About SiteUpwardBot
Last updated: September 29, 2026
SiteUpwardBot is the crawler of SiteUpward. It fetches web pages only when a user asks SiteUpward to audit a specific website. It is not a search engine, does not build a public index, and the pages it fetches are not used to train AI models.
User agent
SiteUpwardBot/1.0 (+https://siteupward.com/bot/)What it does
- It runs only when a user starts an audit of a specific website. Free checks fetch up to 25 pages; full audits up to 500 pages and large audits up to 2,000 pages.
- It reads robots.txt and sitemaps and fetches HTML pages of the same website. It does not follow links to other websites, submit forms, log in or try to access non-public areas.
- It loads a small sample of pages with JavaScript enabled, the way a visitor's web browser does, to check whether content depends on JavaScript. During these few page views, the page's own scripts and resources load as usual – including third-party scripts such as analytics – so they may record a visit. You can filter these visits by the user agent above.
- It briefly requests your homepage with the user agents of well-known AI and search crawlers to test whether your CDN or firewall blocks them (see below). Those test requests come from our servers, not from the AI companies.
How it limits load
- At most 2 requests per second and 2 parallel connections per website, shared by all audits.
- It honours Crawl-delay in robots.txt, up to 10 seconds between requests.
- It slows down when your server answers 429 (Too Many Requests) or 503 (Service Unavailable), and stops after repeated errors.
- Only one audit crawls a website at a time, and it fetches at most 6,000 pages per website per day.
- A website is crawled for a free check at most once every 24 hours; further free checks reuse the recent results. Full audits of the same website start at most once per hour and at most 4 times in 24 hours.
- It downloads at most 5 MB per page and gives up on a page after 15 seconds.
How to opt out
SiteUpwardBot follows robots.txt (RFC 9309) for the product token “SiteUpwardBot”. To stop it from crawling the pages of your site, add:
User-agent: SiteUpwardBot Disallow: /
You can also disallow only certain paths. If your robots.txt cannot be loaded because of a server error, SiteUpwardBot treats the whole website as disallowed.
What still happens after you opt out
A Disallow rule stops SiteUpwardBot from crawling your pages. It does not stop every request: when someone audits your domain, a few requests still reach your site so the report can tell them whether AI crawlers can access it:
- one request for /robots.txt, to read your rules;
- one request for your homepage with a normal web-browser user agent, as a baseline;
- up to 8 requests for your homepage with the user agents of well-known AI and search crawlers (for example OAI-SearchBot, PerplexityBot, Claude-SearchBot and Googlebot), to see whether your CDN or firewall blocks them.
Full exclusion
These test requests come from our servers, are paced by the same per-site rate limits, and happen at most once per audit. If you want no requests at all from us, email support@siteupward.com with your domain and we will get back to you about excluding it from future audits.
What we store
For the audit report we store page addresses, titles, headings, short excerpts and technical measurements, as described in our Privacy Policy. We do not keep full copies of your pages.
Contact
If you think SiteUpwardBot is misbehaving on your site, email support@siteupward.com with your domain and the approximate time, and we will investigate promptly.