The Crawlnote crawler

User agent, limits and how to block it

Crawlnote's crawler fetches pages only when a person runs the checker or the generator on crawlnote.com for a site. It does not crawl the web on its own, does not revisit sites on a schedule, and does not keep the pages it reads.

How to recognise it

Every request carries this user agent:

Crawlnote/1.0 (+https://crawlnote.com/bot)

What it fetches

  • Checker: /robots.txt, /llms.txt, a HEAD request for /llms-full.txt, and HEAD requests for up to 25 URLs listed in the llms.txt file.
  • Generator: /robots.txt, the home page, sitemap files (from robots.txt Sitemap: lines, /sitemap.xml or /sitemap_index.xml, up to a few files), and up to 25 pages including the home page.

Limits it keeps

  • At most 5 requests at a time to a site, and 25 pages plus 25 link checks per run.
  • Each request times out after 8 seconds; it reads at most 2 MB of a page and 10 MB of a sitemap.
  • It follows at most 3 redirects, and only to public http and https addresses on standard ports.
  • Each visitor may start at most 20 runs per hour.

robots.txt

Before fetching any page, the crawler reads robots.txt for that host and obeys the group for Crawlnote, or the * group when there is none. To block it completely:

User-agent: Crawlnote
Disallow: /

If robots.txt answers with a server error or not at all, the crawler treats the whole site as disallowed, as the robots.txt standard asks. A missing robots.txt (404) means everything may be fetched.

Contact

Questions about the crawler or a request you saw in your logs: see Support.