llms.txt vs robots.txt

What each file does, and how they work together with sitemap.xml

Both files sit at the root of a website and both are read by software rather than people, so they are easy to mix up. They do different jobs: robots.txt says what crawlers may fetch; llms.txt suggests what is worth reading. A sitemap is the third file in the set: it lists every page you want found.

Side by side

robots.txtllms.txtsitemap.xml
JobAccess rules: which paths a crawler may fetchA short, curated index of key pages with notesA complete list of the pages you want found
FormatPlain-text directives (User-agent, Disallow, Allow, Sitemap)Markdown: H1, optional summary, H2 sections of linksXML, one <url> per page
Location/robots.txt/llms.txtUsually /sitemap.xml, listed in robots.txt
StatusAn internet standard (RFC 9309)A proposal (llmstxt.org); support variesA long-standing protocol (sitemaps.org)
Can it block access?It asks crawlers not to fetch paths; well-behaved ones complyNo. It only suggests what to readNo
SizeSmallSmall: the pages that matterCan list thousands of pages

llms.txt does not replace robots.txt

Leaving a page out of llms.txt does not hide it, and listing a page there does not grant access to anything robots.txt disallows. If you want a crawler to stay away from part of your site, that rule belongs in robots.txt. If a page is disallowed in robots.txt, it should not appear in llms.txt either; Crawlnote's generator skips such pages.

Do you need all three?

  • robots.txt: yes, if you want to set any crawler rules or point to your sitemap.
  • sitemap.xml: yes, for any site with more than a handful of pages. Crawlnote's generator starts from it.
  • llms.txt: optional. llms.txt is a proposal, not a standard every service follows. Some services read it, others do not, and some major search engines have said they do not use it. Crawlnote makes the file correct; it cannot make anyone read it, and it does not promise rankings or traffic.

How they work together

  1. robots.txt sets the rules and names your sitemap with a Sitemap: line.
  2. The sitemap lists every page.
  3. llms.txt picks the pages that matter from that list and adds a note to each.

That is also the order Crawlnote works in: it reads robots.txt, follows its Sitemap: lines (or /sitemap.xml), and writes llms.txt from the first 25 pages it is allowed to read.