Both files sit at the root of a website and both are read by software rather than people, so they are easy to mix up. They do different jobs: robots.txt says what crawlers may fetch; llms.txt suggests what is worth reading. A sitemap is the third file in the set: it lists every page you want found.
Side by side
| robots.txt | llms.txt | sitemap.xml | |
|---|---|---|---|
| Job | Access rules: which paths a crawler may fetch | A short, curated index of key pages with notes | A complete list of the pages you want found |
| Format | Plain-text directives (User-agent, Disallow, Allow, Sitemap) | Markdown: H1, optional summary, H2 sections of links | XML, one <url> per page |
| Location | /robots.txt | /llms.txt | Usually /sitemap.xml, listed in robots.txt |
| Status | An internet standard (RFC 9309) | A proposal (llmstxt.org); support varies | A long-standing protocol (sitemaps.org) |
| Can it block access? | It asks crawlers not to fetch paths; well-behaved ones comply | No. It only suggests what to read | No |
| Size | Small | Small: the pages that matter | Can list thousands of pages |
llms.txt does not replace robots.txt
Leaving a page out of llms.txt does not hide it, and listing a page there does not grant access to anything robots.txt disallows. If you want a crawler to stay away from part of your site, that rule belongs in robots.txt. If a page is disallowed in robots.txt, it should not appear in llms.txt either; Crawlnote's generator skips such pages.
Do you need all three?
- robots.txt: yes, if you want to set any crawler rules or point to your sitemap.
- sitemap.xml: yes, for any site with more than a handful of pages. Crawlnote's generator starts from it.
- llms.txt: optional. llms.txt is a proposal, not a standard every service follows. Some services read it, others do not, and some major search engines have said they do not use it. Crawlnote makes the file correct; it cannot make anyone read it, and it does not promise rankings or traffic.
How they work together
- robots.txt sets the rules and names your sitemap with a
Sitemap:line. - The sitemap lists every page.
- llms.txt picks the pages that matter from that list and adds a note to each.
That is also the order Crawlnote works in: it reads robots.txt, follows its Sitemap: lines (or /sitemap.xml), and writes llms.txt from the first 25 pages it is allowed to read.