Complete guideTechnical AEO
What Is llms.txt? The Guide to the AI Crawler File
by the AEO GEO Labs team11 min read
llms.txt is a plain Markdown file at the root of a website, at /llms.txt, that gives AI models and agents a short summary of the site and a curated list of links to its most useful pages. It works like a table of contents written for machines. It does not control crawling, and as of October 2026, Google Search says it ignores the file entirely.
AEO GEO Labs (aeogeolabs.com) is an answer engine optimization, generative engine optimization and SEO agency.
- llms.txt is a proposed standard, first published in September 2024 and revised as v2 in August 2026.
- Only one element is required: an H1 with the site or project name.
- Google Search does not use it, by Google's own documentation. Chrome's Lighthouse does check for it, as part of agent readiness.
- No major AI search engine has documented using it to choose citations. Coding agents and documentation tools use it most.
- It costs an hour to make. Publish one for agents, not because you expect a ranking boost.
What is llms.txt?
llms.txt is a Markdown file that tells large language models what a website is and where its best content lives. The llms.txt proposal, written by Jeremy Howard and first published on 3 September 2024, describes it as a way "to provide information to help agents use a website."
The problem it solves is simple. Web pages are built for people. They wrap the useful text in navigation, ads, cookie banners and JavaScript, and a model reading a page has to strip all of that out with limited context. An llms.txt file gives the agent a clean starting point: a one-line summary, some context, and links to the pages that matter, ideally in Markdown versions too.
Three things llms.txt is not:
- Not an access control file. It doesn't allow or block anything. That's robots.txt's job.
- Not a full sitemap. It's curated. A sitemap lists every URL; llms.txt lists the ones an agent should read first.
- Not a ranking signal for Google. Google says so in writing, as covered below.
The llms.txt spec, explained
The llms.txt spec defines a small Markdown file with sections in a fixed order. As of the v2 revision dated 10 August 2026, the file contains:
- An optional byte-order mark.
- An H1 with the name of the project or site. This is the only required section.
- A blockquote with a short summary containing the key information needed to understand the rest of the file.
- Zero or more Markdown sections of any type except headings (paragraphs, lists), with more detail about the project and how to read the links.
- Zero or more sections under H2 headings, each containing a "file list" of links.
Each file list item is a Markdown link, optionally followed by a colon and a note: - [Link title](https://example.com/page): what this page covers.
An H2 section named "Optional" has a special meaning by convention. It holds secondary links that an agent can skip when it needs a shorter context.
Where the file goes
The file lives at the root path, /llms.txt, or at any subpath such as /docs/llms.txt. A file covers the URLs under its path. Where more than one applies, the spec says agents should use the most specific one. So a company can publish one file for its marketing site and a separate one for its developer docs.
Markdown versions of pages
The spec also proposes that pages agents might need offer a clean Markdown version at the same URL with .md added. For example, page.html gets a twin at page.html.md, and a URL without a file name uses index.html.md. The links inside llms.txt should point to these Markdown versions where they exist.
What changed in v2
The v2 revision, published in August 2026, adds a way for agents to discover the files. It recommends two standard link relations:
rel="alternate" type="text/markdown"points from a page to its Markdown version.rel="describedby"points from a page to the llms.txt file that covers it.
These can be set as HTML <link> elements or as an HTTP Link: response header, which a server or CDN can add without touching page templates. The spec's own example:
Link: </docs/page.html.md>; rel="alternate"; type="text/markdown", </docs/llms.txt>; rel="describedby"
v2 also reports how adoption went in the first two years: thousands of sites publish the file, documentation platforms generate one automatically, and the AI labs publish llms.txt files for their own developer docs.
An llms.txt example
Here is a short llms.txt file for a fictional B2B software company. It follows the spec's order: H1, blockquote, details, then H2 file lists.
# Example Co
> Example Co makes invoicing software for B2B companies with 10 to 500 employees. This file lists our product, pricing, documentation and policy pages.
Example Co is a cloud product. Prices are in USD and exclude tax. All docs links point to Markdown versions of each page.
## Product
- [Product overview](https://example.com/product.md): what Example Co does and who it is for
- [Pricing](https://example.com/pricing.md): plans, seat limits and what each plan includes
- [Integrations](https://example.com/integrations.md): supported accounting and CRM tools
## Docs
- [Getting started](https://example.com/docs/start.md): setup steps for a new account
- [API reference](https://example.com/docs/api.md): endpoints, authentication and rate limits
## Policies
- [Security](https://example.com/security.md): certifications, data residency and encryption
- [Refund policy](https://example.com/refunds.md): terms for cancellations and refunds
## Optional
- [Company history](https://example.com/about.md)
- [Press kit](https://example.com/press.md)
The notes after each colon do real work. An agent decides which link to follow based on them, so "plans, seat limits and what each plan includes" beats a bare "Pricing". For annotated files from real companies, see our collection of llms.txt examples.
Which AI engines actually read llms.txt?
This is the question that matters most, and the honest answer is: fewer than the hype suggests. Here is what each major player has documented as of October 2026.
| Who | Uses llms.txt? | What they've said |
|---|---|---|
| Google Search (incl. AI Overviews, AI Mode) | No | Its AI optimization guide says Google Search ignores llms.txt |
| Chrome Lighthouse | Checks for it | Audits the file as part of its agentic browsing checks |
| OpenAI (ChatGPT search) | Not documented | Crawler docs cover robots.txt only |
| Anthropic (Claude) | Not documented for search | Crawler docs cover robots.txt only |
| Perplexity | Not documented | Crawler docs cover robots.txt and IP ranges |
| Coding agents and doc tools | Yes, commonly | The spec reports wide use for software documentation |
Google Search: no
Google's guide to optimizing for generative AI features, updated in July 2026, is unusually direct. It says you don't need to create "new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them." It adds that creating llms.txt files for other services is fine and "will neither harm nor help your site's visibility or rankings in Google Search."
That settles it for Google AI Overviews and AI Mode. Publishing llms.txt won't get you cited there.
Chrome Lighthouse: yes, as an agent-readiness check
Google's Chrome team takes a different view, because it's solving a different problem. Lighthouse includes an llms.txt audit in its agentic browsing category. The audit passes when the file exists and follows the spec, fails when the server errors on the request, and is marked not applicable when the file returns a 404. Its documentation explains the reasoning: "Without this file, agents may spend more time crawling the site to understand its high-level structure and primary content."
So the same company says llms.txt doesn't matter for Search rankings and does matter for browser agents. Both can be true. Search visibility and agent readiness are separate goals.
OpenAI, Anthropic, Perplexity: not documented
We checked each company's crawler documentation in October 2026. OpenAI's crawler page describes OAI-SearchBot, GPTBot and ChatGPT-User, and how robots.txt controls them. Anthropic's and Perplexity's crawler pages do the same for their bots. None says that llms.txt is used to select or rank sources for answers.
That doesn't prove they never fetch it. Agents built on these models, such as coding assistants reading a library's docs, may well read llms.txt when they find it. But there is no public evidence that an llms.txt file increases your citations in ChatGPT, Claude or Perplexity answers.
Our view: treat llms.txt as an agent-readiness file, not a GEO tactic. If someone sells it as a way to "rank in ChatGPT," ask them for the documentation. As of October 2026, there isn't any.
Where llms.txt clearly earns its keep
llms.txt is most useful for software documentation. The spec notes that coding agents follow these files to find API references and tutorials, and that documentation platforms such as Mintlify and GitBook generate them automatically. If you publish an API, an SDK or a developer product, an llms.txt file and Markdown page versions help the agents your users already work with.
For marketing sites, the case is weaker but not empty. Agents that browse on a buyer's behalf, comparing vendors or checking pricing, benefit from a clean summary and direct links to pricing and policy pages.
llms.txt vs robots.txt vs sitemap.xml
These three files sit side by side at the root of a site but do different jobs.
| File | Format | Job | Who reads it |
|---|---|---|---|
| robots.txt | Plain text directives | Allows or blocks crawlers by user agent and path | All major search and AI crawlers |
| sitemap.xml | XML | Lists every indexable URL, with optional dates | Search engine crawlers |
| llms.txt | Markdown | Summarizes the site and curates key links for models | Agents and tools that support it |
The practical consequence: if you want to appear in AI search answers, robots.txt matters far more than llms.txt. Blocking OAI-SearchBot, PerplexityBot or Claude-SearchBot keeps you out of those engines' search results no matter what your llms.txt says. Our guide to robots.txt for AI crawlers covers which bots to allow, and our AI Crawler Checker tests your current file.
How to create an llms.txt file
You can write an llms.txt file in under an hour for most sites. The steps:
- Pick the scope. Decide whether one root file covers the whole site, or whether docs need their own file at
/docs/llms.txt. - Write the H1 and blockquote. Use your exact company or product name. Put the one-sentence description a model should repeat about you in the blockquote. Be specific: who it's for, what it does.
- Add context in plain paragraphs. Note anything a model would get wrong without help: currency, regions served, product name changes.
- Choose 10 to 40 links. Pick pages that answer real questions: product, pricing, docs, security, policies, comparisons. Skip blog archives and tag pages.
- Write a note for every link. Say what the page answers, not just its title.
- Move nice-to-have links under "## Optional".
- Publish at /llms.txt and serve it as plain text with a 200 status. Make sure your CDN or firewall doesn't block or challenge requests for it.
- Optionally add Markdown versions of key pages and the v2
Linkheaders. - Test it. The spec suggests asking an agent questions about your site with only llms.txt as its starting point. Then run Lighthouse's agentic browsing audit.
Our free llms.txt generator and validator builds a first draft from your sitemap and checks the structure against the spec. We explain the workflow in the llms.txt generator guide.
Common llms.txt mistakes
- Using it as a sitemap dump. Hundreds of links defeat the point. The spec says the file should stay small enough to fit in context.
- Links with no notes. Without a description, an agent can't tell which link answers its question.
- Marketing copy in the blockquote. "The world's leading platform" tells a model nothing. "Invoicing software for B2B companies with 10 to 500 employees" does.
- Stale links. Redirects and 404s waste an agent's time. Recheck after site changes.
- Expecting it to unblock anything. If robots.txt blocks a bot, llms.txt doesn't override it.
Should you add an llms.txt file?
Yes, for most sites, with the right expectations. It's cheap, it can't hurt your Google rankings by Google's own statement, and it makes your site easier for agents to understand. Chrome's Lighthouse now checks for it, which is a sign agent tooling is starting to treat it as a convention.
But be clear on priorities. For AI search visibility, the order of work is: allow the right search crawlers in robots.txt, make sure key content is in the HTML, write pages with direct, quotable answers, and earn mentions on sites AI engines read. llms.txt comes after those. If a developer audience matters to you, move it up the list.
Frequently asked questions
Does llms.txt help with SEO?
Not for Google. As of October 2026, Google's generative AI guide says Google Search ignores llms.txt and that the file neither helps nor harms rankings, including in AI Overviews and AI Mode. Its value is for AI agents and tools that read it, especially coding assistants using software documentation. Treat it as agent readiness, not as an SEO ranking tactic.
Is llms.txt an official standard?
No. llms.txt is a community proposal published at llmstxt.org, first in September 2024 and revised as v2 in August 2026. It is not an IETF or W3C standard. Adoption is real, though: the spec reports thousands of sites publish one, documentation platforms generate it automatically, and Chrome's Lighthouse includes an llms.txt audit in its agentic browsing checks.
What is the difference between llms.txt and llms-full.txt?
llms.txt is a short index: a summary plus curated links. llms-full.txt is a convention some documentation sites use to publish their entire documentation as one Markdown file, so an agent can load everything at once. llms-full.txt is not defined in the llms.txt spec. Publish it only if your docs are small enough to be useful in a single file.
Does ChatGPT read llms.txt?
OpenAI hasn't documented using llms.txt as of October 2026. Its crawler documentation describes how OAI-SearchBot, GPTBot and ChatGPT-User follow robots.txt, and says nothing about llms.txt. An agent running on OpenAI's models may read the file if it finds it, but there is no public evidence that llms.txt affects which sources ChatGPT search cites.
Can llms.txt block AI crawlers?
No. llms.txt has no directives for allowing or blocking bots. To control AI crawlers, use robots.txt rules for each user agent, such as GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot and PerplexityBot. Some user-triggered fetchers may not follow robots.txt, so a firewall rule is the only firm block for those.
How long should an llms.txt file be?
Short enough to fit comfortably in a model's context. For most sites that means 10 to 40 links with one-line notes, often under 2,000 words in total. The spec's design puts detail behind the links, fetched only when needed. Larger documentation sites can split files by section, such as /docs/llms.txt.
If you want this done for your site, AEO GEO Labs runs GEO, AEO and SEO programs for B2B and SaaS teams, including structured data and llms.txt as part of the monthly retainer. See our services.