llms.txt is worth adding when your site already has clean crawl access, current sitemaps, entity schema, and pages AI models can quote. It is not a proven ranking or citation signal. Treat it as a low-cost map for agents and documentation-heavy sites, not as the mechanism that makes ChatGPT, Google AI Overviews, or Perplexity recommend your brand.
The honest answer is less exciting than most llms.txt pitches: ship it if the implementation takes an afternoon, but do not let it outrank the fundamentals. Google says its AI features do not require new machine-readable files. OpenAI and Anthropic document crawler control through robots.txt, not llms.txt. Anthropic and Perplexity publish their own llms.txt files, which proves the format is useful as a documentation map, not that it moves citations. Parse tracks AI visibility across ChatGPT, Google AI Overviews, and Perplexity. In that measurement model, llms.txt belongs in the technical hygiene layer. Citation strategy still comes from accessible pages, source authority, entity clarity, and the third-party domains AI systems already trust.
- llms.txt is a proposed Markdown file at
/llms.txtthat points AI systems toward your most important resources. - Google explicitly says no new AI text files or special markup are required for AI Overviews or AI Mode.
- OpenAI and Anthropic give site owners crawler controls through robots.txt user agents, not llms.txt directives.
- Documentation platforms and API companies get the clearest near-term value because agents need clean technical context.
- For most brands, llms.txt should come after crawler access, sitemaps, schema, internal links, and citation gap work.
What is llms.txt supposed to do?
llms.txt is a proposed standard from Jeremy Howard's llms-txt project. The file lives at /llms.txt, uses Markdown, starts with an H1, and lists the site's most important resources with short descriptions. The stated purpose is not to block or allow crawlers. It is to give language models and agentic tools a concise map of what a site is, which pages matter, and where machine-friendly versions of those pages live.
That makes the file closer to a curated table of contents than a ranking signal. The proposal says it is mainly useful at inference time, when an assistant or agent is trying to answer a user's question and needs concise context. That is why the strongest early use case is developer documentation. Anthropic and Perplexity both publish llms.txt files for their documentation. Mintlify automatically generates llms.txt and llms-full.txt for docs sites because agents can use those files to discover the right Markdown page before reading the full docs.
What llms.txt does not do
llms.txt does not replace robots.txt, sitemap.xml, schema, or crawlable HTML. It does not grant permission, revoke permission, force indexing, or tell ChatGPT which pages to cite. Ahrefs' May 2026 assessment is blunt: no major LLM provider has formally adopted llms.txt as a crawler protocol, and there is no public evidence that the file improves AI retrieval, traffic, or model accuracy.
The distinction matters because the name invites a false mental model. robots.txt is an access-control signal that well-behaved crawlers can honor. sitemap.xml lists indexable URLs. Structured data labels visible page facts in a machine-readable way. llms.txt is a self-declared guide. A self-declared guide can help an agent navigate a site, but it cannot make a weak source authoritative. If a comparison page has no citations, no author signal, stale claims, and thin prose, adding its URL to llms.txt only makes the weak page easier to find.
How is llms.txt different from robots.txt and sitemaps?
Use the files for different jobs. The mistake is treating llms.txt as "robots.txt for AI." It is not. The llms-txt proposal itself says the two standards are designed to coexist because robots.txt controls acceptable automated access while llms.txt provides context for allowed content.
llms.txt. Curates the pages, docs, policies, or Markdown resources an agent should read first. It explains what matters, not who may crawl.
robots.txt. Tells crawlers whether they may access paths. OpenAI, Anthropic, and Google document AI crawler controls here.
sitemap.xml. Lists canonical URLs for indexing and discovery. It is broad by design, while llms.txt should be selective.
Schema. Labels visible facts on pages. It helps search systems resolve entities and page types when the markup matches the content.
The practical order is simple: make content crawlable, make URLs discoverable, make entities unambiguous, then add llms.txt as a curated guide.
What do Google, OpenAI, Anthropic, and Perplexity actually support?
Google is the clearest. Its AI features guidance says AI Overviews and AI Mode have no additional requirements beyond Search fundamentals, and that site owners do not need to create new machine-readable files, AI text files, special markup, or special schema to appear. That does not make llms.txt harmful. It means Google does not treat it as an eligibility requirement for its AI surfaces.
OpenAI's crawler documentation separates OAI-SearchBot, GPTBot, and ChatGPT-User, then tells site owners to manage the first two through robots.txt. Anthropic's April 2026 help-center article does the same for ClaudeBot, Claude-User, and Claude-SearchBot, including the warning that disabling search bots can reduce search visibility. Perplexity publishes an llms.txt documentation index, but that is evidence of documentation packaging, not a public commitment that Perplexity's consumer answer engine gives ranking weight to your file. The current provider pattern is: robots.txt controls access; llms.txt may help agents read docs.
If you want to see which sources shape AI answers about your brand, run a free brand check — it takes a minute.
When should a brand implement llms.txt?
Implement llms.txt when the file gives an agent a materially clearer path through your site. That usually means you have developer docs, a large help center, product policy pages, technical references, API docs, or a multi-product site where the canonical resources are easy for humans to browse but noisy for agents to prioritize.
For a B2B SaaS company, a useful file might link to product overview, pricing policy, security page, integrations, API docs, comparison pages, and the canonical "about the company" page. For ecommerce, it might link to buying guides, return policy, shipping policy, product taxonomy, size guides, and review-policy pages. For a services firm, it might link to industry pages, case-study index, methodology, leadership bios, and terms. In each case, the value is orientation. You are telling an agent which pages are canonical and how to describe them, while your main AI visibility work still happens on those pages and the external sources citing them.
What should go in the first version?
Keep the first version small enough that your team will maintain it. Start with one H1, one summary blockquote, and four to six sections. Use absolute URLs. Give every link a plain-language description. Link only to pages that are accurate, current, indexable, and useful if read in isolation.
The starter sections are usually: company facts, product or services, documentation or resources, policies, comparison or category pages, and optional background. Do not stuff every blog post into the file. That turns llms.txt into a worse sitemap. Do not include claims you would not put on a public page. Do not add "preferred answers" or marketing copy that contradicts the linked content. Agents and crawlers can check the page itself, and inconsistency weakens trust. If a page needs a clearer description in llms.txt, it probably needs a clearer H1, meta description, and opening section too. Pair this work with schema markup for AI visibility, because both force you to define the entity cleanly.
How should teams prioritize it against higher-impact work?
llms.txt is rarely the first fix. The first fix is usually crawler access. If OAI-SearchBot, Claude-SearchBot, or Perplexity's crawler cannot fetch your site because of robots.txt, CDN policy, or WAF behavior, llms.txt will not matter. The second fix is page quality: AI systems need passages they can quote, compare, and support with evidence. The third fix is source authority: AI recommendations lean heavily on third-party sources, review sites, communities, publications, and comparison pages (Parse's data on the source domains AI cites most).
Use this priority order for a mid-market site: audit robots.txt and edge access; refresh sitemap coverage; fix indexability and server-rendered text; implement Organization and Person schema where relevant; rewrite the pages AI already finds but does not cite; then publish llms.txt. The ordering is not ideological. It follows the failure modes. If the model cannot fetch the page, a map is useless. If it can fetch the page but the passage is weak, a map accelerates disappointment. For the page-level rewrite pattern, use how to structure content so AI models cite it.
How do you measure whether llms.txt helped?
Measure it as a technical experiment, not a faith-based optimization. Before launch, record whether major AI crawlers hit your site, which pages get fetched, which pages appear as citations, and which prompts show source gaps. After launch, track server logs for requests to /llms.txt, requests to linked Markdown or canonical pages, and changes in citation coverage for prompts tied to those pages.
Do not claim success because the file exists. A better success test is narrower: did an agent or crawler request the file, did it request the linked pages afterward, and did citation coverage improve on the prompts those pages answer? If only the first event happens, you improved packaging, not visibility. If citation coverage changes without corresponding crawl or source changes, treat it as noise until repeated. This is where Parse's source-level view matters: AI citation gap analysis tells you whether your problem is your own site, a missing third-party source, or a competitor-owned citation path.
Who should skip it for now?
Skip llms.txt for now if your site has fewer than 20 meaningful pages, if your robots.txt and sitemap are stale, if important content renders only after client-side JavaScript, or if your category pages lack source-backed claims. In those cases, the file is a distraction from problems that already have proven fixes.
Also skip it if the team will treat it as a second website that drifts from the real one. Drift is worse than absence. A stale llms.txt that points agents to deprecated docs, old pricing pages, or retired product names creates the same entity-confusion problem schema is supposed to solve. If you cannot assign an owner, connect the file to your docs or CMS build, or review it quarterly, leave it out until the maintenance path is clear. The right implementation should reduce ambiguity. A neglected one creates another surface for ambiguity to spread.
The decision framework for marketing leaders
The budget decision is not "is llms.txt real or fake." It is "does this remove a bottleneck in our current AI visibility system." For documentation-heavy companies, the answer is often yes because the file improves agent navigation and costs little. For content-led brands with weak third-party source coverage, the answer is usually no because the bottleneck is authority, not orientation.
Use a three-question gate. First, are our critical pages crawlable, indexable, and visible as server-rendered text? Second, do those pages answer revenue prompts with quotable, source-backed sections? Third, would a curated machine-readable map help an agent choose among multiple valid resources? If any of the first two answers is no, fix that first. If all three are yes, publish llms.txt, monitor logs, and keep expectations modest. The win is operational clarity. Any citation lift is something you prove after, not something you assume before.
Frequently asked questions
Does llms.txt improve AI rankings?
There is no public evidence that llms.txt directly improves rankings, recommendations, or citations in ChatGPT, Google AI Overviews, or Perplexity. Google explicitly says no AI text files are required for its AI search features. Treat llms.txt as a navigation aid for agents and documentation systems, then measure whether crawlers request it and whether linked pages gain citations.
Is llms.txt the same as robots.txt for AI?
No. robots.txt controls crawler access for well-behaved bots. llms.txt provides a curated content map. OpenAI and Anthropic document search and training crawler controls through robots.txt user agents, while llms.txt has no comparable enforcement role. Use both only if each has a clear job.
Should ecommerce brands publish llms.txt?
Only after product pages, return policies, shipping pages, review pages, and buying guides are crawlable and current. Ecommerce llms.txt can help agents find canonical policies and product taxonomies, but it will not compensate for missing Product schema, stale Merchant Center data, weak reviews, or thin category content.
What should be in a B2B SaaS llms.txt file?
Start with the company overview, product pages, pricing policy, security page, integration docs, API docs, support docs, comparison pages, and the canonical about page. Keep descriptions factual and short. Exclude low-value blog archives, duplicate landing pages, and any page you would not want an AI answer to treat as authoritative.
How often should llms.txt be updated?
Review it quarterly and after every major site migration, pricing change, product launch, or docs restructure. If you use a docs platform that auto-generates the file, audit the generated descriptions and section ordering. The file is useful only while it points to current canonical pages.
:::