Read

lms.txt: What the Data Actually Says

There’s a file being sold quite hard at the moment. It’s called llms.txt, it goes at the root of your domain, and the pitch is that it helps AI systems understand your site and therefore recommend you more often.

We add them for clients. We also tell those clients, plainly, that the evidence for the second half of that pitch is close to non-existent.

Here’s why.

What it actually is

llms.txt was proposed in 2024 by Jeremy Howard, co-founder of Answer.AI and fast.ai. It’s a single markdown file at your domain root that summarises what your site is and links to its most important content — the idea being that a language model can orient itself without crawling everything.

That’s it. Two things it is frequently confused with, and isn’t:

  • It is not a robots.txt-style directive. Despite the filename, it controls nothing and blocks nothing.
  • It is not the practice of publishing markdown copies of all your pages. That’s a separate tactic with separate problems.

The “AI visibility” framing came later, attached by the SEO industry on the speculation that AI platforms would reward the file. Not on evidence that they do.

The numbers

In June 2026, Ahrefs published a study of 137,210 domains — server logs and live bot traffic, not guesswork. The findings are worth reading in full, but here are the ones that matter:

  • 28% of those domains publish an llms.txt file. (Ahrefs notes their user base skews technical, so treat that as an upper bound.)
  • 97% of those files received zero requests in May 2026. Not “few requests”. Zero. Nothing fetched them at all.
  • Of the 3% that were fetched, 96% of requests came from bots — and 77% of those bots weren’t AI tools. The top category was SEO audit crawlers.
  • AI retrieval bots — OAI-SearchBot, PerplexityBot, Claude’s search crawler — accounted for 1.1% of requests.
  • Slackbot fetched llms.txt files more often than PerplexityBot did. Link previews in chat apps are generating more llms.txt traffic than the AI search engine the file was ostensibly designed to help.

And the finding that should settle the argument:

Zero requests came from AI bots for llms.txt files that don’t exist.

Ahrefs analysed every request that returned a 404 on /llms.txt. Where valid files drew 96% bot traffic, missing files drew 98% human traffic — SEOs typing the URL to check on competitors. The AI bot share of those 404s was zero.

Nothing goes looking. Publishing the file does not put you on anyone’s radar. Agents fetch it when a link, an index, or a user instruction tells them it exists.

Google is arguing with itself

In late May 2026, Google managed to take both sides of this in under a week.

First, its guide on optimising for generative AI features included a section titled, without irony, “mythbusting” — telling site owners that machine-readable files like llms.txt aren’t neededto appear in generative AI search.

Days later, the Chrome team shipped an llms.txt check inside Lighthouse’s Agentic Browsing audits, with documentation explaining that without the file, agents may spend more time crawling your site to understand its structure.

When Lily Ray pressed John Mueller on the contradiction, his answer was clarifying: llms.txt is “not done for search.” He called it a “temporary crutch, perhaps to save some tokens”for AI coding tools parsing developer documentation — not something non-developer sites need to worry about.

That’s Google’s search relations lead describing the file as a token-saving convenience for coding agents. Which, as it happens, is exactly what the data shows.

Who actually reads it

The largest identifiable AI consumer in the Ahrefs data wasn’t a search bot at all. It was AI agents and agentic infrastructure — 10.5% of requests. Claude-Code, Anthropic’s coding agent, out-fetched every AI retrieval bot, every AI assistant and every AI training crawler.

So there is a real audience. It’s just not the audience the file is being sold on.

If your customers use coding agents to source recommendations, or if agents genuinely act on your site, llms.txt stands a real chance of being read. If you’re a membership body, a professional services firm or a local business hoping to turn up in ChatGPT more often — it won’t do that.

The bit nobody mentions

The Ahrefs study turned up a research crawler identifying itself as prompt-injection-survey/1.0. Someone is systematically studying llms.txt as a prompt injection opportunity — because agents are designed to ingest and trust this file.

That’s worth sitting with. A stale or compromised llms.txt misleads every agent that reads it, and it’s a plain text file that a lot of people treat as a set-and-forget tick box.

If you publish one, treat it like code: version control it, restrict who can edit it, keep the content to plain links and descriptions rather than anything instruction-shaped, only link to resources you control, and review anything a platform auto-generates on your behalf.

So should you have one?

Yes — with clear eyes about why.

Reasons to add it: it takes ten minutes, Lighthouse’s Agentic Browsing audit checks for it, platforms like Wix already generate them automatically and this will likely become a CMS default within a year, and if the agentic web does end up mediating AI search, the file may matter more later than it does now.

Reasons not to oversell it: 97% base rate of zero readership, no measurable effect on AI citations, nothing goes looking for it, and it carries a small but real security consideration.

Our position is straightforward. We’ll add one to your site as part of the technical work, and we’ll link to it properly so it stands some chance of being found. We won’t charge you a strategy fee for it, and we’ll spend your money on the things that actually move: your accessibility tree, your schema markup, your layout stability, and getting your pricing and FAQs out of images and into readable text.

If someone is selling you an llms.txt package as an AI visibility strategy, ask them what percentage of these files get read. If they don’t know, they haven’t looked.


Sources: Ahrefs — We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read · Chrome for Developers — Lighthouse llms.txt audit · llmstxt.org

Want an honest assessment of your site? That’s what our Agent Readiness Audit is for.