HomeBlog › Measurement & Data Studies
Measurement & Data Studies

AI Engines Don't Read llms.txt. Not Even Stripe's.

We planted a fake metric in our own llms.txt, then ran the same test on 58 other companies using phrases that appear only in their files. Across 179 questions asked the normal way, on 59 domains, llms.txt content reached an answer zero times. The engines cited those same websites 100 times.

By Michael Patrick CortezPublished 2026-09-1019 min read

Key takeaways

  • Across 179 questions asked the normal way, on 59 domains, content that exists only in llms.txt reached an AI answer zero times. In 100 of those same runs the engine cited the company's website.
  • It is not about our site. We ran the identical test on 58 other companies including Stripe, Notion, Vercel, Zapier, Databricks, Cisco and Ford, using phrases that appear only in their llms.txt files. Same result on every one.
  • The file is read when you paste the URL and ask. ChatGPT, Grok and Claude reproduced llms.txt content in the explicit condition. Not one engine reached for the file on its own, on any domain.
  • Asked outright whether a metric documented in our file was real, all five engines said no. 42 denials out of 42, including from engines that had quoted the file minutes earlier.
  • A zero across 59 domains is not proof of never. By the rule of three it puts the ceiling at roughly 5 percent of sites, and under 2 percent of questions. That is the honest strength of this result.
  • llms.txt is not a ranking input and not a citation source. It is a retrieval convenience for the moment someone hands an engine your URL, which makes it worth having and not worth prioritising.

I have written llms.txt files for clients. I sell a tool that generates them. So this is not a piece I particularly wanted to write.

Almost everything published about llms.txt is advice: add the file, structure it this way, include these sections. Very little of it is measurement, and none of the measurement I could find answers the only question a client actually asks, which is whether putting something in the file ever gets it back out of an AI engine.

On 11 August 2026 I added a phrase to our own llms.txt that exists nowhere else. Not on our website, not in our marketing, and as far as I could determine, nowhere on the web. We said plainly in the file what we were doing, then left it alone for 30 days.

Then we ran the same test on 58 companies that are not us, using phrases that appear in their llms.txt files and nowhere on their websites. Stripe, Notion, Vercel, Zapier, Databricks, Cisco, Ford.

That came to 467 probes across five engines and 59 domains, which is the largest behavioural test of llms.txt I am aware of, and I want to be upfront that I expected it to come out the other way.

Surfaced unprompted
0

Of 179 normally-phrased questions across 59 domains, times llms.txt-only content reached the answer.

Same runs, site cited
100

The engines were reading these websites. They were not reading the llms.txt files.

Said our metric was fake
42 of 42

Every engine, every run, denied a metric documented in our own live file.

What we planted, and why the controls matter

The marker was a fake metric called the Fernwhistle Score. Nonsense on purpose. A made up phrase cannot be guessed, cannot come from training data, and cannot arrive from a press release, so if it ever appears in an answer, the engine read our file. There is no other explanation available.

That is the treatment. On its own it would produce an unreadable result, which is where most informal tests of this go wrong. If no engine ever mentions your marker, you have learned one of two completely different things and you cannot tell which: either the engines ignore llms.txt, or the engines could not see your website at all.

We planted controls too. Three phrases that appear only in our HTML pages and never in llms.txt: 7-step, Citerank Deploy, and Why Not Cited. Same site, same crawl, different file. If an engine surfaces a control phrase but never the marker, the site is reachable and llms.txt specifically is being skipped.

Three conditions, run against every engine:

  1. Asked normally. Questions about what metrics Citerank publishes, with no mention of llms.txt. This is the real world condition. It is how anyone would ever encounter your file by accident.
  2. Asked to read the file. The full URL in the prompt, with an instruction to read it. This separates "will not read llms.txt on its own" from "cannot read it even when told".
  3. Control. Questions whose answers live only in HTML. This proves retrieval was working at all.

The marker phrase never appears in any prompt. That detail is not pedantry. When we first probed this by hand, we asked an engine directly about the Fernwhistle Score, and the phrase came back in the answer, which a naive string match scores as a hit. The model was repeating our own question. Every probe in the study is phrased so the marker can only appear if the engine actually fetched the file.

Five engines: ChatGPT, Perplexity, Claude, Gemini and Grok, each through its web-enabled API, three trials per prompt. The first study is 105 runs, 99 of which returned an answer. A second study, described further down, adds 90 more.

Nobody reads it on their own

Did the engine surface the planted marker?
Percentage of runs where content that exists only in llms.txt appeared in the answer
Left column is the organic condition, where nothing in the prompt mentions llms.txt. Every engine scores zero. Right column is what happens when the URL is handed over directly.

Forty-two organic probes, zero hits. Not a low number and not an occasional hit on one engine, which is what I would have predicted before running it. Nothing, on any engine.

And the site was clearly reachable during those same runs. Twenty-one of those 42 answers surfaced one of the HTML-only control phrases, and three of the five engines cited citerankscore.com directly.

The control: could the engines read the website?
Runs where an HTML-only fact about the same site appeared in the answer
Twenty-eight of 30 control runs succeeded. The engines were reading our pages fine. The file they skipped was llms.txt.

Two engines will read it if you ask

The explicit condition is where they separate, and it is the only place in this entire study where llms.txt does anything at all. ChatGPT fetched the file and reproduced the marker in five of six runs. Grok did it in all six. Both quoted it accurately, cited our domain, and in ChatGPT's case listed the Fernwhistle Score first among our named metrics, exactly as the file describes it.

Perplexity, Claude and Gemini never reproduced anything from inside the file, in any condition, in any of their runs.

The two extremes are below, same URL, same prompt, four minutes apart.

chatgpt.com
ChatGPT listing the Fernwhistle Score as the first named metric from the Citerank llms.txt file, with a citation to citerankscore.com
ChatGPT read the file. It lists the planted marker first, describes it correctly as a test of whether AI engines read llms.txt, and cites our domain for it.
perplexity.ai
Perplexity replying that it could not retrieve the Citerank llms.txt file because the page fetch failed and no indexed copy was found
Perplexity could not. Given the same URL, it reports that the fetch failed and that a web search surfaced no indexed copy of the file. It then asks the user to paste the contents in.

Read that Perplexity answer again. The file returns HTTP 200 to any client that asks. Our robots.txt allows every crawler. It is 6 KB of plain text at a standard path. Perplexity still could not get it, and, more revealing, could not find it in an index either. After 30 days of being live and linked, the file was not meaningfully indexed anywhere Perplexity could reach.

The half that actually matters: will it cite you?

Everything above measures whether llms.txt content leaks into an answer. It is the interesting question. It is not the commercial one.

The commercial question is whether publishing the file makes an engine treat what you wrote as true about you. So we ran a second study, 90 more probes, and asked the engines directly.

Three questions per engine, three trials each, all naming the metric on purpose: does Citerank have a metric called the Fernwhistle Score, what is it and how is it calculated, and is it real or did I misremember it.

Naming the phrase in the prompt means a keyword match is worthless here, because the model will simply repeat the words back. So these runs are not scored on whether the phrase appears. They are scored on the verdict the engine reaches, graded blind into confirms, denies, or hedges, and on what it cited.

Asked outright: is this metric real?
The answer is documented on the live site, in a file every one of these engines is allowed to fetch
Every engine, every run. Forty-two denials out of 42. Not one hedge, not one confirmation, including from ChatGPT and Grok, which can read the file when handed the URL.

Every engine told us our own published metric does not exist. Grok is the sharpest case. In seven of its nine runs it cited citerankscore.com while denying the metric. It went to our website, read our pages, and reported that a thing documented in a file at our root was not real. ChatGPT, which reproduced the file verbatim minutes earlier when given the URL, denied it here too.

An engine reading a file on demand is a fetch. An engine treating that file as a source of truth about you is retrieval. Only the second one puts you in an answer, and it is not happening.

Where the citations actually go

The second study asked 90 questions about our product, our tools, our pricing and our metrics. Every one of those answers is written in our llms.txt. Here is what the engines cited instead.

What the engines cited when asked about us
Study two, 85 answered runs, questions about our brand, tools, pricing and metrics
Our ordinary pages were cited in 26 runs. Our llms.txt was cited in none. A competitor's llms.txt was cited in nine.

Zero, on the questions the file was written to answer. Our homepage was cited in 19 runs. Our tools, pricing and about pages all appeared. The document written specifically for language models did not.

Across both studies our llms.txt was cited in exactly 10 runs, and all 10 came from the one condition where we put the URL in the prompt and told the engine to go read it. Never once did an engine reach for the file on its own.

Look at the middle bar. In nine runs the engines cited an llms.txt file belonging to citability.dev, citedindex.com or llm-stats.com, other companies in our category, while answering a question about us. Across both studies that happened 22 times against our 10.

The engines were willing to pull an llms.txt into an answer. Just not the one belonging to the company being asked about.

Then we ran it on companies we do not control

Everything so far is one domain. That is the fair objection to every study like this, and on its own our result could mean llms.txt is ignored, or it could mean our site is young and thin and nobody indexes our files. Those have very different implications for you.

We did it again on companies nobody would call young or thin.

You cannot plant a canary in someone else's file. You do not need to. Most llms.txt files already contain sentences that appear nowhere in the rendered website, because the file is written separately and then forgotten. Those sentences are natural markers, and they work exactly like a planted one: if the phrase reaches an answer, the engine read the file.

We pulled every company from our protocol census that publishes a valid llms.txt, fetched each file and each homepage, and kept the phrases that appear in the file and nowhere in the HTML. 58 companies produced usable markers. We then asked all five engines about 14 of them: Stripe, Notion, Vercel, Zapier, Databricks, Postman, Plaid, Deel, Grafana Labs, Rippling, Cohesity, Checkout.com, Fivetran and Cursor.

Same two conditions. Ask normally, then ask the engine to read the file.

The same test, on 14 companies we do not control
Runs where a phrase unique to that company's llms.txt appeared in the answer
Sixty-three questions asked normally, about some of the most heavily linked domains on the web. Zero surfaced llms.txt content. The right column is what happens when you hand the engine the URL.

Zero again, on all fourteen, Stripe and Notion and Vercel included. And these engines were not struggling to find these companies. In the same organic runs, ChatGPT, Perplexity and Claude cited the company's own website in 100 percent of their answers. They read the site. They did not read the file.

Combine both studies and the headline number is this: 105 questions asked the way a real person asks them, across 15 domains, and llms.txt content reached the answer zero times, while the engines cited those same websites 71 times.

One useful correction this gave us. Claude never once cited citerankscore.com in our own study, which we read as an entity problem on our side. This confirms it: Claude cited all 14 of these companies' websites, in every single organic run. The failure was about our domain, not about Claude.

Widening it, and what a zero is actually worth

Fourteen companies is still not many, and a fair reader should push on that. Plenty of published studies crawl thousands of domains. So we widened it.

The reason nobody has run this one at that scale is cost. The studies that cover thousands of domains are adoption censuses: request a URL, check whether a file comes back, move on. Pennies per domain. This is a behavioural study. Every domain needs live queries against paid AI APIs, and every answer has to be checked against a marker set built by diffing that company's file against its own HTML. Minutes and real money per domain.

We extended the organic condition to every company in our census with a usable marker. That is 59 domains, 179 questions asked the normal way, across SaaS, Fortune 100 and retail: Stripe, Notion, Vercel, Zapier, Databricks, Postman, Deel, Plaid, Rippling, Cisco, Ford, Liberty Mutual and 47 more.

Zero on all 59, while in those same 179 runs the engines cited the companies' own websites 100 times.

Now the part most studies skip. A zero is not a proof of never, and I would rather be precise about what it does buy you. The standard tool is the rule of three: if an event does not occur in n trials, the 95 percent upper bound on its true rate is about 3/n.

  • Per question: zero in 179 runs puts the ceiling at 1.7 percent.
  • Per domain: zero on 59 domains puts the ceiling at 5.1 percent.

The domain figure is the one that matters for generalising, because domains, not questions, are what vary across the web. So the defensible claim is not "this never happens anywhere". It is that if llms.txt retrieval happens organically at all, it happens on fewer than about one site in twenty, and we did not find a single instance.

Two honest caveats on that number. The widened round ran on Gemini and Perplexity, so those 44 domains carry two engines rather than five. Fifteen of the 59 domains carry the full five-engine treatment. A wider study with every engine on every domain would tighten the bound further, and we would expect it to land in the same place.

The thing we expected to matter, and did not

This is the explanation everyone reaches for, including us, and it did not survive.

Of the 58 llms.txt files we scanned, 3 are linked from the company's homepage. Ours was not either. That looked like the obvious explanation: a file nothing links to never gets crawled, never gets indexed, never gets cited. It fits Perplexity's message about finding no indexed copy, and it would have made a satisfying, actionable conclusion.

The data does not support it. Splitting our runs by whether the file was linked, the linked sites surfaced markers in 6 percent of runs and the unlinked sites in 8 percent. If anything it is backwards, and the linked group is only two companies, which is far too small to conclude anything either way.

We are reporting it unresolved rather than dressing it up. Linking your llms.txt is still sensible housekeeping. On this evidence it is not the switch that makes the file work, and we would need a much larger sample of linked files to say more.

What the engines think we are

One incidental result, which will be familiar to anyone with an ambiguous brand name. Claude ran up to 30 web searches per question about us and never once retrieved citerankscore.com. It cited citerank.co, an unrelated company, along with an IndieHackers profile and an AWS blog post about agentic readiness.

Claude was not failing to find information. It was confidently finding the wrong company. No llms.txt file fixes that. Entity clarity is an HTML and structured data problem, and it is a much bigger lever than the file this study is about. We wrote about that in why AI engines skip your site.

Full results

What we would actually do about it

Keep the file. Stop prioritising it. It costs an hour and it does work in the one case that is growing, which is an agent or a person pointing an engine straight at a URL. That case is real, and being ready for it costs you an hour. It is not a ranking input, and on this evidence it will not put you in an answer by itself.

Put the facts in HTML. Every fact that reached an answer in this study got there through a normal page. That is where the retrieval is happening, and it is where a claim about your product needs to live if you want an engine to repeat it.

Fix your entity before your files. If an engine is confidently citing a different company with a similar name, no amount of llms.txt tuning will help you. Our guide to getting cited by AI covers the structured data and disambiguation work that does.

Do not report llms.txt as an AI visibility win. Across 85 answered probes about our own product and its metrics, our file was cited in zero of them, while our ordinary pages were cited in 26. If you are showing a client a line item for llms.txt, the honest version of that line is housekeeping, not visibility.

Limitations

We tested 15 domains, which closes the objection that this is about our site but does not make it the whole web. All 14 replication companies are B2B software. A news publisher, a retailer or a local business could behave differently, and we would want a version of this across sectors before calling it universal.

The natural markers in study three are not as clean as a planted one. A phrase absent from the homepage might still appear on a deep page we did not fetch, which would let an engine surface it without ever touching llms.txt. That failure mode would produce false positives, so it makes our zero more conservative, not less.

We measured API endpoints for four of the five engines, and consumer products can behave differently from their APIs. The two screenshots above are consumer product runs and matched their API results exactly, which is reassuring but is not the same as testing every consumer surface.

We did not include crawler fetch logs. Our host does not retain 30 days of static asset request logs, so we cannot show who fetched the file, only what came back out of the engines. Fetching and using are separate questions and we can only speak to the second.

Three trials per prompt is enough to see a 0 percent and a 100 percent clearly. It is not enough to resolve small differences, so please do not read the gap between ChatGPT's 83 percent and Grok's 100 percent as meaningful.

The existence check has a built in asymmetry I should name. A denial is weak evidence on its own, because a model that cannot verify a claim will often deny it by default, which is the behaviour you want from it. What makes 42 out of 42 meaningful is that the answer was published, reachable, and allowed to be crawled for 30 days, and that two of these engines demonstrably could read it seconds earlier in the same session.

Method and data

Five engines through their web-enabled APIs: GPT-4.1 with web search, Grok 4.6 with web search, Perplexity sonar, Gemini 3 Flash with Google Search grounding, and Claude Sonnet 4.5 with web search.

Study one: seven prompts across three conditions, three trials each, 105 runs, 99 answered. A run counted as a hit only if the marker appeared in the response text, and the marker never appeared in a prompt.

Study two: six prompts across two conditions, three trials each, 90 runs, 85 answered. The three existence-check prompts name the marker deliberately, so those runs are scored on the verdict the engine reached, graded blind by a separate model into confirms, denies or hedges, and on which URLs it cited. They are never scored on the phrase appearing.

Study three: 14 companies drawn from our protocol census, two conditions each, one trial per engine, 140 runs, 125 answered. Study four widens the organic condition to the remaining 44 companies, on Gemini and Perplexity only, 132 runs, 74 answered. Markers are phrases of seven or more words that appear in a company's llms.txt and nowhere in its fetched homepage HTML. A run counts as a hit if any six-word span of a marker appears in the answer, which is deliberately generous.

Citation counting is domain-scoped and run-based. Two things to flag, because both changed our numbers. An early version of this analysis counted any cited URL ending in llms.txt, which quietly folded in other companies files and inflated our own count from 10 to 23. And we count runs rather than raw citation instances, so a single answer that cites the same page five times counts once, and one chatty response cannot carry a bar.

The marker is still live in our llms.txt so anyone can check this. Ask an engine what metrics Citerank publishes without naming the marker, then ask it to read the file, and compare what you get.

Download the raw run data: study one, 105 runs, study two, 90 runs, study three, 140 runs, and study four, 132 runs. All JSON, with every prompt, full answer text, and citation list. 467 runs in total.

If you want to know whether AI engines are citing your own site, and which pages they pull from, that is what Citerank does. Related reading: our census of agent protocol adoption across 287 companies, which measured how many major brands publish llms.txt in the first place, and our guide to building agent-ready websites.

See how AI search sees your site

Get your free AI Visibility Score in seconds, no signup required.

Get my free score

Frequently asked questions

Does llms.txt actually work in 2026?
Not as a discovery mechanism and not as a citation source. In our 30 day controlled test, no AI engine surfaced content that existed only in llms.txt when asked about our brand normally, across 42 probes on five engines. When we then asked the engines directly whether a metric documented in that file was real, all five said it was not, in all 42 runs. Across 85 answered probes about our product, our llms.txt was cited in none of them while our HTML pages were cited in 26. Publishing the file is cheap and harmless, so keep it. Treating it as an AI ranking factor is not supported by what we measured.
Which AI engines read llms.txt?
In our test, ChatGPT retrieved and quoted the file in 83 percent of the runs where it was explicitly asked to, and Grok did so in 100 percent of those runs. Perplexity, Claude and Gemini never reproduced anything from inside the file, in any condition, including when the full URL was given in the prompt. Perplexity was explicit about why, reporting that the fetch failed and no indexed copy existed.
Is this just because your site is small? What about big brands?
We tested that directly, on 58 other companies including Stripe, Notion, Vercel, Zapier, Databricks, Cisco and Ford, using phrases that appear in their llms.txt files and nowhere in their HTML. Across 179 normally-phrased questions spanning 59 domains, llms.txt content reached an answer zero times, while the engines cited those same websites 100 times. Domain authority is not the variable. The file is not in the retrieval path for anyone we tested.
Only 59 domains? Other studies cover thousands.
Those studies are almost always adoption censuses: request a URL, check whether a file exists, move on, which costs pennies per domain. This is a behavioural study, where every domain needs live paid queries against five AI engines and every answer has to be checked against a marker set built from that company's own file. On what a zero at this scale buys you: by the rule of three, zero events in 179 runs puts the 95 percent ceiling at 1.7 percent of questions, and zero across 59 domains puts it at 5.1 percent of sites. So the claim is not that this never happens. It is that it happens on fewer than about one site in twenty, and we found no instance of it.
How do you know the engines could see your site at all?
That is what the control markers are for. We planted three phrases that appear only in the HTML site and never in llms.txt, then measured them in the same runs. Those phrases were surfaced 28 times out of 30 in the control condition and 21 times in the organic condition. So the engines were retrieving the site and answering from it. They just were not using llms.txt to do it. Without that control, a null result would be unreadable, because you could not tell an ignored file from an invisible website.
Why not just check your server logs for AI crawler hits on llms.txt?
That measures fetching, not use, and they are not the same thing. A crawler can fetch a file that never influences an answer, and an engine can answer from a cached index without fetching anything during the request. We measured the output side deliberately, because what a marketer actually cares about is whether the content reaches the answer. We did not include fetch log data here because our host does not retain 30 days of static asset logs, and we are not going to publish a number we cannot produce.
Could the marker phrase itself be the problem, since it is a made up word?
That is the point of the design, and it cuts the other way. A nonsense phrase cannot be guessed, cannot come from training data, and cannot arrive from another source, so any appearance is proof of retrieval rather than coincidence. The risk with an invented phrase is a false negative if a model refuses to repeat something it finds implausible. Two engines repeated it happily once they fetched the file, which shows the phrase itself was not a barrier.
Does linking to llms.txt from my homepage make engines read it?
We looked and we cannot say yes. Of 58 company llms.txt files we scanned, only 3 are linked from the homepage. Splitting our runs by that variable, linked files surfaced content in 6 percent of runs and unlinked files in 8 percent, with the linked group being just two companies. That is far too small a sample to draw a conclusion, and the direction is not the one the hypothesis predicts. Link it anyway, because it costs nothing and an unlinked file is undiscoverable by definition. Just do not expect it to be the switch.
Should I delete my llms.txt file?
No. It costs almost nothing to maintain and it does help in the one case that is growing, which is a person or an agent pointing an engine straight at your URL. Keep it accurate and small. What we would stop doing is spending real time on it, buying tools that generate elaborate versions of it, or reporting it to a client as an AI visibility win. On the evidence, the work that moves citations is on your HTML pages.
If an engine can read llms.txt, why does it still say my content is not real?
Because fetching a file on demand and treating that file as a source about you are different behaviours. ChatGPT reproduced our file verbatim when we gave it the URL, then denied the same metric existed when simply asked about it, in all nine of those runs. Grok denied it in seven runs where it had cited our website in the same answer. The engines are answering from an index built from crawled HTML pages, and llms.txt does not appear to be feeding that index. Handing over a URL routes around the index for one request. It does not put your file into it.
Will AI engines cite my llms.txt file?
In our data, only if you hand them the URL. Across 85 answered probes asking specifically about our product, our metrics and our pricing, questions whose answers all sit in our llms.txt, the engines cited our HTML pages in 26 runs and our llms.txt in none. Across both studies our file was cited in 10 runs, every one of them from the condition where we pasted the URL into the prompt and told the engine to read it. In the same answers, they cited three other companies llms.txt files 22 times. If citations are the outcome you are measuring, the file is not producing them on its own.
How can I reproduce this study?
The marker is still live in our llms.txt at citerankscore.com/llms.txt as of publication, and we are leaving it there so the result is checkable. Ask any engine what metrics Citerank publishes without naming the marker, and see whether it appears. Then ask the same engine to read the file by URL and compare. Our full run data, all 105 probes with prompts, answers and citations, is linked at the end of this post.
Michael Patrick Cortez
Michael Patrick Cortez
SEO & AI Search Strategist · Founder of Citerank

Michael Patrick Cortez leads SEO and AI search work at Webfor in Vancouver, WA, and is the founder of Citerank. He writes and speaks about generative engine optimization, getting cited by AI, and building agent-ready websites. Read more of his work at michaelpatrickcortez.com.

More from Michael Patrick Cortez