llms.txt: what it is, an example, best practices, and whether AI actually uses it · Citable

View this page

What an llms.txt file is, what goes in it, a working example, how it differs from robots.txt and a sitemap, how to add one on WordPress or Shopify, and what it does and does not do for AI search.

Price
USD 59
Organisation
Citable

Description

written for teams no score out of 100 no ten-step checklist written for teams no score out of 100 no ten-step checklist written for teams no score out of 100 no ten-step checklist

What is llms.txt?

llms.txt is a plain text file, written in Markdown, that sits at the root of a website — yoursite.com/llms.txt — and tells a language model what the site is and which pages matter most. It was proposed in September 2024 by Jeremy Howard of Answer.AI. The idea is simple: a web page is built for people and is mostly layout, so give a model a short, clean map instead.

It is a proposal, not a standard. Nobody enforces it and no crawler is obliged to look for it. That matters for what you can expect from it, so it is worth keeping in mind through the rest of this page.

An llms.txt example

The format is fixed and short: a title, a one-paragraph summary, then lists of links.

# Acme Bikes

> Acme Bikes sells and services commuter bicycles in Pune.

> Prices, stock and service slots are on the pages below.

## Products

- [City commuter](https://acme.example/bikes/city): 7-speed, ₹24,000, in stock

- [Folding bike](https://acme.example/bikes/fold): 20-inch, ₹31,500

## Policies

- [Returns](https://acme.example/returns): 14 days, unused

- [Warranty](https://acme.example/warranty): 2 years on the frame

## Optional

- [Our story](https://acme.example/about) The # line is the name. The > lines are the summary. Each ## section is a list of links with a short description after the colon. A section called Optional marks links a model can skip when it is short of room.

llms.txt vs robots.txt vs a sitemap

robots.txt is permission: which crawlers may read which paths. See the AI crawler list for the names that go in it.

sitemap.xml is discovery: every URL you want found.

llms.txt is a summary: the few pages that explain you best, with a line about each.

They do different jobs. llms.txt replaces neither of the others.

Does ChatGPT, Claude or Google use llms.txt?

This is the honest part. No major AI search engine has said that its crawler reads llms.txt when it decides what to cite, and Google’s search team has said publicly that it does not use the file. The crawlers that feed AI answers — OAI-SearchBot , PerplexityBot , Claude-SearchBot — ask for your pages themselves, one URL at a time.

Where llms.txt does get read is by agents that are pointed at it: coding assistants reading a product’s documentation, a person pasting the file into a chat, a tool built to look for it. That is real, and it is why documentation sites adopted it first. It is not the same as ranking in ChatGPT.

One sign of where it is heading: Chrome’s Lighthouse now checks for llms.txt among its agentic-browsing audits. It treats the file as optional — a site without one is marked not applicable, not failed — and describes it as an emerging convention for LLMs and AI agents. That is browser tooling preparing for agents that browse on a person’s behalf. It is not a statement that Google Search or Gemini decides what to show by it.

An llms.txt file costs ten minutes and does no harm. It is not what decides whether an AI answer quotes you — the page the crawler actually fetches is.

llms.txt best practices

Keep it short. Ten to thirty links. It is a map, not a second sitemap.

Put facts in the descriptions. A price, a limit, a date — the line after the colon is often all that gets read.

Link to pages that say it in text. A link to a page whose content is drawn by JavaScript sends the model to an empty room.

Keep it true. A file that says ₹24,000 while the page says ₹26,000 is worse than no file.

Serve it as plain text , at the root, without a login or a redirect chain.

How to add llms.txt on WordPress, Shopify and other platforms

On WordPress , the major SEO plugins can generate the file for you, or you can upload one to the root of the site beside robots.txt . On a Next.js or static site it is a file in the public folder, or a route that returns text. On hosted platforms that do not let you place a file at the root — Shopify , Wix, Squarespace — something else has to serve it: an app, or the CDN in front of the site.

A generator can write the first draft from your sitemap. Read it before you publish it: a generated list of two hundred links with no descriptions is the version that helps nobody.

What works on the crawlers that never ask for llms.txt

The problem llms.txt set out to solve is real: a page is hundreds of kilobytes of markup around a few hundred words of fact, and most AI crawlers do not run the scripts that draw the rest. A summary file helps the readers that ask for it. It does nothing for a crawler that asks for the page.

That is the gap Citable closes. When an AI crawler asks for one of your pages, it is answered with the content alone — the same facts, at the same address — while people and Google get the page as it is. No crawler has to know about a special file. You can see what yours are given today with the free scan , and read how the clean copy is served .

The chain, drawn Access, then comprehension, then the choice.

Access — an engine you refuse never reads anything.

Comprehension — the facts have to survive the markup.

Choice — which stays the model’s, whatever anyone sells you.

Read next The rest of the set.

6 min read AI search optimization

How an assistant decides what to say about your site, and which of those steps a team can actually influence.

7 min read What GEO can and cannot do

What generative engine optimisation means, what it borrowed from SEO, and which parts of it are sold as certainty but aren't.

5 min read How AI crawlers read a website

Rendering, fetch budgets, robots.txt, and why a JavaScript-drawn price is invisible to almost every engine reading you.

8 min read Generative engine optimization (GEO)

What it means, how an AI engine chooses its sources, the strategies with evidence behind them, and what GEO tools actually do.

7 min read Answer engine optimization (AEO)

What it means, three before-and-after examples, the practices that matter — and the step that comes before any of them.

6 min read How to get cited by ChatGPT

How ChatGPT picks the pages it cites, and the six things a site can actually do to show up in its answers.

5 min read Can ChatGPT read my website?

How it reads a site, why it does not run JavaScript, the five things that stop it, and how to check in a minute.

5 min read utm_source=chatgpt.com, and tracking AI traffic in GA4

What the tag on your URLs means, why ChatGPT adds it, and how to see every AI engine in Google Analytics.

6 min read Should you block AI crawlers?

Training bots, search bots and live fetch are three different decisions — and the one blocked by accident costs the most.

5 min read GEO vs AEO vs AIO vs LLMO

The full form and meaning of each, where the terms came from, and the four steps underneath all of them.

Or just look at it One page, served both ways.

What survives each trip, with the token count for both.

About Citable Who is behind this page, and what we stand on.

When an AI agent — ChatGPT, Claude, Perplexity or any of the 81 that Citable recognises — opens a page, Citable removes the theme, scripts, navigation and layout and serves only the content: the same facts, at the same URL, on the site's own domain. People and Google still get the full designed page.

Who opens the page What they are given

A person in a browser The full designed page, untouched.

Googlebot, Bingbot and other search engines The full designed page, untouched.

An AI agent — GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot and the rest The content only: no UI, no scripts, the same facts at the same URL.

Measured, not estimated. A real product page of 386.8 KB is served to AI agents as 4.6 KB — 98.8% smaller, with the same facts. Every figure on this site comes from a fetch anybody can repeat.

One list. The 81 AI agents named on these pages are the list the product itself runs on. A page here cannot describe a crawler Citable does not serve.

Search is never touched. Googlebot and Bingbot always get your real page. The clean copy is for AI agents only — at the same address, with the same facts.

Checked, not trusted. A user agent is a claim, and anyone can send one. Where an operator publishes its addresses, every read is checked against them and recorded as verified or not.

Nothing about your visitors. Citable sees which crawler read which page. It does not see who your customers are, and their details never reach us.

Honest about limits. Nobody can promise a citation or a ranking, and we do not. What we show is what an AI agent was given, and whether it sent somebody back.

How it works · The measurements · What it touches and stores · Why Citable · Pricing · Questions, answered plainly · Talk to us

Free · about 20 seconds · no signup Find out whether AI has ever read your site.

Type your address. We fetch a page exactly as GPTBot does, count the tokens, and score fifteen things an answer engine needs. If the log then stays empty for a month, that is worth knowing too.

No contract · cancel any time · Googlebot untouched