sales@simplygj.com WhatsApp +65 8120 0130 25°C • in Singapore
All articles

llms.txt: what it is, who reads it, and whether you need one

1 Oct 2026 8 min read By SimplyGJ

Short answer

llms.txt is a Markdown file at the root of your site that gives AI tools a short summary of who you are and a curated list of your most useful pages. Jeremy Howard proposed it in September 2024, and it’s still a proposal, not a standard. Google says Search ignores it, and OpenAI, Anthropic and Perplexity don’t list it in their crawler documentation as a way to get into their answers. It costs about an hour to write, so add one if you like, but don’t expect it to move your visibility on its own.

llms.txt is a plain Markdown file at yoursite.com/llms.txt that tells AI tools what your site is about and which pages matter. It’s cheap to make. The benefit is unproven.

This guide is for owners and marketing leads in the US, the UK and Australia who’ve heard they “need an llms.txt for AI search”. We’ll cover what the proposal says, what Google and the AI companies have actually published, how to write one in under an hour, and how it differs from robots.txt.

What is llms.txt?

llms.txt is a proposed convention for giving large language models a clean, curated map of a website. It was published at llmstxt.org by Jeremy Howard on 3 September 2024. A second version of the proposal went up on 10 August 2026, which Howard says reflects two years of adoption.

The problem it solves is simple. Web pages are full of navigation, scripts and cookie banners. An AI tool with a limited context window has to dig through all that. A short Markdown file pointing at your best pages saves it the digging.

The proposal is explicit about the use case. Howard writes that he expected the file to be most useful at inference time, when an assistant is answering a question, rather than for training a model. That matters, because it means llms.txt only helps when a tool goes looking for it.

The format, line by line

An llms.txt file is Markdown with a fixed order. Only the first element is required.

  1. An H1 with the name of the site or project. This is the only mandatory part.
  2. A blockquote with a short summary that contains the key facts.
  3. Optional paragraphs or lists with more detail, but no headings.
  4. Zero or more H2 sections, each holding a list of links. Each item is a Markdown link, optionally followed by a colon and a note.
  5. An H2 called “Optional” for secondary links a tool can skip when it’s short on space.

The proposal also suggests clean Markdown copies of key pages, at the same URL with .md added. That suits documentation sites. For most business sites it’s optional.

Who reads llms.txt, and who has said they don’t

This is the part most guides skip. Here’s what each company has published, checked in October 2026.

Company What it has published about llms.txt What actually controls access
Google Search Says you don’t need it, and that Search ignores it Googlebot, robots.txt, normal indexing
Google Chrome (Lighthouse) An optional “agentic browsing” audit checks whether the file exists Not a ranking signal
OpenAI Its crawler docs and publisher FAQ don’t mention it for ChatGPT search OAI-SearchBot via robots.txt
Anthropic Its crawler help article doesn’t mention it ClaudeBot, Claude-User, Claude-SearchBot via robots.txt
Perplexity Its crawler docs don’t mention it PerplexityBot via robots.txt and IP allowlists

Google is the clearest. Its guide to optimising for generative AI features says you don’t need new machine-readable files, AI text files or Markdown to appear in Google Search. A note added in June 2026 goes further. It says creating llms.txt files for other services is fine, but doing so “will neither harm nor help” your visibility in Google Search, “as Google Search ignores them”.

Chrome tells a slightly different story. Lighthouse now has an llms.txt audit in its agentic browsing category. If the file returns a 404, the audit is marked not applicable, “as providing the file is optional at the moment”. It only flags you if the server errors when fetching it. Google’s Chrome team says that without the file, AI agents “may spend more time crawling the site”. That’s a browser-agent point, not a search ranking one.

OpenAI, Anthropic and Perplexity are quieter. Each one publishes an llms.txt for its own developer documentation. But their crawler pages for site owners (OpenAI, Anthropic, Perplexity) talk only about robots.txt user agents and IP ranges. None of them says its search crawler reads llms.txt or uses it to choose sources.

So publishing the file is a convention some AI tools may use. It isn’t a published ranking or citation factor for any major AI search product.

Should you add an llms.txt file?

Yes, if it takes you an hour and you’ll keep it accurate. No, if someone is charging you real money for it or selling it as the fix for AI visibility.

The case for it is cost. It’s a few kilobytes of text. It does no harm in Google, by Google’s own statement. It gives any agent that does look for it a correct summary of your business, written by you.

The case against it is opportunity cost. If your site blocks OAI-SearchBot at the firewall, has thin service pages, or isn’t indexed in Bing, an llms.txt file changes nothing. Those problems decide whether AI tools can find and quote you. We cover them in how to rank in ChatGPT and how to rank in AI Overviews.

Our view: llms.txt is a nice-to-have that sits at the bottom of a GEO checklist. In our experience, the sites that turn up in AI answers got there through crawler access, clear answer-first pages and third-party mentions. We’ve never seen an llms.txt file be the reason a business started getting cited, and we wouldn’t promise a client that it will.

How to write an llms.txt file

You can write one by hand in a text editor. Keep it short and factual.

  1. Write the H1 with your business name.
  2. Write a two or three sentence blockquote: what you do, where, for whom, and any fact you’d want an AI tool to get right, such as prices or locations.
  3. Add an H2 for your main services, linking each service page with a one-line note.
  4. Add an H2 for company pages: about, contact, locations.
  5. Add an H2 for your best guides or resources.
  6. Put anything secondary under “Optional”.
  7. Upload it to the root so it loads at yoursite.com/llms.txt as plain text.

On WordPress, you can upload the file by SFTP, use a plugin, or generate it from your theme. Generating it is better, because a hand-written file goes stale the moment you add a page.

Worked example: a hypothetical accountancy firm in Leeds

This firm is invented. Notice the file is mostly links, with the facts a model might get wrong in the summary.

# Harrow & Lane Accountants

> Harrow & Lane is a chartered accountancy firm in Leeds, UK, working with landlords, sole traders and small limited companies. We handle Making Tax Digital for Income Tax, self assessment, year-end accounts and payroll.

## Services
- [Landlord accounts](https://example.co.uk/landlords): MTD-ready bookkeeping and tax returns for property owners
- [Limited company accounts](https://example.co.uk/limited-companies): year-end accounts, corporation tax, director payroll

## Company
- [About us](https://example.co.uk/about)
- [Contact](https://example.co.uk/contact)

## Optional
- [Blog](https://example.co.uk/blog)

Our own example

simplygj.com publishes one at /llms.txt. It has an H1, a blockquote with our location, team size, the markets we serve and our starting prices, then H2 lists for services, company pages and guides. Our WordPress theme generates it, so a new service or article appears in the file without anyone remembering to update it. That’s the main thing we’d copy.

llms.txt vs robots.txt

robots.txt controls crawling. llms.txt describes content. They do different jobs, and one can’t stand in for the other.

robots.txt is a real standard, published as RFC 9309 in September 2022. Crawlers that follow it read its rules before fetching pages. Google’s robots.txt introduction says it tells crawlers which URLs they can access, mainly to avoid overloading your site. Even RFC 9309 says the rules “are not a form of access authorization”.

llms.txt has no allow or disallow rules at all. You can’t block a bot with it. If you want to stop AI training while staying in AI search, that’s a robots.txt job, using the separate user agents each company publishes.

robots.txt llms.txt sitemap.xml
Purpose Tells crawlers what they may fetch Gives AI tools a curated summary and key links Lists URLs for search engines to discover
Status Internet standard (RFC 9309) Community proposal Widely supported protocol
Format Plain text directives Markdown XML
Can it block a bot? Yes, for bots that comply No No
Read by Google Search? Yes No, per Google Yes

Quick answers

Should I pay an agency to create an llms.txt file?
Not as a standalone job. It’s an hour of work and should come free inside proper GEO work, not as a product.

What should I fix before llms.txt?
Crawler access for OAI-SearchBot, Claude-SearchBot and PerplexityBot, Bing indexing, and service pages that answer buyer questions in plain words.

How do I know if AI tools mention my business?
Ask them, the same questions, every month. Our free AI visibility check runs 10 fixed questions across ChatGPT, Perplexity, Gemini and Google AI Overviews.

If you want the bigger picture, read GEO vs SEO and our guide to LLM SEO. For a site that’s built to be read by people and AI tools from day one, see our website development work, or get in touch if you’re outside Singapore and want to see how we work with international clients.

AI platforms change their guidance often. Statements here were checked in October 2026.

Questions people ask

Does llms.txt help SEO?

Not in Google. Google's AI optimisation guide says Search ignores llms.txt files and they neither help nor harm rankings.

Does ChatGPT read llms.txt?

OpenAI hasn't said it does. Its documentation for site owners says to allow OAI-SearchBot in robots.txt and your firewall, and doesn't mention llms.txt.

What's the difference between llms.txt and llms-full.txt?

llms-full.txt is a community extension that puts the full text of key pages into one file. It isn't in the core llmstxt.org format, and it suits documentation sites far more than business sites.

Can llms.txt stop AI companies using my content?

No. It has no blocking rules. Use robots.txt with the specific AI user agents, and your CDN's bot settings, to control access.

Will an llms.txt file hurt my site?

Not in Google Search, which ignores it. The only real risk is a stale file that gives AI tools out-of-date prices or services.

WhatsApp