Skip to content
Approvalens

Reading room · 6 min read

What Is llms.txt? Format, Examples and Whether You Need One

llms.txt is a proposed Markdown file that gives AI agents a short map of your site. The format, a working example, common errors and what it does not do.

By the Approvalens team

Fixes these report findings

  • llms.txt
  • llms.txt follows the format

llms.txt is a plain Markdown file at /llms.txt that gives AI agents a short, curated overview of your site: a title, a one-paragraph summary and lists of links to your most useful pages. It was proposed by Jeremy Howard in September 2024 and is maintained as an open proposal at llmstxt.org. It is not a crawling permission file like robots.txt, and none of the major AI companies says it's needed to appear in their answers. It's cheap to add, useful for agents that look for it, and safe to skip if you have other priorities.

What it is for

The proposal's reasoning is simple: web pages are built for people, wrapped in navigation, ads and scripts, and an agent with a limited context window does better with a concise, expert-written summary and links to clean text. In the words of the spec, llms.txt information "is instead used on demand, when an agent needs information about a topic while assisting a user". The author's expectation "was that llms.txt would mainly be useful for inference rather than training".

How llms.txt relates to the files you already have:

File Audience Purpose
robots.txt Crawlers What they may fetch
sitemap.xml Search engines Every indexable URL
llms.txt AI agents and assistants A short, curated guide to the site, with links to the pages worth reading

The spec is explicit that a sitemap is no substitute: it lists everything, often too much to fit in a model's context, and doesn't point to clean text versions of pages.

The format

From the llmstxt.org spec, in this order:

  1. An H1 with the site or project name. This is the only required part.
  2. A blockquote with a short summary containing the key facts needed to understand the rest.
  3. Optional paragraphs or lists (no headings) with more detail.
  4. Optional H2 sections containing "file lists": Markdown lists where each item is a link [name](url), optionally followed by : and a note.

An H2 called ## Optional has a special meaning: links an agent can skip when it needs a shorter context.

A working example

# Shrimp Tank Notes

> Practical care guides for freshwater dwarf shrimp (Neocaridina and Caridina):
> water parameters, tank setup, feeding and breeding. Written by one keeper
> since 2019; every guide is tested in our own tanks.

Prices and product links are examples, not recommendations from sponsors.

## Start here

- [Water parameters for Neocaridina](https://shrimptanknotes.com/water-parameters): pH, GH, KH and TDS ranges with the reasons behind them
- [Cycling a new shrimp tank](https://shrimptanknotes.com/cycling): step-by-step, with test results by week

## Care guides

- [Feeding shrimp](https://shrimptanknotes.com/feeding): what, how much and how often
- [Molting problems](https://shrimptanknotes.com/molting): causes of failed molts and what to change

## Optional

- [About the author](https://shrimptanknotes.com/about)
- [Contact](https://shrimptanknotes.com/contact)

Serve it as plain text at the root of the site. Approvalens publishes its own at /llms.txt if you want to see a live one.

Where it can live

Version 2 of the proposal allows the file at the root (/llms.txt) or at any path, such as /docs/llms.txt, covering only the pages below it. When more than one applies, agents "should use the most specific one". The spec also suggests Markdown versions of pages at the same URL plus .md, and Link headers with rel="describedby" pointing to the llms.txt that covers a page.

Common errors

Problem Why it happens Fix
/llms.txt returns your HTML homepage or a "not found" page with status 200 The CMS or framework routes unknown paths to a page Serve a real file, or return a 404 if you don't have one
No H1 on the first line Started with a blockquote or a list Put # Site name first
Links are relative (/guide) Copied from internal navigation Use full URLs, as the spec's examples do
The file is huge, listing every post Generated from the sitemap Pick the pages that answer most questions; the spec wants it small enough to fit in context
Links point to redirected, noindexed or deleted pages File written once and never updated Regenerate it when you publish or remove key pages
Served with Content-Type: text/html Server default for unknown extensions Serve as text/plain or text/markdown
Blocked by robots.txt or a firewall for AI user agents A broad bot block An agent that can't fetch the file can't use it

Does llms.txt help you appear in ChatGPT, Claude or Google?

Honest answer: there's no official statement that it does.

  • Google says for AI Overviews and AI Mode: "You don't need to create new machine readable files, AI text files, or markup to appear in these features" (AI features and your website).
  • OpenAI, Anthropic and Perplexity document their crawlers and robots.txt tokens (OpenAI, Anthropic, Perplexity). Their crawler pages don't mention llms.txt as a signal.
  • llmstxt.org reports that thousands of sites publish one, that documentation platforms generate it, and that the AI labs publish llms.txt files for their own developer docs. That shows adoption by publishers, not that a search crawler ranks you by it.

So treat llms.txt as a courtesy to agents that read it, mostly coding assistants and tools that fetch a site on a user's behalf. It doesn't replace an open robots.txt, a firewall that lets AI search bots through, or content that's in the HTML. Those decide whether AI tools can read your site at all; the allow AI crawlers in robots.txt guide covers them.

Generating one

llmstxt.org lists tools that create the file for you, including the Yoast SEO and AIOSEO plugins for WordPress, and Wix, which generates one for every Wix site. On a custom site, generate it from the same data as your sitemap or navigation so it can't drift out of date; that's how ours is built.

llms.txt and AdSense

They're separate. Nothing in Google's documentation connects llms.txt to AdSense; the review looks at your pages, your policies and whether its own crawler can reach the site. A free Approvalens scan covers what the AdSense review does look at, and the full report checks up to 20,000 pages.

Check your llms.txt

The free AI crawler checker fetches /llms.txt, tells you whether it exists and whether it follows the basic llmstxt.org format, and shows whether AI bots can reach your pages at all. Note that our fetch uses a simulated user agent from our server, so a firewall that verifies bots by IP may answer real AI agents differently.

FAQ

Is llms.txt a standard?

It's a proposal published on llmstxt.org with a public GitHub repository and community input, now in a second version. It isn't an IETF or W3C standard like robots.txt (RFC 9309).

Can llms.txt block AI training?

No. It has no allow or disallow rules. Use robots.txt tokens such as GPTBot, ClaudeBot and Google-Extended for that.

What is llms-full.txt?

A convention some sites and tools use for a single file containing the full text of the linked pages. It isn't part of the core format described on llmstxt.org, so name and contents vary by generator.

Will a missing llms.txt hurt my rankings?

Not according to any search engine's documentation. Google says no AI text files are needed for its AI features.

Spotted something out of date or wrong? Tell us and we'll correct it.

Read this guide in Turkish →

Check it on your own site. Free, no sign-up.

Free tools for this

Free scan

Check your own site

Free scan: readiness score and every issue, usually in a few minutes.

Free scan · score and every problem found · no sign-up

All guides →