Skip to content
← Field notes

Field notes

What is llms.txt? How to write one for a store site, and how to check it is really there

llms.txt is a note for AI: “this is what the site is, read these pages.” You can do without it. If you publish one, make sure AI actually sees it.

Published llms-txtai-crawlerec

What llms.txt is

A text file named llms.txt at the root of your site. It is written in Markdown: the site's name, a short summary, and links to the pages you want AI to read. Jeremy Howard proposed it in 2024; v2 of the proposal is at llmstxt.org.

According to the spec, it is used mostly at inference time — when an AI is looking something up for a user — rather than for training. The AI reads llms.txt, then follows the links to the pages it needs, because HTML pages are built for people and pulling clean text out of navigation and ads is hard.

How it differs from robots.txt

The two get mixed up, so first: robots.txt says where a crawler may go; llms.txt says what matters. Writing something in llms.txt cannot block an AI crawler. To block one, use robots.txt (see how to check whether robots.txt blocks GPTBot).

robots.txtllms.txt
JobWhere crawlers may and may not goWhat the site is and which pages to read
FormatUser-agent, Disallow and Allow linesMarkdown (headings and link lists)
If missingEverything may be readNo problem (optional)

How to write one (a store example)

The order is fixed. Only the first heading is required; the rest is optional.

  1. One # heading with the site or store name. This is the only required part
  2. A > blockquote with a short summary
  3. Paragraphs without headings, if you need them
  4. ## sections, each a list of links: one [name](URL) per line, optionally followed by : and a short note
# 山田茶園 オンラインショップ

> 静岡の自社茶園の煎茶とほうじ茶を売っています。注文は会員登録なしでもできます。

送料は全国一律、5,000円以上で無料です。

## 商品

- [煎茶の一覧](https://example.com/collections/sencha): 産地と収穫時期つき
- [ほうじ茶の一覧](https://example.com/collections/hojicha)

## ご注文について

- [送料とお届け日数](https://example.com/pages/shipping)
- [返品と交換](https://example.com/pages/returns)

## Optional

- [茶園の紹介](https://example.com/pages/about)

(The example is a Japanese tea shop: a sencha and hojicha store, shipping and returns pages, and an “about” page.) The final ## Optional section has a set meaning: links an AI can skip when it needs a shorter context. For a store, put links that answer pre-purchase questions — product lists, shipping, returns — near the top.

You will also see llms-full.txt. That name is not in the spec itself; documentation platforms like Mintlify use it for a single file with the full text of every page. A store site is fine with llms.txt alone.

Check that it is really there

  1. Open https://your-domain/llms.txt in a browser
  2. If the Markdown you wrote shows up as plain text, it is there. Check that the # heading is the first line (after any blank lines)
  3. If you see your homepage or a “page not found” screen, it is not there — see trap 1 below

Three ways an uploaded llms.txt counts as missing

1. /llms.txt returns an HTML page

Some servers answer a missing file with the homepage or an error screen and a success status (200). Something shows up, so it looks uploaded, but the body is an HTML document, not llms.txt. Our free check also counts it as missing when the body starts with <!DOCTYPE html> or <html>. Some cart and site builders do not let you put arbitrary files at the root; check your platform's help for how to add a root file.

2. It lives only under a subpath

The spec allows a file under a path such as /docs/llms.txt, which then describes only the pages under that path. For a site-wide guide, use the root /llms.txt; Chrome's Lighthouse docs also name the root (https://example.com/llms.txt) as the place to put it. Our free check also looks only at the root /llms.txt and /llms-full.txt.

3. Using llms.txt to say “no”

Writing “do not read this site” in llms.txt blocks nothing. Allowing and blocking is robots.txt's job. llms.txt is a guide for the AI that does come to read.

You can do without it

llms.txt is optional for now. Chrome's Lighthouse has an llms.txt audit, but a missing file (404) is marked not applicable; it is flagged when fetching the file hits a server error. Check first that robots.txt is not blocking AI crawlers and that your page has no noindex. llms.txt can come after that.

Check all three at once

Paste your store's URL into the form at the bottom of this page. It shows whether the root llms.txt and llms-full.txt exist (and are not an HTML page), whether eight AI crawlers can read your homepage, and whether that page has noindex — right away. No signup, no email. It reads only public files and the page itself.

Try it on your own site

Paste a URL and see right away: whether AI crawlers are let in, whether there is a file written for AI, and whether the page is marked to stay out of search — using public pages only.