Field notes
How to check whether robots.txt really blocks GPTBot, and three ways it looks blocked but isn't
“We opted out of AI training, I think.” Checking takes one look at robots.txt — as long as you know the one rule that makes a block look real when it isn't.
There is exactly one place to look: robots.txt at the root of your store's domain. Open https://your-domain/robots.txt in a browser; anyone can read it.
Read it in three steps
- Open
https://your-domain/robots.txt. If it does not exist (404), nothing is blocked: under the robots.txt standard (RFC 9309), a crawler may read any page when robots.txt is unavailable. - Search the page for
GPTBot. If there is aUser-agent: GPTBotline, theDisallowandAllowlines right below it are GPTBot's rules. - If
GPTBotis not there, read theUser-agent: *group. Crawlers that are not named follow*.
A typical block looks like this:
User-agent: GPTBot
Disallow: /Disallow: / means the whole site. A narrower rule such as Disallow: /admin/ leaves every other page readable.
Looks blocked, isn't
1. Blocking everything with `*`, then adding a GPTBot group
User-agent: *
Disallow: /
User-agent: GPTBot
Allow: /products/This reads like “block everything, but let GPTBot see product pages”. It does not work that way. Once a group names GPTBot, GPTBot stops looking at the * group. RFC 9309 says a crawler uses the group that matches its name and falls back to * only when none does. Here GPTBot is blocked from nothing, including the homepage.
To block everything except product pages, put Disallow: / in the GPTBot group too:
User-agent: GPTBot
Disallow: /
Allow: /products/Now only paths starting with /products/ are readable. When two rules match a path, the longer one wins.
2. An empty `Disallow:`
User-agent: GPTBot
Disallow:An empty Disallow blocks nothing. It looks like a block and means the opposite. The / tends to get lost when copying an example.
3. A misspelled crawler name
Names are case-insensitive, so gptbot works. A different spelling such as GPT-Bot or ChatGPT names some other crawler, and GPTBot falls back to *. The two names OpenAI lists for robots.txt control are GPTBot (training) and OAI-SearchBot (ChatGPT search).
Blocking GPTBot does not remove you from ChatGPT search
OpenAI says the GPTBot and OAI-SearchBot settings are independent. Disallowing GPTBot signals that your content should not be used for training; allowing OAI-SearchBot keeps you in ChatGPT search results. To leave search as well, disallow OAI-SearchBot by name.
ChatGPT-User is different. It visits a page because a ChatGPT user asked, and OpenAI notes that since the action is user-initiated, robots.txt rules may not apply (OpenAI's crawler docs).
One more thing: robots.txt is a request, not a lock. Only crawlers that read it obey it. Changes also take time; for search, OpenAI says about 24 hours after you update robots.txt.
If reading it by hand is a chore
Paste your store's URL into the form at the bottom of this page. It lists eight AI crawlers, GPTBot included, and whether each can read your homepage — and whether that is because you named it or because it falls under *. No signup, no email. It reads only public files such as robots.txt and the page itself.
How each cart's default robots.txt treats AI crawlers is in our measurement of six platforms.
robots.txt decides what AI may read. llms.txt is the file that tells it what you would like it to read. An example for a store and how to check it is really there are in what is llms.txt.
Try it on your own site
Paste a URL and see right away: whether AI crawlers are let in, whether there is a file written for AI, and whether the page is marked to stay out of search — using public pages only.