Field notes
Japanese EC platforms name zero AI crawlers in their default robots.txt (six hosts measured)
Can a store owner decide whether AI reads their shop? I measured the stores each platform publishes itself to find out.
Short version: the default store setup names no AI crawler at all. Neither allowed nor disallowed. If an owner does nothing, everything is readable — for training and for AI-search citation alike.
Method: I ran each URL through our Tier 0 check (no signup, no email, public pages only) and looked at how robots.txt treats eight AI crawlers (GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Google-Extended, Applebot-Extended). Measured on September 21, 2026. The six hosts are stores and corporate sites each company publishes itself; no merchant's own shop was measured.
What the six hosts returned
| Host | The eight AI crawlers | llms.txt | sitemap |
|---|---|---|---|
| official.base.shop (BASE official store) | none named (fall under *, readable) | could not fetch | could not fetch |
| shop.stores.jp (STORES official store) | none named | absent | could not fetch |
| demo.shop-pro.jp (Color Me demo store) | none named | absent | absent |
| hydrogen.shop (Shopify official demo) | none named | absent | present |
| www.makeshop.jp (MakeShop corporate site) | GPTBot, ClaudeBot, Google-Extended and Applebot-Extended disallowed by name; the other four allowed | absent | present |
| www.future-shop.co.jp (futureshop corporate site) | none named | absent | present |
“Could not fetch” is not “absent”. The BASE official store rate-limited us (429), and the STORES official store returned 403 for sitemap.xml. Neither can be called missing, so both stay as “could not fetch”. The Color Me demo store redirects its page to err.shop-pro.jp/403.htm, so only its robots.txt is used as evidence. futureshop redirects co.jp to future-shop.jp, so the robots.txt we read is the one at the destination.
SEO crawlers get named. The AI ones do not
SEO crawlers are a different story. The STORES official store disallows MJ12bot and Amazonbot by name; the Color Me demo store disallows eight, including AhrefsBot, SemrushBot and PetalBot. Naming crawlers is already in use here.
None of the eight we checked appear on those lists (some people count Amazonbot as an AI crawler; STORES disallows that one). Those blocklists accumulated over years, and GPTBot only appeared in August 2023 (OpenAI's bot docs).
Can an owner change it? That differs by cart. Shopify lets you add a robots.txt.liquid template to your theme and append your own rules to the defaults (Shopify developer docs). For the other four (BASE, STORES, Color Me, futureshop) we could not confirm whether an owner can edit it — some help centers block automated access — so that stays unverified here.
One site separates training from citation
The only host of the six that names AI crawlers is MakeShop's corporate robots.txt — their own service site, not the store default.
It disallows training crawlers (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended among others) while allowing the ones that fetch for AI-search citation (OAI-SearchBot, ChatGPT-User, Claude-User, PerplexityBot). The file states that policy in Japanese comments, with a last-updated date.
I am not claiming that split is the right answer. But “don't train on us” and “do cite us in AI search” are compatible wishes, and they can be written separately at the robots.txt level. That did not click for me until I saw it in the wild.
llms.txt: absent on all five we could read
I also looked for llms.txt, the file that summarizes a site for AI readers. All five hosts we could fetch returned absent. The sixth could not be fetched at all, so we are not calling it missing.
That reads less like a decision and more like an early stage. We publish one on our own site, and so far I have felt no difference in inquiries.
Measure your own store
The same check runs from the form at the bottom of this page. No signup, no email. It reads only public files (robots.txt, llms.txt, sitemap.xml and the page itself) and never walks the purchase path. The result gets a shareable URL you can hand to your agency or your cart's support desk.
If the result shows something being read that you would rather keep out, the next question is whether your cart lets you edit robots.txt. On Shopify you can add rules yourself. Elsewhere, pasting the result URL to your platform's support desk is the quicker route.
Once you have edited it, how to confirm the rule actually blocks, and three ways it silently does not, are in how to check whether robots.txt really blocks GPTBot.
Try it on your own site
Paste a URL and see right away: whether AI crawlers are let in, whether there is a file written for AI, and whether the page is marked to stay out of search — using public pages only.