KoreLens research · 16 July 2026
UK stores aren’t blocking AI. They’re invisible to it.
We audited 249 verified UK Shopify and WooCommerce storefronts the way an AI shopping system reads them. Almost none block AI crawlers — and almost none give them any product data to read.
389
UK store URLs seeded from public roundups
249
platform-verified stores (231 Shopify · 18 Woo)
80
other platforms — excluded
60
unreadable by our identified crawler
Sample bias, stated up front: seeds come from curated “best of” lists of UK stores, so this sample skews toward well-run stores. The wider population is unlikely to look better than these numbers.
Finding 1
AI shopping systems get almost no product data at the front door
77% of stores carry some structured data — organisation, navigation, search. But Product markup with prices and availability is essentially absent from the homepage, the page where crawlers and agents land first. Apparel, beauty, food & drink, jewellery and sports & outdoors all measured 0% on their homepages; the best sectors managed about 7%.
Scope: a homepage measurement. Stores may mark up individual product pages — that per-store crawl is the natural follow-up. This finding is about the front door.
Finding 2
The “stores are blocking AI” story is mostly a myth
Just 7 of 248 stores block any of the 21 named AI crawlers in robots.txt — mostly the same handful blocking everything. GPTBot, ClaudeBot, Google-Extended, CCBot and Amazonbot are each blocked by 2.4%. UK stores are not shutting AI out on purpose. They’re open — and illegible. The problem isn’t permission. It’s product data.
Finding 3
Two-thirds of homepages give AI no returns-policy signal
AI shopping surfaces treat returns, shipping and support information as trust data. Only a third of stores state or link a returns policy where a crawler can find it from the homepage. Gifts & stationery lead at 70%; food & drink is worst at 16.2% among sectors with ten or more stores.
Finding 4
llms.txt “adoption” is a platform artifact
The split between platforms points to auto-generation, not merchant action. Per our methodology, llms.txt is experimental — reported as adoption trivia, and never scored.
Finding 5
15% of stores couldn’t be read by an identified bot at all
These sites refused or timed out for our transparent, self-identified crawler. We claim only that our crawler was refused — whatever blocks an identified auditor may or may not block a given AI system. But a storefront that default-blocks unknown bots should at least know it’s doing so.
Sector league table
Sectors with ten or more verified stores, ranked by median readiness score. Smaller sectors are in the dataset but too small to rank honestly.
| Sector | Stores | Median /100 | Structured data | Product data on homepage | Returns signal | Blocks any AI bot |
|---|---|---|---|---|---|---|
| Health & fitness | 14 | 81 | 85.7% | 7.1% | 28.6% | 0% |
| Home & garden | 54 | 77 | 83.3% | 7.5% | 35.8% | 1.9% |
| Gifts & stationery | 10 | 77 | 80.0% | 0% | 70.0% | 0% |
| Food & drink | 37 | 76 | 70.3% | 0% | 16.2% | 2.7% |
| Beauty | 23 | 76 | 73.9% | 0% | 34.8% | 4.3% |
| Jewellery | 11 | 75 | 90.9% | 0% | 36.4% | 0% |
| Apparel | 59 | 74 | 72.9% | 0% | 34.5% | 3.4% |
| Sports & outdoors | 21 | 72 | 76.2% | 0% | 28.6% | 9.5% |
Overall median readiness score: 75/100.
What we did not find, and will not claim
- No claim that any store is losing sales or traffic — readiness, never outcome.
- Product field quality (prices, stable IDs, availability) was measurable on only 5 homepages — too few to publish honestly. That measurement needs a product-page crawl.
- We never rank or single out individual stores — the league table is by sector only. The raw per-store observations are published below as open data so anyone can verify every number; each row is a dated, factual reading of public pages, and any store can re-run the same check free at any time.
The raw data, published
All 389 seeded observations, including unreachable sites and non-Shopify/Woo platforms — the full working set, not just the rows behind the headline numbers, so the exclusions are checkable too.
Licensed CC BY 4.0 — cite the study and reuse freely. Each row is a dated, point-in-time reading of public pages; if your store is in it and you’ve since fixed something, the free check below shows today’s reality.
Methodology
One pass per site on 16 July 2026: robots.txt, homepage, sitemap.xml, llms.txt. Public pages only, 8-second timeouts, at most 4 sites in flight, one attempt, transparent user-agent (KoreLensAudit/1.0, linking this methodology). Platform verified from each site’s own HTML markers, never assumed from the seed list. Checks are the same deterministic checks as the free KoreLens audit. A site we could not read is excluded from metrics, never scored. JS-render risk is a text-length heuristic, not a certain diagnosis. Nothing is simulated.
The full measurement contract is on the methodology page.
These are the same checks as the free audit.
Run them on your own store — score, gaps, and the word-for-word answer an AI engine gives about you today. In about a minute, no signup.
Check my store — free