KoreLens research · 16 July 2026

UK stores aren’t blocking AI. They’re invisible to it.

We audited 249 verified UK Shopify and WooCommerce storefronts the way an AI shopping system reads them. Almost none block AI crawlers — and almost none give them any product data to read.

389

UK store URLs seeded from public roundups

249

platform-verified stores (231 Shopify · 18 Woo)

80

other platforms — excluded

60

unreadable by our identified crawler

Sample bias, stated up front: seeds come from curated “best of” lists of UK stores, so this sample skews toward well-run stores. The wider population is unlikely to look better than these numbers.

Finding 1

AI shopping systems get almost no product data at the front door

2%of verified stores (5 of 247)

77% of stores carry some structured data — organisation, navigation, search. But Product markup with prices and availability is essentially absent from the homepage, the page where crawlers and agents land first. Apparel, beauty, food & drink, jewellery and sports & outdoors all measured 0% on their homepages; the best sectors managed about 7%.

Scope: a homepage measurement. Stores may mark up individual product pages — that per-store crawl is the natural follow-up. This finding is about the front door.

Finding 2

The “stores are blocking AI” story is mostly a myth

2.8%of stores (7 of 248)

Just 7 of 248 stores block any of the 21 named AI crawlers in robots.txt — mostly the same handful blocking everything. GPTBot, ClaudeBot, Google-Extended, CCBot and Amazonbot are each blocked by 2.4%. UK stores are not shutting AI out on purpose. They’re open — and illegible. The problem isn’t permission. It’s product data.

Finding 3

Two-thirds of homepages give AI no returns-policy signal

33.6%of stores (of 247)

AI shopping surfaces treat returns, shipping and support information as trust data. Only a third of stores state or link a returns policy where a crawler can find it from the homepage. Gifts & stationery lead at 70%; food & drink is worst at 16.2% among sectors with ten or more stores.

Finding 4

llms.txt “adoption” is a platform artifact

95.7% vs 11.1%Shopify vs WooCommerce

The split between platforms points to auto-generation, not merchant action. Per our methodology, llms.txt is experimental — reported as adoption trivia, and never scored.

Finding 5

15% of stores couldn’t be read by an identified bot at all

60 of 389seeded sites (15.4%)

These sites refused or timed out for our transparent, self-identified crawler. We claim only that our crawler was refused — whatever blocks an identified auditor may or may not block a given AI system. But a storefront that default-blocks unknown bots should at least know it’s doing so.

Sector league table

Sectors with ten or more verified stores, ranked by median readiness score. Smaller sectors are in the dataset but too small to rank honestly.

SectorStoresMedian /100Structured dataProduct data on homepageReturns signalBlocks any AI bot
Health & fitness148185.7%7.1%28.6%0%
Home & garden547783.3%7.5%35.8%1.9%
Gifts & stationery107780.0%0%70.0%0%
Food & drink377670.3%0%16.2%2.7%
Beauty237673.9%0%34.8%4.3%
Jewellery117590.9%0%36.4%0%
Apparel597472.9%0%34.5%3.4%
Sports & outdoors217276.2%0%28.6%9.5%

Overall median readiness score: 75/100.

What we did not find, and will not claim

  • No claim that any store is losing sales or traffic — readiness, never outcome.
  • Product field quality (prices, stable IDs, availability) was measurable on only 5 homepages — too few to publish honestly. That measurement needs a product-page crawl.
  • We never rank or single out individual stores — the league table is by sector only. The raw per-store observations are published below as open data so anyone can verify every number; each row is a dated, factual reading of public pages, and any store can re-run the same check free at any time.

The raw data, published

All 389 seeded observations, including unreachable sites and non-Shopify/Woo platforms — the full working set, not just the rows behind the headline numbers, so the exclusions are checkable too.

Licensed CC BY 4.0 — cite the study and reuse freely. Each row is a dated, point-in-time reading of public pages; if your store is in it and you’ve since fixed something, the free check below shows today’s reality.

Methodology

One pass per site on 16 July 2026: robots.txt, homepage, sitemap.xml, llms.txt. Public pages only, 8-second timeouts, at most 4 sites in flight, one attempt, transparent user-agent (KoreLensAudit/1.0, linking this methodology). Platform verified from each site’s own HTML markers, never assumed from the seed list. Checks are the same deterministic checks as the free KoreLens audit. A site we could not read is excluded from metrics, never scored. JS-render risk is a text-length heuristic, not a certain diagnosis. Nothing is simulated.

The full measurement contract is on the methodology page.

These are the same checks as the free audit.

Run them on your own store — score, gaps, and the word-for-word answer an AI engine gives about you today. In about a minute, no signup.

Check my store — free