The US Catalogue Truth Index · 2026 edition · 11 August 2026
US stores aren’t blocking AI. They’re invisible to it.
We audited 160 verified, readable US Shopify and WooCommerce storefronts the way an AI shopping system reads them. Almost none block AI crawlers on purpose — almost none give them product data to read — and one in four of the stores we seeded couldn’t be read by an identified crawler at all.
238
US store URLs seeded from public roundups
160
readable platform-verified stores (159 Shopify · 1 Woo)
16
verified but gated (bot challenge / script-only shell) — excluded
17
other platforms — excluded
45
unreachable by our identified crawler
Sample bias, stated up front: seeds come from curated “best of” lists of US stores, so this sample skews toward well-run, design-conscious DTC brands. The wider population of US stores is unlikely to look better than these numbers.
Finding 1
AI shopping systems get almost no product data at the front door
83.1% of stores carry some structured data — organisation, navigation, search. But Product markup with prices and availability is essentially absent from the homepage, the page where crawlers and agents land first: exactly one of the 160 stores we could read exposes any. Apparel (0 of 49), beauty, food & drink, health & fitness and jewellery all measured 0% on their homepages.
Scope: a homepage measurement. Stores may mark up individual product pages — that per-store crawl is the natural follow-up. This finding is about the front door.
Finding 2
The “stores are blocking AI” story is a myth in the US too
Just 4 of 159 stores with readable robots.txt block any of the 21 named AI crawlers — Google-Extended and CCBot lead at 2.5% each, GPTBot and ClaudeBot at 1.3%. Like the UK, US stores are not shutting AI out on purpose. They’re open — and illegible. The problem isn’t permission. It’s product data.
Finding 3
A quarter of US stores couldn’t be read or scored by an identified crawler
45 of 238 seeded sites refused or timed out for our transparent, self-identified crawler, and a further 16 verified stores answered with a bot challenge or a script-only shell our checker refuses to score. That is markedly heavier bot protection than the 15.4% we measured for UK stores in July. We claim only that our crawler was stopped — whatever blocks an identified auditor may or may not block a given AI system — but a storefront that default-blocks unknown bots should at least know it’s doing so.
Finding 4
Six in ten homepages give AI no returns-policy signal
AI shopping surfaces treat returns, shipping and support information as trust data. Only 39.4% of stores state or link a returns policy where a crawler can find it from the homepage. Food & drink leads at 57.1%; jewellery and health & fitness sit at 25% among the sectors we measured.
Finding 5
llms.txt “adoption” is a platform artifact, again
Near-universal presence on a sample that is 159 parts Shopify points to auto-generation, not merchant action — the same pattern the UK edition measured. Per our methodology, llms.txt is experimental: reported as adoption trivia, and never scored.
Sector league table
Sectors with ten or more readable stores, ranked by median readiness score. Smaller sectors are in the dataset but too small to rank honestly.
| Sector | Stores | Median /100 | Structured data | Product data on homepage | Returns signal | Blocks any AI bot |
|---|---|---|---|---|---|---|
| Apparel | 49 | 78 | 79.6% | 0% | 38.8% | 2.0% |
| Home & garden | 38 | 78 | 78.9% | 2.6% | 42.1% | 2.6% |
| Food & drink | 14 | 77 | 85.7% | 0% | 57.1% | 7.1% |
| Beauty | 10 | 77 | 70.0% | 0% | 50.0% | 0% |
Overall median readiness score: 79/100.
What we did not find, and will not claim
- No claim that any store is losing sales or traffic — readiness, never outcome.
- Product field quality (prices, stable IDs, availability) was measurable on only 1 homepage — far too few to publish. That measurement needs a product-page crawl.
- The 16 gated stores are reported as unreadable, never as “no schema” — a page we could not read well enough to score contributes nothing to any percentage.
- We never rank or single out individual stores — the league table is by sector only, and only for sectors with ten or more readable stores. The raw per-store observations are published below as open data so anyone can verify every number; each row is a dated, factual reading of public pages, and any store can re-run the same check free at any time.
The raw data, published
All 238 seeded observations, including unreachable sites, gated stores, and non-Shopify/Woo platforms — the full working set, not just the rows behind the headline numbers, so the exclusions are checkable too.
Licensed CC BY 4.0 — cite the study and reuse freely. Each row is a dated, point-in-time reading of public pages; if your store is in it and you’ve since fixed something, the free check below shows today’s reality.
Methodology
One pass per site on 11 August 2026: robots.txt, homepage, sitemap.xml, llms.txt. Public pages only, 8-second timeouts, at most 10 sites in flight (never more than one request set per site), one attempt, transparent user-agent (KoreLensAudit/1.0, linking this methodology). Platform verified from each site’s own HTML markers, never assumed from the seed list. Checks are the same deterministic checks as the free KoreLens audit. A site we could not read is excluded from metrics, never scored — including verified stores that answered with a bot challenge or a script-only shell. JS-render risk is a text-length heuristic, not a certain diagnosis. Nothing is simulated.
The full measurement contract is on the methodology page.
About The US Catalogue Truth Index
An annual measurement of how legible US ecommerce catalogues are to the systems that read them. Published annually. If the method changes materially between editions we will say so rather than present the numbers as a like-for-like trend.
This page is the 2026 edition and keeps its address permanently — its numbers will never be edited. To link to whichever edition is current, use /index/us-catalogue-truth. The UK sibling lives at /index/uk-catalogue-truth.
- 2026 edition — 160 verified and readable US Shopify and WooCommerce storefronts, crawled 11 August 2026
These are the same checks as the free audit.
Run them on your own store — score, gaps, and the word-for-word answer an AI engine gives about you today. In about a minute, no signup.
Check my store — freeJust want the product-markup finding on one page? Run the product schema checker — it reads the JSON-LD on a single product URL and shows the value found for each field.