Methodology

How every number is measured.

This page defines how KoreLens measures AI readiness and citation evidence, exactly what each number means, and what may — and may not — be claimed from it. It is the reference to hand your analyst when they ask “how do you know?”.

Written and maintained by Yaseen Mohammed, founder of KoreLens. Changes to how anything is measured appear on this page before they appear in any number.

What the audit measures deterministically

The URL audit is a set of deterministic checks over your actual pages and files — the same input produces the same result, and anything that cannot be verified is reported as not checked, never scored.

  • Crawler access

    Your robots.txt is checked against the access rules of 21 named AI crawlers (GPTBot, OAI-SearchBot, Google-Extended, ClaudeBot and more). Blocked means blocked — we report the exact rule.

  • Readable business and product facts

    The visible text and structured data (JSON-LD) on your page are read the way an AI system reads them, and checked against each other. Facts that disagree or are missing are listed individually.

  • Product feed validation, field by field

    Your product data file is validated against the published OpenAI product feed spec and Google Merchant Center field requirements. The date those specifications were last verified is recorded in the output itself.

  • Agent actionability

    A synthetic probe checks whether an AI assistant could actually find and use your contact, booking, and checkout paths. What cannot be checked without a real browser is reported as not checkable — never guessed.

  • Real AI answers, stored verbatim

    Tracked questions are asked of real AI engines through their APIs. Every stored answer records the exact model id, word-for-word text, and whether the engine used live search. Failed calls are stored as failures.

Which number is the score

One number is Readiness, out of 100. It is the same figure in your dashboard, in a shared client report, and in an email — if you see a readiness score anywhere, it is that one. It measures how much of what AI and search systems need to read about a business is actually present on the site. It is not a traffic, ranking, or sales figure.

Two other numbers exist and are deliberately NOT merged into it, because they measure different things: Shop readiness (out of 100) is about product data being complete enough to build a feed from, and Record completeness (out of 100) is about how much of your saved KoreLens record is filled in. Each is shown with its own name, its own denominator, and a line saying what it means — and you will never see two different scales for the same thing on one page.

How much of a site we read — and what “missing” means

The audit reads your homepage and up to four more pages found by following your own links and sitemap: a product page, a policy page, a collection, and a contact page. It stays on your domain, obeys your robots.txt, stops at a fixed page cap and a time budget, and reads one page per role — so a shop with 400 products still costs one product page. We also cap how often any single site is read, across everyone: a site checked several times in quick succession gets the homepage-only check, and the report says so. That limit protects sites from being pointed at by whoever pastes their address into our free tool — including sites that never asked us to look.

This matters for what a finding means. “Not on the homepage” is a weak claim and was often simply wrong — a returns policy one click away was reported as absent. Every negative finding now states its scope, and the report lists the exact URLs we read, so “not found on any of the pages we read” is something you can check rather than take on trust. Anything the cap, the clock or your robots.txt kept us from reading is named too.

Reading more pages can only remove a false negative. A fact found on a policy page can clear a check the homepage failed, and it is labelled with the page it was found on — but nothing found elsewhere can ever fail a check your homepage passed, so a wider read cannot lower your score.

Not every check applies to every business

A street address and opening hours describe a business customers physically visit. Asking a software company for its opening hours is noise, and scoring it down for not having a shop counter is worse than noise. So when a saved record is set to “an online business or organization”, those two checks are not run, not counted, and named in the not-checked list — the score is out of the checks that apply to you, and it says which ones those were.

This switch follows your own declaration on the record, never our inference. A free audit of a URL we know nothing about runs every check, because leaving a check out is a claim that it does not apply — and we would rather ask a question that turns out to be irrelevant than quietly stop telling a real shop that its address is missing.

How we measure AI answers — and what we cannot measure

Every AI answer KoreLens shows was obtained by calling a named model through its API, with live web search on, sending the tracked question word for word. This table renders from the same configuration the calls use — it cannot drift from what production sends.

EngineModel we callRetrieval
Claude (Anthropic)claude-haiku-4-5Live web search (provider server tool; real per-call search count stored)
ChatGPT (OpenAI)gpt-4o-search-previewLive web search (search-enabled model; provider does not report a per-call count)
PerplexitysonarSearches the web natively (provider does not report a per-call count)
Gemini (Google)gemini-3.6-flashLive Google Search grounding (per-call query count stored)
  • One measurement is one answer. Weekly tracking asks each prompt once per engine per week and stores the answer verbatim with the exact model id, the date, and whether live search really ran on that call. AI answers vary between runs, so a single stored answer is a dated observation — evidence of what this instrument returned that day, not “the” answer. Where a claim needs tighter precision, we run repeated samples of the same prompt and report the rate with an interval instead of a single answer.
  • One answer is an observation; a rate needs a sample. Where product-level claims need more than a single dated answer, KoreLens runs repeated samples of the same buyer question and reports an appearance rate with a 95% confidence interval (Wilson), never a “ranking” from one run. Two rates are called different only when their intervals do not overlap — otherwise the honest verdict is “no verdict yet”, and that is what the product says. Sample size and interval are shown wherever a rate is shown.
  • If retrieval fails, we say so. A check that could not run with live search is stored and labelled “no live web search”, and is excluded from measured mention rates — an answer from the model’s stored knowledge is a different instrument, and we do not pool instruments. A comparison whose two sides ran under different retrieval states carries a method-change warning on the comparison itself.
  • What we cannot measure: ChatGPT (chatgpt.com), Gemini (gemini.google.com), Claude (claude.ai), Microsoft Copilot — the consumer apps have no supported way to be measured programmatically. You (or your agency) can record what you saw there, word for word. Recorded answers are stored labelled “entered by you”, shown with that label everywhere, and are never counted in measured rates or before/after evidence. Measured and recorded answers are never blended — this is enforced in code, not just promised.

What API measurement can tell you

  • Whether a named model, with live search on, mentioned or cited your business when asked a specific question on a specific date — word for word, stored.
  • How that same question, on that same engine, under the same settings, answered differently between two dates — with any change in our own method flagged on the comparison itself.
  • Aggregate mention rates over many checks, always with the denominator shown, and only above minimum sample sizes.

What it cannot tell you

  • What the consumer apps answered. The models are related and search the same web, but chatgpt.com is not the OpenAI API model, and gemini.google.com is not the Gemini API. We name exactly which instrument we measured; we never imply it was the app.
  • What any specific buyer saw. AI answers are non-deterministic and can be personalised — the same prompt, minutes apart, can produce different answers. No canonical answer exists for anyone to measure: not for us, and not for tools that read answers out of a browser.
  • A "position" or "rank". A claim like “you rank #3 in ChatGPT” needs a stable ordering that these systems do not produce — the order can change between two runs with identical settings. We do not sell position claims at all.

Exactly what we send — nothing hidden

The tracked question is sent word for word as the user message. Around it we set a short system prompt that frames the model as a general-purpose assistant and tells it to search before answering and to say so rather than guess when it cannot find reliable information — the same kind of frame every consumer surface applies, made visible. Answers are capped at 600 tokens; temperature is the provider default (the consumer apps do not pin it either); no location or personalisation parameters are sent, so answers reflect the providers’ default search context, not any specific buyer’s. The frame is versioned (presence-v1); if it ever changes, results measured under different versions are never compared as if the method held.

The system prompt, verbatim
You are a general-purpose AI assistant answering an everyday question about a business or place, the way you would for any user. If web search is available, search for the business before answering. Answer naturally in 2–4 sentences based on what you find or reliably know. If you cannot find reliable information about this specific business, say so plainly instead of guessing or inventing details.

Why we do not sell “position” — and why held-method measurement is the stronger claim

Some tools read answers out of the consumer interfaces and report where a brand “ranks”. Whatever is doing the reading — an API or a browser — is sampling a system that is non-deterministic and can be personalised. A browser-based tool is sampling its own signed-in account’s view, shaped by that account’s settings: a single web-search toggle has been shown to move a brand between position 1 and position 8 on the same prompt. That is not a flaw in one tool — it is what happens when a rank is claimed over a distribution that reshuffles between runs. No tool, ours included, can measure “the” answer a buyer sees, because no such single answer exists.

So we sell the two things that survive that fact. First: whether your data is machine-readable and verifiable — deterministic checks over your actual pages, feeds, and structured data, where the same input gives the same result and a re-crawl proves a fix landed. No sampling problem exists there at all. Second: change in the same prompt on the same engine between two dates, with the method held constant — same model id, same retrieval state, same frame, and the method printed on the comparison. For measuring change, holding the instrument steady matters more than matching any one person’s personalised view; that is how measurement works everywhere else, too.

Single answers are shown as dated exhibits, word for word — claims about what was said on a date, which stay true regardless of variance. Rates and before/after deltas are stated only with their denominators and only above minimum sample sizes, and repeated-sample runs carry intervals. A tool claiming your client “ranks #3” is making a weaker claim than either: it needs a canonical ranking that does not exist. Ours are built to not need one.

The metric: cited-answer share

Cited-answer share = of the successful sampled answers in a period, the fraction that mention the business.

  • A sampled answer is one real API call to a named model — the exact model id is stored with every check. Answers are stored verbatim; there is no fabrication, templating, or estimation path anywhere in the pipeline.
  • Checks run search-grounded, because consumer surfaces search before answering. Every stored check records whether live search really ran; only grounded, API-measured checks count in this metric. Ungrounded checks and owner-recorded answers are excluded from the rate, counted, and shown labelled — never pooled.
  • Only self-questions (questions about this business) count toward its share. Competitor questions measure rivals and are excluded.
  • A failed call is stored as a failure and excluded from the denominator. It is never counted as “not mentioned”.

Citation Velocity: before/after around a change

Citation Velocity is the percentage-point change in cited-answer share between a fixed window of ISO weeks before a recorded change and the same-length window after it. Six design rules, all enforced in code:

  1. The change week is excluded

    It is partially before and partially after the change; counting it would smear the boundary in whichever direction flatters the number.

  2. Fixed windows

    Default 4 ISO weeks per side, always the weeks nearest the change. Distant history cannot be cherry-picked into the baseline.

  3. Sample-size gate

    Fewer than 6 successful checks on either side → insufficient sample, and no delta is stated. 6–11 per side → medium confidence; 12+ → high. Confidence is computed once and never upgraded downstream.

  4. Fixed question set

    Baseline weeks retain their own stored checks, so adding flattering questions after a change does not rewrite the baseline.

  5. Deterministic event identity

    Every mirrored analytics event carries a UUID derived from its identifying facts, so retries and cron overlaps cannot double-count evidence.

  6. Attribution is correlational

    A before/after delta around a change is evidence of association, not proof of causation. Confounders (model updates, seasonality, press, competitor changes) are not controlled in a single-business measurement, and every emitted event carries that caveat.

What may be claimed

  • “In our sampled measurements, answers on {engine} mentioned {business} in {after}% of checks in the 4 weeks after the fix, up from {before}% in the 4 weeks before (n={n_before}/{n_after}, confidence: {confidence}).”
  • “Across N businesses that deployed fixes in {period}, the median change in cited-answer share was +X points within 4 weeks (methodology attached).”
  • “IndexNow submission accepted (HTTP 202)” — a receipt, nothing more.

What may never be claimed

  • “Guaranteed citations / rankings / traffic / sales.”
  • “#1 cited source” — no ranking ground truth exists for consumer AI surfaces.
  • “Schema confirmed ingested by {provider}” — no provider issues ingestion receipts; we can only sample answers.
  • “+X% in 24 hours” — below crawl-and-index latency; our own dashboard would falsify it.
  • Any number whose window failed the sample-size gate.

The Fix Trail: what a closed loop actually records

One data defect followed all the way through: found, fixed, confirmed by an independent re-read, and — where the fix is tied to a tracked buyer question — that question asked again, with the verdict reported either way.

  1. 1 · Detected. A deterministic check found the defect — sources disagreeing, a product changed from the record you accepted, or a readiness check failing on a live page. The values and the dates that triggered it are stored, not a model’s summary of them.
  2. 2 · Fixed. The correction is recorded with who authorised it and when. Where a fix is applied outside KoreLens — most are — the trail says so rather than implying we did it.
  3. 3 · Confirmed by an independent re-read. We read the data back from the source and check it actually changed. A successful API response is not a confirmation, and we do not treat it as one; only the re-read can move this stage.
  4. 4 · Re-measured, where there is something to re-measure. If the fix is tied to a tracked buyer question, that question is asked again under the frozen conditions of the first measurement, and the two appearance rates are compared with their intervals. If it isn’t tied to one, the trail says that plainly instead of showing a stage that can never arrive.

A stage that hasn’t happened reads as not yet reached, never as a failure. Every stage carries its date and where the fact came from. And the verdict is reported whatever it is — see below.

Prompt coverage: how many buyer questions surface you at all

Prompt coverage is the share of tracked discovery questions — realistic buyer questions that do not name the business — whose latest answered check surfaced the business. The denominator is answered questions only: a question no engine has answered yet is stated as unanswered, never counted as a miss. This is the same number the dashboard has always computed for discovery; the name is the one practitioners use for it. It measures whether the business appears at all — not position, not preference, not traffic.

Cited share: a percentage, not a ranking

For each competitor domain, cited share is the share of successful checks in the window whose stored answer cited that domain — with the sample size and a 95% interval, on the same sample gate as every other rate. The denominator is the same eligible check set the mention rate uses: API-measured, live-search answers that passed the name-collision gate. Aggregators, press pages and known name-collision domains are excluded from the table using the same classification that files them under “also cited” — and the surface says the exclusion happened. It is never presented as a ranking: being cited says an engine used the domain, nothing about position or preference, and no “#1 cited source” claim exists in the product.

Sentiment: how answers characterise you — a model judgment, labelled as one

Sentiment is the least deterministic number in this product, and it is built so that fact is visible rather than smoothed over. For each stored answer that passed the name-collision gate and actually mentions the business, a judge model classifies how the answer characterises it: favourable, neutral, unfavourable — or factually wrong, a separate outcome that is never folded into “unfavourable”, because a flattering description of the wrong company is not positive sentiment. “Wrong” means: contradicts your saved record, which is the only authority the judge is given.

  • Every value names its instrument. Each judgment stores the judge model and rules version that produced it, and links to the verbatim answer it was derived from. A judgment whose supporting sentence cannot be found verbatim in the stored answer is rejected, never stored.
  • Excluded answers get no value at all. Answers the name-collision gate excluded or flagged are not judged — not given a neutral score, not counted in the denominator. No characterisation of someone else can describe you.
  • Same sample gate as every other rate. Sentiment is reported as a rate with its sample size and a 95% interval. Below the sample gate there is no sentiment verdict, and the product says so instead of showing one.
  • Never blended. Sentiment is a fourth, separate scale. It is never merged into the readiness score or any other number, and deterministic checks are never averaged with model judgments.

What may be claimed from it: “of the N judged answers, X were classified unfavourable by the named judge model, and here is each supporting sentence.” What may not: that sentiment predicts traffic, rankings or sales; that a model’s classification is a fact about your business; or any figure at all when the sample is below the gate.

What counts as a change worth telling you about

KoreLens emails you when something it measures actually moves. Because an alert interrupts you, the bar for sending one is a published rule rather than a judgement call. An AI answer is treated as changed only when the finding changes: whether you were named, which sources were cited, or whether the question drew a usable answer at all.

  • Rewording alone is never an alert. Engines rewrite the same answer constantly. If you are still named, the same sources are cited, and the viability verdict is unchanged, nothing is sent — the stored answers still differ, and you can still read both side by side.
  • A week we could not measure is never a movement. If a check was refused or the page could not be read, there is no score for that week — so there is no fall to report. Gaps stay gaps.
  • Answers about a different company are excluded. Where our checks indicate an engine described a similarly-named business rather than yours, that answer is left out of alerts for the same reason it is left out of your rates. It is surfaced separately, as its own finding.
  • Changed conditions are not a changed answer. If one check ran with live retrieval and the other did not, the difference is our method, not their answer, and the two are never compared.
  • A newly cited domain is filtered first. Marketplaces, press and reference pages, and sites whose name merely resembles yours are demoted before anything is sent. A domain we cannot categorise is still shown — refusing to guess must not become a way of hiding a real competitor. On a client’s first measured week nothing is sent at all: with no earlier week to compare against, every cited domain would be “new” only because we had not looked before.
  • Small score movement is noise, and you can ask for less of it. A readiness move under five points out of 100 is check-level noise and is never recorded as a change. Above that, you can raise your own threshold so only larger moves email you — that setting decides what reaches your inbox, never what we record. The measurement and your quietness preference are kept apart on purpose, so the stream stays a complete account of what we observed.

Every alert states the change itself and links to the dated evidence for it. None of them explains why the change happened: we report that two measurements differ and show you both. Readiness movement is a measure of what AI and search systems can read on a site — never a traffic, ranking, or sales figure.

What we won’t claim

Four things KoreLens is built not to say. Each one is a design decision, not a limitation we ran into — and each costs us a claim we could otherwise make.

  • A verdict when the numbers don’t support one

    When we re-measure a buyer question after a fix, we compare two appearance rates with 95% confidence intervals. If those intervals overlap, we report “no verdict” — the intervals overlap, so we won’t call it — and we say what would narrow them. A higher percentage with overlapping intervals is not an improvement; it is noise that happens to point the right way. We would rather hand your client an honest “not yet” than a number that falls apart when their technical lead checks it. Inconclusive results appear in the product and in client reports exactly as prominently as movements do.

  • That a tracked question is still measuring something — prompt viability

    A question can stop working. An engine changes how it answers, a prompt drifts out of relevance, and what comes back is “I don’t have information about that business” — a real answer to a question that is no longer measuring anything. Left alone it quietly enters the denominator as a miss and drags a rate down for no reason. KoreLens classifies every stored answer for whether it actually engaged with the business, flags a question whose answers no longer do, and stops counting it until it is rewritten. The flag is visible in the product, and a flagged question never appears in a client report as evidence.

  • A readiness score for a page we could not read

    Bot protection, a CAPTCHA wall, or a page whose content only appears after scripts run all produce the same thing on our side: a response we cannot read. Every content check then “fails”, and the result is a low score with a long list of red flags — describing our own blocked request, not the site. So a score is now something a page has to earn: if an identified auditor was refused, or we received less readable content than the published threshold, KoreLens returns no readiness score, no shop-readiness score, no missing-field list and no fix list, and shows what we tried, what happened, and how to let us in. We claim only that our checker was turned away — never that AI systems are blocked, because we are not those crawlers and cannot speak for them. The live AI answers alongside are gathered independently and often describe a brand accurately even when the site refuses us.

  • That we know which of your sources is right

    Your product feed, your product page, and your structured data disagree more often than anyone expects, and any AI or shopping system reading the wrong one describes the product wrongly. When KoreLens finds a disagreement it shows every source’s stated value with the date it was read, and stops. We do not pick a winner, apply a heuristic, or quietly prefer the newest record — the merchant knows which is correct and we do not. Once it is corrected at source, the next read confirms the sources agree, and that confirmation is what we report.

The common thread: a measurement that cannot decline to answer is not a measurement. We would rather be the tool your client’s technical lead can’t catch out.

Data lineage

Every number can be traced back through this chain — each stage keeps its own honesty property.

StageStoreHonesty property
Question askedtracked_questionsOwner-editable, deterministic templates
Answer sampledcitation_checksVerbatim answer, exact model id, grounded flag, failures stored as failures
Change recordedchange_eventsOnly real accepted actions (verified IndexNow receipts, declared fixes)
Velocity computedin code, pure functionGates and exclusions above, unit-tested
Evidence mirroredPostHog citation_velocity_recordedDeterministic UUID, caveat attached, measured results only

The shortest honest summary: we measure readiness and record evidence. We never guarantee a citation, a ranking, traffic, or a sale.