Methodology

How every number is measured.

This page defines how KoreLens measures AI readiness and citation evidence, exactly what each number means, and what may — and may not — be claimed from it. It is the reference to hand your analyst when they ask “how do you know?”.

Written and maintained by Yaseen Mohammed, founder of KoreLens. Changes to how anything is measured appear on this page before they appear in any number.

What the audit measures deterministically

The URL audit is a set of deterministic checks over your actual pages and files — the same input produces the same result, and anything that cannot be verified is reported as not checked, never scored.

  • Crawler access

    Your robots.txt is checked against the access rules of 21 named AI crawlers (GPTBot, OAI-SearchBot, Google-Extended, ClaudeBot and more). Blocked means blocked — we report the exact rule.

  • Readable business and product facts

    The visible text and structured data (JSON-LD) on your page are read the way an AI system reads them, and checked against each other. Facts that disagree or are missing are listed individually.

  • Product feed validation, field by field

    Your product data file is validated against the published OpenAI product feed spec and Google Merchant Center field requirements. The date those specifications were last verified is recorded in the output itself.

  • Agent actionability

    A synthetic probe checks whether an AI assistant could actually find and use your contact, booking, and checkout paths. What cannot be checked without a real browser is reported as not checkable — never guessed.

  • Real AI answers, stored verbatim

    Tracked questions are asked of real AI engines through their APIs. Every stored answer records the exact model id, word-for-word text, and whether the engine used live search. Failed calls are stored as failures.

How we measure AI answers — and what we cannot measure

Every AI answer KoreLens shows was obtained by calling a named model through its API, with live web search on, sending the tracked question word for word. This table renders from the same configuration the calls use — it cannot drift from what production sends.

EngineModel we callRetrieval
Claude (Anthropic)claude-haiku-4-5Live web search (provider server tool; real per-call search count stored)
ChatGPT (OpenAI)gpt-4o-mini-search-previewLive web search (search-enabled model; provider does not report a per-call count)
PerplexitysonarSearches the web natively (provider does not report a per-call count)
Gemini (Google)gemini-3.6-flashLive Google Search grounding (per-call query count stored)
  • One measurement is one answer. Weekly tracking asks each prompt once per engine per week and stores the answer verbatim with the exact model id, the date, and whether live search really ran on that call. AI answers vary between runs, so a single stored answer is a dated observation — evidence of what this instrument returned that day, not “the” answer. Where a claim needs tighter precision, we run repeated samples of the same prompt and report the rate with an interval instead of a single answer.
  • If retrieval fails, we say so. A check that could not run with live search is stored and labelled “no live web search”, and is excluded from measured mention rates — an answer from the model’s stored knowledge is a different instrument, and we do not pool instruments. A comparison whose two sides ran under different retrieval states carries a method-change warning on the comparison itself.
  • What we cannot measure: ChatGPT (chatgpt.com), Gemini (gemini.google.com), Claude (claude.ai), Microsoft Copilot — the consumer apps have no supported way to be measured programmatically. You (or your agency) can record what you saw there, word for word. Recorded answers are stored labelled “entered by you”, shown with that label everywhere, and are never counted in measured rates or before/after evidence. Measured and recorded answers are never blended — this is enforced in code, not just promised.

What API measurement can tell you

  • Whether a named model, with live search on, mentioned or cited your business when asked a specific question on a specific date — word for word, stored.
  • How that same question, on that same engine, under the same settings, answered differently between two dates — with any change in our own method flagged on the comparison itself.
  • Aggregate mention rates over many checks, always with the denominator shown, and only above minimum sample sizes.

What it cannot tell you

  • What the consumer apps answered. The models are related and search the same web, but chatgpt.com is not the OpenAI API model, and gemini.google.com is not the Gemini API. We name exactly which instrument we measured; we never imply it was the app.
  • What any specific buyer saw. AI answers are non-deterministic and can be personalised — the same prompt, minutes apart, can produce different answers. No canonical answer exists for anyone to measure: not for us, and not for tools that read answers out of a browser.
  • A "position" or "rank". A claim like “you rank #3 in ChatGPT” needs a stable ordering that these systems do not produce — the order can change between two runs with identical settings. We do not sell position claims at all.

Exactly what we send — nothing hidden

The tracked question is sent word for word as the user message. Around it we set a short system prompt that frames the model as a general-purpose assistant and tells it to search before answering and to say so rather than guess when it cannot find reliable information — the same kind of frame every consumer surface applies, made visible. Answers are capped at 600 tokens; temperature is the provider default (the consumer apps do not pin it either); no location or personalisation parameters are sent, so answers reflect the providers’ default search context, not any specific buyer’s. The frame is versioned (presence-v1); if it ever changes, results measured under different versions are never compared as if the method held.

The system prompt, verbatim
You are a general-purpose AI assistant answering an everyday question about a business or place, the way you would for any user. If web search is available, search for the business before answering. Answer naturally in 2–4 sentences based on what you find or reliably know. If you cannot find reliable information about this specific business, say so plainly instead of guessing or inventing details.

Why we do not sell “position” — and why held-method measurement is the stronger claim

Some tools read answers out of the consumer interfaces and report where a brand “ranks”. Whatever is doing the reading — an API or a browser — is sampling a system that is non-deterministic and can be personalised. A browser-based tool is sampling its own signed-in account’s view, shaped by that account’s settings: a single web-search toggle has been shown to move a brand between position 1 and position 8 on the same prompt. That is not a flaw in one tool — it is what happens when a rank is claimed over a distribution that reshuffles between runs. No tool, ours included, can measure “the” answer a buyer sees, because no such single answer exists.

So we sell the two things that survive that fact. First: whether your data is machine-readable and verifiable — deterministic checks over your actual pages, feeds, and structured data, where the same input gives the same result and a re-crawl proves a fix landed. No sampling problem exists there at all. Second: change in the same prompt on the same engine between two dates, with the method held constant — same model id, same retrieval state, same frame, and the method printed on the comparison. For measuring change, holding the instrument steady matters more than matching any one person’s personalised view; that is how measurement works everywhere else, too.

Single answers are shown as dated exhibits, word for word — claims about what was said on a date, which stay true regardless of variance. Rates and before/after deltas are stated only with their denominators and only above minimum sample sizes, and repeated-sample runs carry intervals. A tool claiming your client “ranks #3” is making a weaker claim than either: it needs a canonical ranking that does not exist. Ours are built to not need one.

The metric: cited-answer share

Cited-answer share = of the successful sampled answers in a period, the fraction that mention the business.

  • A sampled answer is one real API call to a named model — the exact model id is stored with every check. Answers are stored verbatim; there is no fabrication, templating, or estimation path anywhere in the pipeline.
  • Checks run search-grounded, because consumer surfaces search before answering. Every stored check records whether live search really ran; only grounded, API-measured checks count in this metric. Ungrounded checks and owner-recorded answers are excluded from the rate, counted, and shown labelled — never pooled.
  • Only self-questions (questions about this business) count toward its share. Competitor questions measure rivals and are excluded.
  • A failed call is stored as a failure and excluded from the denominator. It is never counted as “not mentioned”.

Citation Velocity: before/after around a change

Citation Velocity is the percentage-point change in cited-answer share between a fixed window of ISO weeks before a recorded change and the same-length window after it. Six design rules, all enforced in code:

  1. The change week is excluded

    It is partially before and partially after the change; counting it would smear the boundary in whichever direction flatters the number.

  2. Fixed windows

    Default 4 ISO weeks per side, always the weeks nearest the change. Distant history cannot be cherry-picked into the baseline.

  3. Sample-size gate

    Fewer than 6 successful checks on either side → insufficient sample, and no delta is stated. 6–11 per side → medium confidence; 12+ → high. Confidence is computed once and never upgraded downstream.

  4. Fixed question set

    Baseline weeks retain their own stored checks, so adding flattering questions after a change does not rewrite the baseline.

  5. Deterministic event identity

    Every mirrored analytics event carries a UUID derived from its identifying facts, so retries and cron overlaps cannot double-count evidence.

  6. Attribution is correlational

    A before/after delta around a change is evidence of association, not proof of causation. Confounders (model updates, seasonality, press, competitor changes) are not controlled in a single-business measurement, and every emitted event carries that caveat.

What may be claimed

  • “In our sampled measurements, answers on {engine} mentioned {business} in {after}% of checks in the 4 weeks after the fix, up from {before}% in the 4 weeks before (n={n_before}/{n_after}, confidence: {confidence}).”
  • “Across N businesses that deployed fixes in {period}, the median change in cited-answer share was +X points within 4 weeks (methodology attached).”
  • “IndexNow submission accepted (HTTP 202)” — a receipt, nothing more.

What may never be claimed

  • “Guaranteed citations / rankings / traffic / sales.”
  • “#1 cited source” — no ranking ground truth exists for consumer AI surfaces.
  • “Schema confirmed ingested by {provider}” — no provider issues ingestion receipts; we can only sample answers.
  • “+X% in 24 hours” — below crawl-and-index latency; our own dashboard would falsify it.
  • Any number whose window failed the sample-size gate.

Data lineage

Every number can be traced back through this chain — each stage keeps its own honesty property.

StageStoreHonesty property
Question askedtracked_questionsOwner-editable, deterministic templates
Answer sampledcitation_checksVerbatim answer, exact model id, grounded flag, failures stored as failures
Change recordedchange_eventsOnly real accepted actions (verified IndexNow receipts, declared fixes)
Velocity computedin code, pure functionGates and exclusions above, unit-tested
Evidence mirroredPostHog citation_velocity_recordedDeterministic UUID, caveat attached, measured results only

The shortest honest summary: we measure readiness and record evidence. We never guarantee a citation, a ranking, traffic, or a sale.