Skip to content

Methodology

How KeenTag measures. Every step, published.

Short answerHow does KeenTag measure AI visibility?

We ask ChatGPT, Perplexity and Gemini the questions your buyers ask, several times each, and read every answer for your name. We report the share of answers that named you with its 95% Wilson range, never a score or a rank. After you ship a fix, a re-check says "changed beyond noise" only when the old and new ranges stop overlapping.

Engines
3
ChatGPT, Perplexity and Gemini
Samples
3–4
per question, per engine, per check
Range
95%
Wilson interval on every rate
Scores or ranks
0
A rate and its range instead

The questions

Short answerWhich questions do you ask?

The questions a buyer in your category would type into an AI assistant. The questions we suggest never contain your brand name, so being named has to come from the engine. You review and edit them before tracking starts, and we ask the same wording every time.

  • When you add a brand, we read your site and suggest questions for your category, for example “best invoicing tool for freelancers”. They are suggestions: you edit, add or remove them before anything is tracked.
  • Suggested questions never mention your brand. Keep it that way: a question that names you gets answers that name you, and the rate stops meaning anything.
  • We ask the same wording on every check, so this week’s answers are comparable with last week’s.
  • Plans track 25 (Solo), 60 (Pro) or 150 (Agency) questions per brand. The free check asks 3 questions, which you can edit before it runs.

The engines

Short answerWhich engines do you ask, and how?

ChatGPT, Perplexity and Gemini, on every plan. We ask through each engine's API with web search switched on, so answers reflect what the engine finds today rather than old training data. That is close to, but not the same as, what a signed-in person sees in the app.

  • ChatGPTChatGPTAPI, web search on, every plan
  • PerplexityPerplexityAPI, web search on, every plan
  • GeminiGeminiAPI, web search on, every plan

Google AI Overviews is planned for Pro and Agency. It is not live yet, and no rate includes it until it is. A request that errors or takes longer than 60 seconds is recorded as a no-answer and left out of the rate.

Samples per plan

Short answerHow many times do you ask each question?

On a plan, each question goes to each engine 3 times (4 on Agency) on every check. The free check asks once per engine. Every sample is a fresh request: we never count one answer twice.

Samples and answers per plan
PlanSamples per question, per engineAnswers per question, each checkQuestions per brandMost answers in one brand’s checkHow often
Solo39 = 3 × 325225Weekly tracking
Pro39 = 3 × 360540Daily tracking
Agency412 = 3 × 41501,800Daily tracking
Free check13 = 3 × 139Once, when you run it

Why sample more than once? The same question can get a different answer a minute later. One answer per engine is a reading, not a rate. More answers make the range narrower, which is why a full check over many questions is far more precise than any single question. Compare plans

What counts as named

Short answerWhat counts as being named?

An answer names you when your brand name appears as a whole word. The rate is the answers that named you divided by the answers that came back. Failed requests are left out; unclear matches count as answers but not as named.

  • Named

    Your brand name appears as a whole word, not inside another word. If your name is also an everyday word (like Linear or Notion), it must appear capitalised as a name, or the answer must cite your own domain.

    Counts as named, and as an answer

  • Not named

    An answer came back without your name. We note which of your competitors it named instead.

    Counts as an answer, not as named

  • Unclear

    An everyday-word name appears only in lower case with nothing else pointing to you, so it is probably the ordinary word.

    Counts as an answer, not as named. Never treated as a gap to fix

  • No answer

    The engine returned an error or did not answer within 60 seconds.

    Left out of the rate entirely

rateanswers that named youanswers that came back

Across questions and engines we add up the named answers and the answers, then compute one range from the totals. We don’t average rates, so a question with more answers carries more weight, as it should. We keep each full answer, so you can check any count yourself.

The 95% range

Short answerHow is the 95% range computed?

With the Wilson score interval at 95% confidence (z = 1.96), a standard method for a share measured from few answers. It never goes below 0% or above 100%, and it stays honest when you are named in none or all of the answers.

Wilson score interval

p = named ÷ answers, n = answers, z = 1.96centre = (p + z²/2n) ÷ (1 + z²/n)half-width = z × √(p(1 − p)/n + z²/4n²) ÷ (1 + z²/n)range = centre ± half-width

Try it. This calculator runs the same function the product uses to compute every range we report.

Examples
no-answers left out

With these numbers

p = 3 / 9 = 0.333centre = (p + z²/2n) / (1 + z²/n) = 0.383half = z·√(p(1−p)/n + z²/4n²) / (1 + z²/n) = 0.263range = 0.121 to 0.646, z = 1.96
Mention rate with its 95% range

Named in 3 of 9 answers: a rate of 33%, likely between 12% and 65%.

Same rate, more answers
AnswersRate95% rangeWidth
933%12–65%53 pts
3633%20–50%30 pts
14433%26–41%15 pts

How a re-check decides

Short answerHow does a re-check decide if something changed?

We re-ask that one question on all 3 engines at your plan's sample count and compare it with the question's answers from the previous 90 days. If the two 95% ranges don't overlap, the result is "changed beyond noise". If they overlap, it is "no detectable change".

  1. Baseline. Every answer to that question from the 90 days before the re-check, pooled into one rate and range.
  2. Re-check. A fresh run of that one question on ChatGPT, Perplexity and Gemini, at your plan’s sample count (3 or 4 per engine).
  3. Compare. If the ranges don’t overlap: changed beyond noise, and we say whether it went up or down. If they overlap: no detectable change.
  4. Record.We save the before and after with their counts. A re-check that got no answers at all is not recorded, so it can’t show up later as a fall.

“No detectable change” doesn’t mean the fix failed. It means these answers can’t tell a real change from noise yet; tracking keeps collecting. And when ranges do separate, we show the before and after rather than claim the fix caused it: engines also change on their own.

Examples
Baseline, last 90 days
Re-check, after the fix
Baseline12 of 117 answers
10%
Re-check7 of 9 answers
78%
Changed beyond noise: up

Baseline 6–17%, re-check 45–94%. The re-check's range sits entirely above the baseline's. The difference is larger than run-to-run noise.

What we keep, and for how long

Short answerWhat do you keep, and for how long?

For tracked questions we keep each full answer, the engine, when it was asked, whether you were named, who was named instead and the sources cited, for as long as your account is active. A free check without an account is deleted after 30 days. The privacy policy is the full, binding version.

What we keep and for how long
WhatHow long
Your account and tracking data: brands, questions, competitors and check history, including every answerWhile your account is active. Erased when you delete your account; it leaves our database backups within 30 days.
A free check without an account: the domain and brand name, suggested category and competitors, your questions, the answers with their sources, and a salted one-way hash of your IP address (never the address itself)Deleted after 30 days. Backups roll off within their own retention window.
Billing and invoice recordsHeld by Paddle, our merchant of record, for the period tax law requires (up to 6 years).
Optional product analytics (only with your consent)12-month retention setting, deleted earlier on request or when you delete your account.
Hosting request logsNo more than 90 days.

Your questions are sent to OpenAI, Perplexity and Google to get the answers. You can export your data or delete your account yourself at any time. The privacy policy is the binding version; if this summary and the policy ever differ, the policy wins.

What this can't tell you

Short answerWhat can't this measure?

Personalised answers, apps and engines we don't query, and anything an engine says between checks. A rate tells you how often you are named for these questions, not your traffic or your sales.

  • Personal answers. A signed-in person can get answers shaped by their history, settings and location. We ask through the API, with no personal history.
  • Between checks. A rate describes the answers we collected, on the days we collected them.
  • Why an engine changed. Engines update their models and sources without notice. A rate can move with nothing changing on your side.
  • Business outcomes.We don’t predict traffic, leads or revenue, and we don’t guarantee any result.
  • A score or a rank.We don’t compute one. A single number hides how sure it is; a rank between brands named a few times each is mostly noise.

Changes to this page

  1. First published.

When the method changes, we update this page, its date and this log. Questions about the method? Write to us.