← Back to blog

How Many Prompts Should You Actually Track for AI Visibility? The Sample-Size Math Most GEO Dashboards Skip

Run the same prompt through the same model twice, minutes apart, and it will repeat only about 40% of the domains it cited the first time. Same model, same question, no change in the world between the two runs. In a paired-run test of 20 prompts against claude-sonnet-5 with web search enabled, mean overlap of cited domains was 0.40, with per-prompt overlap ranging from 0.06 to 1.00. (cloro, August 2026)

That is a small experiment and it should be read as one, but the instability it demonstrates is not in doubt. And it is the single most important fact about the GEO dashboard your team just started paying for, because almost every AI visibility number a lean B2B team looks at is built on far fewer samples than the noise requires.

Most teams get the shape of the investment backwards. They spend the budget on prompt coverage, read the chart daily, and treat a screenshot of ChatGPT naming a competitor as evidence of something. The math says to do close to the opposite.

Why one AI answer proves nothing

Citation is probabilistic per answer. Each response is one draw from a pool of plausible sources, which means a brand can be genuinely visible on a prompt and still be missing from any individual answer.

Model tier moves the numbers further. Across the same prompt set, claude-haiku-4-5 averaged 4.0 sources per answer, claude-sonnet-5 averaged 6.5, and claude-opus-5 averaged 8.8, with cross-tier domain overlap falling to 0.22 to 0.25. (cloro, August 2026) Bigger models cite more sources, which opens more citation slots, and they draw from measurably different source pools. Your visibility can change because the user upgraded their plan.

There is a structural reason this instability is not going away either. Ahrefs found that 28.3% of ChatGPT's most-cited pages have zero organic visibility in Google, and fewer than 10% of the sources cited across ChatGPT, Gemini and Copilot rank in the top 10 Google organic results for the same query. (2026 GEO statistics roundup) The candidate pool these systems draw from is not your rank report. You cannot derive it, so you have to sample it.

The practical consequence: presence is a rate over many draws. Screenshots are anecdotes. A dashboard that shows you today's answer is showing you one coin flip.

The math that sizes your prompt set

AI visibility sample size is a confidence-interval problem, and the arithmetic is unforgiving in a useful way.

A visibility rate near 25% measured over 72 answers carries a 95% confidence interval of roughly plus or minus 10 percentage points. Over about 300 answers it tightens to roughly plus or minus 5. To claim with standard statistical power that visibility moved from 20% to 30%, you need about 294 answers in each period. (cloro, August 2026)

One prompt checked daily produces 7 answers a week. That resolves nothing at all.

The good news is that answers multiply fast, because the volume is prompts times engines times days. Fifty prompts across six engines sampled once a day is 2,100 answers a week. At that volume, a week-over-week comparison is statistically sound. A day-over-day comparison at the same volume still is not, because a single day is only 300 answers spread across engines that behave differently from one another.

That one sentence is worth internalising before your next reporting cycle: at a normal lean-team setup, weekly reads are real and daily reads are noise.

How many prompts, really

Between 50 and a few hundred distinct prompts covers most B2B teams. The constraint is distinctness, not count.

Ten phrasings of "best rank tracker" are one intent sampled ten times. That is genuinely useful as extra samples, and it is not extra coverage. Build the set from intents you actually sell into: category questions, competitor comparisons, problem phrasings, and how-to questions around your product. Then add phrasing variants deliberately and tag them as replicas so your reporting treats them as what they are.

The market has converged on the same order of magnitude. Peec's published plans cap tracked prompts at 50, 150 and 350. (cloro, August 2026) That is not a coincidence; it is roughly where the statistics stop paying you back.

A cap changes strategy rather than just budget. A 50-prompt plan is a sample of your intent space, not the space itself. Spend it where a citation is worth money, and accept explicitly that the long tail goes unmeasured until the budget grows. Writing that sentence down is what stops someone from later reporting the untracked tail as a decline.

Two more choices worth making on purpose:

  • Write prompts the way buyers type them. Full questions, not keyword strings. Engines rewrite prompts into their own search queries, and conversational phrasing survives that rewrite better.
  • Verify retrieval before you commit budget. Run a candidate prompt a few times and look at what the engines actually pull. A prompt whose answers never touch your category cannot measure you, no matter how well it is phrased.

And when the trade-off is more prompts versus more daily runs of the same prompt, take more prompts. Variance between prompts is larger than variance within one, so breadth buys more statistical information than depth. (cloro, August 2026)

Four rules that keep the number honest

1. Weekly aggregates are the smallest sound unit. Compare periods, not days. If someone asks what happened yesterday, the correct answer is that yesterday is not a measurable unit at this sample size.

2. Never pool engines into one rate. Each engine has its own citation behaviour, so a single blended number describes none of them. Read per-engine rates and let each carry its own interval. A pooled "AI visibility score" is the metric most likely to move for reasons nobody can explain.

3. Hold the prompt set frozen inside a measurement window, and date every change. Adding prompts mid-window changes the denominator and manufactures a trend. Batch prompt changes. The prompt list is your instrument, and recalibrating an instrument mid-experiment invalidates the read.

4. Size the read to the claim. Detecting a 2-point move needs several times the volume of a 10-point move. When the dashboard cannot supply the volume, the honest report is "no detectable change," in those words. The discipline of saying so is what makes the bigger claims credible later.

A four-week plan a lean team can actually run

Week 0: freeze the instrument. Write the prompt set, say 60 prompts across your real intents. Pick the engines. Fix the competitor panel. Date the file. Nothing in it changes for four weeks.

Weeks 1 and 2: baseline, in silence. Run daily and report nothing. Two weeks at 60 prompts across six engines is roughly 5,000 answers, enough to put a plus-or-minus-5-point interval on rates near 25%. Resist reading the day charts. They are precisely the noise the plan exists to average over.

Week 3: first read. Compute per-engine weekly rates and publish the denominators alongside them. Store the raw answers behind the number, not just the rate, because every later comparison is against this one.

Week 4: first comparison. Two real points, one delta. If the delta is smaller than the interval, report "no detectable change."

After week 4: intervene one variable at a time. Ship the content change, the PR push, or the new pages. Date the intervention. Judge it against the frozen baseline no earlier than two weeks after it lands. An intervention inside the baseline window restarts the baseline, and that rule has no exceptions that leave the data readable.

The cost of the data itself is trivial at this scale. The expensive part is the month of restraint.

Alert design, or how not to get paged about nothing

Alert on runs of misses long enough that chance cannot explain them, never on a single missing answer.

The variance sets the bar. If your brand normally appears in half of a prompt's answers, three consecutive daily misses on that one prompt still happens by chance more than one day in ten. As an alert, that pages you weekly about nothing and trains the team to ignore the channel. (cloro, August 2026)

Three rules make alerting survivable:

  • Alert on windows, not answers. "Mention rate across this prompt group below 15% for 7 consecutive days" fires on real disappearance. "We were not in this morning's answer" fires on the same 40% domain churn measured between runs minutes apart.
  • Scale the window to the prompt's volume. A brand tracked on one prompt needs a longer run of misses before the alert means anything than a brand tracked on twenty, because twenty prompts a day is twenty draws and the aggregate stabilises faster.
  • Route engine-wide drops differently. If every brand's rate falls on one engine at once, the engine changed its behaviour. That alert belongs to whoever owns the measurement, not to the brand team.

What to do once the measurement is trustworthy

Only then is it worth arguing about tactics. For what it is worth, the original Princeton-led GEO research, published in November 2023 and presented at KDD 2024, found that optimization methods lifted visibility in generative engine responses by up to 40% across a 10,000-query benchmark, with quotations (up to +41%), statistics (+30% to +40%) and cited sources (+30%) leading, while keyword stuffing performed worse than doing nothing.

Those are effect sizes worth chasing. They are also exactly the size of effect that a badly instrumented dashboard will fail to detect, or will invent in a week where nothing happened.

Being in the engine's candidate pool across many related queries beats optimising for one prompt's answer, and you can only tell the difference between those two outcomes if the measurement holds still long enough to show you.

The takeaway

Your AI visibility number is a rate with an interval around it, and the interval is almost certainly wider than the change your last report claimed. Fifty to a few hundred distinct prompts, every engine read separately, a frozen set inside a dated window, weekly aggregates, and the willingness to say "no detectable change" out loud. That is the whole discipline.

This is the kind of measurement rigour Nukipa builds into a lean team's GTM system from the start, so the visibility number in your board deck survives the first person who asks how it was calculated. If you want a second set of eyes on how your AI visibility is currently being measured, test Nukipa.

  1. AI Visibility Sample Size: How Many Prompts and Runs (cloro, August 14 2026)
  2. 70+ Generative Engine Optimization (GEO) Statistics for 2026 (Peec AI)

Related