Parse
Pricing
Sign inCheck your brandOpening…
  1. Research
  2. How many prompts do you need to measure AI visibility?
ResearchHow many prompts do you need to measure AI visibility?

How many prompts do you need to measure AI visibility?

Ten reruns of one prompt only tell you as much as two different prompts. Holding a brand's AI visibility to plus or minus 5 points took about 145 prompts asked once.

2.0
ten reruns of one prompt carried the information of about two different prompts
  • The finding
  • How we measured
  • Sources
  • More like this

Ten reruns of one prompt are worth about two prompts

A market is a set of buyer questions about one product category. A brand-market series is one brand's visibility on one engine in one market. Its visibility rate is the share of AI answers naming the brand, averaged over the market's prompts. A tracking design is k prompts asked r times each.

Reruns of the same prompt are correlated, because the prompt fixes much of the answer. Measured across 12,053 series, ten reruns of one prompt carried the information of 2.02 prompts asked once, and three reruns carried 1.60 prompts' worth. Two independent computations agree: 2.02 from the measured variance components and 2.02 from the observed disagreement between random samples. On the 72 market-engine pairs with at least 20 prompts the figure was 2.10.

The practical form: three different prompts asked once had a typical error of 19.97 points, and one prompt asked ten times had 24.28 points. Typical error is the root-mean-square gap, in percentage points, between a design's reading and the brand's full-window visibility rate.

2.0
prompts' worth of information in ten reruns of one prompt (2.02 from components, 2.02 observed)

Takeaway

A rerun of a prompt you already track buys far less information than one prompt you do not.

At the same answer count, prompt-heavy designs are about twice as accurate

Ten answers can be one prompt asked ten times or ten prompts asked once. The ten-prompt design had a typical error of 11.28 points; the one-prompt design had 24.28, 2.2 times larger. Thirty answers as 10 prompts × 3 runs gave 8.77 points; as 3 prompts × 10 runs, 14.07, 1.6 times larger. Three answers as 3 prompts × 1 run gave 19.97; as 1 prompt × 3 runs, 27.30.

The 10-prompt designs are measured on the 72 market-engine pairs with at least 20 prompts, so that two teams can hold disjoint 10-prompt sets. On that same panel, one prompt asked ten times gave 24.47 points and 3 × 10 gave 14.16, so the comparison holds on matched markets.

Typical error by design at equal answer counts

  • 10 prompts × 1 run (10 answers)11.28 pts
  • 1 prompt × 10 runs (10 answers)24.28 pts
  • 10 prompts × 3 runs (30 answers)8.77 pts
  • 3 prompts × 10 runs (30 answers)14.07 pts
  • 3 prompts × 1 run (3 answers)19.97 pts
  • 1 prompt × 3 runs (3 answers)27.30 pts
10 prompts × 1 run (10 answers) is in color · grey is the other rows. Points are percentage points of visibility rate; lower is better.

Takeaway

Spend the next unit of budget on a new prompt until the prompt set is large; only then add runs.

The full error table: 1 to 10 prompts by 1 to 10 runs

The table gives the typical error and the 95% band for every measured design. The 95% band is ±1.96 times the typical error: the range a single reading falls in 19 times out of 20. The run-only error is the disagreement between two teams that track the same prompts on different runs; it isolates rerun noise.

Ten prompts asked once (10 answers) had a 95% band of ±22.1 points. Ten prompts asked ten times (100 answers) narrowed it to ±15.1. One prompt asked ten times stayed at ±47.6.

Typical error and 95% band by design

11134.567.6325.86
13327.353.5114.92
31319.9739.1314.93
51515.4530.2711.59
33915.8231.018.63
1101024.2847.68.17
1011011.2822.118.53
531512.1923.896.68
3103014.0727.574.73
103308.7717.24.9
5105010.8321.223.65
10101007.715.092.7
Two disjoint random samples per market and engine, 20 draws each; 12,053 series (1,303 for 10-prompt designs).

Takeaway

Ten prompts asked once already beat any single-prompt design; ten prompts asked ten times are still ±15 points.

Prompt choice, not rerun noise, drives most of the disagreement

Split each series' answer-level variation into a between-prompt part and a run-to-run part. The between-prompt part was 43.8% of the total when pooled over all 12,053 series (median series 40.2%). This is the prompt share of variance: how much of whether a brand appears in an answer is settled before the engine runs.

Reruns average out only the run-to-run part. For five-prompt designs, prompt choice accounted for 43.7% of the disagreement between two teams at one run, 70.0% at three runs, and 88.6% at ten runs. The more you rerun, the more of what remains is the prompt set.

Share of two-team disagreement driven by prompt choice, 5-prompt designs

  • 1 run43.7%
  • 3 runs70.0%
  • 10 runs88.6%
10 runs is in color · grey is the other rows. 1 minus the run-only disagreement variance divided by the total disagreement variance.

Takeaway

After three runs, most of the remaining error is prompt choice, and only new prompts can reduce it.

Binomial margin tables understate the error by up to 2.3 times

Vendor guides compute a margin as 1.96 × √(p(1−p)/n), with n the number of answers. That formula assumes every answer is an independent draw. Across the measured designs the observed 95% band was 1.06 times the formula at one run per prompt, 1.44 to 1.45 times at three runs, and 2.28 to 2.36 times at ten runs.

Example: 5 prompts × 10 runs is 50 answers. The formula promises ±9.1 points; the observed band was ±21.2. With one run per prompt the formula is close. Each added run widens the gap, because the added answers repeat the same prompts.

Binomial formula vs observed 95% band

10 prompts × 10 runs1006.6315.092.276
10 prompts × 3 runs3012.117.21.421
10 prompts × 1 run1020.9622.111.055
5 prompts × 10 runs509.0621.222.343
5 prompts × 3 runs1516.5423.891.444
5 prompts × 1 run528.6430.271.057
Formula band averaged over each series' own visibility rate.

Takeaway

Use the binomial formula only with n equal to the number of prompts, not the number of answers.

What common tracking setups deliver

Using the measured variance components, the table computes the 95% band for setups practitioners recommend. 15 prompts × 5 runs, a common persona setup, gives ±12.9 points from 75 answers. 40 prompts asked once gives ±10.7 from 40 answers. The rule that about 300 answers resolve a rate to ±5 points holds when those are 300 different prompts asked once (±3.9). As 30 prompts × 10 runs, the same 300 answers give ±8.7; as 100 prompts × 3 runs, ±5.3.

The binomial column shows what the independent-answers formula promises for the same design. The two columns agree only when each prompt is asked once.

95% band for common tracking setups

1 prompt × 10 runs1047.4920.25
10 prompts × 1 run1021.3620.25
20 prompts × 1 run2015.114.32
40 prompts × 1 run4010.6810.13
15 prompts × 5 runs7512.947.4
100 prompts × 1 run1006.756.41
50 prompts × 3 runs1507.555.23
150 prompts × 1 run1505.515.23
30 prompts × 10 runs3008.673.7
50 prompts × 6 runs3006.973.7
100 prompts × 3 runs3005.343.7
300 prompts × 1 run3003.93.7
50 prompts × 10 runs5006.722.86
500 prompts × 1 run5003.022.86
Computed from the measured between-prompt and run-to-run variance components, pooled over 12,053 series.

Takeaway

Judge a tracking plan by its prompt count first; the answer count flatters rerun-heavy designs.

Getting to ±5 points takes about 145 prompts asked once

For each series we computed the prompt count at which the 95% band reaches ±5 points. The median series needed 144.5 prompts asked once, 92.3 prompts asked three times (277 answers), 71.4 asked ten times (714 answers), or 65.0 asked thirty times (1,950 answers). For ±10 points the median needs were 36.1, 23.1, 17.9, and 16.2 prompts.

A quarter of series needed more than 247.9 prompts asked once to reach ±5 points. These figures extend the measured components beyond ten prompts; the components reproduced the observed error within 0.1 points at all twelve measured designs.

Median prompts needed for a ±5-point 95% band, by runs per prompt

  • 1 run144.5 prompts
  • 3 runs92.3 prompts
  • 10 runs71.4 prompts
  • 30 runs65.0 prompts
1 run is in color · grey is the other rows. Median across 12,053 series; the 75th percentile at one run is 247.9 prompts.

Takeaway

A 20-to-40-prompt set reports a brand's visibility to roughly ±10 to ±15 points, not ±5.

Reruns cannot buy their way past the prompt floor

With unlimited reruns, rerun noise vanishes and only the between-prompt part remains. That floor is ±44.7 points for one prompt, ±14.1 for 10 prompts, ±10.0 for 20, ±6.3 for 50, and ±4.5 for 100. Ten runs get most of the way to the floor: at 10 prompts the typical error is 10.90 points at one run, 8.62 at three, 7.66 at ten, and 7.21 with unlimited runs.

The curves below give the typical error by prompt count at one, three, ten, and unlimited runs, computed from the measured components.

Typical error by prompt count and runs per prompt

Percentage points; the unlimited-runs curve is the between-prompt floor.

Takeaway

The prompt count sets the floor; runs only decide how close to it you get.

Mid-visibility brands need the most prompts

Error in percentage points is largest when a brand appears in some answers and not others. Series with 30–60% visibility needed a median of 382.4 prompts asked once to reach ±5 points; 15–30% needed 257.1; brands above 60% needed 330.2; brands at 5–15% needed 112.4. At 5 prompts × 3 runs, the typical error ran from 9.85 points in the 5–15% band to 17.82 in the 30–60% band.

The prompt share of variance was similar in every band, 42.4% to 46.4%, so the extra prompts are a matter of scale, not a different mechanism.

Median prompts asked once for a ±5-point 95% band, by visibility band

  • 5–15%112.4
  • 15–30%257.1
  • 30–60%382.4
  • 60%+330.2
30–60% is in color · grey is the other rows. Bands by full-window visibility rate; 8,001 / 2,262 / 1,414 / 376 series.

Takeaway

A challenger brand at 30–60% visibility is the hardest case to measure; budget prompts accordingly.

ChatGPT Search and Google AI Mode need the same number of prompts

Median prompts asked once for ±5 points: 143.9 on ChatGPT Search and 145.0 on Google AI Mode. The prompt share of variance was 46.0% on ChatGPT Search and 41.7% on Google AI Mode, so reruns average out slightly more on Google AI Mode: at ten runs the medians were 74.8 and 68.8 prompts. Typical error at 5 prompts × 3 runs was 12.41 and 11.97 points.

The two engines are scored on the same markets and prompts, so the near-identical figures are not a coverage artifact.

143.9 vs 145.0
median prompts asked once for ±5 points, ChatGPT Search vs Google AI Mode
46.0% vs 41.7%
prompt share of variance
74.8 vs 68.8
median prompts at ten runs for ±5 points
12.41 vs 11.97
typical error at 5 prompts × 3 runs (pts)

Takeaway

One prompt-count rule works for both engines; track them separately but size them the same.

Even 100 answers name the same market leader only about half the time

A market leader is the brand with the highest visibility rate in a team's reading; a leader is clear when no other brand ties it. With 10 prompts × 10 runs, both teams had a clear leader in 90.0% of 1,440 market samples, and when both did they named the same brand 53.6% of the time. At 5 prompts × 3 runs (15 answers), both were clear in 57.3% of 11,080 samples and agreed 45.3%; at 3 prompts × 1 run, both were clear in 12.8% and agreed 37.1%.

Leader identity depends on the gap between the top brands. In many markets several brands sit within a few points of each other, and a 10-prompt reading cannot separate them.

Market-leader agreement between two teams, by design

1111.6%24.0%
13312.3%16.3%
31312.8%37.0%
51525.1%43.1%
33940.4%37.3%
1101042.9%18.4%
1011046.9%49.9%
531557.3%45.3%
3103071.9%37.1%
1033073.6%52.5%
5105082.8%45.7%
101010090.0%53.5%
11,080 market samples per design (1,440 for 10-prompt designs); a leader is clear when no other brand ties it.

Takeaway

Report the leader with its margin over #2, and not from fewer than 10 prompts.

In named markets, two 15-answer readings rarely agreed on the leader

The table lists the 14 active markets with the most prompts in the panel. For each, it shows the leader and its visibility rate on each engine over the full window, and the share of 40 paired samples (20 per engine) at 5 prompts × 3 runs in which both teams had a clear leader and it was the same brand.

Agreement ran from 60.0% in online therapy platforms, where Talkspace led at 68.9% on ChatGPT Search and 69.7% on Google AI Mode, to 7.5% in pet insurance, where Pets Best led at 50.9% and 45.6% with several brands close behind. A dominant brand does not guarantee agreement: Pipedrive led CRM and sales pipeline software at 92.4% on ChatGPT Search, yet the 15-answer readings agreed only 30.0% of the time, because other brands tied it at the top of small samples.

Leader agreement at 5 prompts × 3 runs in 14 named markets

Online Therapy & Psychiatry Platforms32Talkspace logoTalkspace68.9%Talkspace logoTalkspace69.7%60.0%
Hard Money Lending Services40Lima One Capital50.9%Easy Street Capital62.3%50.0%
Business VoIP Phone Systems35RingCentral logoRingCentral64.4%RingCentral logoRingCentral56.4%42.5%
Online Stock & Options Brokers28Charles Schwab logoCharles Schwab77.4%Charles Schwab logoCharles Schwab69.1%40.0%
Payroll and HRIS Software29Gusto logoGusto80.8%Gusto logoGusto79.2%37.5%
Online Mattress Shopping and Comparison47Saatva logoSaatva54.4%Helix Sleep logoHelix Sleep35.2%30.0%
CRM and Sales Pipeline Software25Pipedrive logoPipedrive92.4%Pipedrive logoPipedrive82.3%30.0%
AI Customer Support Chatbots28Intercom logoIntercom51.3%Intercom logoIntercom36.7%22.5%
Construction Industry Software and Services30Procore logoProcore30.0%Procore logoProcore28.0%17.5%
Online Sports Betting Apps38FanDuel logoFanDuel91.2%FanDuel logoFanDuel83.7%15.0%
Meal Kit Delivery Services25HelloFresh logoHelloFresh65.7%Green Chef logoGreen Chef46.6%15.0%
Insurance Quote Comparison Tools25State Farm logoState Farm42.8%The Zebra logoThe Zebra61.0%12.5%
Pet Insurance Plans46Pets Best logoPets Best50.9%Pets Best logoPets Best45.6%7.5%
Personal Injury & Accident Lawyers36Morgan & Morgan9.1%Morgan & Morgan8.2%7.5%
Leader and rate are computed on every prompt and run in the window; agreement is over 40 paired samples.

Takeaway

The live version of each market's full-window leaderboard is on the public Parse index; a small prompt sample is not.

What this study counts, and what it leaves out

The panel keeps markets with at least 10 prompts that have 20 or more answers per engine in the window: 277 of 1,698 markets and 4,048 of 17,970 organic prompts. Series need at least 5% visibility. Raising that threshold to 10% leaves 6,113 series with a prompt share of 43.5% and a 5 × 3 typical error of 14.52 points, so the headline does not depend on the threshold.

Two disjoint samples from the same market give an unbiased error at any prompt count, because the finite-population correction cancels in their difference. The variance components predicted the observed typical error within 0.1 points at every design (5 × 3: 12.19 predicted, 12.19 observed; 1 × 10: 24.23 vs 24.28). Runs were drawn at random across 13 weeks, so run-to-run error includes drift inside the window. Brands resolve to their owning company, which is why one market whose leaders resolved to Alphabet and Meta was left out of the named table. The visibility rate is prompt-weighted, matching designs with equal runs per prompt.

277 of 1,698
markets kept: at least 10 prompts with 20+ answers per engine
6,113
series at a 10% threshold; prompt share 43.5%
0.1 pts
largest gap between predicted and observed typical error across 12 designs
20
random prompt-and-run samples per market and engine

Takeaway

The design was checked two ways; the numbers move by tenths of a point under the alternative choices.

The GEO takeaway

Size the prompt set first. Forty prompts asked once (±10.7 points) beat 15 prompts asked five times (±12.9) with about half the answers. Add runs only after the prompt set is as large as the budget allows; three runs capture most of the rerun benefit, and ten runs reach within a point of the floor.

Report every visibility rate with its 95% band from the prompt count, not the answer count, and treat a change smaller than the band as no change. Two teams with 10 prompts × 10 runs agreed on the market leader about half the time, so a leader claim needs a margin over the next brand, not just a rank.

Takeaway

Prompts first, runs second, and a band on every number.

How we measured

In the Parse index, we analyzed 186,216 AI answers to 4,048 organic prompts in 277 markets on ChatGPT Search and Google AI Mode from June 2 through August 31, 2026, covering 958,454 brand-answer pairs and 12,053 brand-market series, and drew 20 random prompt-and-run samples from every market on each engine.

2.0
prompts' worth of information in ten reruns of one prompt
43.8%
of answer-level variation in a brand's visibility is fixed by the prompt
145
prompts asked once to hold a brand's visibility rate to ±5 points
12,053
brand-market series analyzed

Get the data

Dataset CSVThe metrics behind every figure in this report.

Sources

These are the pages this study used.

  1. cloro — AI visibility sample size: how many prompts and runs (Ricardo Batista, August 14, 2026) · accessed September 7, 2026
  2. MaxAEO — How many prompts to test AI visibility? Sample size and prompt-run math (Chris Han, July 6, 2026) · accessed September 7, 2026
  3. Nick Lafferty — AI visibility metrics: formulas, benchmarks and sample sizes (updated August 20, 2026) · accessed September 7, 2026
  4. Search Engine Land — How to make prompt tracking much more accurate (Kevin Indig, June 10, 2026) · accessed September 7, 2026
  5. SE Ranking — How to choose prompts to track for AI visibility (Yevheniia Khromova, April 23, 2026) · accessed September 7, 2026

More like this

Which AI engine changes its answers the most?
Nearly a tie. Ask the same question again and ChatGPT Search keeps 40.5% of its brand list while Google AI Mode keeps 41.5% — and the less stable engine changes month to month.
Does AI recommend the same brand when you ask again?
Only about six in ten times. The top recommendation stayed the same in 90,817 of 161,023 consecutive same-prompt, same-engine answer pairs, or 56.4%.
Does question wording change which brands AI names?
Yes, a lot. Two phrasings of the same buyer question shared 11.7% of named brands on average; the identical phrasing re-asked one to three days later shared 38.0%.
AI citation volatility by industry: one-shot checks miss the signal
Ask the same AI question again and the cited sources usually change. Across repeated runs, ChatGPT answers shared only about 21% of their cited sources, and every industry showed high churn.
AI recommendation movers: how sticky is the top?
Two snapshots of where AI ranks brands, five months apart, on the same set of brands. The top is sticky but not frozen: four in ten of the top-100 brands churned out. A one-time AI-visibility reading is a weak signal.

About this research

Dimitry Apollonsky

Founder, Parse

I built Parse to track where AI answers really come from: the sources they cite and the brands they name. DM me on LinkedIn to talk shop.

See your brand's visibility rate with its error band, computed on a full prompt set.

Run a free check against live AI answers — no account needed.

Parse

See where your brand stands in AI recommendations.

Products

  • Brands
  • Markets
  • Integrations
  • Work with us
  • Pricing
  • MCP

Resources

  • Research
  • Methodology
  • Blog

© 2026 Parse. All rights reserved.

LegalPrivacy PolicyTerms of Service