How many prompts do you need to measure AI visibility?
Ten reruns of one prompt only tell you as much as two different prompts. Holding a brand's AI visibility to plus or minus 5 points took about 145 prompts asked once.
Ten reruns of one prompt are worth about two prompts
A market is a set of buyer questions about one product category. A brand-market series is one brand's visibility on one engine in one market. Its visibility rate is the share of AI answers naming the brand, averaged over the market's prompts. A tracking design is k prompts asked r times each.
Reruns of the same prompt are correlated, because the prompt fixes much of the answer. Measured across 12,053 series, ten reruns of one prompt carried the information of 2.02 prompts asked once, and three reruns carried 1.60 prompts' worth. Two independent computations agree: 2.02 from the measured variance components and 2.02 from the observed disagreement between random samples. On the 72 market-engine pairs with at least 20 prompts the figure was 2.10.
The practical form: three different prompts asked once had a typical error of 19.97 points, and one prompt asked ten times had 24.28 points. Typical error is the root-mean-square gap, in percentage points, between a design's reading and the brand's full-window visibility rate.
Takeaway
At the same answer count, prompt-heavy designs are about twice as accurate
Ten answers can be one prompt asked ten times or ten prompts asked once. The ten-prompt design had a typical error of 11.28 points; the one-prompt design had 24.28, 2.2 times larger. Thirty answers as 10 prompts × 3 runs gave 8.77 points; as 3 prompts × 10 runs, 14.07, 1.6 times larger. Three answers as 3 prompts × 1 run gave 19.97; as 1 prompt × 3 runs, 27.30.
The 10-prompt designs are measured on the 72 market-engine pairs with at least 20 prompts, so that two teams can hold disjoint 10-prompt sets. On that same panel, one prompt asked ten times gave 24.47 points and 3 × 10 gave 14.16, so the comparison holds on matched markets.
Typical error by design at equal answer counts
- 10 prompts × 1 run (10 answers)11.28 pts
- 1 prompt × 10 runs (10 answers)24.28 pts
- 10 prompts × 3 runs (30 answers)8.77 pts
- 3 prompts × 10 runs (30 answers)14.07 pts
- 3 prompts × 1 run (3 answers)19.97 pts
- 1 prompt × 3 runs (3 answers)27.30 pts
Takeaway
The full error table: 1 to 10 prompts by 1 to 10 runs
The table gives the typical error and the 95% band for every measured design. The 95% band is ±1.96 times the typical error: the range a single reading falls in 19 times out of 20. The run-only error is the disagreement between two teams that track the same prompts on different runs; it isolates rerun noise.
Ten prompts asked once (10 answers) had a 95% band of ±22.1 points. Ten prompts asked ten times (100 answers) narrowed it to ±15.1. One prompt asked ten times stayed at ±47.6.
Typical error and 95% band by design
| 1 | 1 | 1 | 34.5 | 67.63 | 25.86 |
| 1 | 3 | 3 | 27.3 | 53.51 | 14.92 |
| 3 | 1 | 3 | 19.97 | 39.13 | 14.93 |
| 5 | 1 | 5 | 15.45 | 30.27 | 11.59 |
| 3 | 3 | 9 | 15.82 | 31.01 | 8.63 |
| 1 | 10 | 10 | 24.28 | 47.6 | 8.17 |
| 10 | 1 | 10 | 11.28 | 22.11 | 8.53 |
| 5 | 3 | 15 | 12.19 | 23.89 | 6.68 |
| 3 | 10 | 30 | 14.07 | 27.57 | 4.73 |
| 10 | 3 | 30 | 8.77 | 17.2 | 4.9 |
| 5 | 10 | 50 | 10.83 | 21.22 | 3.65 |
| 10 | 10 | 100 | 7.7 | 15.09 | 2.7 |
Takeaway
Prompt choice, not rerun noise, drives most of the disagreement
Split each series' answer-level variation into a between-prompt part and a run-to-run part. The between-prompt part was 43.8% of the total when pooled over all 12,053 series (median series 40.2%). This is the prompt share of variance: how much of whether a brand appears in an answer is settled before the engine runs.
Reruns average out only the run-to-run part. For five-prompt designs, prompt choice accounted for 43.7% of the disagreement between two teams at one run, 70.0% at three runs, and 88.6% at ten runs. The more you rerun, the more of what remains is the prompt set.
Share of two-team disagreement driven by prompt choice, 5-prompt designs
- 1 run43.7%
- 3 runs70.0%
- 10 runs88.6%
Takeaway
Binomial margin tables understate the error by up to 2.3 times
Vendor guides compute a margin as 1.96 × √(p(1−p)/n), with n the number of answers. That formula assumes every answer is an independent draw. Across the measured designs the observed 95% band was 1.06 times the formula at one run per prompt, 1.44 to 1.45 times at three runs, and 2.28 to 2.36 times at ten runs.
Example: 5 prompts × 10 runs is 50 answers. The formula promises ±9.1 points; the observed band was ±21.2. With one run per prompt the formula is close. Each added run widens the gap, because the added answers repeat the same prompts.
Binomial formula vs observed 95% band
| 10 prompts × 10 runs | 100 | 6.63 | 15.09 | 2.276 |
| 10 prompts × 3 runs | 30 | 12.1 | 17.2 | 1.421 |
| 10 prompts × 1 run | 10 | 20.96 | 22.11 | 1.055 |
| 5 prompts × 10 runs | 50 | 9.06 | 21.22 | 2.343 |
| 5 prompts × 3 runs | 15 | 16.54 | 23.89 | 1.444 |
| 5 prompts × 1 run | 5 | 28.64 | 30.27 | 1.057 |
Takeaway
What common tracking setups deliver
Using the measured variance components, the table computes the 95% band for setups practitioners recommend. 15 prompts × 5 runs, a common persona setup, gives ±12.9 points from 75 answers. 40 prompts asked once gives ±10.7 from 40 answers. The rule that about 300 answers resolve a rate to ±5 points holds when those are 300 different prompts asked once (±3.9). As 30 prompts × 10 runs, the same 300 answers give ±8.7; as 100 prompts × 3 runs, ±5.3.
The binomial column shows what the independent-answers formula promises for the same design. The two columns agree only when each prompt is asked once.
95% band for common tracking setups
| 1 prompt × 10 runs | 10 | 47.49 | 20.25 |
| 10 prompts × 1 run | 10 | 21.36 | 20.25 |
| 20 prompts × 1 run | 20 | 15.1 | 14.32 |
| 40 prompts × 1 run | 40 | 10.68 | 10.13 |
| 15 prompts × 5 runs | 75 | 12.94 | 7.4 |
| 100 prompts × 1 run | 100 | 6.75 | 6.41 |
| 50 prompts × 3 runs | 150 | 7.55 | 5.23 |
| 150 prompts × 1 run | 150 | 5.51 | 5.23 |
| 30 prompts × 10 runs | 300 | 8.67 | 3.7 |
| 50 prompts × 6 runs | 300 | 6.97 | 3.7 |
| 100 prompts × 3 runs | 300 | 5.34 | 3.7 |
| 300 prompts × 1 run | 300 | 3.9 | 3.7 |
| 50 prompts × 10 runs | 500 | 6.72 | 2.86 |
| 500 prompts × 1 run | 500 | 3.02 | 2.86 |
Takeaway
Getting to ±5 points takes about 145 prompts asked once
For each series we computed the prompt count at which the 95% band reaches ±5 points. The median series needed 144.5 prompts asked once, 92.3 prompts asked three times (277 answers), 71.4 asked ten times (714 answers), or 65.0 asked thirty times (1,950 answers). For ±10 points the median needs were 36.1, 23.1, 17.9, and 16.2 prompts.
A quarter of series needed more than 247.9 prompts asked once to reach ±5 points. These figures extend the measured components beyond ten prompts; the components reproduced the observed error within 0.1 points at all twelve measured designs.
Median prompts needed for a ±5-point 95% band, by runs per prompt
- 1 run144.5 prompts
- 3 runs92.3 prompts
- 10 runs71.4 prompts
- 30 runs65.0 prompts
Takeaway
Reruns cannot buy their way past the prompt floor
With unlimited reruns, rerun noise vanishes and only the between-prompt part remains. That floor is ±44.7 points for one prompt, ±14.1 for 10 prompts, ±10.0 for 20, ±6.3 for 50, and ±4.5 for 100. Ten runs get most of the way to the floor: at 10 prompts the typical error is 10.90 points at one run, 8.62 at three, 7.66 at ten, and 7.21 with unlimited runs.
The curves below give the typical error by prompt count at one, three, ten, and unlimited runs, computed from the measured components.
Typical error by prompt count and runs per prompt
Takeaway
Mid-visibility brands need the most prompts
Error in percentage points is largest when a brand appears in some answers and not others. Series with 30–60% visibility needed a median of 382.4 prompts asked once to reach ±5 points; 15–30% needed 257.1; brands above 60% needed 330.2; brands at 5–15% needed 112.4. At 5 prompts × 3 runs, the typical error ran from 9.85 points in the 5–15% band to 17.82 in the 30–60% band.
The prompt share of variance was similar in every band, 42.4% to 46.4%, so the extra prompts are a matter of scale, not a different mechanism.
Median prompts asked once for a ±5-point 95% band, by visibility band
- 5–15%112.4
- 15–30%257.1
- 30–60%382.4
- 60%+330.2
Takeaway
ChatGPT Search and Google AI Mode need the same number of prompts
Median prompts asked once for ±5 points: 143.9 on ChatGPT Search and 145.0 on Google AI Mode. The prompt share of variance was 46.0% on ChatGPT Search and 41.7% on Google AI Mode, so reruns average out slightly more on Google AI Mode: at ten runs the medians were 74.8 and 68.8 prompts. Typical error at 5 prompts × 3 runs was 12.41 and 11.97 points.
The two engines are scored on the same markets and prompts, so the near-identical figures are not a coverage artifact.
Takeaway
Even 100 answers name the same market leader only about half the time
A market leader is the brand with the highest visibility rate in a team's reading; a leader is clear when no other brand ties it. With 10 prompts × 10 runs, both teams had a clear leader in 90.0% of 1,440 market samples, and when both did they named the same brand 53.6% of the time. At 5 prompts × 3 runs (15 answers), both were clear in 57.3% of 11,080 samples and agreed 45.3%; at 3 prompts × 1 run, both were clear in 12.8% and agreed 37.1%.
Leader identity depends on the gap between the top brands. In many markets several brands sit within a few points of each other, and a 10-prompt reading cannot separate them.
Market-leader agreement between two teams, by design
| 1 | 1 | 1 | 1.6% | 24.0% |
| 1 | 3 | 3 | 12.3% | 16.3% |
| 3 | 1 | 3 | 12.8% | 37.0% |
| 5 | 1 | 5 | 25.1% | 43.1% |
| 3 | 3 | 9 | 40.4% | 37.3% |
| 1 | 10 | 10 | 42.9% | 18.4% |
| 10 | 1 | 10 | 46.9% | 49.9% |
| 5 | 3 | 15 | 57.3% | 45.3% |
| 3 | 10 | 30 | 71.9% | 37.1% |
| 10 | 3 | 30 | 73.6% | 52.5% |
| 5 | 10 | 50 | 82.8% | 45.7% |
| 10 | 10 | 100 | 90.0% | 53.5% |
Takeaway
In named markets, two 15-answer readings rarely agreed on the leader
The table lists the 14 active markets with the most prompts in the panel. For each, it shows the leader and its visibility rate on each engine over the full window, and the share of 40 paired samples (20 per engine) at 5 prompts × 3 runs in which both teams had a clear leader and it was the same brand.
Agreement ran from 60.0% in online therapy platforms, where Talkspace led at 68.9% on ChatGPT Search and 69.7% on Google AI Mode, to 7.5% in pet insurance, where Pets Best led at 50.9% and 45.6% with several brands close behind. A dominant brand does not guarantee agreement: Pipedrive led CRM and sales pipeline software at 92.4% on ChatGPT Search, yet the 15-answer readings agreed only 30.0% of the time, because other brands tied it at the top of small samples.
Leader agreement at 5 prompts × 3 runs in 14 named markets
| Online Therapy & Psychiatry Platforms | 32 | 68.9% | 69.7% | 60.0% | ||
| Hard Money Lending Services | 40 | Lima One Capital | 50.9% | Easy Street Capital | 62.3% | 50.0% |
| Business VoIP Phone Systems | 35 | 64.4% | 56.4% | 42.5% | ||
| Online Stock & Options Brokers | 28 | 77.4% | 69.1% | 40.0% | ||
| Payroll and HRIS Software | 29 | 80.8% | 79.2% | 37.5% | ||
| Online Mattress Shopping and Comparison | 47 | 54.4% | 35.2% | 30.0% | ||
| CRM and Sales Pipeline Software | 25 | 92.4% | 82.3% | 30.0% | ||
| AI Customer Support Chatbots | 28 | 51.3% | 36.7% | 22.5% | ||
| Construction Industry Software and Services | 30 | 30.0% | 28.0% | 17.5% | ||
| Online Sports Betting Apps | 38 | 91.2% | 83.7% | 15.0% | ||
| Meal Kit Delivery Services | 25 | 65.7% | 46.6% | 15.0% | ||
| Insurance Quote Comparison Tools | 25 | 42.8% | 61.0% | 12.5% | ||
| Pet Insurance Plans | 46 | 50.9% | 45.6% | 7.5% | ||
| Personal Injury & Accident Lawyers | 36 | Morgan & Morgan | 9.1% | Morgan & Morgan | 8.2% | 7.5% |
Takeaway
What this study counts, and what it leaves out
The panel keeps markets with at least 10 prompts that have 20 or more answers per engine in the window: 277 of 1,698 markets and 4,048 of 17,970 organic prompts. Series need at least 5% visibility. Raising that threshold to 10% leaves 6,113 series with a prompt share of 43.5% and a 5 × 3 typical error of 14.52 points, so the headline does not depend on the threshold.
Two disjoint samples from the same market give an unbiased error at any prompt count, because the finite-population correction cancels in their difference. The variance components predicted the observed typical error within 0.1 points at every design (5 × 3: 12.19 predicted, 12.19 observed; 1 × 10: 24.23 vs 24.28). Runs were drawn at random across 13 weeks, so run-to-run error includes drift inside the window. Brands resolve to their owning company, which is why one market whose leaders resolved to Alphabet and Meta was left out of the named table. The visibility rate is prompt-weighted, matching designs with equal runs per prompt.
Takeaway
The GEO takeaway
Size the prompt set first. Forty prompts asked once (±10.7 points) beat 15 prompts asked five times (±12.9) with about half the answers. Add runs only after the prompt set is as large as the budget allows; three runs capture most of the rerun benefit, and ten runs reach within a point of the floor.
Report every visibility rate with its 95% band from the prompt count, not the answer count, and treat a change smaller than the band as no change. Two teams with 10 prompts × 10 runs agreed on the market leader about half the time, so a leader claim needs a margin over the next brand, not just a rank.
Takeaway
How we measured
In the Parse index, we analyzed 186,216 AI answers to 4,048 organic prompts in 277 markets on ChatGPT Search and Google AI Mode from June 2 through August 31, 2026, covering 958,454 brand-answer pairs and 12,053 brand-market series, and drew 20 random prompt-and-run samples from every market on each engine.
Get the data
Sources
These are the pages this study used.
- cloro — AI visibility sample size: how many prompts and runs (Ricardo Batista, August 14, 2026) · accessed September 7, 2026
- MaxAEO — How many prompts to test AI visibility? Sample size and prompt-run math (Chris Han, July 6, 2026) · accessed September 7, 2026
- Nick Lafferty — AI visibility metrics: formulas, benchmarks and sample sizes (updated August 20, 2026) · accessed September 7, 2026
- Search Engine Land — How to make prompt tracking much more accurate (Kevin Indig, June 10, 2026) · accessed September 7, 2026
- SE Ranking — How to choose prompts to track for AI visibility (Yevheniia Khromova, April 23, 2026) · accessed September 7, 2026