There is no universal category prompt count. Across 595 active categories in the Parse public index, the median category contains 9 active organic prompts. Only 17 categories, or 2.9%, contain at least 25, and none contains 50. That does not mean nine prompts are statistically sufficient for a campaign. It means a fixed quota like "track 50 prompts per category" confuses category coverage with repeated measurement.
- The 595 active categories contain 6,475 eligible primary prompt assignments.
- The median category contains 9 prompts and the average contains 10.9.
- 270 categories contain at least 10 prompts, 90 contain at least 15, and 36 contain at least 20.
- Only 17 categories contain 25 or more prompts; none contains 50.
- Campaign design should separate prompt breadth, engine coverage, repeat runs, and review capacity instead of treating one prompt-count target as the whole sample plan.
Most categories sit between 5 and 14 prompts
The center of the distribution is clear. Of 595 active categories, 317 contain 5 to 9 eligible prompts and another 180 contain 10 to 14. Those two bands cover 497 categories, or 83.5% of the index.
| Prompt band | Active categories | Share |
|---|---|---|
| 0 | 4 | 0.7% |
| 1 to 4 | 4 | 0.7% |
| 5 to 9 | 317 | 53.3% |
| 10 to 14 | 180 | 30.3% |
| 15 to 19 | 54 | 9.1% |
| 20 to 24 | 19 | 3.2% |
| 25 to 49 | 17 | 2.9% |
| 50 or more | 0 | 0.0% |
The median of 9 is descriptive, not prescriptive. It tells us how many distinct category questions have survived Parse's active public-index selection. It does not tell a brand how many repeated runs are needed to estimate a stable share of model. Those are different denominators.
This distinction matters because AI visibility vendors sell volume in several units: prompts, engines, checks, responses, and refresh frequency. A plan with 10 prompts across 5 engines every day can produce more observations than a plan with 50 prompts checked once a month. Buying the larger prompt number does not automatically buy the better measurement.
Why a fixed 25 or 50 prompt rule breaks
Some categories naturally contain more distinct buying decisions. Online mattress shopping has 47 active prompts in the index. Pet insurance has 46. Hard money lending has 40. Sports betting apps have 38. Personal injury lawyers have 36. These categories combine product or service choice with price, location, risk, feature, eligibility, and trust questions.
Other categories are narrower. A buyer may still ask many phrasings, but only a small set changes the relevant shortlist or source evidence. Adding synonyms until the set reaches 50 can make the dashboard larger without making it more representative.
The existing guide to building an AI visibility prompt set is about a full brand monitoring program across buyer stages, competitor questions, and source-sensitive diagnostics. This benchmark measures something narrower: distinct primary questions inside one active public category. A company can track 50 prompts overall while using 9 category questions, 12 comparison questions, 10 branded diagnostics, and repeated variants for priority personas. The numbers are compatible once the units are named.
Prompt breadth and sampling depth are separate choices
Campaign planning has four independent levers:
- Breadth. How many distinct buyer decisions are in scope?
- Surfaces. Which AI engines matter to the audience?
- Repetition. How many responses are collected for each prompt and surface?
- Cadence. How often is the same design rerun?
Prompt count controls only the first lever. It cannot compensate for one response per prompt on a probabilistic system. It also cannot fix a set that ignores the buyer's actual constraints.
A smaller category set with repeated measurements can answer whether a brand's position is changing. A broad one-time scan can answer which questions and brands exist. The first is monitoring. The second is discovery. Many procurement mistakes happen because teams buy one and expect the other.
Ahrefs explicitly separates broad search-backed discovery from custom tracked prompts. Semrush also separates visibility discovery from prompt tracking. Parse uses the public index to expose the category and custom monitoring to follow a brand's chosen questions. The AI visibility tools comparison explains the product models; this study shows why the underlying units should not be collapsed into one volume claim.
If you want to know when AI changes its answer about your brand, start with a free brand check — it takes a minute.
A practical category coverage method
Start with decisions, not a quota. List the buyer questions that could change which brand is recommended. Group them into a compact matrix:
| Decision layer | Planning question |
|---|---|
| Category entry | Which options belong in the market? |
| Use case | Which option solves the buyer's job? |
| Fit | Which option suits the persona, size, location, or workflow? |
| Trust | Which option satisfies risk, review, or compliance requirements? |
| Commercial | Which option fits price, contract, or implementation constraints? |
| Displacement | Which option wins a comparison or alternative question? |
Not every category needs one prompt in every cell. Price is absent from most established category sets, as the price-prompt benchmark shows. A free consumer product may need no contract question. A regulated platform may need several trust questions. The count follows the decision map.
Once the map is complete, add engines and repeat runs as separate dimensions. Do not duplicate a prompt merely to make a pricing tier feel fully used. Every added prompt should name a decision that the existing set cannot answer.
How to choose a starting range
The distribution suggests three planning modes:
- Focused category audit: 5 to 9 distinct questions. This matches the largest category band and is useful for a first discovery pass.
- Operating category set: 10 to 19 questions. This covers multiple decision layers without requiring a large review team.
- Complex category map: 20 to 49 questions. This fits categories with several personas, commercial models, risk gates, or product subtypes.
These are coverage modes, not confidence thresholds. Each prompt still needs enough engine and repeat observations for the decision being made. If the goal is a controlled experiment, use the statistical design in how to run an AI visibility experiment. If the goal is a weekly operating review, use a stable set the team can inspect rather than the largest set the plan allows.
The public index provides a useful starting map. /rankings shows the categories and buyer questions already tracked. /brands shows whether the current questions produce a meaningful competitive field. /sources shows whether the set reaches different evidence types or keeps retrieving the same source layer.
What to ask before buying a prompt allowance
A vendor's prompt cap matters only after the unit is clear. Ask these questions:
- Does one prompt across three engines count once or three times?
- Does a daily refresh consume more allowance than a weekly refresh?
- Are repeat runs stored separately, averaged, or discarded?
- Can prompts be grouped by category, persona, and buyer decision?
- Can inactive or duplicate prompts be retired without losing history?
- Does the discovery index count against custom monitoring volume?
Then model the plan as observations, not prompts alone. A campaign with 15 category prompts, 3 engines, and 4 repeated checks per reporting period contains 180 prompt-engine observations. That is a more useful procurement number than "15 prompts."
How we measured this
We queried Niche, NichePromptMembership, and Prompt in production through Cosmo. The session used SET default_transaction_read_only = on, a repeatable-read transaction, and a 60-second statement timeout. The snapshot was taken on August 29, 2026.
We kept niches where isActive is true and publishState is active. We counted only primary niche memberships attached to organic, active, visible, unpaused, and unarchived prompts. That produced 595 categories and 6,475 assignments.
The primary query calculated the median and thresholds. An independent query returned the complete prompt-count histogram. Summing the histogram produced 595 categories; summing bins at 25 or more produced 17. The largest category contained 47 prompts, confirming that none reached 50.
Niche assignment is a curated and model-assisted taxonomy. A prompt has one primary niche for this analysis, but a buyer question can relate to more than one market. The result describes current Parse category coverage, not a universal minimum sample size.
How many AI visibility prompts should one category have?
There is no universal count. In Parse's 595 active categories, the median is 9 prompts and 83.5% contain 5 to 14. Use the number of distinct buyer decisions in the category, then plan engines, repeats, and cadence separately.
Does this mean nine prompts are enough for reliable measurement?
No. Nine is the median category-coverage count, not a statistical confidence threshold. Reliable measurement also depends on how many engines, repeat responses, and reporting periods you collect for each prompt.
Why do some AI visibility tools recommend 25 to 50 prompts?
That range often describes a whole brand program containing category, comparison, branded, persona, and diagnostic prompts. This study counts primary organic questions inside one public-index category. The units serve different jobs.
Should I use every prompt included in my tool plan?
No. Add a prompt only when it represents a buyer decision the current set cannot answer. Filling a quota with synonyms increases review work and can overweight one intent without improving coverage.