A defensible AI visibility prompt set has four categories (buyer-intent, category, competitive, and source-sensitive) and starts at roughly 50 prompts before expanding toward 200. Buyer-intent prompts dominate the first batch because they drive pipeline; category and competitive prompts come next because they set the baseline and benchmark. Anything under 25 prompts is too noisy to act on. Anything over 300 without structure becomes a dashboard nobody reads.
Why prompt selection is the highest-leverage decision in AI visibility
Every decision downstream of this one, from which platforms matter to which sources to close to what to report to the CEO, is filtered through the prompts you chose. Pick the wrong ones and you will track variance, not signal. SparkToro found a less than 1% chance that ChatGPT returns the same brand list when asked the same question 100 times (SparkToro, 2025), and BrightEdge measured only 17% of queries producing the same brands across ChatGPT, Google AI Overviews, and AI Mode (BrightEdge, 2025). That volatility is not a bug in the tools; it is the baseline behavior of generative retrieval. Parse's data on how long an AI citation lasts shows the cited sources turning over run to run as well. The only way to get stable measurement is to choose prompts that actually matter to your buyer and to measure them across enough runs to cancel the noise. Parse tracks AI visibility across ChatGPT, Google AI Overviews, and Perplexity, covering millions of indexed prompt responses, but the starting input is still a deliberate prompt list, not a grab bag.
What a "prompt" means in AI visibility monitoring
A prompt is a single natural-language query a real buyer could plausibly type into an AI model to evaluate options in your category. It is not a keyword, a search query string, or a clever hack to force your brand into the response. "best helpdesk software for small teams" is a prompt. "helpdesk" is not. The distinction matters because AI models decompose the prompt into sub-queries: ALM Corp's retrieval research showed 89.6% of ChatGPT prompts trigger two or more follow-up fan-out searches, and 95% of those fan-out queries have zero traditional search volume (ALM Corp, 2025). Your prompt is the seed; the model generates the rest. Short, abstract keywords do not behave this way and do not reflect how buyers actually interact with AI chatbots.
The four prompt categories
A complete set has four categories. Most teams start by covering only one, then wonder why their dashboard does not explain why they are losing deals.
- Buyer-intent prompts. How a buyer asks about solving their problem. "best tool for X," "how do I solve Y," "what software helps with Z for [persona]." These are closest to pipeline.
- Category prompts. How the market is described. "AI visibility platforms," "B2B CRM options," "HR software for 500-person companies." These set your category baseline.
- Competitive prompts. How buyers compare options. "[competitor] alternatives," "[you] vs [competitor]," "companies like [competitor] for [use case]." These measure displacement.
- Source-sensitive prompts. Queries where the answer is known to lean heavily on specific sources: G2 for software, Wikipedia for established companies, Reddit for consumer products. These tell you where your citation work is actually landing.
If any category is missing, your prompt set is optimizing for a partial picture.
How many prompts is enough
The right number is whatever your team can actually review each week, weighted by commercial value. Public tool docs and vendor reports converge on three ranges: 25 to 50 is the starter set most teams run for 30 to 60 days, 50 to 200 is where serious tracking lives, and 200 to 500 is enterprise scale with dedicated analysts. Below 25 the noise floor is too high; BrightEdge's platform disagreement data implies you need at least 30 to 40 prompts to see a stable cross-platform picture (BrightEdge, 2025). Above 500 without segmentation, marginal prompts do not change any decision you would have made at 300. Start at 50, expand in batches of 25 tied to a specific question ("do we show up in agency buyer prompts?"), and stop adding when you cannot name the decision a new prompt would inform.
How to find your starter prompts in one afternoon
Three methods, run in sequence, get a first-draft list of 50 in a few hours.
- Rewrite your keyword list as natural questions. Take your top 30 organic keywords and phrase each as a question a buyer would type into ChatGPT. "small business crm" becomes "what is the best CRM for small businesses under 50 employees." The AI Overviews data showing 48% of Google searches now trigger AI Overviews (Stackmatix via Conductor, 2026) means most informational keywords already have natural-language analogs.
- Mine sales and support transcripts. Pull the last 50 discovery calls and CS tickets and extract the actual language prospects used to describe their problem. This is where buyer-intent prompts come from and it is almost always missing from SEO-derived lists.
- Harvest from competitor content. Take each competitor's top five pages by estimated AI citations and note the questions those pages answer. If a competitor owns a buyer-intent question, that question goes in your set.
Deduplicate, group into the four categories, and cut anything that does not connect to a revenue decision.
Mapping prompts to the buyer journey
Prompts decompose along the journey, not a flat list. A clean map has three tiers so you can spot which stages you are winning and which you are invisible in.
| Journey stage | Prompt style | What it measures | Example |
|---|---|---|---|
| Problem-aware | "how do I solve X" | Category entry and education | "how do I reduce churn in SaaS under 1000 employees" |
| Solution-aware | "best X for Y," "X software comparison" | Inclusion in the consideration set | "best customer success platform for B2B SaaS" |
| Vendor-aware | "[you] vs [competitor]," "[competitor] alternatives" | Head-to-head displacement | "Gainsight alternatives for mid-market" |
Buyer-intent research from G2 shows 50% of B2B software buyers now start vendor research in an AI chatbot (G2, 2025), and earned-media research points to third-party coverage as the place brand mentions most often surface in AI answers. Both data points push the same conclusion: vendor-aware prompts are where displacement shows up first, and they should be at least a quarter of your tracked set.
Prompts that map to revenue, not ego
A common failure mode is stocking the set with prompts where your brand wins easily. "Best [your category] made by [your exact positioning]" is a vanity prompt that reassures the team and tells you nothing. Replace those with prompts where the decision is contested. Good buyer-intent prompts name a persona, a constraint, and a use case: "best project management tool for a 200-person professional services firm that bills hourly" is worth tracking. "best project management tool" is not, because it is too broad for AI to return a stable answer and too disconnected from any buyer you can reach. BCG's agentic-commerce guidance is useful here: the visibility metric that matters is how often AI systems surface your brand across category-relevant prompts, not whether you appear in a contrived one-in-a-million query (BCG, 2025).
If you want to know when AI changes its answer about your brand, start with a free brand check — it takes a minute.
Source-sensitive prompts and why they close the loop
Some queries are dominated by a single source type regardless of topic. TryProfound's platform analysis shows Wikipedia as ChatGPT's top source at 7.8% of citations, Reddit as the top source in Google AI Overviews and Perplexity, and G2 commanding the majority of B2B software citations (TryProfound, 2025). Ahrefs found 67% of ChatGPT's top 1,000 citation sources are effectively off-limits to marketers: Wikipedia, government sites, established publishers (Ahrefs, 2025). Parse's own ranking of the source domains AI cites most shows the same handful of platforms dominating. Source-sensitive prompts let you check whether the citation work you are doing is landing on prompts that actually depend on those sources. Example: if you are investing in G2 reviews, track five prompts in the form "best X for Y" where the buyer-intent is comparison-heavy. If G2 coverage does not move those prompts, the investment is not paying where you expected.
Platform variants: track the same prompt across models
The same question produces meaningfully different answers across platforms. BrightEdge measured 62% disagreement on brand recommendations between ChatGPT and Google AI Overviews; AI Overviews averages 6.02 brands per query, ChatGPT averages 2.37 (BrightEdge, 2025). Do not build three separate prompt sets. Build one canonical set and run every prompt across every platform you care about. Default coverage for most mid-market brands: ChatGPT, Google AI Overviews, Perplexity. Add Claude if you sell to developer-heavy buyers, Gemini if your category leans enterprise IT. Running a single prompt across three platforms counts as three data points, not one prompt.
Geography, language, and persona variants
One prompt set often needs to branch. A B2B SaaS company selling in the US and UK should run the core set in both locales; the model's retrieval shifts regionally. Persona variants matter when the same category has materially different buyer questions: "best EHR for solo practice" and "best EHR for a 500-bed hospital" are not the same prompt even though they share a category. Keep variants as a second axis, not a duplicate set. A practical rule: if a prompt would get answered with a completely different brand list when a persona or geography changes, it is a variant worth tracking. If it would return the same list with minor ranking shuffles, it is not.
Prompts to avoid
Certain prompt styles waste monitoring slots and generate misleading reports.
- Branded prompts. "Is [your brand] good?" returns self-referential noise. Your brand will usually appear in its own branded query.
- Hallucination bait. "Who invented [your niche product]?" often returns confident fiction; tracking it tells you more about model quirks than buyer behavior.
- News-dependent prompts. "Latest funding round in [space]" drifts based on model training date, not your visibility.
- Over-specific prompts. "Best CRM for a fintech company in Omaha with Salesforce integration and under $50K ACV" is so narrow it returns near-random results.
- Prompts that duplicate keywords. Tracking "CRM software" and "CRM platforms" and "CRM tools" as separate prompts inflates the set without adding signal.
Expanding the set: quarterly review, not ad hoc
Prompt sets should change with the business, but not on a whim. A healthy cadence is a quarterly review tied to these four triggers: a new product or segment launched, a new competitor entered the category, a buyer question surfaced repeatedly in sales calls, or a platform shift (new AI model, changed citation pattern) changed which sources get cited. Between reviews, the set should be stable so trends mean something. Omnius reported that 94% of marketing leaders plan to increase GEO spend in 2026 (Omnius, 2025), and GoodFirms found 65% of marketers call AI search changes their single biggest challenge (GoodFirms, 2026). Both of those findings mean investment is real and measurement has to be serious. Weekly prompt churn destroys historical comparability.
Sample starter sets by industry
A rough 50-prompt starter split across the four categories. Adjust ratios to your business; the structure is the point.
| Industry | Buyer-intent | Category | Competitive | Source-sensitive |
|---|---|---|---|---|
| B2B SaaS | 20 ("best X for [persona]") | 10 ("X software," "X platforms") | 15 ("[competitor] alternatives," "[you] vs [competitor]") | 5 (G2- and Reddit-heavy) |
| Ecommerce | 25 ("best X for [use case]") | 10 (category-level) | 5 (brand vs brand) | 10 (review-site and Reddit-heavy) |
| Agency / services | 15 ("best X agency for [vertical]") | 10 ("X agencies," "X consultants") | 15 ("[competitor] alternatives") | 10 (Clutch, review-site, PR-source heavy) |
| Regulated (health, finance, legal) | 15 ("best X for [compliance constraint]") | 15 (category + subcategory) | 10 (competitive) | 10 (authority-publisher heavy) |
How to instrument the set for measurement
A prompt set is only as useful as the scoring rules behind it. Four fields per prompt make the set auditable: the prompt text, the category, the platforms it runs on, and the decision it informs. Without the last field, the prompt becomes dashboard clutter. Then define what "a win" means at the prompt level: some teams use appear-at-all, some use top-three, some weight by position. Parse surfaces position and citation data per prompt, but the scoring rule is yours to set and document. Once the set is instrumented, treat the overall visibility score as a weighted sum across the four categories, not a flat average. Buyer-intent prompts should carry the most weight because they are closest to revenue.
How this fits the rest of AI visibility work
The prompt set is the input layer. Everything else (citation gap analysis, content restructuring, source work, reporting cadence) runs on top of it. See how to build an AI visibility scorecard for how the scorecard consumes prompt data, and the weekly AI visibility review for the operating rhythm that turns prompt-level signals into team action. If your prompt set is wrong, the scorecard and review will be wrong. If it is right, the downstream work compounds quarterly.
FAQ
How many AI visibility prompts should I track as a starter?
Start with about 50 prompts, split across buyer-intent, category, competitive, and source-sensitive categories. Below 25, the inherent variance in AI recommendations is higher than the signal you are trying to measure. Above 200 without a segmentation plan, new prompts do not change decisions. Expand in batches of 25 tied to a specific question rather than by volume target.
Are keywords and prompts the same thing?
No. Keywords are short search strings optimized for classical ranking. Prompts are natural-language questions that AI models fan out into sub-queries. ALM Corp's research shows 89.6% of ChatGPT prompts trigger multiple follow-up searches and 95% of those fan-outs have zero traditional search volume. Track prompts, and use your keyword list only as a seed for rewriting into buyer-style questions.
Should I track branded prompts?
Usually not. Branded prompts like "is [your brand] good?" return self-referential answers that tell you nothing about your reach. The exception is when you have a known entity confusion problem and need to track whether AI models are correctly distinguishing your brand from a similarly named competitor. That is a troubleshooting use case, not part of the core tracking set.
How often should I change the prompt set?
Quarterly, tied to named triggers: a new product or segment, a new competitor, repeated buyer questions from sales, or a major platform shift. Between reviews, keep the set stable so trends are comparable over time. Weekly prompt churn destroys the ability to tell signal from noise.
How do I weight prompts in an overall visibility score?
Weight by the business decision each prompt informs. Buyer-intent prompts usually carry the highest weight because they are closest to pipeline. Category prompts set the baseline. Competitive prompts measure displacement. Source-sensitive prompts are diagnostics for citation work. A flat average hides which categories are moving; a weighted sum shows it.
Ready to run a disciplined prompt set?
A good prompt set changes which AI visibility decisions are legible, which is the whole point of measurement. Parse lets you define the set, run it across ChatGPT, Google AI Overviews, and Perplexity, and score every prompt against position, citations, and competitive presence, so the reporting up the chain reflects what buyers are actually asking.