Real AI buyer prompts are full questions, not dressed-up keywords. Parse analyzed 14,977 active organic prompts in its public index and found a median length of 15 words. Fully 81.1% contain at least 10 words, while only 15 prompts are four words or shorter. A campaign built from head terms like "payroll software" or "pet insurance" is measuring a different language from the one buyers use when they ask AI for a recommendation.
- The 14,977 prompts contain 223,079 words, or 14.9 words per prompt on average.
- The median prompt is 15 words; 12,141 prompts, or 81.1%, contain at least 10 words.
- Only 15 prompts contain four words or fewer.
- The three most common opening words are "what," "I," and "which." Together they open 9,290 prompts, or 62.0% of the panel.
- A useful AI visibility campaign should preserve the buyer's use case, constraint, and decision, not reduce the prompt to a keyword label.
The median buyer prompt is 15 words
The distribution is unusually tight for a corpus this broad. The average is 14.9 words and the median is 15, which means the headline is not being pulled upward by a small set of long questions. The middle half of the panel runs from 11 to 18 words. Four out of five prompts contain at least 10 words.
| Prompt-length measure | Result |
|---|---|
| Active organic prompts | 14,977 |
| Total words | 223,079 |
| Average words per prompt | 14.9 |
| Median words per prompt | 15 |
| Prompts with 10 or more words | 12,141, or 81.1% |
| Prompts with 4 or fewer words | 15, or 0.1% |
That shape changes what counts as a monitoring unit. A short label names a market. A buyer prompt describes a decision inside that market. "CRM" is a market label. "Which CRMs route property inquiries across a team and report response time, conversion, and source quality?" is a decision. The second version gives the model a job, a workflow, and three evaluation criteria. It can return a competitive set that reflects a real sale.
The practical test is simple: if removing the last half of a prompt would not change which brand should win, the prompt may be padded. If reducing it to a two-word keyword would erase the buyer's constraint or use case, the detail is doing real measurement work.
Buyers ask from inside the problem
The opening words show how different this language is from a keyword list. "What" opens 4,097 prompts, "I" opens 3,324, and "which" opens 1,869. Together those three account for 62.0% of the entire panel. "We," "how," "as," "for," and "where" follow.
| Opening word | Prompts | Share of panel |
|---|---|---|
| What | 4,097 | 27.4% |
| I | 3,324 | 22.2% |
| Which | 1,869 | 12.5% |
| We | 1,328 | 8.9% |
| How | 878 | 5.9% |
| As | 646 | 4.3% |
These are not cosmetic sentence starters. "I" and "we" place the buyer inside a situation. "Which" asks for a choice. "What" often asks the engine to define the relevant set before choosing. A prompt library made only from category nouns loses those decision signals, even if every noun is commercially important.
This is why the existing guide to building an AI visibility prompt set starts with buyer language rather than search volume. The new evidence adds a benchmark to that rule: a representative public-index prompt is not a keyword with a question mark. It is usually a 10 to 20 word description of the choice the buyer is trying to make.
Why keywords still help, but cannot be the final input
Keyword research remains useful for finding categories, products, and demand themes. Ahrefs describes its AI visibility corpus as search-backed, combining keyword data, People Also Ask questions, and semantic expansion. Semrush likewise recommends tracking prompts tied to a brand, category, and competitors. Those are good discovery inputs. The mistake is stopping before they are rewritten in buyer language.
A keyword can seed a prompt in three steps:
- Name the decision. Is the buyer discovering a category, comparing vendors, validating trust, or checking fit?
- Add the constraint that changes the shortlist. That can be team size, price, integration, compliance, location, or a required feature.
- State what a useful answer must do. Ask for a recommendation, comparison, tradeoff, or implementation choice.
For example, "business phone system" becomes "Which business phone systems support a distributed sales team, integrate with our CRM, and report call quality by office?" The category stays the same. The measurable competitive question becomes much sharper.
The point is not to make every prompt long. It is to keep every word that changes the answer. Our separate benchmark on AI visibility prompts per category shows why depth should follow category structure rather than an arbitrary global quota.
If you want to know when AI changes its answer about your brand, start with a free brand check — it takes a minute.
A campaign needs prompt families, not isolated phrases
One long prompt is still one observation. The campaign unit should be a small family of buyer decisions that share a commercial job. A useful family might include:
- a category-entry question that asks which options exist;
- a feature or workflow question that tests practical fit;
- a price or contract question when cost is genuinely part of the category;
- a trust or compliance question;
- a comparison or alternative question that exposes displacement risk.
This structure prevents two common failures. First, it stops a team from filling the set with slight wording variants that return the same brands. Second, it stops a tool default from treating every category as if buyers evaluate it the same way.
Use /rankings to inspect the questions already associated with a category. Use /brands to see which brands appear across those questions. Then use /sources to trace the evidence behind the answers. That sequence moves from buyer language to competitive outcome to source strategy without pretending a head term can stand in for all three.
What this changes in tool evaluation
The competitor audit found prompt selection everywhere. Profound, Peec, Scrunch, Semrush, and Ahrefs all explain prompts, monitoring, visibility, and sources. The open procurement question is not whether a tool supports prompts. It is whether the prompt model represents the questions your buyers ask.
Ask a vendor to export a sample before you buy. Then check four things:
- Are prompts full buyer questions or short search-derived labels?
- Do they include use cases and constraints that can change the winning brand?
- Can you separate discovery prompts from comparison, price, trust, and workflow prompts?
- Can you preserve a stable custom set while still discovering new public-index questions?
A large prompt database is valuable for discovery. A custom prompt set is valuable for campaign measurement. The right product can explain which job each layer performs. The existing AI visibility tools comparison covers that procurement boundary; this study supplies a concrete language benchmark for evaluating the sample.
How to audit your own prompt set
Export the prompt list and add four columns: word count, buyer voice, decision type, and constraint. Do not optimize toward 15 words mechanically. Use the distribution as a warning signal.
If most prompts contain fewer than five words, the set probably describes markets rather than decisions. If every prompt contains 25 words, the set may be overfitted to edge cases. If none begins with "I" or "we," the set may be written from the brand's taxonomy instead of the buyer's situation. If every prompt asks "best," it may measure list inclusion while missing workflow, trust, and fit.
Review the exceptions manually. A short prompt can be valuable when the category itself is the decision. A long prompt can be valuable when a regulated or technical purchase really has several hard constraints. The benchmark should start a review, not replace judgment.
Once the set passes the language check, freeze it long enough to measure change. The weekly AI visibility review explains how to turn stable prompt evidence into an operating backlog. Changing the questions every week makes a clean trend impossible.
How we measured this
We queried the production Prompt table through the Cosmo worker path with SET default_transaction_read_only = on, a repeatable-read transaction, and a 60-second statement timeout. The snapshot was taken on August 29, 2026. We kept prompts that were organic, active, visible to users, not paused, and not archived. That produced 14,977 prompts.
Word counts split trimmed prompt text on whitespace. The primary query computed the count, average, median, and length thresholds. An independent query reconstructed the headline from total words and direct conditional counts. Both returned 14,977 prompts, 223,079 words, 12,141 prompts with at least 10 words, and 15 prompts with four words or fewer.
The panel reflects Parse's active public index, not every prompt people ask every AI system. It is curated category coverage, not a random survey of global usage. The result should guide campaign design, not be read as market share for a phrase or opener.
How long is a real AI buyer prompt?
In Parse's active public index, the median buyer prompt is 15 words and the average is 14.9. The middle half runs from 11 to 18 words. Length is not a target by itself, but it shows that real recommendation questions usually carry a use case, constraint, or decision that a short keyword omits.
Are AI prompts the same as SEO keywords?
No. Keywords are useful discovery inputs, but a buyer prompt states the decision an AI answer must support. "Payroll software" names a category. "Which payroll tools handle multi-state tax for a 200-person team?" asks for a recommendation under a constraint.
Should every tracked prompt contain at least 10 words?
No. The 81.1% figure is a corpus benchmark, not a hard writing rule. Keep the words that change the likely answer, remove padding, and preserve short prompts when the broad category question is itself commercially important.
How do I know whether my prompt set is too keyword-like?
Audit word count, buyer voice, decision type, and constraints. A set dominated by two-to-four-word category labels, with no first-person questions and no fit criteria, is likely measuring market vocabulary rather than real buyer decisions.