How to Build an AI Visibility Prompt Set That Is Not Just Your Keyword List in a Fake Mustache
A useful prompt set samples how people choose, not how marketers label pages. Start with buyer situations, vary the constraints and wording, keep branded prompts out of the score, and version the panel so a trend means something.
A keyword list is not a buyer panel
The fastest way to build a bad AI visibility study is to export an SEO keyword list, put “What is the best” in front of every row, and call the result buyer research. It looks tidy. It also measures a narrow dialect spoken mostly by marketers.
Search keywords compress intent because people learned to talk to Google in fragments. Conversations with assistants expand it again. A buyer is more likely to describe a mess: the location, the stakes, the budget, what has already gone wrong, and the trade-off they care about. “Divorce lawyer NYC” and “I own a business with my spouse and think assets are being moved before we separate” may belong to the same practice area, but they do not invite the same recommendation.
The unit of design should therefore be a buyer situation, not a keyword. Keywords still help discover demand. They just should not be allowed to write the questionnaire.
Build the matrix before writing the prompts
Start with a small matrix of dimensions that genuinely change fit. For a service firm, that usually means service need, geography, buyer type, urgency, complexity, budget posture, and one or two decision criteria such as discretion, speed, specialization, or willingness to litigate. Not every combination deserves a prompt. The matrix exists to show what the panel covers and, just as importantly, what it does not.
Then choose cells that represent distinct commercial situations. A useful legal panel might separate a routine uncontested matter from a high-asset dispute, a local resident from an out-of-state owner, and a buyer seeking a specialist from one seeking an affordable first consultation. A useful agency panel might separate a funded startup launch from an enterprise migration. If changing a constraint would reasonably change the shortlist, it probably deserves separate coverage.
Do not quietly overweight the product’s favorite use case. Assign weights explicitly: by qualified lead volume when you have it, by revenue opportunity when that is the business question, or evenly when the goal is category research. “We weighted all prompts equally” is a methodology. “We happened to write twelve prompts about our best service” is a thumb on the scale.
Use paraphrases as a stress test, not decoration
Prompt wording matters. A 2026 study of commercial recommendations found that natural paraphrases of the same buying intent produced much less overlap in recommended brands than exact-prompt reruns. The precise percentages belong to that study’s models and categories, not to the universe, but the design lesson travels: one sentence is a brittle proxy for an intent.
Give important situations two or three natural forms. Vary register and information order, not the underlying need. “Who handles complex partnership disputes in Boston?” and “My co-founder has locked me out of the company accounts; which Boston firms work on cases like this?” probe the same neighborhood from different streets. A constraint-adding prompt is not a paraphrase; it is a new situation and should be labeled that way.
Paraphrases also expose suspicious wins. If a firm appears only when the prompt repeats the exact wording of its title tag, that is not broad authority. It is a narrow retrieval match wearing a medal.
Keep discovery prompts separate from diagnostic prompts
An unbranded discovery prompt asks the assistant to choose without planting a candidate: “Which firms should I consider?” That is the right material for share of voice. A branded diagnostic prompt asks what the assistant knows about a named firm, how it compares, or whether it fits a scenario. Those prompts are useful, but they answer a different question.
Mixing the two inflates visibility because the brand has already been supplied. The same goes for prompts copied from your own positioning. If every question contains your preferred phrase, the study is partly measuring whether the assistant can echo your website.
Maintain separate scores for unbranded recommendation presence, branded factual accuracy, source citation, and competitive comparison. One blended number is convenient for a dashboard and disastrous for interpretation.
Freeze the instrument, then version it honestly
A tracking panel needs stable prompts, stable model labels, a recorded run date, and a fixed sampling rule. Otherwise a higher score can come from editing the test rather than improving the market position. Save the exact prompt string, not a human-friendly summary of it. Record whether web search was used and preserve the answer and its cited URLs.
Panels still need maintenance. Buyer language changes, services change, and assistants change how they search. Use versions: keep a frozen core for longitudinal comparison and add an exploratory edge for new situations. When a prompt retires, do not splice its replacement into the old time series as if nothing happened.
The result should resemble a research instrument more than a content calendar. If another analyst cannot rerun it and understand why each prompt exists, the score is theater with decimals.
A compact quality check
Before running the expensive part, ask five questions. Does every prompt map to a documented buyer situation? Are material constraints represented without combinatorial nonsense? Are the important intents tested in more than one natural phrasing? Are branded and unbranded prompts reported separately? Can the exact panel be rerun next month?
Then read the prompts aloud. If they sound like a committee of SEO tools trying to order lunch, rewrite them. Real buyers are specific, occasionally emotional, and rarely kind enough to use your taxonomy.
Key takeaways
- Design around buyer situations and fit-changing constraints, not keyword rows.
- Use multiple natural phrasings for important intents and label constraint changes as new situations.
- Keep unbranded discovery, branded diagnostics, citations, and comparisons as separate measurements.
- Freeze a longitudinal core, version changes, and preserve exact prompts and answers.
Sources and further reading
Primary documentation and research used for this field note. Product behavior changes; check the linked source before treating any implementation detail as permanent.
Next step
Atlas shows the public map. A Viclaro audit turns that map into the prompts your firm is losing and the page edits most likely to change the next scan.
See how Atlas measures a category.