Back to blog
Research design 9 min read

Why AI Recommendations Change Between Runs (and When a Change Is Real)

An assistant is not a laminated directory. Search results move, generated wording varies, and a tiny prompt edit can change the shortlist. The answer is not to give up on measurement. It is to stop treating one answer like a closing stock price.

The screenshot is a sample, not the truth

Ask an assistant for a recommendation twice and the lists may differ. That does not automatically mean the market changed, the model broke, or your latest heading tag worked. It means you observed two outputs from a system with several moving parts.

A web-enabled answer can vary at retrieval: the assistant may rewrite the question into different searches, receive a changed result set, or select different pages. It can vary at generation: given similar evidence, the model may name, order, or describe firms differently. It can also vary because the product changed models or ranking logic between your runs.

OpenAI documents that ChatGPT Search may rewrite one prompt into one or more targeted queries. Google describes its AI features as using a “query fan-out” technique across subtopics and data sources. You do not need access to either internal trace to see the implication: one buyer question can create more than one retrieval path.

There are four kinds of movement worth separating

First is rerun variance: identical prompt, model, settings, and short time window; different answer. Second is phrasing variance: the same intended need expressed differently. Third is context variance: location, prior conversation, account state, search mode, or other product context changes. Fourth is temporal change: the web index, cited pages, model, or business evidence genuinely changes.

Those are not interchangeable error bars. Reruns estimate output stability under a narrow protocol. Paraphrases test whether the result survives normal buyer language. Context variants test the markets or product states you actually care about. Repeated snapshots test movement over time.

A 2026 commercial-recommendation study found paraphrase sensitivity larger than same-prompt rerun variance in its tested settings. A separate study comparing Google Search, AI Overviews, and Gemini reported low overlap among retrieved sources and lower consistency for AI Overviews across reruns and minor query edits. Treat those findings as warnings about design, not universal conversion factors.

Hold still what you can

Save the exact prompt, full answer, cited URLs, model identifier, run timestamp, search mode, and declared location. Use clean sessions when measuring first-turn discovery. Do not compare a fresh chat with a ten-message conversation and call the difference ranking movement.

Run the same frozen panel on a schedule and preserve per-model results. When a provider silently changes a model alias, annotate the series if you can detect it. When you deliberately change the prompt set, create a new panel version and overlap old and new long enough to estimate the break.

Perfect control is impossible in consumer AI products. The goal is not laboratory cosplay. It is enough protocol discipline that you can list plausible causes of a change instead of crediting the last thing the marketing team shipped.

Use repeated observations, then look for breadth

One appearance is evidence of possibility. Repeated appearances estimate propensity. Multiple models and prompt variants show breadth. That is why a useful result includes a numerator and denominator: named in 7 of 24 eligible runs is interpretable; “ranked number three in AI” is mostly costume jewelry.

When volume allows, report an interval around the rate. When it does not, show the raw count and resist decimals. A score of 33.3% from one mention in three runs did not become scientific because the dashboard found a tenths place.

Look for coherent movement: more prompts, more models, repeated snapshots, and source changes that make sense. A single jump isolated to one model and one phrasing belongs in the investigation queue, not the victory email.

Retest on two clocks

Use a regular clock for trend measurement and an event clock for interventions. The regular snapshot might run monthly or weekly depending on category volatility and cost. Its job is comparability. The event retest follows a meaningful change: a new evidence page, corrected business identity, major third-party coverage, crawler access fix, or model release.

Do not retest five minutes after publishing and declare failure. Discovery, indexing, retrieval, and model behavior operate on schedules you do not control. First verify that the page is public, crawlable, canonical, linked internally, and available to the relevant search crawler. Then allow a predeclared observation window and rerun the same protocol.

Also keep a no-change comparison group when possible. If every firm in the category rises on one assistant, your edit is not the only candidate explanation. Competitors are a useful control group even when they have not agreed to participate in your experiment.

What counts as a real win

A credible win has direction, durability, and mechanism. Direction means the target metric improves. Durability means it survives reruns or a later snapshot. Mechanism means the changed answers or sources connect plausibly to the evidence you added.

For example: a firm publishes a specific, crawlable page about a buyer situation; two assistants begin citing it; recommendation presence rises across several prompts in that situation; the gain remains in the next scheduled snapshot. That is not perfect causal proof, but it is a much stronger story than “we changed schema Tuesday and ChatGPT named us Wednesday.”

Good measurement makes marketing slightly less cinematic and much more useful. The aim is not to eliminate uncertainty. It is to know how much uncertainty you bought with the sample.

Key takeaways

  • Separate identical-prompt rerun variance from phrasing, context, and true change over time.
  • Preserve exact prompts, answers, citations, model labels, timestamps, search mode, and location.
  • Treat one appearance as possibility; require breadth and repeated observations before calling movement.
  • Retest on a fixed schedule and after meaningful interventions, with comparison firms when possible.

Sources and further reading

Primary documentation and research used for this field note. Product behavior changes; check the linked source before treating any implementation detail as permanent.

  1. 1. ChatGPT Search: query rewriting, location, and sources — OpenAI
  2. 2. AI features and your website — Google Search Central
  3. 3. How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews — arXiv
  4. 4. Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation — arXiv

Next step

Atlas shows the public map. A Viclaro audit turns that map into the prompts your firm is losing and the page edits most likely to change the next scan.

See how Viclaro measures repeated recommendations.