Skip to content

Blog

Proving a change moved an AI answer

An AI answer is a rate, not a reading. How to run a before and after that survives the noise: how many runs, what to record, and what to hold back.

By Alex Cloudstar9 min read

The short answer

Measure it as a rate, not as a reading. Pick the prompts you care about, run each one ten times before you touch anything, and record how often you get named. Ship one change. Wait three to four weeks. Run the same prompts, worded identically, the same number of times, and compare the two rates. The result counts only if a second set of prompts you deliberately did not act on stayed flat over the same window.

That is the method. Everything below is why each of those constraints is there and what breaks when you drop one.

It is worth saying what the usual advice gives you instead, because it is not this. The published guides on measuring AI visibility, Ahrefs’ included, teach you to read a dashboard: mentions, citations, impressions, share of voice, benchmarked against competitors and re-checked monthly. Checked on 4 September 2026, their AI visibility audit walks through eight steps of that and never reaches attribution: no before and after design, no holdout, nothing on how many observations a claim needs, and no guidance on how long to wait. Those numbers tell you where you stand. They do not tell you whether the thing you shipped is the reason you moved, and those are separate questions.

A ranking is a reading, a citation is a rate

Position 7 on Google is a fact. You look it up once and you have it, and if you look again an hour later it will almost certainly still be 7.

An assistant does not work that way. The answer is assembled fresh each time you ask. The model samples its own output, and where it searches the web first, the pages it gets back are not guaranteed to be the same set. Brands sitting near the edge of the answer drop in and out for no reason you did anything about. Run one prompt ten times this afternoon and count: on any prompt where you are close to the boundary, you will not get ten identical lists.

So “does ChatGPT name us for this prompt” is not a yes or a no. It is a probability, and one check is one sample from it. If you ran the prompt once in March, rewrote the page, and ran it once in April, you have two coin flips and a story.

The common way to get this wrong is not measuring the wrong thing. It is measuring the right thing once.

How many runs it takes

Which raises the obvious question, and it has an answer you can check rather than a rule of thumb. Score each run as named or not named and compare the before count against the after count with Fisher’s exact test, which is the standard tool for two small counts like these. Two-sided, at the usual p below 0.05, going from never named to named in:

  • 5 runs each side: you need 4 of the 5 after. Three of five is not enough.
  • 10 runs each side: 5 of the 10.
  • 20 runs each side: 5 of the 20.

Read those three lines again, because they say something you can act on. Doubling the runs from 10 to 20 halves the improvement you need in order to see it at all, from half the answers to a quarter. Adding prompts does nothing for this: 40 prompts checked once each is 40 unreliable readings, not one reliable one. Runs buy you sensitivity. Prompts buy you coverage. They are not interchangeable, and the tools that meter you on prompt count quietly push you toward the one that cannot settle an argument.

This is also cheaper than it sounds. Twenty prompts, ten runs each, before and after, on one assistant, is 400 answers. Bought through a data API, an answer with a live web search behind it runs around two cents once the search is paid for, so the whole experiment is under ten dollars in fees. What you are really spending is the three weeks.

Record the answer, not a score

One row per run. The date, the prompt verbatim, which assistant, the model version if it is shown to you, whether your brand appears in the text, which competitors appear, the order they appear in, and every URL cited.

The order matters more than it looks. A prompt that named three competitors and now names four with you last is a smaller result than being named first, and both beat not appearing. Record only a yes or a no and you cannot tell those apart, which throws away most of what a second rewrite would aim at.

Named and cited are two different outcomes

Your brand appearing in the answer text and one of your pages being cited as a source are separate events, and treating them as one is where a lot of confusion starts. You can be named without being cited: the assistant knows the brand from somewhere else, usually a roundup on a site you do not own, and never touches your pages. You can be cited without being named: your page supplies the explanation and the answer credits the topic, or credits the competitor whose brand happens to sit next to the same claim.

Give them two columns, because they move for different reasons. Rewriting your own page can change whether you get cited. Whether you get named often depends on what other people’s pages say about you, which is a slower and largely separate project.

The control set is the part everyone skips

Split your prompts before you start. Some you will act on. Some you will deliberately leave alone, and you will not touch a page relevant to those for the length of the experiment. Ten and ten is a fine first split.

Run both sets on the same days, with the same number of runs. Then the comparison has a second axis and starts meaning something:

  • Treated group up, held-out group flat: that is your evidence. It is the only pattern that points at you.
  • Both groups up: something moved that is not you. A model update, a competitor page dropping out, a change in how the assistant searches.
  • Treated group up, held-out group down: be more suspicious, not more pleased. Something is shifting under the whole set and you happened to be on the right side of it this week.
  • Both flat: your change did not move this prompt. That is a result. Write it down, because it is the one nobody records and it is why people repeat the same fix for a year.

This is also the specific failure mode of a share-of-voice number. Ahrefs defines share of voice as your percentage of impressions against the other brands you chose to track, so it rises when a competitor falls and you did nothing at all. As a picture of the market that is useful. As the number you judge a rewrite by, it is measuring your competitors.

How long to wait, and which prompts can move at all

Weeks, not days. Rewrite on Tuesday, re-check on Wednesday, and you have measured your own impatience. Three to four weeks is a sane first window, and the reason is worth understanding, because the same reason tells you which prompts belong in the experiment.

Look at whether the original answer showed sources. An answer that cited pages was assembled from a live search, so it can change as soon as the index behind that search holds your new page. That is a crawl-and-refresh problem, and it genuinely runs on the order of weeks. An answer that cited nothing leaned on the model’s weights instead, and no amount of rewriting your page moves that until the model is trained again, which you cannot schedule and will not be told about.

So sort your prompts by that before you start. The sourced ones are the experiment. The unsourced ones are the slower project about getting mentioned on pages you do not control, and mixing them into the same measurement is how a fix that worked ends up looking like a fix that failed.

One change at a time, and a list of things that are not you

Rewrite the title, add a comparison table and add an FAQ block in the same week, watch the prompt start naming you, and you have learned that something worked. Next time you will be guessing again.

Before you start, write down what else could move the number inside your window, so you check those at the end instead of discovering them in an argument three months later:

  • The assistant shipped a new model, or changed how it searches. Record the model version on every run: it is the cheapest confounder to rule out and the one most likely to have fired.
  • A competitor published, or lost, the page that was being cited. Open the cited URLs from your baseline at the end and check they still exist and still say the same thing.
  • You shipped something else. A migration, a template change, a robots rule, a CDN default that started treating agent fetches as scrapers. A page an assistant can no longer fetch is indistinguishable from a rewrite that did not land.
  • The wording drifted. Re-typing a prompt from memory and adding one word makes it a different prompt. Copy and paste from the spreadsheet, every time.

Numbers that cannot answer this question

  • A single visibility score. It folds named, cited, position and prompt coverage into one figure, so two changes that push it opposite ways cancel out and you see a flat line where two real things happened.
  • Traffic. Assistants answer most questions without sending a click, and the analytics that exist miss most of what is left.
  • A one-off audit. A scan with no second scan cannot answer a question about a change, whatever it cost.
  • Anything measured on a prompt you did not freeze. If the wording moved, the two runs are not comparable and no amount of arithmetic fixes that afterwards.

On the traffic point specifically, The AI traffic GA4 cannot see goes through what GA4’s AI Assistant channel actually counts and what it cannot see by construction. It is a shorter list than most people expect.

Doing this yourself, starting today

None of this needs a tool for the first round. It needs a spreadsheet and the patience to run one prompt ten times.

  • Pick 10 prompts you currently lose and 10 you will not touch. If you do not have a prompt list yet, Ranking is not being named covers writing one that describes the market rather than flattering you.
  • Run all 20, ten times each, on one assistant. One assistant measured properly beats four measured once.
  • Record the full row for every run: date, prompt, model version, named, order, competitors, cited URLs.
  • Change one thing on the pages behind the first ten. Note the date you shipped it.
  • Wait three weeks. Run all 20 again, same wording, same number of runs, and compare the two rates within each group.

The tedious part is real and there is no way around it by hand: 400 rows, entered twice, three weeks apart, with the discipline to not touch the held-out pages in between. Most people who start this stop at the baseline, which is exactly why so much of the advice in this category has never been checked against anything.

A claim that a rewrite worked is worth what its second measurement is worth. If there was no second measurement, the claim is a guess with a date on it.

Running that loop on a schedule is the whole of what we built. How it works sets out what gets pulled, how often the prompts run, and what a rewrite changes and in what order. If you would rather not keep the spreadsheet, start the free trial. Starter is free for 14 days, and the baseline lands on the first check. The moved answer takes the three to four weeks this piece is about, which is after the trial, not inside it.