Measurement
Did your change move the answer?
Most numbers an AI visibility dashboard shows you cannot answer that. Here is the one comparison that can, and how to tell in advance whether your site is even the right thing to change.
The short answer
Measure how often you get named, not whether you showed up. Run the prompts you care about 10 times each before you touch anything and record the rate. Ship one change. Wait three to four weeks. Run the same prompts, worded identically, the same number of times, and compare the two rates. That comparison, on a prompt set you froze, is the only figure on this page that answers the question.
Before you spend three to four weeks on it, check one thing. Open the answer you want to change and look at whether it cited any sources. If it did, the assistant searched the web to write it, and a page of yours can end up in that search. If it cited nothing, it answered from what the model already held, and no rewrite of your pages moves it until that model is trained again, which you cannot schedule and will not be told about.
That fork decides whether your website is the right thing to be changing at all, and it costs you one look at an answer you have already got.
Sort the prompts before you measure any of them
Losing a prompt is one outcome with three different causes, and only one of them is about the pages you own. Split them first, because a single number averaged over all three is how a change that worked ends up looking like a change that failed.
- The answer cited sources, and one of them is the kind of page you have or could write. Your pages are the lever here. These are the prompts worth experimenting on.
- The answer cited sources and every one of them belongs to somebody else: a directory, a review site, a forum thread, a roundup on a publication. Rewriting your own pages does not reach into those. The work is getting named on them, and it runs on a slower clock.
- The answer cited nothing. Whatever you publish this quarter, this prompt is not going to notice. Keep tracking it, but do not run your experiment on it.
The second case is the one most people are actually in and the one the category is quietest about. If an assistant knows your competitor from four independent pages that mention them and knows you from none, the fix is not a better title tag. Ranking is not being named covers where being ranked and being named come apart, which is the same split seen from the other side.
Which number settles it
An AI visibility tool will show you most of the rows below on the same screen, with nothing marking out which one answers a question about a change you made. This is that marking.
| The number | What it answers | What it cannot |
|---|---|---|
| Named rate on a frozen prompt, before against after | Whether this prompt moved. | Whether you moved it. That takes a second set of prompts you deliberately left alone. |
| Cited rate: one of your URLs appears as a source | Whether your page is being read at all. | Whether the brand got the credit. You can be cited in the sources and absent from the sentence. |
| Where you land in the list of names | Whether you went from an afterthought to a first pick. | Anything, if you only recorded a yes or a no. The order has to be written down at the time. |
| Share of voice against tracked competitors | Where you sit in the market this week. | Whether anything about you changed. It rises when a competitor falls and you did nothing. |
| A single visibility score | One line for a report. | Which two opposite movements cancelled underneath it to produce a flat line. |
| Sessions from assistant referrals | That some people clicked through. | Almost everything else. Most answers resolve without a click, and the analytics miss much of what is left. |
Two of those deserve saying plainly. Share of voice is your percentage of the mentions among the brands you chose to track, so it goes up when a rival goes down and reports that as your progress. And referral traffic is the wrong instrument entirely for this: The AI traffic GA4 cannot see goes through what the analytics can and cannot see by construction.
The comparison, in five lines
- Take the prompts a change could plausibly reach, and run each one 10 times before you touch anything. Record every run.
- Split them. Half you will act on, half you will not touch for the length of the experiment. Ten and ten is a fine first split.
- Ship one change to the pages behind the first half, and write down the date you shipped it.
- Wait three to four weeks. A page has to be crawled and indexed before it can be cited, and that alone is usually days.
- Run all of them again, same wording, same count, and compare the before and after rates within each half separately.
The held-out half is what turns a number into evidence. Treated prompts up and held-out prompts flat points at you. Both halves up points at something else: a model update, a competitor page disappearing, a change in how the assistant searches. How many runs a claim of that shape actually needs, and the test that settles it, are worked through in Proving a change moved an AI answer.
What has to stay frozen
A before and after is only a comparison if the only thing that changed between them is the thing you changed. Five of these get broken routinely, and every one of them invalidates the result quietly rather than loudly.
- The wording. Copy and paste the prompt from wherever you stored it, every single time. Retyping it from memory and adding one word makes it a different prompt.
- The number of runs. Five before and twenty after is not a comparison between two rates, it is a comparison between a rate and a guess.
- The assistant, and the model version if it is shown to you. Record the version on every run: it is the cheapest confounder to rule out and the one most likely to have fired.
- The account state. An assistant carrying memory of your previous chats is answering a personalised question, and yours is the one account in the world guaranteed to have been reading about your brand. Run both rounds signed out, or in a fresh session, and never mix the two.
- Everything else you ship. A migration, a template change, a robots rule, a CDN that started treating agent fetches as scrapers. A page an assistant can no longer fetch looks exactly like a rewrite that did not land.
Reading a result that has not moved yet
The check the week after you ship will usually look identical to the one before it, and that is not a failure. It is what three to four weeks means.
Movement that shows up before the rate does
- Your URL starts appearing in the sources while your brand is still absent from the text. The page is being read. What is missing is the brand sitting next to the claim, so the citation has something to attach to.
- You go from unmentioned to named last. A smaller result than being named first, and a real one, which is why the order belongs in the record rather than a yes or a no.
- The set of competitors named starts changing even though you are not in it yet. The answer is being assembled from a different set of pages than it was, which is the thing that has to happen before your page can be one of them.
And a flat result after the full window is a finding, not a null. Write it down. It is the one nobody records, and it is why people repeat the same fix for a year.
Doing this by hand, and when to stop
None of this needs a tool for the first round. Twenty prompts, 10 runs each, twice, is a spreadsheet and an afternoon on each end. What it also is, is 400 rows entered three weeks apart with the discipline not to touch the held-out pages in between, and most people who start it stop at the baseline.
Running that loop on a schedule is the whole of what Beseen is. How it works sets out what gets pulled, how often the prompts run, and what the re-check compares against. The FAQ answers what a flat citation rate after four weeks actually means. If you would rather not keep the spreadsheet, start free: the first two checks cost nothing and the baseline lands on the first one. The moved answer takes the three to four weeks this page is about, which is further out than your second check, and no tool shortens it.