Skip to content

Blog

Your AI visibility score is hiding lost prompts

An AI visibility score can rise while buyers stop seeing you where it matters. Track prompts, assistants, names and citations before you trust the average.

By Alex Cloudstar10 min read

The score is an average of different jobs

An AI visibility score is useful only after you can see the prompt results underneath it. A score can rise because your company was cited in a broad educational answer while it disappears from the comparison question that a ready buyer asks. The average looks better. The business outcome does not.

This is not a complaint about arithmetic. An average does exactly what it says: it combines results. The problem starts when a team treats that combined result as proof that it is being recommended for the questions that matter. A dashboard should help you find the losing questions, not give you permission to ignore them.

If you track AI visibility, keep the full answer beside the number. Record the exact prompt, the assistant, every product named, its position, and any cited source. That is the evidence a product marketer can review. A score is only the index to it.

Ask for the prompts behind the score. If a tool cannot show the response that made a number move, it cannot show what needs to change.

One company can win the score and lose the buyer

Imagine a payroll product tracks ten questions. On six broad prompts such as “what is payroll software?” its help centre is linked as a source. On four purchase prompts such as “what payroll software works for a UK startup with contractors in two countries?” the assistant names three competitors and leaves it out.

A blended metric can still report a decent result. It may even improve after the company publishes another strong explainer. But the second group is where a buyer is choosing. Treating both groups as equivalent turns a useful warning into a cheerful chart.

The same trap appears across assistants. A product can be named by ChatGPT and absent from Claude. It can be a source link in Perplexity and never appear in the recommendation sentence. It can rank well in Google while no assistant selects it as an answer. Ranking is not being named explains why a search position and a product choice are separate observations.

Do not make a citation carry a recommendation's weight

A citation answers a page question: did the response visibly use this URL or domain as support? A recommendation answers a choice question: did the response put this product forward for the buyer's stated need? Both are useful. Neither proves the other.

Keep them in separate columns even when the same response contains both. A citation may reveal a page worth studying. A recommendation tells you that the product was selected. When they move in different directions, that difference is the finding.

Start with the questions that make a decision

A keyword list is a poor substitute for a prompt set. Buyers do not always ask an assistant for “project management software.” They ask for a tool that has a specific permission model, moves data from a particular system, costs less than an incumbent, or works for a team that has no administrator.

Build the first set from evidence you already own: sales-call notes, demos that stalled, support requests, comparison pages, search terms that lead to high-intent pages, and the reasons a customer says it chose you. Preserve the wording. A prompt is not a keyword with extra words. Its constraints decide which products an assistant considers relevant.

  • Category prompts: who belongs on a shortlist?
  • Alternative prompts: what replaces a named incumbent?
  • Constraint prompts: which product works with a specific budget, team size, region, or requirement?
  • Implementation prompts: what connects, migrates, or supports the workflow a buyer already has?
  • Trust prompts: which product is suitable when the buyer cares about reliability, privacy, or support?

Label the group as well as the prompt. If recommendation rate drops, you need to know whether the loss is in category discovery or a single integration question. One score cannot make that distinction for you.

The minimum evidence in an AI visibility review

A useful weekly review does not need a complicated model. It needs a record that lets another person see what happened without repeating your interpretation. Save one row per prompt and assistant run.

  • The prompt, unchanged and dated.
  • The assistant and the model or mode used, when the product exposes it.
  • Your company's result: recommended, mentioned, cited, or absent.
  • Every competitor named and the order in which it appeared.
  • The surrounding sentence, so a casual example is not counted as an endorsement.
  • Every visible citation, including the cited page rather than only its domain.
  • A saved response or screenshot that a teammate can inspect later.

This makes the dashboard answer questions a headline number cannot. Which competitor appears whenever a buyer asks about migrations? Which of your own pages is used for educational support but never accompanies a recommendation? Which assistant has stopped naming you? The response is where those answers live.

A metric needs a denominator before it needs a chart

“We appeared 40 times this month” says very little on its own. Forty appearances in forty comparable runs is different from forty in four hundred. A recommendation rate makes the sample explicit: recommendations divided by the runs for that prompt group and assistant.

Keep the denominator narrow enough to be honest. Do not combine a new exploratory prompt with a set you have run for months, then call the result a trend. Do not mix ChatGPT and Claude into one rate unless you still show each assistant's result. And do not let an educational question outweigh a decision question because it happens more often in the sample.

The repeat runs matter too. Assistant responses vary. One result can be a change, an outlier, or a different retrieval path. ReadProving a change moved an AI answer before declaring that a content change caused a movement. The short version is simple: retain the old prompts, repeat them, and compare like with like.

Write down the counting rule before results arrive

A visibility score is only comparable when its counting rule stays put. Teams often discover halfway through a review that one person counted a product only when the assistant said it was a good fit, while another counted every appearance of the name. Both can be useful counts. They are not the same metric.

Define the events before you look at a result. For example, call an answer a recommendation only when the assistant presents the product as a suitable choice for the stated request. Call it a mention when the product name appears without that selection. Call it a citation when the response visibly links to or attributes information to a page. If the assistant says a product is a poor fit, save that too. An appearance is not automatically a positive result.

The rule needs to handle lists as well. A brand named first with a reason is a different result from a brand listed fifth after “other options include.” You do not need to invent a false precision score for the difference. Record the order and the sentence. That gives the team enough context to decide later whether position deserves its own metric for this prompt group.

Publish the rule beside the dashboard. It stops a quarterly chart changing its meaning every time someone new joins the review. It also keeps product, content, and growth teams from arguing about the same saved response with three private definitions of “visibility.”

Keep the baseline stable and explore separately

A prompt set has two jobs that pull in opposite directions. The baseline set needs to stay stable so that week-to-week movement has a meaning. The discovery set needs to grow as sales calls, competitor pages, and new buyer language reveal better questions. Put those jobs in separate lists.

The baseline can be small. Ten carefully chosen, purchase-minded prompts are more useful than a hundred loose category queries. Retain their wording, assistant settings, and grouping. If a prompt stops matching the market, do not quietly rewrite history. Retire it with a date, state why, and add its replacement as a new series. The old results remain evidence of what buyers were asking at the time.

Use the discovery set for the messier work. Add the exact phrase from a sales call. Add a competitor's claimed specialty. Add a constraint that appeared in support. These prompts can uncover a blind spot quickly, but they should not be merged into last month's denominator simply because the answer is interesting.

This split also makes changes easier to interpret. A stable baseline falls after a release, while a newly added prompt is already missing you. Those are two facts with different next steps. A single expanding score would collapse them into one downward line and remove the reason it moved.

What to do when the score and the prompt disagree

Do not start by changing every page on the site. Open the losing response. First, check that the assistant understood the product category and the buyer's constraint. Then look at the companies it chose. Their names, positions, and supporting sources tell you more than a generalized optimization recommendation.

If the response misunderstands what you do, your public language may be too vague or split across pages. If it understands the category but prefers another product for a named condition, inspect whether your site plainly addresses that condition. If it cites a third-party page that omits you, that is a different distribution problem from an on-page explanation.

Make one scoped change tied to the evidence. Add a missing implementation detail. Clarify an unsupported comparison claim. Publish a direct answer to the question if you can do so accurately. Then rerun the same prompt set. A score may tell you that something changed. The saved response tells you whether the change helped the question you meant to solve.

Give each team the question it can answer

An AI visibility score often becomes a content-team KPI because the output has links in it. That is too narrow. A product team may own the capability that makes an assistant rule you out. Sales may recognize a buyer constraint that no page explains. Support may know the objection that turns a generic recommendation into a wrong one. The saved answer gives each group something concrete to assess.

Content can check whether the product is described plainly and whether the answer has a useful page to cite. Product marketing can compare the assistant's wording with the buyer narrative it intends to own. Product can confirm whether a missing feature is actually missing, merely hard to find, or described in language that does not match how buyers ask. Growth can decide whether a recommendation loss is worth a campaign, a comparison page, or a closer look at where competitors are being discussed.

This is why the prompt and response belong in the report. A score delegates nothing. “AI visibility declined by three points” has no owner. “Claude stopped naming us for the migration question after it began citing a competing integration guide” gives a team a place to start.

Run a review that ends with a decision

The practical review is short. Start with the prompt groups that connect to an active product line, a current campaign, or a sales objection. Sort for changes in recommendation status, then read the full answers rather than a chart. Flag a result only after you can show what was present before and what is present now.

For each meaningful change, write one sentence that names the evidence and one next action. For example: “We are now cited as a source for the category explainer but are not recommended in the pricing-constrained prompt. Review the pricing page and the three competitors named there.” The action may be research, a factual product clarification, a content update, or no action at all. A negative result is still useful if it prevents a team from chasing a vague score.

Keep a small change log beside the prompt results. Note the date a page changed, a product capability launched, a comparison page went live, or a prompt entered the baseline. Then wait for enough repeated runs to see whether the answer changed in the direction you expected. The goal is not a dashboard that moves every day. It is a record that can distinguish a real change from ordinary answer variation.

The output of an AI visibility review should be a decision with evidence attached, not a score that asks the next meeting to make sense of it.

AI referral traffic is a third measurement

Traffic matters, but it is not the score either. A buyer can see a recommendation and act later, choose another route to your site, or never click a link. An assistant can cite a page without naming your product at all. Referral reporting is useful for the visits it receives, but it starts after the answer has already been generated.

That is why traffic, citations, and recommendations deserve their own lines in a review. The AI traffic GA4 cannot see covers the attribution holes that make referral data incomplete. Use it to validate visits. Use prompt evidence to see the buyer conversation that may never result in one.

A good score is a drill-down, not a verdict

Keep a summary if it helps a team watch the whole system. Just make the summary clickable in practice: every number should lead to the prompt group, assistant, dated answer, named competitors, and source pages that created it. That is the difference between reporting visibility and investigating it.

Beseen's report starts from the same premise. It shows the searches behind a result, the AI citations and product names it found, and the competitors present in those answers. Run a a free report to see the evidence before you decide what should count as progress.