How Often Should You Re-Test Your AI Visibility
Testing too rarely means you find out about a problem a quarter after it started. Testing too often means drowning in variance that looks like signal and reacting to noise. Both failures are common and the second is more expensive, because it produces work.
This section sounds procedural and it is the foundation of everything after it. A prompt set quietly edited between runs makes every trend line in the document meaningless, and it is the easiest way to manufacture improvement without doing anything.
A third response, attempting to manipulate the review platform, fails for mechanical as well as ethical reasons. Fabricated accounts tend to be uniform in language and timing, which is the pattern that gets discounted, and platforms enforce against it with increasing effectiveness.
If you must change the prompt set, add new prompts as a separate cohort and keep the original series running unchanged. Editing the instrument retrospectively destroys the comparison you have been building.
It is also worth asking for the report a day before the meeting rather than seeing it in the room. A document presented live is experienced as a narrative and approved on the strength of the delivery. The same document read beforehand is experienced as evidence, and the questions that occur to you reading it alone are usually the ones worth asking.
You are unlikely to read all of it, and its presence changes the incentives entirely. An agency that knows the raw evidence ships with the report writes a different summary than one that knows it will not be checked.
Show the Cheap Failures First Before asking for a programme, ask for permission to check whether you are readable. Crawler access, rendering without JavaScript, listing accuracy on the sources your prompts cited.
Results Split by Intent, With Run Counts Not one number. Mention rate reported as a fraction with the run count visible, broken out by prompt tier, so buying intent is never blended with definitional questions.
One preparation step is worth the effort. Before the meeting, check whether anyone in the business has already noticed something relevant: a customer who mentioned an assistant, a support ticket citing wrong information, a salesperson who was asked about a competitor comparison they had not seen. Internal anecdote carries disproportionate weight because nobody can dismiss it as vendor material.
Handle the Statistics Carefully Numbers circulate in this field faster than anyone checks them, and using an unsourced one is the fastest way to lose a room. Attach the provenance to everything you cite:
The Prompt Set, Unchanged The report opens with the prompt set used, versioned and dated, and a statement that it is identical to last month's. If it changed, the change is listed explicitly with a reason, and the previous series is kept alongside so comparisons remain honest.
Overclaiming here is the main risk to your own standing. A proposal that promises a channel shift and delivers a corrected directory listing will be remembered. One that promised a baseline and delivered a baseline plus some unexpected fixes will be renewed. llm seo
One test separates a report written to inform from one written to reassure. Read it and try to write down a question it does not answer. In a good report you will find several, because it contains enough specifics to make new questions obvious. In a padded one you will struggle, not because everything is covered but because there is nothing specific enough to interrogate.
One scheduling detail improves comparability more than it should. Run on roughly the same date each month rather than whenever somebody remembers. Retrieval behaviour and the freshness of competing sources both vary over a month, and a series taken at irregular intervals introduces variation that looks like a trend.
The condition is that it has to be honest. A comparison where every row favours you is transparent to readers and produces nothing quotable as an impartial claim. Name real competitors, use concrete axes, and state plainly where somebody else is the better choice.
Search traffic includes everybody at every stage, including a large volume of people gathering background information with no intention of buying anything. Assistant referrals skip most of that, because the informational portion was satisfied before the click.
A useful way to think about the sequence is that each stage moved a task from the user to the interface. First the fact, then the summary, and now the comparison. Each move removed a reason to visit a website, and each was followed by an industry insisting the change had been overstated. It is reasonable to expect the pattern to continue rather than to stop at a convenient point.
Watch the source list as closely as the mention rate, because it usually moves first. New citations from a directory you corrected are a leading indicator, and they typically appear a month or two before any change in whether you are recommended.