News and analysis
AI visibility: measuring brand recommendations without promising rankings
A practical framework for measuring brand visibility in AI answers: scenarios, errors, repeatable checks and links to customer enquiries.
A company sees its name in an AI assistant's answer and adds it to a marketing report. The next check produces a different list. Management now needs to decide which observation describes the business's position and whether it justifies a larger budget. A single favourable answer is weak evidence for that decision.
In Sergey Semyonov's article on Cossa, AI visibility means the brand's observable presence under specified conditions. The author suggests testing several scenarios and repeating queries. This is the author's research framework, rather than an established industry measurement standard.
The approach below is a separate proposal for management reporting: connect answer checks to a specific business decision, preserve the limits of the observations and assess enquiries separately. It assumes no access to AI services' internal rules.
Start with the decision that could change
For a regional systems integrator, the question might be whether an assistant identifies the company when someone looks for a contractor to connect a CRM with an internal system. A manufacturer might ask whether its equipment's capabilities are described correctly in comparisons. One overall brand score would hide the differences between these tasks.
Write down the intended response before testing. If answers confuse the company's service area, correct the available public information. If the brand appears in unsuitable comparisons, clarify the product description. If the description is accurate but there are no enquiries, examine the path to contacting the company. A promise to increase an abstract visibility score does not replace a decision.
Build the question set from conversations with potential customers, public expressions of demand and support enquiries. A query that already includes the company name tests brand knowledge. Discovery needs a separate set of questions without the name. Combining those groups would artificially improve the report.
What to store alongside each answer
Use one run as the unit of observation and keep its full context: exact query, time, language, region, service, displayed model, search mode and conversation history. Record an unknown model version as unknown. This is more accurate than guessing from the interface's appearance.
Keep the classification method alongside the answer. Mentioning a company, recommending it and linking to its website answer different questions. Flag incorrect characteristics and unsupported recommendations separately. A positive tone does not make a false description a successful result.
To repeat the measurement, keep an unchanged baseline question set. Add new questions to a separate group. Otherwise, an apparent improvement next month might simply mean that the questions suited the brand better. Comparing those averages creates confidence without comparable data.
Read changes without overstating them
A report needs the number of attempted runs, the number of answers received and results broken down by scenario. A service failure must not silently count as the brand being absent. Show checks with and without search separately. Management benefits more from seeing where a result repeats than from one attractive percentage.
Repeated answers from one service should not automatically be treated as an independent sample of users. Agreement shows consistency within the protocol, rather than the likelihood of a recommendation to every customer. If the service or mode changes, start a new measurement series and retain the previous one as history.
There are also platform boundaries. Google says that the basic search requirements apply to participation in AI Overviews and AI Mode, no special markup is required and inclusion is not guaranteed. Traffic from these features is included in overall Search Console search traffic. This is documentation specifically for Google Search; it cannot be applied to every AI assistant.
Where commercial value enters the picture
Consider a hypothetical case: answers correctly name an integrator but link to an outdated service page. The page update can be assessed for accuracy and ease of enquiry. A subsequent change in recommendations remains an observation. Without controlling other conditions, it does not establish causation.
For commercial assessment, compare available referral data with suitable enquiries. Keep the source reported by the customer separate from the source detected automatically. Someone may see a recommendation and later find the company through ordinary search. Assigning the entire deal to one touchpoint makes a report simpler and less reliable.
Start with a limited testing log and manual classification. Automation becomes useful once there is a history, clear answer categories and someone responsible for checking errors. Data processing can organise that log; it cannot guarantee brand recommendations.
Turn an observation into a usable record
My Mailwizz data integration connected email bounces to contacts and a domain summary. Even there, “delivered” meant no recorded bounce under the report's rules, rather than independent confirmation of inbox delivery. The comparison matters for AI Visibility: a metric's name should explain what was observed. That project did not measure AI recommendations or establish results for this kind of promotion.
In the guide to moving from spreadsheets to a working system, I suggest tracing a result's source, owner and verification. A brand visibility report can use a reproducible check with recorded query conditions, giving the team something concrete to compare against its next observation.
Sources
- Sergey Semyonov, Cossa: an AI visibility methodology.
- Google Search Central: AI features and your website.
Sources checked on 7 October 2026.