
A company asks ChatGPT one question, sees its name and declares its AI optimisation successful. A week later, the same question returns different competitors and sources.
That is not measurement. It is one observation in a variable system.
Answers may vary with wording, language, location, time, user context, platform and retrieved sources. Visibility therefore cannot be evaluated through one screenshot or an opaque universal "AI score".
A repeatable method is required that separates brand mention, site citation, competitor list position, factual accuracy, real referrals and impact on enquiries and sales.
Define the outcome
AI visibility is not a single metric.
| Outcome | Question |
|---|---|
| Mention | Is the brand named? |
| Citation | Is the domain used as a source? |
| Recommendation | Is the company recommended? |
| Position | Where does it appear in a list? |
| Accuracy | Is it described correctly? |
| Referral | Did a user visit the site? |
| Conversion | Did AI influence an enquiry? |
Citation, mention and recommendation are not interchangeable.
Build a question set
A traditional keyword is often short: "SEO agency Riga". An AI user may ask: "Which digital agency in Latvia can run a technical SEO audit and then implement fixes on a three-language website?"
Include at least five categories in the set:
- Category discovery;
- specific customer situations;
- option comparisons;
- evidence of sector or technology experience;
- questions about the brand itself.
Begin with 10-20 questions per important service, based on sales conversations, objections, Search Console and the semantic core.
Start with commercially significant decisions rather than hundreds of random prompts.
Control the test
Record:
- the exact prompt;
- platform and mode;
- language;
- target market;
- date and time;
- session type;
- answer, sources and visible URLs;
- the model name if visible.
Measure LV, RU and EN separately when each language represents a market: a translation can retrieve different sources and competitors.
Repeat observations
A generated answer is not a fixed SERP position. Each question must be repeated and evaluated over a period.
A practical starting design:
- 30 priority questions;
- 3 platforms;
- 3 repetitions;
- one full monthly measurement;
- separate language dashboards.
This creates 270 observations per language rather than one favourable screenshot.
Core calculations
Mention Rate
answers naming the brand / valid answers × 100
Citation Rate
answers citing the domain / valid answers × 100
Recommendation Rate
genuine recommendations / commercial questions × 100
Share of Answer
brand mentions / mentions of predefined competitors × 100
The competitor set must be defined in advance. Changing the comparison group each month makes the trend meaningless.
Factual Accuracy
Classify each claim as correct, partly correct, outdated, unsupported or wrong. A wrong recommendation can be worse than absence.
Separate sources from recommended brands
AI may cite your guide while recommending a competitor. The information is useful, but the commercial association is weak.
Conversely, an external directory may support a brand recommendation, revealing the value of off-site reputation.
Report cited domains and recommended brands separately.
Referral traffic and zero-click influence
OpenAI adds `utm_source=chatgpt.com` to ChatGPT Search links. Track landing pages, engagement, forms, calls and assisted conversions.
Create an AI referral segment and monitor sessions, landing pages, engagement and enquiries.
However, users may later search the brand, visit directly or enquire on another device. Also review branded search, direct traffic and "How did you hear about us?" responses.
Establish a baseline
Complete one measurement cycle before major changes.
Maintain a log of new pages, technical fixes, external mentions, profile updates, structured-data changes and crawler rules.
This does not establish perfect causation, but it helps distinguish systematic work from random fluctuation.
Executive reporting
Executives do not need a 200-prompt spreadsheet. A monthly report should show:
- visibility across priority questions;
- change from the previous period;
- Share of Answer against 3-5 competitors;
- most frequently cited sources;
- factual errors and reputation risk;
- referral and assisted conversions;
- the next three actions.
Common mistakes
- testing only branded prompts;
- running one query;
- confusing citation with recommendation;
- changing prompts between periods;
- merging different languages into one metric;
- ignoring factual errors;
- evaluating traffic only;
- publishing a "visibility score" without disclosing its formula.
If the company is currently absent, start with Why Don't ChatGPT and Google AI Mention Your Company?. Our AI SEO, GEO and AEO guide explains the strategic work behind the metrics.
Conclusion
AI visibility is measurable, but not with one magic number. It requires a stable question sample, repetitions and separate analysis of mentions, citations, recommendations, accuracy and business outcomes.
The system must answer three questions: can AI find us, does it describe and recommend us correctly, and does that visibility help create demand.
Juice can establish an initial benchmark through its SEO services and connect the findings to technical, content and reputation priorities.
Frequently asked questions
How often should AI search visibility be measured?
For most businesses, one complete monthly measurement and a lighter weekly check of the highest-priority questions is sufficient. Increase the frequency after major site changes, a product launch or a reputation risk.
How many questions are required for an AI visibility benchmark?
For a small company, 20-30 commercially relevant questions per primary market or service group is a practical starting point. A representative sample and consistent repetition method matter more than a very large volume.
Does one ChatGPT screenshot prove a result?
No. One answer is a single observation because results can vary with wording, session, time, language, location and retrieved sources. Conclusions require repetitions and a comparable measurement period.
Does AI referral traffic show the complete result?
No. Referral traffic measures clicks only. An AI recommendation may later produce a branded search, direct visit, phone call or enquiry on another device, so assisted conversions and changes in branded demand should also be considered.
What is the difference between citation rate and mention rate?
Citation rate measures how often AI uses the company's domain as a source. Mention rate measures how often the brand itself is named in the answer. A company may be cited without being mentioned or mentioned based on another source.
Sources
Google generative AI optimisation guide: developers.google.com/search/docs/fundamentals/ai-optimization-guide
OpenAI publisher FAQ: help.openai.com/en/articles/12627856-publishers-and-developers-faq
OpenAI crawler documentation: developers.openai.com/api/docs/bots


