Most "AI visibility" numbers are too flattering to be useful. They are usually produced by asking a model some questions, searching the response for your brand name, and counting the hits. That method reliably reports progress that is not happening.
If you are going to measure this, measure it in a way that can tell you bad news.
The three false positives
1. The mention-in-denial
You ask "what are the best GEO tools?" and the model replies: "Tools like Visibilitas are sometimes mentioned, but I would not recommend it for enterprise teams." A substring match scores that as a citation. It is closer to the opposite. Any honest measurement has to classify sentiment and role, not just presence.
2. Source-list bleed
Perplexity-style answers append a sources list. If your domain appears there, a naive scraper counts a mention, even when the prose never mentions you and the source was used for one incidental statistic. Being in the bibliography is not being in the answer.
3. The generic-name collision
If your brand is a common word, you will "appear" constantly. Scores built on substring matching are noise for any brand not lucky enough to be named something unusual.
What to measure instead
- Citation rate: of N tracked prompts, in how many are you cited in the prose, not the footnotes?
- Recommendation rate: of those, how many are positive recommendations rather than neutral mentions or denials? This is the number that correlates with pipeline.
- Share of answer: when competitors are named too, what proportion of the named set are you? Visibility is relative.
- Position within the answer: first named beats last named, for the same reason position one beats position ten.
- Which URL was used: tells you which page is actually doing the work, which is what you scale.
Choose the prompts like a buyer, not a marketer
The most common measurement mistake after naive counting is tracking the wrong prompts. Teams track their brand name, "what is Visibilitas?", and celebrate a strong result. Of course it is strong: the query names you.
The prompts that matter are the ones where you are not named:
- Category: "best tools for tracking AI search visibility"
- Problem: "how do I find out if ChatGPT recommends my competitor"
- Comparison: "alternatives to [competitor]"
- Qualified: "GEO tool for a 5-person marketing team on a budget"
Those are the queries with commercial intent, and they are the honest test of whether you are visible.
Sample properly
Generative answers are non-deterministic. One run tells you almost nothing; the same prompt can name you today and not tomorrow. Meaningful measurement means repeated sampling over time, across engines, with the variance made visible. A single-shot check dressed up as a score is a coin flip with a progress bar.
Close the loop with the boring data
Finally, connect it to reality. AI referral traffic is still small for most sites, but it is measurable, and it converts unusually well, because the assistant has already done the qualifying. Watch referrals from assistant domains in your analytics, and watch Search Console for the query shifts that AI Overviews cause.
The goal is not a bigger number on a dashboard. It is knowing, honestly, whether the machines that increasingly mediate your category are telling people to use you.


