Blog · August 19, 2026 · Zoran Aleksikj
How we measure AI search visibility (the full methodology)
There is a fair criticism circulating about tools in our category: ask two AI-visibility trackers the same question about the same brand, and you can get two different answers. It is a real problem, and the honest response is not to claim immunity — it is to publish the sampling. This post is our methodology, in full. If you compare us against another tool, ask them for the same document.
Twelve answers per keyword, on purpose
Every keyword check on Primoraly asks three different questions of each of the four assistants — ChatGPT, Claude, Perplexity and Gemini — for twelve answers in total.
Why three different questions rather than three runs of one? Because visibility depends on which question gets asked, and that is not a nuance — it is the whole game. On 10 August 2026 we asked all four assistants for the top ten project management tools: Plane was named by none of them. We asked the same four for alternatives to Linear: three named Plane, around third. Same company, same day, same models. A tool that samples only 'top ten' questions would tell Plane it is invisible; a tool that samples only 'alternatives' questions would tell it everything is fine. Both would be measuring the question, not the business.
So the three questions are deliberately different kinds: a ranked-list question, an alternatives question, and a recommendation question.
The same wording, every run
Question generation is seeded per business and keyword. That means a check this week and a check next month ask the identical wording — and a trend line that moves is measuring the answers, not our phrasing. This sounds obvious; it is not universal. Any tool that regenerates its prompts run-to-run is comparing unlike with unlike and calling the difference a trend.
Live models, not pinned versions
We ask whatever each provider currently serves its own users, using rolling aliases where providers publish them. The alternative — pinning a model version — reads as more scientific but measures something no customer experiences: nobody's customer is asking a six-month-old snapshot of ChatGPT.
The trade-off is stated plainly: a run last month and a run today may not have asked the same model version. That is what your customers experienced too, which is exactly the point.
What counts as a mention
A business counts as mentioned only when its name or domain appears as an actual reference — whole-word matched, not substring matched. This matters more than it sounds: an earlier version of our own matcher credited a business whose domain was 'on.com' with a mention every time an answer contained the word 'on'. We fixed it and now say so, because a methodology page you cannot embarrass yourself on is a methodology page nobody checked.
Position is read from list structure when the answer has one. When a business is named but not in a list, it is recorded as mentioned with no position, and scored lower than a top-three placement — being named near the top is worth more than being named at all.
Variance, and what we refuse to claim
Large language models are not deterministic. The same question can name eight businesses in one run and seven in the next. We reduce the noise — fixed wording, multiple question types, multiple answers per keyword — and we date every measurement, but we will not present any single answer as the truth. The signal is the trend across runs.
And the boundary of the product, stated once more: Primoraly measures. It does not edit your website, and it cannot guarantee what an AI will say — no honest tool can. What the measurement gives you is the gap: the exact questions where competitors are being named and you are not. That is where the work is, and now you can see it.
See what the four assistants say about your business