How to Track Brand Presence in AI Search: By Hand or With a Tool
Summarize with AI

There are two honest ways to measure whether AI engines recommend your business, and this covers both.
You can do it by hand with a spreadsheet, a fixed list of prompts and about ninety minutes a month. Or you can hand the running to software and read the report. Neither is the beginner option: they answer the same question with different amounts of your time, and the choice depends on how many locations and languages you are measuring, not on how serious you are.
What does not change between them is the part before the measurement. The prompt set, the vocabulary and the formula are the same either way, and if you get those wrong the tool will be wrong faster than you would have been by hand.
So: the full manual method first, because it is the one nobody publishes in full, then what a tracker actually does with it. With the real numbers both produced for a nail salon in Hialeah.
What you are actually counting
Before the formula, the vocabulary, because the two numbers that matter get used interchangeably and they are not the same thing.
A mention is the engine naming your business in the body of its answer. A citation is the engine linking to a page as a source for that answer. You can be mentioned without being cited, which means the engine knows you exist but is sourcing the claim from somewhere else. You can be cited without being mentioned, which means your page fed the answer and a competitor got named in it.
That second case is the one worth catching, and it is invisible in any tool that reports a single visibility score.
Here is what that looks like with real numbers. This is a Semrush AI Search snapshot for Live Love Beauty Salon & Spa, a salon in Hialeah, measured across four engines in the same period.
| Engine | Mentions | Cited pages | What the gap says |
|---|---|---|---|
| Gemini | 25 | 23 | Balanced. Names it and sources it from its own site. |
| Google AI Overview | 22 | 15 | Names it often, cites it less. The recommendation is coming from third-party sources. |
| ChatGPT | 18 | 31 | Cites the site heavily but names it less. The pages are feeding answers that credit someone else. |
| Google AI Mode | 13 | 17 | The weakest engine of the four, and the one to work on next. |
Read the ChatGPT row again. Eighteen mentions against thirty-one cited pages. The site is doing the work of informing the answer almost twice as often as it is getting named in it. That is a specific, fixable problem, and the aggregate score of 24 says nothing about it.
Why one blended score is the wrong unit
Look at the spread in that table. Gemini 25, AI Mode 13. Same business, same week, same category. One engine names it nearly twice as often as the other.
Now imagine the only number you had was the aggregate: AI Visibility 24. What would you do with it? You cannot tell which engine to work on, whether the problem is being named or being sourced, or whether last month's work moved anything. The average is arithmetically true and operationally useless.
Most tools open on a single blended score because it fits on a card. Treat it as a headline, not a diagnosis. Every decision you make comes from the per-engine breakdown underneath it.
This is not a quirk of one business. Engines differ in how many sources they pull and how much weight each one carries, so the same brand routinely lands in a different band on each. Any report that pools them is averaging away the only signal you can act on.
Build the prompt set before you measure anything
The measurement is only as good as the prompts, and this is where most manual attempts fall apart.
Prompt unbranded. If you ask ChatGPT "what do you think of Live Love Beauty Salon," it will tell you, and you will have learned nothing. You did the recommending. The prompts that matter are the ones a customer types when they do not yet know who you are: they describe a need, a place and a constraint, and the engine chooses.
Prompt the way people actually ask. Not keywords. Sentences with the awkward specifics real people include.
These are nine prompts that were run against this salon. They are worth reading closely, because the shape is the lesson.
| Prompt | Engine | Language |
|---|---|---|
| premium nail salon in hialeah | Google AI Mode | EN |
| bilingual nail salon in hialeah | Google AI Mode | EN |
| full services nail salon in hialeah | Google AI Mode | EN |
| salon de uñas en hialeah full services donde hablen español | Google AI Mode | ES |
| salón de uñas premium en hialeah | Google AI Overview | ES |
| modern, long-lasting gel nails with non-toxic products | ChatGPT | ES |
| which salon do you recommend for my daughter quinceañera in hialeah | ChatGPT | ES |
| recommend a bilingual nail salon in hialeah | Claude | ES |
| premium nail salon in hialeah | Gemini | EN |
Notice the quinceañera prompt. No keyword tool will ever hand you that, and it is the one closest to how a real customer decides. Your prompt set should be roughly a third category prompts, a third constraint prompts (bilingual, non-toxic, open Sunday, near me) and a third occasion prompts.
Every prompt gets asked in both languages, and the results get counted separately. The engines answer differently in English and Spanish, and the set of competitors they name changes with the language. Pooling them gives you a number that describes neither market.
Twenty to thirty prompts is enough to be stable. Fewer and one odd answer swings the whole thing.
The formula, calculated per engine
The math is not the hard part, and it is already public. Share of voice is:
AI SOV (%) = (your brand's citations ÷ total category citations) × 100
The discipline is in the denominator and the scope. Total category citations means every brand named across your whole prompt set on that engine, not just the ones you consider competitors. And you calculate it once per engine, never pooled.
So a twenty-prompt set run across four engines gives you four share-of-voice numbers, not one. That is the deliverable. If you finish with a single percentage, you did the arithmetic right and the study wrong.
Route one: the workflow by hand
Write the prompt set and freeze it
Twenty to thirty unbranded prompts, split between category, constraint and occasion. Freeze the list. The moment you change prompts between runs, you lose the ability to compare months, which is the entire point of measuring.
Run every prompt in a clean session
Log out, or use a temporary chat. A logged-in session carries your history and personalizes the answer toward you, which quietly inflates every number you are about to record. Run each engine separately: ChatGPT, Gemini, Perplexity, Google AI Mode and AI Overviews all answer differently.
Record two columns, not one
For every prompt and every engine, log whether your brand was MENTIONED in the answer and whether your site was CITED as a source. They are separate columns because they are separate problems. Log every other brand named in the same answer while you are there, because that is your denominator.
Screenshot the answers you will want to prove later
Answers change. A claim about last quarter with no capture behind it is unverifiable in six weeks, and screenshots are what turn a spreadsheet into something a client believes.
Calculate share of voice per engine, then read the gaps
Four engines, four numbers. Then compare mentions against citations within each engine. A high-citation, low-mention engine is treating your site as a source while naming someone else, which is a different fix from being invisible entirely.
Repeat monthly with the same prompts
The absolute number matters less than its direction. One run is a photograph; three runs are a measurement.
Route two: what a tracker actually does
A tracker does not know anything you do not. It runs the same prompts you would have run, records the same two columns, and does the arithmetic. What you are buying is the running, and the running is the expensive part: thirty prompts across five engines in two languages is three hundred sessions a month, and that stops being a spreadsheet job.
Three things it does that a person realistically will not:
It runs on schedule, forever. The value of this measurement is the trend, and the trend dies the first month you are busy. Software is not busy.
It keeps clean sessions by default. The most common way a manual study goes wrong is running prompts while logged in, which personalizes the answers toward you and inflates every number. A tracker has no history to contaminate it.
It captures the answer at the moment it ran. AI answers change. Six weeks later you cannot reconstruct what ChatGPT said, and a claim with no capture behind it is not a claim.
Per-engine numbers, mentions and citations as separate columns, your own prompt set rather than a generated one, and the raw captures. A tool that shows you one blended score and no evidence has taken the running off your hands and thrown away the finding.
Which route fits you
| By hand | With a tracker | |
|---|---|---|
| Cost | Free | A subscription |
| Time per month | About 90 minutes | Minutes, reading the report |
| Best for | One business, one market, establishing a baseline | Several locations or languages, or a monthly client report |
| Where it breaks | The month you are too busy to run it | Never, which is the point |
| What you learn | Everything, because you saw every answer | The numbers, unless the tool shows you the captures |
If you are measuring one business in one market and you have never done this before, do it by hand first. Not on principle — because reading the answers yourself is what teaches you what the number means, and you only need to learn that once.
If you are past that, or you are reporting to clients every month, the running is not a good use of anyone's time. That is what AI Signal Rank is: our own tracker, built because we hit this ceiling ourselves running measurements for local businesses in two languages. It exists whether or not you use it, and the method above works exactly the same if you would rather keep the spreadsheet.
The prompt set. No tool can write it for you, because it depends on how your customers describe their problem, in their language, in your market. Build it by hand once and it keeps working whatever you plug it into later.
Common questions about tracking brand presence in AI search
What is a good AI visibility score?
There is no published benchmark, and be skeptical of anyone quoting one. The scores are vendor-specific and calculated differently, so a 24 on one platform does not mean what a 24 means on another. Use your own first measurement as the baseline and compete against it.
How often should I measure my AI brand presence?
Monthly is enough for most businesses. AI answers shift constantly, so weekly measurement mostly records noise, and quarterly is too slow to connect a change to the work that caused it.
Do I need to be cited to be recommended?
No, and that gap is the useful finding. An engine can name your business while sourcing the claim from a directory or a review site. It means the recommendation is resting on third-party evidence rather than on your own pages.
Do I need a tool to measure AI brand presence?
No. The manual method in this article is complete and produces the same numbers. A tool buys you the running, not the knowledge, and it becomes worth paying for when the run count crosses a few hundred sessions a month.
Does this replace tracking Google rankings?
No. It is a second surface, not a replacement. The same business can be strong in AI answers and weak in classic results, and the two are diagnosed separately.
Why do my results change every time I run the same prompt?
Because the engines are non-deterministic and personalize on session history. Use clean sessions, keep the prompt set frozen, and read the trend across runs rather than any single answer.
How is this different from optimizing for AI search?
Measurement tells you where you stand; optimization is the work that moves it. If you are ready for that side, it is covered in SEO in the age of AI.
What the numbers are for
The point of measuring brand presence in AI search is not the score. It is the sentence you can say afterwards: this engine names us but does not cite us, that one cites us but names a competitor, and the Spanish results look nothing like the English ones.
None of those sentences survive being averaged into a single number, and every one of them tells you what to do next.
Run it by hand once, whichever route you end up on. You will never read a blended visibility score the same way again, and you will know what to ask of the tool that eventually runs it for you — ours or anyone's.
And if you would rather not run it at all, that measurement is where our AI SEO work starts.




