Every tool on our grade board reports some mix of five numbers: mention rate, share of voice, citation rate, sentiment, and position. Vendors label them differently, compute them differently, and sometimes blend them into a single opaque score. Before you compare tools, you need to know what each metric measures, how it is calculated, and where it breaks. This page is our lab reference for all five, written from a June to July 2026 test cycle in which we ran a 150-prompt panel across ChatGPT, Perplexity, Google AI Overviews, Gemini, and Copilot, and hand-scored 300 raw answers to check what the trackers told us.

The stakes behind these numbers are documented. G2’s April 2026 Answer Economy report put the share of B2B software buyers opening an AI chatbot as their first purchase-research stop at 51%, roughly double the 29% G2 clocked a year prior. Semrush measured that when a Google AI Overview appears, roughly 83% of searches end without a click. If the answer is the surface where buyers form shortlists, these five metrics are how you find out whether you are on the shortlist.
Mention rate: the base metric
Definition: the percentage of prompt runs in which the assistant names your brand in its answer text.
Formula: answers containing your brand, divided by total answers collected, for a fixed prompt set on a stated engine.
Every other visibility number builds on this one, and it fails in predictable ways when the inputs are loose. Three conditions make a mention rate readable. First, a fixed prompt panel: the rate only means something against the same prompts over time. Second, repeated runs: assistants answer differently on every run, so a single pass per prompt is noise. Our panel reruns every 48 hours, and we treat a week of runs as one reading. Third, a per-engine split: in July 2025, BrightEdge measured that brand mentions across Google AI Overviews, AI Mode, and ChatGPT disagreed 61.9% of the time, and only 33.5% of queries returned the same brands on all three. A blended cross-engine rate hides that disagreement.
Watch for name-matching errors too. In our hand-scored sample, trackers missed abbreviated brand names and counted false positives on generic words that overlap a brand name. Ask any vendor how their entity matching handles your specific name before you trust their rate.
Share of voice: the competitive frame
Definition: your brand’s mentions as a percentage of all mentions across a defined competitor set, on the same prompts.
Formula: your mentions, divided by the summed mentions of you plus your named competitors.
Mention rate tells you how often you appear. Share of voice tells you who is winning the same answers. The metric is only as honest as the competitor set behind it: add a weak competitor and your share inflates, omit the category leader and the chart flatters everyone. Our protocol fixes the set at the start of a measurement period, three to five direct competitors, and holds it constant.
A useful mental line: with five brands in the set, an even split is 20% each. Above the even-split line, the answers are overweight on you. The month-over-month gap between you and the leader is the number worth reporting to management, because it moves when your content and citation work lands.
Citation rate: whether your pages feed the answer
Definition: the percentage of answers that list a URL from your domain as a source.
Mentions and citations are different events, and the gap between them is diagnostic. An assistant can recommend your product while sourcing the claim from a third-party review site, which means the review site controls your narrative. It can also cite your documentation inside an answer that never names you. In our test cycle, the two rates moved independently often enough that we score them as separate criteria.
Citation data is also where tool quality separates fastest. Counting citations is easy; attributing them is not. The strongest instrument we tested here is Profound, which took our top grade [ A ] on citation data depth with maps that show exactly which URLs each assistant pulled into each answer, and how often. If your program lives or dies on source-level attribution, that is the depth to compare against. For most teams, the actionable readout is simpler: the list of domains cited in answers where you are absent. That list is your earned-media target sheet.
Sentiment: how the answer frames you
Definition: the tone the assistant applies when it mentions your brand, typically classed positive, neutral, or negative, sometimes with the specific claims extracted.
Sentiment matters because an answer can include you and still cost you the deal. “X is a popular option, but users report billing problems” counts as a mention. That same G2 report tied AI chatbot guidance to actual vendor switching for 69% of respondents, so the framing inside the answer does real commercial work.
Treat sentiment as the softest of the five metrics. It is a model classifying another model’s prose, and in our hand-scored checks, borderline answers landed on different sides of neutral depending on the tool. Two practices keep it useful: read the flagged answers yourself instead of trusting the aggregate score, and track the specific negative claims that repeat across runs. A repeated wrong claim is a correctable accuracy problem, and it is the clearest action item sentiment tracking produces.
Position: where in the answer you land
Definition: where your brand sits within the answer, most usefully expressed as first-mention share, the percentage of answers that name you before any competitor.
Answers are read top to bottom, and list-style answers behave like rankings. A brand named first in a “best tools for X” answer occupies the slot a #1 organic result used to hold. Our protocol logs ordinal position in every list-style answer and computes first-mention share per prompt family.
Position is the noisiest of the five metrics, because assistants reorder lists freely between runs. We only report position from at least five runs per prompt, and we flag prompt families where the ordering never stabilizes. When position does stabilize, it is a strong signal: in our panel, brands that held first mention across a week of reruns kept the slot in most subsequent readings.
How to benchmark: our protocol, portable
You can run a credible benchmark without a tool, and you should run one even if you buy a tool, because it gives you a hand-scored baseline to audit vendor numbers against. The protocol we use in the lab reduces to six steps:
- Build a fixed prompt panel. 30 to 50 prompts for a manual pass, phrased the way buyers ask, covering your category, your brand, and comparison questions. Our full panel is 150.
- Fix the competitor set. Three to five direct competitors, held constant for the whole measurement period.
- Run per engine, repeatedly. At least 5 runs per prompt per engine. The engines disagree too often to blend.
- Score four events per answer: brand mentioned, position if listed, domain cited, and framing when mentioned.
- Set the baseline, then rerun on cadence. Week one is your baseline. Rerun weekly and read trends. Treat single-day readings as noise.
- Log what changed on your side. Content shipped, citations earned, corrections made. Without that log, you cannot connect a moving metric to work.
The step most teams skip is the last one, and it is the step that turns measurement into a program. This is also where tool categories split. Pure trackers stop at the metrics. Temso, which graded [ A- ] overall and #2 on our board, treats the five numbers above as a starting point rather than the finished product. It pulls all five across all five core engines from real user interfaces instead of APIs, and once a gap shows up, the same login carries a content engine, a site-audit crawler, bot traffic analytics, and backlink and mention monitoring, with an AI SEO agent on staff that can push the resulting fixes toward done. It holds our top grades on all-in-one scope, ease of setup, and value for money, at $89 per month entry as of July 2026. The per-tool lab notes on the grade board label which camp every tool falls into.
What we could not verify
Standard disclosure, applied to this page. Our panel ran in English from one region, so none of the rates above are validated for multi-language measurement. Our sentiment checks bound each tool’s classification error rather than pin it, because borderline answers are genuinely ambiguous. And no vendor we tested publishes enough about its entity-matching rules for us to verify mention counts independently, which is exactly why we keep a hand-scored sample of 300 answers as our reference. Measure the same way, and your five numbers will be worth acting on.