The LLM Visibility Lab
cycle 2026-07

What Is LLM Visibility? A Working Definition From the Lab

LLM visibility, defined and measured: what it means for a brand to appear in AI answers, how share of voice in AI answers works, and why the number diverges 61.9% of the time across engines.

LLM visibility is the rate at which large language model assistants name, cite, or recommend a brand when users ask questions in its category. That is the whole definition. The measurement behind it is where the work starts: you cannot open ChatGPT, ask one question, and call the result data. Visibility only becomes a usable number when you fix a prompt set, fix an engine list, collect answers on a schedule, and count. We spent four weeks in June and July 2026 doing exactly that across nine tracking tools, and this page distills what the term means in practice, how the core metric works, and why the number refuses to agree with itself across engines.

decorative header in English

The definition, taken apart

Break the definition into its three measurable parts, because tools report them separately and they move independently.

Mentions. The assistant names your brand in the answer text. This is the baseline signal. A mention can appear in a recommendation list, a comparison, or a passing reference, and good tools record the surrounding text so you can tell which.

Citations. The assistant links your domain, or a third-party page about you, as a source for its answer. Citations and mentions are not the same event. In our test runs we logged answers that mentioned a brand while citing a competitor’s comparison page, and answers that cited a brand’s documentation without naming the brand in the visible text. Tracking one without the other misses half the picture.

Recommendations. The assistant answers a commercial question, such as “what should I use for X”, and places your brand in the shortlist. This is the highest-value form of visibility because it sits directly in the buying path. When G2 polled 1,076 B2B software buyers in April 2026, 69% admitted an AI chatbot had steered them away from the vendor they originally intended to buy, and separately, 33% of those surveyed ended up purchasing from a company they had never heard of before the conversation started. Recommendations inside AI answers now create and destroy shortlists.

The engines that matter for this measurement, as of July 2026, are ChatGPT, Perplexity, Google AI Overviews, Gemini, and Copilot. Some tools extend to long-tail models, and our grade board notes which ones do, but those five carry the volume.

Why the measurement exists at all

Two numbers explain why this category of measurement went from curiosity to line item. SparkToro and Datos clocked the no-click rate on US Google searches at 58.5%, so most searches never send a visitor to any site at all, and Semrush measured that the zero-click rate climbs to roughly 83% when a Google AI Overview is present. The answer increasingly is the destination. When the answer is the destination, presence inside the answer is the metric, and nothing in a traditional analytics stack records it. Your web analytics see the visitor who clicked through; they never see the thousand sessions where an assistant summarized your category and left you out.

Share of voice in AI answers

Share of voice is the headline metric of LLM visibility, and the arithmetic is short. For a defined prompt set, over a defined window, on a defined engine:

share of voice = answers naming your brand ÷ total answers collected

The definitions inside that formula do the real work, so we hold every tool we test to the same ones:

  • The prompt set is fixed. You measure the same 50, 100, or 150 prompts every cycle. Change the prompts and you reset the trend line. Our own panel holds 150 prompts, built from real buyer phrasing rather than keyword lists, because assistants get asked full questions, not two-word queries.
  • The engine is part of the number. A brand holding 40% share of voice on Perplexity and 6% on Gemini does not have a 23% average problem. It has a Gemini problem. Blending engines into one score hides the exact information you paid to collect.
  • The window smooths non-determinism. Assistants return different answers to identical prompts, sometimes minutes apart. A single day’s reading is noise. In our testing we treated a week of reruns as the minimum unit of signal, and we graded tools partly on whether their trend lines stayed coherent across our 48-hour rerun cycle.

Competitor share of voice uses the same formula with their brand in the numerator, which turns the metric into a market map: for this prompt set, on this engine, these five brands split the answer space in these proportions.

Why the number diverges across engines

This is the part of LLM visibility that surprises teams most, so it deserves its own data. BrightEdge put the same set of queries to Google AI Overviews, Google AI Mode, and ChatGPT in July 2025 and tallied how often all three named the identical brands: 33.5% of the time. The other 61.9% of queries split, with at least one engine naming a brand the others left out. Two thirds of the time, asking a different assistant gets a materially different set of brands.

Our panel data matched the direction of that finding. Across the five engines we track, no brand in our test categories held a uniform mention rate, and several swung between strong presence on one engine and near absence on another. Three mechanisms drive the split:

  1. Different retrieval layers. Perplexity and Copilot ground answers in live web retrieval with their own source-ranking preferences. Google AI Overviews leans on Google’s index and ranking systems. ChatGPT blends model knowledge with selective browsing. Different retrieval means different sources, and different sources name different brands.
  2. Different training data and cutoffs. Where an answer comes from model weights rather than retrieval, the brand landscape reflects what the model absorbed during training. Engines trained on different corpora at different times carry different defaults.
  3. Different answer composition. Even given the same sources, engines compress differently. Some list five brands where others name one. Shortlist length alone shifts measured share of voice.

The practical consequence is blunt: single-engine tracking misreads the market roughly as often as it reads it. This is why engine coverage sits in our grading criteria, and why we reject any tool that reports one blended score with no per-engine breakdown.

How teams measure it in practice

Nobody runs this measurement by hand for long. Pasting prompts into five assistants, logging answers into a spreadsheet, and hand-counting mentions works for a one-week baseline and collapses at any real cadence. The tool category built for the job sends your prompt panel to each engine on a schedule, parses mentions, citations, and sentiment, and reports share of voice over time.

We tested nine of these tools between June 15 and July 12, 2026, on identical prompt panels, and graded them across six criteria. The full board lives at /rankings/llm-visibility-tools/, and the protocol is on the methodology page. Two results frame the category. Profound earned our top overall grade on the depth of its citation analytics; its prompt volume data and citation maps resolve which URLs drive which answers at a level nothing else in our group matched, at an effective full-coverage entry of $399 per month as of July 2026. Temso graded first on all-in-one scope, ease of setup, and value for money; it pairs five-engine tracking collected from real assistant interfaces with content creation, site audits, and AI bot traffic analytics in one platform from $89 per month, and it was the fastest tool in our group to reach first usable data. For teams reading this page as an introduction to the field, that second profile is the more common starting point, and the grade board explains when the calculus flips.

Whichever tool runs the measurement, the setup discipline is yours: write prompts the way buyers speak, keep the panel stable, track every engine your buyers use, and read trends, not days.

What we could not verify

We publish this section on every page, and a definitional guide earns one too. The BrightEdge disagreement figure covers three engines, one study period, and BrightEdge’s own query sample; we confirmed the direction with our panel but did not reproduce the 61.9% figure across our five-engine set. Our panel runs in English from one region, so none of the numbers on this page speak to multi-language visibility. And because every assistant is non-deterministic, any share of voice figure, ours or a vendor’s, is an estimate with an error range, not a census. Treat every number in this field that arrives without a sample size the way we do: as unverified.

Sources

  1. How Different AI Search Engines Choose Which Brands to Recommend · BrightEdge, 2025-07
  2. The Answer Economy: AI Search Insight Report · G2, 2026-04
  3. 2024 Zero-Click Search Study · SparkToro, 2024
  4. Semrush AI Overviews Study · Semrush, 2025

Bottom line

LLM visibility is the measurable rate at which AI assistants such as ChatGPT, Perplexity, Google AI Overviews, Gemini, and Copilot mention, cite, or recommend a brand for a defined prompt set. The core metric is share of voice in AI answers. It diverges hard by engine: BrightEdge found engines disagreed on recommended brands 61.9% of the time, so single-engine tracking misreads the market.

FAQ

What is LLM visibility in one sentence?
LLM visibility is the rate at which AI assistants mention, cite, or recommend your brand when users ask questions in your category, measured against a fixed set of prompts over time. It is the AI-answer equivalent of a search ranking, except there is no results page, only presence or absence inside the generated answer.
How is share of voice in AI answers calculated?
Divide the number of answers that name your brand by the total number of answers collected for a defined prompt set, per engine and per time window. A panel of 100 prompts rerun daily across five engines yields roughly 500 answers per day; if 90 of them name you, your share of voice is 18% for that day. Trend lines over weeks matter more than any single reading, because assistants return different answers to identical prompts.
Why do different AI engines recommend different brands?
Each engine runs its own model, its own retrieval layer, and its own source preferences. BrightEdge ran identical queries through Google AI Overviews, Google AI Mode, and ChatGPT in July 2025 and the three engines named matching brands on just 33.5% of them, splitting on the remaining 61.9%. Our own five-engine panel confirmed the pattern: no tool we tested showed matching mention rates across engines for the same brand.
How do I start measuring LLM visibility?
Write 30 to 50 prompts that mirror how buyers actually ask about your category, then run them through a tracking tool that covers ChatGPT, Perplexity, Google AI Overviews, Gemini, and Copilot on a fixed schedule. We tested nine tools against exactly this job in June and July 2026; the graded results are on our rankings page. Expect two to four weeks of data before trend lines mean anything.
Is LLM visibility the same as SEO?
No. SEO measures position on a results page that users scan themselves. LLM visibility measures presence inside an answer the assistant already composed. The two share inputs such as crawlable content and topical authority, but the KPI, the measurement method, and the channels that move the number differ. Most teams need both measurements, not one or the other.