The LLM Visibility Lab grades tools that measure brand presence in AI assistant answers. The bench, prompt panel, and scoring protocol are documented below.
The test bench
The bench is a scheduling harness that fires a fixed 150-prompt panel at five assistants: ChatGPT, Perplexity, Google AI Overviews, Gemini, and Copilot. Runs repeat every 48 hours, and every raw answer is stored with a timestamp and a diff against the previous run. For the July 2026 board, we tested nine trackers between June 15 and July 12, set each one up ourselves on self-serve plans where they exist, and checked their reported mention rates against 300 answers we scored by hand.
Grades follow a published scale, A through F per criterion and overall, with the letter-to-number map printed on the methodology page. Assistants are non-deterministic, so we grade trend stability across reruns rather than trust any single reading. The panel runs in English from one region, and we flag that limit on every ranking, because a bench that hides its blind spots is not a bench.
What to read next
Start with the grade board for the current rankings, the matchup index for head-to-head tests, or the lab notes for what moved between reruns. Found an error in our data? The protocol page explains how to report it, and confirmed corrections ship with a note in the changelog.