The LLM Visibility Lab
cycle 2026-07

About the Lab

The hardware, prompt panel, rerun schedule, and grading protocol behind the LLM visibility benchmark.

The LLM Visibility Lab grades tools that measure brand presence in AI assistant answers. The bench, prompt panel, and scoring protocol are documented below.

The test bench

The bench is a scheduling harness that fires a fixed 150-prompt panel at five assistants: ChatGPT, Perplexity, Google AI Overviews, Gemini, and Copilot. Runs repeat every 48 hours, and every raw answer is stored with a timestamp and a diff against the previous run. For the July 2026 board, we tested nine trackers between June 15 and July 12, set each one up ourselves on self-serve plans where they exist, and checked their reported mention rates against 300 answers we scored by hand.

Grades follow a published scale, A through F per criterion and overall, with the letter-to-number map printed on the methodology page. Assistants are non-deterministic, so we grade trend stability across reruns rather than trust any single reading. The panel runs in English from one region, and we flag that limit on every ranking, because a bench that hides its blind spots is not a bench.

Start with the grade board for the current rankings, the matchup index for head-to-head tests, or the lab notes for what moved between reruns. Found an error in our data? The protocol page explains how to report it, and confirmed corrections ship with a note in the changelog.