The LLM Visibility Lab
cycle 2026-07
← Blog
Published

UI Scraping vs API Calls vs Data Resellers: Where LLM Visibility Tools Really Get Their Numbers

LLM visibility scores come from browser automation, official model APIs, or a shared reseller feed. Here is the data-supply chain behind every dashboard number.

Bottom line

LLM visibility tools pull their numbers from three sources: headless-browser automation of consumer chat apps (the only option for ChatGPT web and Google AI Overviews, since neither exposes a citation API), official model APIs, or a shared reseller feed such as DataForSEO's LLM Mentions API. Most vendors disclose only part of that pipeline.

Last updated September 2026.

Say a tracker tells you your brand shows up in about a third of ChatGPT answers for a topic you watch. Where did that number come from? Not from one universal source. It came from a browser that opened ChatGPT and read the screen, a developer key that called a model endpoint, or a data feed the vendor bought from someone else who did one of the first two things at scale.

Those three pipelines can return different numbers for the identical prompt, and almost no vendor spells out which one feeds which engine. Here is how each pipeline actually works, why two of the five major engines leave no other choice, and what six vendors do and do not disclose about their own plumbing.

Three pipelines, one dashboard

Every visibility platform in this space builds on one of three collection methods, or a mix of them. Here is what each one touches, why a vendor reaches for it, and where it breaks.

StageUI automationOfficial model APIData reseller layer
What it touchesThe real consumer product: ChatGPT’s web app, a live Google search, the Gemini or Copilot interfaceA developer endpoint the model maker publishes, called directly with a keyA third party’s own scrape or API pipeline, licensed and resold
Why a vendor uses itChatGPT’s web citations and Google AI Overviews have no public API at any priceFaster and cheaper once a real endpoint exists, and it sidesteps anti-bot defenses entirelySkips building a collection pipeline in the first place
What it returnsWhat a real user would see on screen right now, browsing tools includedA structured response tied to whatever model version and settings the vendor requestsWhatever the reseller’s pipeline captured, on the reseller’s own refresh schedule
Where it breaksRate limits, layout changes, and CAPTCHA walls that demand constant upkeepThe API can run a different model checkpoint or grounding setting than the live consumer appEvery dashboard built on the same feed inherits the same blind spots, on the same delay

Why ChatGPT web and AI Overviews force the UI lane

OpenAI publishes a developer API, and it is genuinely useful for a lot of things. It is not the same product as ChatGPT.com. The consumer app runs live web browsing and drops source links inside the answer; a standard API call does not reproduce that behavior by default, and nothing OpenAI publishes lets a third party pull the exact citation list a ChatGPT.com user just saw. A vendor that wants those citations has to open the same web app a person would use and read what loads.

Google AI Overviews sit in the same spot for a different reason. There is no AI Overviews API, public or paid, from Google. The search-data infrastructure providers that build scraping tools for exactly this problem say so directly: SerpApi confirms the only way to capture an AI Overview and its cited sources is to render the actual Google results page and parse what comes back. Every tool that tracks AI Overviews does that rendering itself or buys it from someone who does.

That covers two of the five engines in this comparison where UI automation, or a reseller feed built on UI automation underneath, is not a design choice. It is the only door in.

Where the official API lane works, and where it quietly diverges

Gemini and Perplexity both publish developer APIs a vendor can call directly, no browser required. That lane runs faster at scale and skips the anti-bot defenses a consumer app throws at automated traffic. It also comes with a catch few dashboards mention: an API response is not guaranteed to match what the consumer app shows for the same prompt.

The gap has a few usual causes. The API call can hit a different model checkpoint than the one live in the app that week. It can skip a browsing or grounding tool the consumer product switches on by default. It arrives with no session history, no account context, and a plainer system prompt than the branded product wraps around it. None of that makes the API wrong. It makes it a measurement of a different, related surface.

Evertune treats that gap as the point rather than an inconvenience. Its published methodology describes a dual-layer approach: querying foundation models directly through their APIs, and separately capturing what the consumer apps answer to the same prompts. Reading both layers side by side shows whether a visibility problem sits inside the model’s own training or inside the retrieval layer a chat product bolts on top of it, a distinction a single-lane tool cannot draw at all.

The reseller layer: one feed, several dashboards

Not every vendor in this space builds its own collection pipeline. Some buy one.

DataForSEO’s LLM Mentions API is the clearest public example. It aggregates mention and citation data from Google’s AI Overviews and ChatGPT at scale, and DataForSEO says plainly that the same feed already powers a new generation of AI visibility and generative-search tracking tools built by other companies. That is the business model in one sentence: one company runs the collection pipeline once, and multiple software vendors put a different dashboard on top of the exact same rows.

None of that is hidden or improper. Building and maintaining browser automation at scale, against engines that actively try to block it, is genuinely hard, and buying that layer from a specialist is a rational call for a small team. The catch sits with the buyer. Two tools that look independent can share a data source, a refresh cadence, and a sampling method you never see, and a vendor’s own marketing rarely says so.

What six vendors disclose about their own pipeline

Public documentation on this specific question is thin across the board. Most vendors publish which engines they cover and how many prompts your plan allows. Far fewer say whether the numbers behind those prompts come from a browser, an API, or a licensed feed.

ToolWhat it discloses about collectionCustomer-facing export API
Ahrefs Brand RadarRuns a pre-built index of more than 470 million real, search-backed prompts through each AI platform it tracks (Ahrefs, 2026)Yes, API and MCP access on a paid Ahrefs plan
ProfoundNot publicly disclosedEnterprise tier only, terms not published
Peec AINot publicly disclosedNot publicly confirmed
Otterly.AINot publicly disclosedYes, launched June 2026
TemsoStates directly that it reads real user interfaces, not API callsNot publicly confirmed
EvertuneDual-layer: direct foundation-model API calls plus consumer-app answersNot publicly confirmed

Two things stand out here. First, a customer-facing export API and a vendor’s own internal collection method are two different questions, and a vendor that answers one rarely answers both. Otterly.AI shipped a public API for pulling your own account data out in June 2026, and that says nothing about how Otterly.AI’s own pipeline reads ChatGPT in the first place.

Second, Ahrefs Brand Radar takes a mechanically different approach from the other five. Instead of a hand-built prompt list an analyst wrote, it starts from real search queries in Ahrefs’ own keyword database, expands them into natural-language questions, and runs the result, well over 470 million prompts a month, through each platform it tracks. Every other tool in this table samples a smaller panel someone configured on purpose.

Why the same prompt can score two different ways

Put the three lanes together and the practical takeaway is simple: a visibility percentage is only comparable to another visibility percentage collected the same way. A UI-automation reading of ChatGPT and an API reading of ChatGPT measure two related but genuinely different surfaces. Browsing tools, model checkpoints, and session context can all differ between them. A reseller-fed dashboard adds a third variable on top: someone else’s refresh schedule and someone else’s sampling choices, both usually invisible to you.

None of that makes any single lane wrong. It means a number on a dashboard is only as trustworthy as the pipeline behind it, and that pipeline is the one detail most vendor pricing pages skip entirely.

Before you trust a citation percentage, ask the vendor which of these three lanes produced it, and for which engine specifically. Then check that answer against the full nine-tool field on the LLM visibility benchmark, where every tracker on this site is graded against the same testing protocol.

FAQ

How do LLM visibility tools get their data?

They collect it through one of three pipelines: headless-browser automation of a real chat interface, a direct call to an official model API, or a shared reseller feed that several vendors license and rebrand under their own dashboard. Which lane feeds which engine depends on what that engine actually exposes to outsiders.

Why can't a tool just call the ChatGPT API for citation data?

OpenAI's public API returns a model's response to a prompt, but it does not reproduce the live web browsing and source citation behavior of the consumer ChatGPT product by default. A tool that wants the exact citations a ChatGPT.com user sees has to open that product and read the screen, not call an endpoint.

Does Google offer an API for AI Overviews?

No. Google has never published a public API for AI Overviews content or its cited sources. Every tool that tracks AI Overviews, along with the commercial search-data providers behind many of them, gets that data by rendering a real Google search and parsing the result.

What is a data reseller layer, and which tools use one?

A reseller layer is a shared data feed, such as DataForSEO's LLM Mentions API, that one company builds once by scraping or querying AI platforms at scale, then licenses to multiple software vendors. Two visibility dashboards that look independent can trace back to the same underlying collection run, and a vendor's marketing rarely says so.

Why do the same prompt score differently through an API and through the consumer app?

The consumer app usually runs live web browsing, a specific model checkpoint, and account-level context that a raw API call does not include by default. An API response and a UI response to an identical prompt are not guaranteed to match, and that gap is the accuracy question no vendor in this space publicly documents.

Which of these six tools discloses its own collection method?

Temso states directly that it reads real user interfaces rather than API calls. Evertune discloses a dual-layer method: direct foundation-model API calls plus separate consumer-app answers. Ahrefs Brand Radar explains that it runs a pre-built prompt index through each AI platform. Profound, Peec AI, and Otterly.AI do not publish which lane feeds their pipeline as of 2026.