Research reports
How often do answer engines cite you?
We sampled 180,000 prompts across four engines to see which sources get quoted, and what they have in common.
We run a maintained prompt panel — roughly 180,000 prompts spanning twelve verticals — against Google AI Overviews, ChatGPT, Perplexity and Gemini on a fixed weekly schedule, and record which domains are cited, in what position, and with what framing.
This report covers Q1 2026. Method and limitations are at the end; please read them before quoting the numbers.
Headline findings
- 68% of tracked queries returned a generated answer, up from 54% a year earlier
- 3.2 sources cited per answer on average, with wide variance by vertical
- 41% of citations went to a page already ranking in the top three organically
- 12% week-on-week churn in cited sources, median across categories
Ranking predicts citation, but does not determine it
The 41% overlap is the number most people misread in both directions.
It is not "ranking is all that matters" — a clear majority of citations went to pages *outside* the top three, including a meaningful share from pages ranking below ten.
Nor is it "ranking is irrelevant". Pages that ranked nowhere were cited rarely. Retrieval has to find you first.
The honest reading: ranking gets you considered, structure gets you quoted.
What cited pages have in common
Comparing cited against non-cited pages for the same prompts, three properties separated them most:
- A direct answer early. Cited pages answered in the opening lines far more often than uncited pages at the same rank.
- Entity coverage. Cited pages named more of the concepts an expert would expect.
- Self-contained claims. Sentences that survive extraction without their surrounding paragraph.
Schema presence correlated, but weakly, and mostly disappeared once rank was controlled for.
Variation by vertical
- Health and finance cited fewer sources, weighted heavily toward institutional domains
- Technical queries increasingly cited official documentation over community forums
- Retail comparison prompts showed the highest source churn of any vertical
Method and limitations
Prompts are sampled from real question phrasing, not keyword lists, and are held stable so week-on-week comparison is meaningful. Each engine is queried without personalisation from a fixed region.
Limitations worth stating plainly: results vary by region and personalisation in ways a fixed panel cannot capture; engines change without notice, and a shift in our numbers may reflect a change in the engine rather than in the sites we track. Treat the direction as more reliable than the absolute level.