The 2025 Controlled Experiment: What Makes AI Cite a Brand?
Content with explicit entity markup and structured context achieved an 81% citation rate across 63 prompts and 4 major LLMs — versus a 22% baseline for standard, well-written blog posts.
Abstract
Generative engines do not rank pages — they retrieve chunks. The practical question for any brand is whether content structure measurably changes retrieval outcomes. In 2025, RAG Signal ran a controlled experiment to answer it: identical subject matter, two content treatments, 63 buyer prompts, 4 major LLMs.
The treatment group — content with explicit entity markup, standalone claims, and structured context — was cited in 81% of prompts. The control group — the same information written as standard, well-structured blog posts — was cited in 22% of prompts. The 59-point gap is not a stylistic nuance; it is the difference between being a source and being invisible.
Methodology
Design. Controlled A/B comparison. The same core information was written in two treatments: (A) control — standard, well-written blog posts with conventional keyword optimization; (B) treatment — the same content restructured with explicit entity definitions, standalone answer-shaped claims, and machine-addressable context. All other variables (domain, publishing frequency, internal linking) were held constant.
Prompt set. 63 prompts across 7 topic categories: research excellence, specific departments, notable alumni, recent breakthroughs, rankings, funding, and industry partnerships. Prompts were designed to mimic real buyer and researcher queries.
Models. GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and Perplexity Pro — the four major generative engines at the time of testing.
Measurement. Each prompt was run three times per model to account for output variability; the average citation rate was recorded. A prompt counts as a citation when the model names the brand or source explicitly in its answer.
Reliability. A test-retest analysis over a two-week window produced a correlation of 0.94 between rounds, indicating high measurement consistency.
Findings
Primary Result: Structure Determines Citation
| Treatment | Prompts citing brand | Citation rate |
|---|---|---|
| Control — standard blog posts | 14 / 63 | 22% |
| Treatment — entity markup + structured context | 51 / 63 | 81% |
Secondary Result: Backlink Authority Does Not Predict AI Citations
In a companion audit of the same 63-prompt framework, brands with zero traditional backlink authority but content structured for RAG retrieval were cited in 34% of AI responses. Brands with Domain Authority scores above 70 but poorly structured content were cited in only 12% of responses. In B2B software queries, fewer than 12% of cited sources came from domains with a DR above 70 — the majority were mid-tier niche blogs, documentation sites, and community threads.
| Profile | AI citation rate |
|---|---|
| High DA (70+) — poorly structured content | 12% |
| Low/no backlink authority — RAG-ready content | 34% |
| Treatment condition (entity markup, full framework) | 81% |
Applied Results
The same framework applied in client engagements produced consistent outcomes: a DevOps SaaS engagement (12 landing pages + 4 whitepapers restructured) saw a 34% increase in citation rate across 63 tracked prompts in 90 days, and the Filmfolk engagement went from 0% to 81% in 90 days.
Interpretation
Retrieval is chunk-level, not page-level. Generative engines retrieve passages and score them for semantic closeness to the query. A page can satisfy Google's crawler and fail the vector test — which explains why top-3 Google rankings produced AI citations only 34% of the time in our companion research on Google rankings vs AI citations.
Entity markup is the highest-leverage change. The treatment differed most in explicit entity definitions and standalone claims. Models retrieve entities, not prose. When a brand is defined as an entity — with clear relationships to products, people, and categories — it becomes attachable to an answer.
Authority transfers differently in AI. Backlinks build PageRank; they do not automatically build retrieval trust. Source trust in AI is earned through consistency, structure, and third-party reference — a different asset class.
Limitations
- Results reflect the model versions available at testing time (2025–2026). Model updates can shift retrieval behavior; the retainer monitoring exists precisely because of this.
- The prompt set (63) is small relative to the full space of buyer queries, though it is representative of high-intent B2B research behavior.
- Citation rate measures brand mention, not sentiment or placement within the answer.
- Controlled treatments were limited to content structure; other retrieval signals (third-party citations, temporal freshness) were evaluated separately in the companion audits cited above.
Run this experiment on your brand
Get your Citation Baseline — the same 63-prompt framework applied to your category, with a full report in 5 business days.
Get Your Citation Baseline →