RAG Signal
Original Research

The 2025 Controlled Experiment: What Makes AI Cite a Brand?

Content with explicit entity markup and structured context achieved an 81% citation rate across 63 prompts and 4 major LLMs — versus a 22% baseline for standard, well-written blog posts.

Bora Kurum Published: August 2026 Apache 2.0 Peer-reviewable

Abstract

Generative engines do not rank pages — they retrieve chunks. The practical question for any brand is whether content structure measurably changes retrieval outcomes. In 2025, RAG Signal ran a controlled experiment to answer it: identical subject matter, two content treatments, 63 buyer prompts, 4 major LLMs.

The treatment group — content with explicit entity markup, standalone claims, and structured context — was cited in 81% of prompts. The control group — the same information written as standard, well-structured blog posts — was cited in 22% of prompts. The 59-point gap is not a stylistic nuance; it is the difference between being a source and being invisible.

Methodology

Design. Controlled A/B comparison. The same core information was written in two treatments: (A) control — standard, well-written blog posts with conventional keyword optimization; (B) treatment — the same content restructured with explicit entity definitions, standalone answer-shaped claims, and machine-addressable context. All other variables (domain, publishing frequency, internal linking) were held constant.

Prompt set. 63 prompts across 7 topic categories: research excellence, specific departments, notable alumni, recent breakthroughs, rankings, funding, and industry partnerships. Prompts were designed to mimic real buyer and researcher queries.

Models. GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and Perplexity Pro — the four major generative engines at the time of testing.

Measurement. Each prompt was run three times per model to account for output variability; the average citation rate was recorded. A prompt counts as a citation when the model names the brand or source explicitly in its answer.

Reliability. A test-retest analysis over a two-week window produced a correlation of 0.94 between rounds, indicating high measurement consistency.

Findings

Primary Result: Structure Determines Citation

Treatment Prompts citing brand Citation rate
Control — standard blog posts 14 / 63 22%
Treatment — entity markup + structured context 51 / 63 81%

Secondary Result: Backlink Authority Does Not Predict AI Citations

In a companion audit of the same 63-prompt framework, brands with zero traditional backlink authority but content structured for RAG retrieval were cited in 34% of AI responses. Brands with Domain Authority scores above 70 but poorly structured content were cited in only 12% of responses. In B2B software queries, fewer than 12% of cited sources came from domains with a DR above 70 — the majority were mid-tier niche blogs, documentation sites, and community threads.

Profile AI citation rate
High DA (70+) — poorly structured content 12%
Low/no backlink authority — RAG-ready content 34%
Treatment condition (entity markup, full framework) 81%

Applied Results

The same framework applied in client engagements produced consistent outcomes: a DevOps SaaS engagement (12 landing pages + 4 whitepapers restructured) saw a 34% increase in citation rate across 63 tracked prompts in 90 days, and the Filmfolk engagement went from 0% to 81% in 90 days.

Interpretation

Retrieval is chunk-level, not page-level. Generative engines retrieve passages and score them for semantic closeness to the query. A page can satisfy Google's crawler and fail the vector test — which explains why top-3 Google rankings produced AI citations only 34% of the time in our companion research on Google rankings vs AI citations.

Entity markup is the highest-leverage change. The treatment differed most in explicit entity definitions and standalone claims. Models retrieve entities, not prose. When a brand is defined as an entity — with clear relationships to products, people, and categories — it becomes attachable to an answer.

Authority transfers differently in AI. Backlinks build PageRank; they do not automatically build retrieval trust. Source trust in AI is earned through consistency, structure, and third-party reference — a different asset class.

Limitations

  • Results reflect the model versions available at testing time (2025–2026). Model updates can shift retrieval behavior; the retainer monitoring exists precisely because of this.
  • The prompt set (63) is small relative to the full space of buyer queries, though it is representative of high-intent B2B research behavior.
  • Citation rate measures brand mention, not sentiment or placement within the answer.
  • Controlled treatments were limited to content structure; other retrieval signals (third-party citations, temporal freshness) were evaluated separately in the companion audits cited above.

Run this experiment on your brand

Get your Citation Baseline — the same 63-prompt framework applied to your category, with a full report in 5 business days.

Get Your Citation Baseline →