RAG Signal
Methodology

Citation Engineering: MAP → BUILD → WEIGHT → REINFORCE → MEASURE

A structured, repeatable 5-step methodology — not guessing, not content spraying. Proven across 7 industries with 77.1% measured attribution rate. Published as open source →

Citation Engineering is RAG Signal's proprietary 5-step methodology. Unlike standard SEO — which optimizes for rankings — Citation Engineering optimizes for retrieval: the signals that influence whether AI models cite your brand. Every step is measurable. Every outcome is auditable. Read the full whitepaper for the algorithm behind it.

MAP prompt map BUILD Brand Memory WEIGHT signal weights REINFORCE across models MEASURE Citation Delta same prompt set Change the set and the comparison with the previous measurement is gone.
The process is a loop, not a list. The final phase re-runs the exact prompt set fixed in the first — that repetition is what makes the measurement comparable.
Operated by the RAG Signal team

Prompt Discovery Algorithm

The platform makes the evidence and decision flow our team manages in every client engagement visible.

  1. 01 Collect Evidence Evidence first
  2. 02 Discover Candidates Find opportunities
  3. 03 Measure Visibility Establish the baseline
  4. 04 Prioritize Focus the work
  5. 05 Expand Candidates Test the adjacent questions
  6. 06 Build Roadmap Set the delivery sequence
01

MAP — Discover Your Prompt Landscape

Find every real buyer question where your brand should appear.

We don't guess what prompts matter. We mine Google Search Console, competitor citation data, LLM query patterns, and category-specific buyer journeys to build a comprehensive prompt map. For Filmfolk, we discovered 63 distinct prompts — from "best corporate video production London" to "video agency for enterprise internal comms."

Tracked in Google 8 branded keywords Asked by buyers 63 prompts 4 branded 59 non-branded
One dot, one prompt. The 8 keywords the brand was tracking and the 63 questions buyers were asking are not the same set — the decision gets made inside the 59 nobody was watching.

Before MAP: Filmfolk was tracking 8 branded keywords in Google. After MAP: 63 prompt map — only 4 were branded. The other 59? Where buyers actually make decisions.

02

BUILD — Construct Brand Memory

Build structured knowledge that AI models can retrieve and cite.

AI models retrieve from their training data and context windows. We build structured Brand Memory: entity definitions, relationship maps, factual assertions, and citation anchors that models can retrieve. This is the structured knowledge infrastructure that helps make your brand a consistent retrieval target.

Filmfolk Organization factual assertions 35+ enterprise clients 97% client retention headquartered in London founded 2016 corporate video · internal comms · events outside corroboration 12 industry publication citations 3 award references 40+ case study pages
Brand Memory is not an abstraction but a countable structure: one entity, the assertions attached to it, and the outside sources that corroborate them. An assertion sitting on your own site is not enough to be retrieved — something else has to point at it.

Entity: Filmfolk (Organization)
Assertions: "serves 35+ enterprise clients," "97% client retention rate," "headquartered in London, UK," "founded 2016," "specializes in corporate video, internal comms, and event coverage"
Signals: 12 citations from industry publications, 3 award references, 40+ case study pages

03

WEIGHT — Score and Prioritize Signals

Not all signals are equal. Our RAG Scoring Algorithm weights them by retrieval impact.

The Signal Scoring Engine evaluates each signal across multiple dimensions: source authority, factual consistency, cross-model persistence, temporal freshness (with 180-day hard cutoff), entity linkage density, citation frequency, and competitive differentiation. Each dimension contributes to a composite score — published in full in our open-source whitepaper.

Source Authority 25% Factual Consistency 20% Entity Linkage 15% Cross-Model Persistence 12% Temporal Freshness 12% Citation Frequency 10% Competitive Diff 6% total 100
The seven signals do not carry equal weight. The first three account for 60% of the score on their own, which is also the order the work follows.

40+ platform modules track signal weight in real time. A claim on your own site scores lower than the same claim cited by an industry publication — unless your site has high EEAT signals. The algorithm handles this automatically.

04

REINFORCE — Deploy Across Models

Push weighted Brand Memory into the retrieval paths of each target model.

Each LLM has different retrieval mechanics. Based on observable citation patterns, ChatGPT tends to favor recency and authority domains. Perplexity weights real-time web signals. Claude emphasizes document structure and factual consistency. We deploy model-specific reinforcement: llms.txt for Claude, structured data for Perplexity's web index, entity-rich content for ChatGPT's knowledge base, and cross-model signal amplification for Gemini.

Model Observed tendency What gets deployed ChatGPT recency · authority domains entity-rich content Perplexity real-time web signals structured data Claude document structure · factual consistency llms.txt Gemini consistency across sources cross-model amplification Tendencies are inferred from observed citation patterns; no provider publishes them.
Putting the same thing in four places the same way does not work, because the four are not looking at the same thing. The form of the signal changes per model — which is also why the measurement is taken per platform.

Filmfolk, Day 45: ChatGPT citation rate: 42% → 71%. The reinforcement wasn't "more content." It was structured entity deployment + llms.txt optimization + 14 high-authority citation anchors placed in model training pipelines.

05

MEASURE — Track Citation Delta

Citation rate isn't a vanity metric. It's the only metric that matters in AI.

At 30, 60, and 90 days, we re-run every prompt across all 4 models and measure: Is your brand cited? In what position? With what framing? Against which competitors? The Citation Delta Report shows exactly what moved — and what didn't. Retainer clients get continuous monitoring with anomaly detection.

0% 25% 50% 75% 100% baselineday 30day 60day 90 Perplexity — baseline: 0% Perplexity — day 30: 38% Perplexity — day 60: 67% Perplexity — day 90: 86% Perplexity 86% ChatGPT — baseline: 0% ChatGPT — day 30: 42% ChatGPT — day 60: 71% ChatGPT — day 90: 84% ChatGPT 84% Claude — baseline: 0% Claude — day 30: 31% Claude — day 60: 59% Claude — day 90: 78% Claude 78% Gemini — baseline: 0% Gemini — day 30: 28% Gemini — day 60: 52% Gemini — day 90: 76% Gemini 76%
No platform reached its end point in the first 30 days; the movement was gradual and the four did not finish in the same place. The figures are in the table below.

Filmfolk Citation Delta Report, Day 90:

ChatGPT
0% → 84%
+84pp
Perplexity
0% → 86%
+86pp
Claude
0% → 78%
+78pp
Gemini
0% → 76%
+76pp
Why Not DIY?

"Can't I just use ChatGPT for this?"

It's the most common question we hear. Here's why the answer is no.

Can't Self-Measure

AI models hallucinate when asked about themselves. You cannot audit your own citation rate by prompting.

Prompting ≠ Engineering

Brand Memory requires structured data, entity definitions, and knowledge graphs — built with code, not chat.

Model Updates Can Reset Citations

A single model update can shift citation patterns significantly overnight. Without monitoring, you won't know for weeks.

Future Dataset Prep

We position your Brand Memory for the next training run. You can't optimize for a dataset that doesn't exist yet.

Download the Methodology PDF

The 7-dimension Signal Scoring Algorithm with scoring criteria, weightings, and interpretation guidelines — for your team or your audit checklist.

Ready to Engineer Your Brand's AI Signals?

Start with a baseline audit. We'll map your current citation rate across every major model — no commitment. Or read the full methodology whitepaper first.

Get Your Citation Baseline →