DiscoverAI Citation Accuracy Benchmark: Methodology
A frozen protocol for testing whether AI research answers are correct, fully supported, precisely cited, and willing to abstain.

Bottom line
This benchmark uses 12 synthetic source documents, 18 answer-keyed questions, conflicting and outdated evidence, and six deliberately unanswerable prompts. Products receive no score until preserved outputs pass blinded item-level adjudication.
Editorial accountability
Who checked this guide
- Evaluation type
- Research-based verification
- Last materially checked
- Evidence
- 4 listed sources
Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.
Editorial basis
What this guidance is based on
- Editorial basis
- Source-led analysis
- Primary references
- 4
- Products covered
- 4
- Last checked
- 2026-10-02
Important limits
- • This page preregisters a method; it is not a product leaderboard.
- • No product result is published in protocol version 2026.10-v1.
- • Product access, connectors, file limits, and citation interfaces may prevent matched runs and will be disclosed.
In this guide
Short answer
The DiscoverAI Citation Accuracy Benchmark tests a configured product on a frozen, synthetic evidence packet. It separates answer correctness from citation quality and measures whether every material claim is entailed by the cited source, whether all claims that need support receive it, whether the link lands on the precise supporting span, and whether the system abstains when the packet cannot answer the question.
No product has passed merely because its review is eligible. A public result requires preserved inputs and outputs, product and model configuration, two independent reviewers, adjudication, and a reproducibility rerun.
Frozen dataset
Version 2026.10-v1 contains 12 synthetic documents about a fictional service organization. The packet includes dated policies, meeting notes, tables, a superseded memo, two sources that disagree, a statistic with an easily lost denominator, and passages that share vocabulary without supporting the same claim. Because every entity and event is fictional, the public fixtures contain no participant or customer data.
The question set contains 18 tasks: 12 answerable and six where the correct behavior is to say the evidence does not establish an answer. Tasks cover direct lookup, multi-source synthesis, chronology, numeric comparison, conflict disclosure, quotation, and policy application.
Matched run conditions
Each product receives the same files, filenames, question wording, turn order, and time budget. Reviewers record product, plan, model, retrieval mode, connector state, web access, date, region, system instructions, file limits, retries, latency, and output. Web search is disabled for the closed-book track. If a product cannot ingest the packet or expose citations, that capability boundary is reported rather than converted into a zero.
Item-level scoring
Claim correctness asks whether each atomic claim agrees with the answer key. Citation entailment asks whether the cited passage actually supports that exact claim. Citation completeness checks whether all externally verifiable material claims are cited. Span precision checks whether the destination exposes the supporting passage rather than an entire undifferentiated document. Abstention rewards declining the six unanswerable tasks without fabricating support.
Scores publish as numerators and denominators, not one decorative star rating. Severe errors are separate: fabricated quotations, nonexistent sources, material numeric distortion, concealed source conflict, or a confident answer to a safety-relevant unanswerable question.
Human adjudication
Two reviewers independently label atomic claims and citation links without seeing a product leaderboard. Disagreements are adjudicated against the frozen answer key and source span. We publish agreement, exclusions, configuration differences, failures, and rerun variance alongside any product result.
Catalog-wide run status
Every DiscoverAI tool review now receives a citation-benchmark applicability check. Eligible means the reviewed workflow includes source-grounded research, retrieval, or document answers. It does not mean the product was tested or passed. Unrelated products are marked not applicable so absence is never misread as failure. The live coverage ledger is available in the [Research Benchmark Center](/research-insights/benchmarks).
Publication threshold
A result ships only when the dataset version is frozen, output artifacts are retained, configuration is reproducible, both reviewers finish, disagreements are adjudicated, severe errors are reviewed, and a second run confirms that the first was not a lucky sample. Until then the status remains eligible · not yet run.
Sources and verification
Product details and claims were checked against the following primary sources.
Frequently asked questions
What does the citation benchmark measure?
It measures atomic claim correctness, citation entailment, citation completeness, supporting-span precision, appropriate abstention, and severe citation failures.
Does an eligible review have a benchmark score?
No. Eligibility only means the workflow fits the protocol. A score requires a controlled product run and human adjudication.
Why use synthetic documents?
Synthetic fixtures can be public, frozen, permission-safe, and fully answer-keyed without exposing customer or research-participant data.
Why not publish one overall score?
A single average can hide fabricated quotes or unsupported safety-relevant claims. Component results and severe errors remain visible.
Tools mentioned in this article
ChatGPT
The general-purpose AI assistant that started it all
OpenAI's flagship conversational AI model, powering everything from casual chat to complex reasoning, coding, and creative work.
Claude
Anthropic's thoughtful, safety-focused AI with exceptional long-form reasoning
Claude excels at deep analysis, long-form writing, and nuanced reasoning. Built by Anthropic with a focus on safety and helpfulness.
Google Gemini
Google's deeply integrated AI assistant with unmatched access to Google's ecosystem
Gemini combines powerful AI with Google's vast data ecosystem — Search, Gmail, Docs, YouTube, and more — for a uniquely integrated experience.
Perplexity AI
AI-powered search engine with real-time citations and research capabilities
Perplexity combines AI chat with real-time web search, delivering cited, verifiable answers. Think Google Search meets ChatGPT.
Read next
Recommended for you

Audit AI Citations Before the Answer Leaves Your Team
A clickable citation is a route to evidence—not proof that the evidence is current, authorized, complete, or correctly interpreted.
Use a repeatable source audit for AI research, financial analysis, legal work, and client materials. Check identity, time, rights, support, and transformation.
Read guide
Best AI Writing Tools in 2026: 9 Picks by Use Case, Budget, and Workflow
NotebookLM Review 2026: Is Google's Research Assistant Worth Using?
ChatGPT Review 2026: The AI Assistant That Defined a Category, Thoroughly Tested