GuideUpdated 2026-10-02

DiscoverAI Citation Accuracy Benchmark: Methodology

A frozen protocol for testing whether AI research answers are correct, fully supported, precisely cited, and willing to abstain.

By DiscoverAI Editorial TeamReviewed by DiscoverAI Editorial Review3 min readContent & SearchHow we evaluate
Paper-cut source documents linked to exact highlighted evidence spans and verified claim cards on a calibrated citation benchmark
Original DiscoverAI editorial illustration. Editorial illustration: every answer claim must connect to a precise supporting source span.

Bottom line

This benchmark uses 12 synthetic source documents, 18 answer-keyed questions, conflicting and outdated evidence, and six deliberately unanswerable prompts. Products receive no score until preserved outputs pass blinded item-level adjudication.

Editorial accountability

Who checked this guide

Meet the editorial team →
Evaluation type
Research-based verification
Last materially checked
Evidence
4 listed sources

Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.

Editorial basis

What this guidance is based on

Editorial basis
Source-led analysis
Primary references
4
Products covered
4
Last checked
2026-10-02

Important limits

  • • This page preregisters a method; it is not a product leaderboard.
  • • No product result is published in protocol version 2026.10-v1.
  • • Product access, connectors, file limits, and citation interfaces may prevent matched runs and will be disclosed.
In this guide
  1. Short answer
  2. Frozen dataset
  3. Matched run conditions
  4. Item-level scoring
  5. Human adjudication
  6. Catalog-wide run status
  7. Publication threshold

Short answer

The DiscoverAI Citation Accuracy Benchmark tests a configured product on a frozen, synthetic evidence packet. It separates answer correctness from citation quality and measures whether every material claim is entailed by the cited source, whether all claims that need support receive it, whether the link lands on the precise supporting span, and whether the system abstains when the packet cannot answer the question.

No product has passed merely because its review is eligible. A public result requires preserved inputs and outputs, product and model configuration, two independent reviewers, adjudication, and a reproducibility rerun.

Frozen dataset

Version 2026.10-v1 contains 12 synthetic documents about a fictional service organization. The packet includes dated policies, meeting notes, tables, a superseded memo, two sources that disagree, a statistic with an easily lost denominator, and passages that share vocabulary without supporting the same claim. Because every entity and event is fictional, the public fixtures contain no participant or customer data.

The question set contains 18 tasks: 12 answerable and six where the correct behavior is to say the evidence does not establish an answer. Tasks cover direct lookup, multi-source synthesis, chronology, numeric comparison, conflict disclosure, quotation, and policy application.

Matched run conditions

Each product receives the same files, filenames, question wording, turn order, and time budget. Reviewers record product, plan, model, retrieval mode, connector state, web access, date, region, system instructions, file limits, retries, latency, and output. Web search is disabled for the closed-book track. If a product cannot ingest the packet or expose citations, that capability boundary is reported rather than converted into a zero.

Item-level scoring

Claim correctness asks whether each atomic claim agrees with the answer key. Citation entailment asks whether the cited passage actually supports that exact claim. Citation completeness checks whether all externally verifiable material claims are cited. Span precision checks whether the destination exposes the supporting passage rather than an entire undifferentiated document. Abstention rewards declining the six unanswerable tasks without fabricating support.

Scores publish as numerators and denominators, not one decorative star rating. Severe errors are separate: fabricated quotations, nonexistent sources, material numeric distortion, concealed source conflict, or a confident answer to a safety-relevant unanswerable question.

Human adjudication

Two reviewers independently label atomic claims and citation links without seeing a product leaderboard. Disagreements are adjudicated against the frozen answer key and source span. We publish agreement, exclusions, configuration differences, failures, and rerun variance alongside any product result.

Catalog-wide run status

Every DiscoverAI tool review now receives a citation-benchmark applicability check. Eligible means the reviewed workflow includes source-grounded research, retrieval, or document answers. It does not mean the product was tested or passed. Unrelated products are marked not applicable so absence is never misread as failure. The live coverage ledger is available in the [Research Benchmark Center](/research-insights/benchmarks).

Publication threshold

A result ships only when the dataset version is frozen, output artifacts are retained, configuration is reproducible, both reviewers finish, disagreements are adjudicated, severe errors are reviewed, and a second run confirms that the first was not a lucky sample. Until then the status remains eligible · not yet run.

Sources and verification

Product details and claims were checked against the following primary sources.

Frequently asked questions

What does the citation benchmark measure?

It measures atomic claim correctness, citation entailment, citation completeness, supporting-span precision, appropriate abstention, and severe citation failures.

Does an eligible review have a benchmark score?

No. Eligibility only means the workflow fits the protocol. A score requires a controlled product run and human adjudication.

Why use synthetic documents?

Synthetic fixtures can be public, frozen, permission-safe, and fully answer-keyed without exposing customer or research-participant data.

Why not publish one overall score?

A single average can hide fabricated quotes or unsupported safety-relevant claims. Component results and severe errors remain visible.

Free content AI buyer checklist

Choose tools that improve accepted work—not output volume.

Get a checklist for accuracy, edit time, sourcing, brand fit, and cost per accepted asset—plus one briefing a week.

Free · one email a week · unsubscribe any timePreview the checklist →

Tools mentioned in this article

ChatGPT

The general-purpose AI assistant that started it all

4.6

OpenAI's flagship conversational AI model, powering everything from casual chat to complex reasoning, coding, and creative work.

FreemiumChatbotsWriting

Claude

Anthropic's thoughtful, safety-focused AI with exceptional long-form reasoning

4.5

Claude excels at deep analysis, long-form writing, and nuanced reasoning. Built by Anthropic with a focus on safety and helpfulness.

FreemiumChatbotsWriting

Google Gemini

Google's deeply integrated AI assistant with unmatched access to Google's ecosystem

4.2

Gemini combines powerful AI with Google's vast data ecosystem — Search, Gmail, Docs, YouTube, and more — for a uniquely integrated experience.

FreemiumChatbotsProductivity

Perplexity AI

AI-powered search engine with real-time citations and research capabilities

4.4

Perplexity combines AI chat with real-time web search, delivering cited, verifiable answers. Think Google Search meets ChatGPT.

FreemiumChatbotsData Analysis

Read next

More on Content & Search →