DiscoverAI Thematic Analysis Benchmark: Methodology
A human-reviewed protocol for testing whether AI can organize qualitative evidence without flattening disagreement or inventing themes.

Bottom line
This benchmark uses 16 synthetic interview transcripts, a frozen reference codebook, contradictory cases, negation, and minority concerns. It measures evidence-backed themes and attribution—not how polished a summary sounds.
Editorial accountability
Who checked this guide
- Evaluation type
- Research-based verification
- Last materially checked
- Evidence
- 4 listed sources
Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.
Editorial basis
What this guidance is based on
- Editorial basis
- Source-led analysis
- Primary references
- 4
- Products covered
- 4
- Last checked
- 2026-10-02
Important limits
- • This page preregisters a method; it does not report comparative product performance.
- • The synthetic fixture cannot represent every culture, language, research question, or interpretive tradition.
- • Product results will apply only to the declared plan, model, configuration, and dataset version.
In this guide
Short answer
The DiscoverAI Thematic Analysis Benchmark tests whether a configured AI product can code a frozen set of synthetic interviews, recover supported themes, preserve minority and contradictory evidence, attribute quotations correctly, and keep a stable codebook across repeated runs. It does not reward eloquent summaries that cannot be traced to participants.
No product is scored from feature pages or demonstrations. Eligible reviews remain not yet run until their actual outputs are preserved and adjudicated.
Frozen dataset
Version 2026.10-v1 contains 16 synthetic interviews about a fictional team's adoption of an internal knowledge system. The set includes varied roles, tenure, access needs, positive and negative experiences, conditional approval, negation, sarcasm, one contradictory participant, and three low-frequency but decision-important concerns. The reference model contains eight themes, permitted code merges, exact evidence spans, negative cases, and three minority themes.
Required tasks
Products receive the same transcripts and instructions to create an initial codebook, code every transcript, propose themes, attach participant-level evidence, identify contradictory cases, summarize limitations, and export the analysis. A second run adds two held-out transcripts to test whether the codebook evolves transparently instead of silently rewriting history.
Scoring
Evidence-backed theme recall measures recovery of reference themes with supporting spans. Unsupported theme rate penalizes themes that lack sufficient evidence. Quote attribution checks words, participant, and context. Minority-theme retention checks whether rare safety, access, or trust concerns survive clustering. Codebook stability measures whether labels and definitions remain interpretable across the held-out update.
We also report contradictory-case handling, over-merging, duplication, correction minutes, export completeness, reviewer agreement, run-to-run variance, and cost. Severe errors include fabricated quotations, wrong participant attribution, disclosure of hidden identifiers, erasure of a preregistered minority concern, or a conclusion that reverses the evidence.
Human reference and adjudication
The reference codebook is not treated as metaphysical truth. Two human coders independently code the fixtures, discuss disagreements, preserve defensible alternate labels, and freeze the adjudicated map before product outputs are opened. Product themes may use different words and still receive credit when their definition and evidence match.
Catalog-wide run status
Every DiscoverAI tool review now receives a thematic-benchmark applicability check. Reviews covering interviews, surveys, transcripts, feedback, repositories, or qualitative insight are eligible for a future controlled run. Other workflows are marked not applicable. See the complete [Research Benchmark Center ledger](/research-insights/benchmarks).
Publication threshold
Results require a frozen dataset, declared configuration, preserved exports, independent coding, adjudication, severe-error review, and a reproducibility run. We publish component measures and evidence examples with privacy-safe fixture IDs. We do not turn “not applicable,” inaccessible, or not yet tested into a zero.
Sources and verification
Product details and claims were checked against the following primary sources.
Frequently asked questions
Can AI perform thematic analysis?
AI can accelerate coding, retrieval, and clustering, but researchers remain responsible for interpretation, sampling limits, contradictory evidence, privacy, and conclusions.
How are different theme labels compared fairly?
Reviewers compare definitions and supporting evidence, not exact wording. Defensible alternate labels can match the frozen reference concept.
Why measure minority themes separately?
Low-frequency concerns can be operationally or ethically important. An average recall score can hide their erasure.
Does benchmark eligibility mean a tool passed?
No. It means the tool's reviewed workflow is relevant. Only a controlled, adjudicated run can produce a result.
Tools mentioned in this article
ChatGPT
The general-purpose AI assistant that started it all
OpenAI's flagship conversational AI model, powering everything from casual chat to complex reasoning, coding, and creative work.
Claude
Anthropic's thoughtful, safety-focused AI with exceptional long-form reasoning
Claude excels at deep analysis, long-form writing, and nuanced reasoning. Built by Anthropic with a focus on safety and helpfulness.
Google Gemini
Google's deeply integrated AI assistant with unmatched access to Google's ecosystem
Gemini combines powerful AI with Google's vast data ecosystem — Search, Gmail, Docs, YouTube, and more — for a uniquely integrated experience.
Read next
