GuideUpdated 2026-10-02

DiscoverAI Thematic Analysis Benchmark: Methodology

A human-reviewed protocol for testing whether AI can organize qualitative evidence without flattening disagreement or inventing themes.

By DiscoverAI Editorial TeamReviewed by DiscoverAI Editorial Review2 min readHow we evaluate
Paper-cut interview evidence flowing through a reviewed codebook into theme clusters while a minority finding remains visible
Original DiscoverAI editorial illustration. Editorial illustration: themes stay connected to participant evidence, including rare contradictory findings.

Bottom line

This benchmark uses 16 synthetic interview transcripts, a frozen reference codebook, contradictory cases, negation, and minority concerns. It measures evidence-backed themes and attribution—not how polished a summary sounds.

Editorial accountability

Who checked this guide

Meet the editorial team →
Evaluation type
Research-based verification
Last materially checked
Evidence
4 listed sources

Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.

Editorial basis

What this guidance is based on

Editorial basis
Source-led analysis
Primary references
4
Products covered
4
Last checked
2026-10-02

Important limits

  • • This page preregisters a method; it does not report comparative product performance.
  • • The synthetic fixture cannot represent every culture, language, research question, or interpretive tradition.
  • • Product results will apply only to the declared plan, model, configuration, and dataset version.
In this guide
  1. Short answer
  2. Frozen dataset
  3. Required tasks
  4. Scoring
  5. Human reference and adjudication
  6. Catalog-wide run status
  7. Publication threshold

Short answer

The DiscoverAI Thematic Analysis Benchmark tests whether a configured AI product can code a frozen set of synthetic interviews, recover supported themes, preserve minority and contradictory evidence, attribute quotations correctly, and keep a stable codebook across repeated runs. It does not reward eloquent summaries that cannot be traced to participants.

No product is scored from feature pages or demonstrations. Eligible reviews remain not yet run until their actual outputs are preserved and adjudicated.

Frozen dataset

Version 2026.10-v1 contains 16 synthetic interviews about a fictional team's adoption of an internal knowledge system. The set includes varied roles, tenure, access needs, positive and negative experiences, conditional approval, negation, sarcasm, one contradictory participant, and three low-frequency but decision-important concerns. The reference model contains eight themes, permitted code merges, exact evidence spans, negative cases, and three minority themes.

Required tasks

Products receive the same transcripts and instructions to create an initial codebook, code every transcript, propose themes, attach participant-level evidence, identify contradictory cases, summarize limitations, and export the analysis. A second run adds two held-out transcripts to test whether the codebook evolves transparently instead of silently rewriting history.

Scoring

Evidence-backed theme recall measures recovery of reference themes with supporting spans. Unsupported theme rate penalizes themes that lack sufficient evidence. Quote attribution checks words, participant, and context. Minority-theme retention checks whether rare safety, access, or trust concerns survive clustering. Codebook stability measures whether labels and definitions remain interpretable across the held-out update.

We also report contradictory-case handling, over-merging, duplication, correction minutes, export completeness, reviewer agreement, run-to-run variance, and cost. Severe errors include fabricated quotations, wrong participant attribution, disclosure of hidden identifiers, erasure of a preregistered minority concern, or a conclusion that reverses the evidence.

Human reference and adjudication

The reference codebook is not treated as metaphysical truth. Two human coders independently code the fixtures, discuss disagreements, preserve defensible alternate labels, and freeze the adjudicated map before product outputs are opened. Product themes may use different words and still receive credit when their definition and evidence match.

Catalog-wide run status

Every DiscoverAI tool review now receives a thematic-benchmark applicability check. Reviews covering interviews, surveys, transcripts, feedback, repositories, or qualitative insight are eligible for a future controlled run. Other workflows are marked not applicable. See the complete [Research Benchmark Center ledger](/research-insights/benchmarks).

Publication threshold

Results require a frozen dataset, declared configuration, preserved exports, independent coding, adjudication, severe-error review, and a reproducibility run. We publish component measures and evidence examples with privacy-safe fixture IDs. We do not turn “not applicable,” inaccessible, or not yet tested into a zero.

Sources and verification

Product details and claims were checked against the following primary sources.

Frequently asked questions

Can AI perform thematic analysis?

AI can accelerate coding, retrieval, and clustering, but researchers remain responsible for interpretation, sampling limits, contradictory evidence, privacy, and conclusions.

How are different theme labels compared fairly?

Reviewers compare definitions and supporting evidence, not exact wording. Defensible alternate labels can match the frozen reference concept.

Why measure minority themes separately?

Low-frequency concerns can be operationally or ethically important. An average recall score can hide their erasure.

Does benchmark eligibility mean a tool passed?

No. It means the tool's reviewed workflow is relevant. Only a controlled, adjudicated run can produce a result.

Free AI tool buyer checklist

Make the next AI subscription earn its place.

Get the printable buyer checklist now, plus one useful five-minute AI briefing each week.

Free · one email a week · unsubscribe any timePreview the checklist →

Tools mentioned in this article

ChatGPT

The general-purpose AI assistant that started it all

4.6

OpenAI's flagship conversational AI model, powering everything from casual chat to complex reasoning, coding, and creative work.

FreemiumChatbotsWriting

Claude

Anthropic's thoughtful, safety-focused AI with exceptional long-form reasoning

4.5

Claude excels at deep analysis, long-form writing, and nuanced reasoning. Built by Anthropic with a focus on safety and helpfulness.

FreemiumChatbotsWriting

Google Gemini

Google's deeply integrated AI assistant with unmatched access to Google's ecosystem

4.2

Gemini combines powerful AI with Google's vast data ecosystem — Search, Gmail, Docs, YouTube, and more — for a uniquely integrated experience.

FreemiumChatbotsProductivity

Read next