GuideUpdated 2026-09-22

Build an AI Evaluation Set From Real Work

A useful test set reflects the work that actually arrives—including ambiguity and failure—not a handful of polished demo prompts.

By DiscoverAI Editorial TeamReviewed by DiscoverAI Editorial Review2 min readContent & SearchHow we evaluate
Paper-cut editorial illustration of representative work samples sorted into ordinary, ambiguous, edge, failure, and refusal trays before a blind human-scored model comparison
Original DiscoverAI editorial illustration. Editorial illustration: representative work samples sorted into ordinary, ambiguous, edge, failure, and refusal trays before a blind human-scored model comparison.

Bottom line

Build an AI evaluation set by sampling real, permission-safe work across common cases, important segments, edge cases, known failures, and should-refuse requests. Remove unnecessary personal or confidential data, define expected qualities before running the model, double-label a subset, freeze a holdout, and version every test with the prompt, model, tools, and data configuration.

Editorial accountability

Who checked this guide

Meet the editorial team →
Evaluation type
Research-based verification
Last materially checked
Evidence
4 listed sources

Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.

Editorial basis

What this guidance is based on

Editorial basis
Source-led analysis
Primary references
4
Products covered
3
Last checked
2026-09-22

Important limits

  • Company telemetry and model claims may not generalize to other populations or workloads.
  • Availability, policy, pricing, and product behavior can change.
In this guide
  1. Short answer
  2. Sample the workload, not the demo
  3. Write rubrics before seeing outputs
  4. Separate development from decision evidence
  5. What readers should do

Short answer

Build an AI evaluation set by sampling real, permission-safe work across common cases, important segments, edge cases, known failures, and should-refuse requests. Remove unnecessary personal or confidential data, define expected qualities before running the model, double-label a subset, freeze a holdout, and version every test with the prompt, model, tools, and data configuration.

Free content AI buyer checklist

Choose tools that improve accepted work—not output volume.

Get a checklist for accuracy, edit time, sourcing, brand fit, and cost per accepted asset—plus one briefing a week.

Free · about 5 minutes · one email a week · unsubscribe any time

Free · one email a week · unsubscribe any timePreview the checklist →

Sample the workload, not the demo

Start with 50 to 200 cases drawn from the real distribution. Stratify by language, complexity, source quality, customer type, risk, and frequency. Add rare severe failures deliberately so an average score cannot hide them.

Write rubrics before seeing outputs

Define required facts, allowed sources, format, tone, abstention, tool permissions, and critical-failure rules. Use binary gates for non-negotiable constraints and scaled judgments only where nuance is genuinely needed.

Separate development from decision evidence

Use one set while improving prompts and a frozen holdout for the release decision. Track disagreements, reviewer time, cost, latency, and failures by segment. Promote confirmed production incidents into regression cases without continuously tuning to the same small sample.

What readers should do

Begin with 60 cases: 30 ordinary, ten ambiguous, ten edge cases, five known failures, and five should-refuse cases. Have two reviewers label 15 independently, resolve disagreements, run the current workflow as a baseline, and set release thresholds before testing a replacement.

Claims were checked against the linked sources on September 22, 2026. Vendor measurements and company announcements are attributed evidence, not independent guarantees.

Sources and verification

Product details and claims were checked against the following primary sources.

Frequently asked questions

How many cases should an AI evaluation set contain?

Start with 50 to 200 representative cases; consequence and workload diversity matter more than a universal minimum.

Can production data be used in an AI evaluation?

Only with appropriate permission and controls; remove unnecessary personal, confidential, and regulated data or use faithful synthetic cases.

What belongs in an AI test set?

Common tasks, important segments, ambiguity, edge cases, known failures, adversarial inputs, and requests the system should refuse.

Why keep a holdout set?

It reduces the risk of tuning prompts or workflows to the same examples used to claim improvement.

Free content AI buyer checklist

Choose tools that improve accepted work—not output volume.

Get a checklist for accuracy, edit time, sourcing, brand fit, and cost per accepted asset—plus one briefing a week.

Free · one email a week · unsubscribe any timePreview the checklist →

Recommended tool

Use ChatGPT if this workflow fits your team

It has one of the clearest workflow fits in its category and is easier to recommend than tools that only look impressive in demos.

Tools mentioned in this article

ChatGPT

The general-purpose AI assistant that started it all

4.6

OpenAI's flagship conversational AI model, powering everything from casual chat to complex reasoning, coding, and creative work.

FreemiumChatbotsWriting

Claude

Anthropic's thoughtful, safety-focused AI with exceptional long-form reasoning

4.5

Claude excels at deep analysis, long-form writing, and nuanced reasoning. Built by Anthropic with a focus on safety and helpfulness.

FreemiumChatbotsWriting

Langfuse

Open-source tracing, evaluation, prompt management, and metrics for LLM applications

4.0

Langfuse unifies traces, costs, prompts, datasets, and evaluation with cloud and self-hosted options, but telemetry sensitivity, retention, operational load, and fast-rising plan costs demand a scoped pilot.

FreemiumCodeAnalytics

Read next

More on Content & Search