DiscoverAI AI Support Resolution Benchmark: Methodology
A preregistered protocol for future results—published before any customer-service AI receives a score.

Bottom line
This is DiscoverAI's public protocol for comparing AI support systems with identical cases, knowledge, permissions, and review rules. It measures durable accepted resolution, not containment or vendor-reported automation, and publishes no product ranking until collection and adjudication are complete.
Editorial accountability
Who checked this guide
- Evaluation type
- Research-based verification
- Last materially checked
- Evidence
- 4 listed sources
Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.
Editorial basis
What this guidance is based on
- Editorial basis
- Source-led analysis
- Primary references
- 4
- Products covered
- 4
- Last checked
- 2026-09-26
Important limits
- • No product scores or independent benchmark observations have been collected for publication yet.
- • Results from one case mix, configuration, language, region, or review window may not transfer to another support operation.
- • Products, models, pricing, reporting definitions, and capabilities can change after the verification date.
In this guide
Short answer
This is DiscoverAI's public protocol for comparing AI support systems with identical cases, knowledge, permissions, and review rules. It measures durable accepted resolution, not containment or vendor-reported automation, and publishes no product ranking until collection and adjudication are complete.
Free workflow pilot checklist
Test the workflow before you buy the tool.
Get the buyer checklist, including task, owner, approval, fallback, and time-saved fields—plus one useful briefing a week.
Research question
For a fixed support operation, which configured system correctly and durably resolves eligible customer needs, escalates the rest with useful context, follows policy, minimizes customer and reviewer effort, and does so at a reproducible total cost? The benchmark evaluates a product, model, knowledge base, integrations, settings, and operating process together—not a timeless model or a polished demonstration.
Frozen case set
Build a versioned set from de-identified historical cases and authored edge cases. Stratify by intent, channel, language, customer segment, complexity, risk, required action, and expected disposition. Include ambiguous requests, missing identity, outdated knowledge, conflicting policy, multiple intents, emotional customers, prompt injection, duplicate contacts, revoked credentials, rate limits, and downstream outages. Keep a hidden holdout set to detect tuning to the public examples.
Each case receives a reference packet: permitted knowledge, applicable policy, required evidence, allowed actions, acceptance criteria, correct escalation destination, severity, and reopen window. Remove personal data and secrets unless a controlled test explicitly requires synthetic equivalents.
Matched conditions
Give every system the same eligible cases, knowledge snapshot, integration fixtures, customer history, permissions, time budget, and escalation staff. Record product, plan, model, version, region, channel, configuration, prompts, tools, retrieval sources, identity state, approvals, timestamps, usage, and failures. Differences that cannot be normalized remain visible limitations rather than disappearing into one score.
Run three phases. Offline replay tests answers and proposed actions without customer impact. Shadow mode observes real work while humans remain authoritative. A small limited-live cohort follows only after the system clears the preregistered offline and shadow gates; consequential actions stay approval-gated.
Primary metric: durable accepted resolution
A case counts as a durable accepted resolution only when the answer or action is correct, supported by the allowed evidence, compliant with policy, authorized, complete, and not reopened for the same need inside the declared window. Abandonment, silence after an answer, duplicate closure, unsupported claims, hidden transfers, partial completion, and incorrectly closed cases do not qualify.
Publish the numerator, denominator, eligibility exclusions, confirmed versus inferred outcomes, and reopen window. Never substitute containment, answer rate, deflection, conversations handled, or a vendor's billable unit for independent acceptance.
Secondary measures
Report factual and evidential correctness, action correctness, policy compliance, privacy and authorization, appropriate abstention, escalation timing, destination and context quality, customer effort, latency, correction time, reviewer minutes, reopen rate, abandonment, total operating cost, and cost per durable accepted resolution.
Severe events remain separate from averages: privacy exposure, unauthorized action, material financial or account change, unsafe advice, discrimination, evasion of escalation, corrupted records, or unrecoverable customer harm. One severe event cannot be averaged away by many easy successes.
Human review and adjudication
Two trained reviewers independently score a stratified sample and every severe or ambiguous case using the frozen rubric. Blind product identity where the transcript and action record permit it. Reviewers cite the source, policy, or system state supporting each judgment. Disagreements receive documented adjudication; agreement and unresolved cases are published.
Customer feedback is useful but not sufficient: a positive rating cannot make an incorrect or unauthorized action acceptable, and silence cannot prove resolution. Automated graders may assist triage only after validation against human decisions and must not be the sole judge of their own class of system.
Cost and statistical reporting
Total cost includes platform fees, billable outcomes or conversations, model and tool usage, implementation amortization, integrations, monitoring, human review, corrections, support, and incident response. Report volume, eligible-case rate, uncertainty intervals, medians where appropriate, and results by intent, channel, risk, and language. Do not publish a composite score unless every weight is declared before collection and all component results remain available.
Release gate and reproducibility
No leaderboard or winner ships until every included system completes the declared sample or is clearly labeled incomplete; every accepted resolution has inspectable evidence; severe events and reviewer disagreements are adjudicated; costs reconcile to logs or invoices; and the limitations review is complete.
The results package should expose a versioned rubric, case metadata without customer content, configuration sheet, observation schema, exclusions, aggregate calculations, conflicts, corrections, and change log. Raw transcripts or private operational data are released only when consent, security, and licensing permit. Until that gate is met, this page is methodology—not performance evidence.
How teams can use the protocol now
Teams can adapt the rubric for procurement without waiting for DiscoverAI results. Freeze representative cases, define acceptance and severe failures, establish a human baseline, give finalists matched conditions, and progress from offline replay to shadow mode to a limited live cohort. Choose from durable accepted outcomes, customer impact, reviewer burden, recovery, and full cost—not the vendor with the highest self-reported automation rate.
Sources and verification
Product details and claims were checked against the following primary sources.
Frequently asked questions
Does this benchmark already rank AI support agents?
No. This page preregisters the method and withholds product scores until matched collection, review, cost reconciliation, and the release gate are complete.
What is durable accepted resolution?
It is a correct, supported, policy-compliant, authorized, complete answer or action that is not reopened for the same need during a declared window.
Why not use containment or deflection rate?
Those metrics can count silence, abandonment, incorrect closure, or hidden transfers. The protocol requires independent evidence that the customer's need was actually resolved.
Can vendors sponsor inclusion or improve their score?
No. Commercial relationships remain outside eligibility, case selection, scoring, adjudication, and release decisions; conflicts are disclosed.
Recommended tool
Use Sierra if this workflow fits your team
Sierra is a serious shortlist for large customer-experience teams that want an AI agent to resolve—not merely summarize—customer requests across several channels. Its distinctive commercial idea is outcome-based pricing. That alignment is valuable only when the contract defines a successful outcome, exclusions, reversals, quality thresholds, and disputed attribution with unusual precision.
Tools mentioned in this article
Sierra
Sierra is a serious shortlist for large customer-experience teams that want an AI agent to resolve—not merely summarize—customer requests across several channels
Sierra is a serious shortlist for large customer-experience teams that want an AI agent to resolve—not merely summarize—customer requests across several channels. Its distinctive commercial idea is outcome-based pricing. That alignment is valuable only when the contract defines a successful outcome, exclusions, reversals, quality thresholds, and disputed attribution with unusual precision.
Cresta
Cresta is worth evaluating for large contact centers that want AI agents, real-time human-agent guidance, and conversation intelligence on one enterprise platform
Cresta is worth evaluating for large contact centers that want AI agents, real-time human-agent guidance, and conversation intelligence on one enterprise platform. The opportunity is a shared learning loop across automated and human conversations. The risk is optimizing a vendor score or containment rate while customer outcomes, consent, fairness, or escalation quality deteriorate.
Pylon
Pylon is a strong shortlist for B2B software companies that support customers across Slack, Teams, email, chat, phone, and shared operational systems
Pylon is a strong shortlist for B2B software companies that support customers across Slack, Teams, email, chat, phone, and shared operational systems. Its distinctive opportunity is turning those fragmented conversations into account context that humans and agents can use. The tradeoff is a broad data and action surface, sales-led core pricing, and a still-evolving Agentic Support credit model.
Intercom Fin
A practical AI tool for customer support workflows
Intercom Fin helps professionals improve customer support workflows with AI-assisted drafting, automation, analysis, or production features.
Read next
