ReviewUpdated 2026-09-01

Promptfoo Review 2026: LLM Testing, Red Teaming, Pricing, and Fit

A research-based Promptfoo review covering capabilities, pricing, privacy, limitations, and a fair buyer test.

By DiscoverAI Editorial TeamReviewed by DiscoverAI Editorial Review2 min readContent & SearchHow we evaluate
Paper-cut illustration of adversarial prompt fragments meeting an AI shield and a governed continuous-testing pipeline
Original DiscoverAI editorial illustration. LLM security tests become useful when calibrated findings have risk levels, owners, regressions, and remediation deadlines.

Bottom line

Promptfoo brings repeatable evals, model comparisons, assertions, red teaming, and CI gates to AI development, but test-set quality, judge calibration, sensitive traces, and remediation ownership determine its value.

Editorial accountability

Who checked this guide

Meet the editorial team →
Evaluation type
Hands-on evaluation
Last materially checked
Evidence
4 listed sources

Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.

Review evidence

What this guidance is based on

Editorial basis
Current first-party product, pricing, documentation, privacy, security, and license material
Review type
Research-based product assessment
Material review date
September 1, 2026
Buyer test
Controlled workflow test with evidence, correction, cost, permission, privacy, and ownership checks

Important limits

  • DiscoverAI did not complete the proposed long-term paid deployment for this research-based review.
  • Features, prices, limits, security controls, privacy terms, licensing, and provider data paths can change; verify the linked first-party pages before purchase.
In this guide
  1. Short answer
  2. Best for
  3. Look elsewhere if
  4. What Promptfoo verifiably does
  5. Important limitations
  6. Pricing snapshot
  7. A fair buyer test
  8. Final verdict

Short answer

Promptfoo is worth evaluating for teams moving from manual prompt checks to versioned, repeatable quality and security gates. It can compare providers and exercise adversarial cases before deployment. The tool is only as strong as the expected outputs, rubrics, attack policy, judge calibration, and people assigned to fix failures.

Best for

  • Engineering teams adding AI tests to CI
  • Security teams red-teaming LLM applications
  • Builders comparing prompts and providers

Look elsewhere if

  • Teams without labeled acceptance criteria
  • Buyers seeking one universal quality score
  • Sensitive tests with unreviewed cloud sharing

What Promptfoo verifiably does

Official documentation covers declarative test configuration, provider and prompt comparison, deterministic and model-graded assertions, datasets, caching, RAG and agent evaluation, red-team plugins and strategies, vulnerability reports, CI integration, local execution, sharing, and enterprise deployment or governance.

Important limitations

Model graders can be biased, inconsistent, correlated with the system under test, or vulnerable to manipulated content. Synthetic attacks do not represent every user or threat. Test data and traces may contain secrets or personal information. Broad suites consume model cost and can create noisy gates unless failures map to owners and risk levels.

Pricing snapshot

Promptfoo's open-source CLI can run without a software subscription, while hosted collaboration and enterprise security offerings use current plan limits or sales-led pricing that should be confirmed directly. Evaluation model calls, target systems, CI compute, storage, and red-team generation create separate costs. Reviewed September 1, 2026.

A fair buyer test

Take 100 production-like cases and 30 human-labeled failures. Compare deterministic checks, two model judges, and blinded human review. Add prompt injection, data exposure, harmful action, tool abuse, and policy cases. Measure judge agreement, false passes, false blocks, reproducibility, runtime, cost, secret handling, and remediation time.

Final verdict

Promptfoo earns a shortlist for engineering and security teams that want evals in source control and CI. Begin with a small human-labeled suite, favor deterministic assertions where possible, and treat red-team results as a triage input rather than a compliance certificate.

This is a research-based assessment, not a claim of hands-on product testing. Product, pricing, privacy, security, licensing, and usage claims were checked against the first-party sources below on September 1, 2026. Verify current terms and run the proposed test with approved data before adoption.

Sources and verification

Product details and claims were checked against the following primary sources.

Frequently asked questions

Is Promptfoo free?

Yes. The core CLI is open source; model calls, infrastructure, hosted collaboration, and enterprise services can add cost.

Can Promptfoo test for prompt injection?

Yes. Its red-team framework includes prompt-injection and other vulnerability strategies, which should be supplemented with application-specific attacks.

Does Promptfoo work in CI?

Yes. Tests can run in CI and gate changes using configured assertions and thresholds.

Can Promptfoo prove an AI app is secure?

No tool can provide that guarantee. Its results cover the configured targets, attacks, models, and test period and require human triage and remediation.

Continue exploring

A useful next step

View topic →
Paper-cut illustration of varied AI outputs passing through calibrated metric lenses into a comparison notebook
ReviewBuild, Design & Govern

Ragas Review 2026: RAG and Agent Evaluation, Cost, and Fit

A research-based Ragas review covering capabilities, pricing, privacy, limitations, and a fair buyer test.

Ragas helps teams replace informal AI vibe checks with datasets, experiments, custom metrics, and model-assisted evaluation, but metric validity, judge alignment, token cost, and human labels remain essential.

Read guide

Paper-cut illustration of AI outputs crossing test lanes with calibrated gauges
ReviewContent & Search

Braintrust AI Review 2026: Evals, Observability, Pricing, and Fit

A research-based Braintrust review covering capabilities, pricing, privacy, limitations, and a fair buyer test.

Braintrust connects production traces, datasets, experiments, scorers, prompts, and human review, but judge validity, sensitive logs, retention, score volume, and release-gate design require calibration.

Read guide

Paper-cut illustration of conversation fragments becoming a time-aware knowledge graph
ReviewBuild, Design & Govern

Zep Review 2026: Agent Memory, Pricing, Security, and Fit

A research-based Zep review covering capabilities, pricing, privacy, limitations, and a fair buyer test.

Zep turns conversations and business events into time-aware agent memory, but extraction quality, stale facts, deletion, credit usage, and the deployment trust boundary need controlled evaluation.

Read guide

Paper-cut illustration of model streams entering a protected routing hub
ReviewWork & Operations

Portkey AI Review 2026: Gateway, Pricing, Guardrails, and Fit

A research-based Portkey review covering capabilities, pricing, privacy, limitations, and a fair buyer test.

Portkey centralizes model access, fallbacks, caching, guardrails, keys, budgets, and traces, but a gateway becomes a critical data and availability boundary that needs failure testing.

Read guide

The five-minute weekly AI briefing

One useful change, workflow, and decision—already filtered.

Stay current without tracking every launch. Built for lean teams weighing budget, privacy, and implementation effort.

Recommended tool

Use Promptfoo if this workflow fits your team

Open-source local workflow

Tools mentioned in this article

Promptfoo

Open-source evaluation and security testing for prompts, models, RAG systems, and agents

4.0

Promptfoo brings repeatable evals, model comparisons, assertions, red teaming, and CI gates to AI development, but test-set quality, judge calibration, sensitive traces, and remediation ownership determine its value.

FreemiumCodeResearch

Langfuse

Open-source tracing, evaluation, prompt management, and metrics for LLM applications

4.0

Langfuse unifies traces, costs, prompts, datasets, and evaluation with cloud and self-hosted options, but telemetry sensitivity, retention, operational load, and fast-rising plan costs demand a scoped pilot.

FreemiumCodeAnalytics

Helicone

An open-source AI gateway with request monitoring, cost tracking, caching, fallbacks, prompts, and evaluations

4.0

Helicone combines multi-provider routing and observability behind a familiar API, but proxy trust, logged payloads, retention, usage-based costs, and gateway dependency need careful architecture review.

FreemiumCodeAnalytics

Ragas

An open-source framework for systematic evaluation of RAG, prompts, workflows, and agents

4.0

Ragas helps teams replace informal AI vibe checks with datasets, experiments, custom metrics, and model-assisted evaluation, but metric validity, judge alignment, token cost, and human labels remain essential.

FreemiumCodeResearch