Promptfoo Review 2026: LLM Testing, Red Teaming, Pricing, and Fit
A research-based Promptfoo review covering capabilities, pricing, privacy, limitations, and a fair buyer test.

Bottom line
Promptfoo brings repeatable evals, model comparisons, assertions, red teaming, and CI gates to AI development, but test-set quality, judge calibration, sensitive traces, and remediation ownership determine its value.
Editorial accountability
Who checked this guide
- Evaluation type
- Hands-on evaluation
- Last materially checked
- Evidence
- 4 listed sources
Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.
Review evidence
What this guidance is based on
- Editorial basis
- Current first-party product, pricing, documentation, privacy, security, and license material
- Review type
- Research-based product assessment
- Material review date
- September 1, 2026
- Buyer test
- Controlled workflow test with evidence, correction, cost, permission, privacy, and ownership checks
Important limits
- • DiscoverAI did not complete the proposed long-term paid deployment for this research-based review.
- • Features, prices, limits, security controls, privacy terms, licensing, and provider data paths can change; verify the linked first-party pages before purchase.
In this guide
Short answer
Promptfoo is worth evaluating for teams moving from manual prompt checks to versioned, repeatable quality and security gates. It can compare providers and exercise adversarial cases before deployment. The tool is only as strong as the expected outputs, rubrics, attack policy, judge calibration, and people assigned to fix failures.
Best for
- Engineering teams adding AI tests to CI
- Security teams red-teaming LLM applications
- Builders comparing prompts and providers
Look elsewhere if
- Teams without labeled acceptance criteria
- Buyers seeking one universal quality score
- Sensitive tests with unreviewed cloud sharing
What Promptfoo verifiably does
Official documentation covers declarative test configuration, provider and prompt comparison, deterministic and model-graded assertions, datasets, caching, RAG and agent evaluation, red-team plugins and strategies, vulnerability reports, CI integration, local execution, sharing, and enterprise deployment or governance.
Important limitations
Model graders can be biased, inconsistent, correlated with the system under test, or vulnerable to manipulated content. Synthetic attacks do not represent every user or threat. Test data and traces may contain secrets or personal information. Broad suites consume model cost and can create noisy gates unless failures map to owners and risk levels.
Pricing snapshot
Promptfoo's open-source CLI can run without a software subscription, while hosted collaboration and enterprise security offerings use current plan limits or sales-led pricing that should be confirmed directly. Evaluation model calls, target systems, CI compute, storage, and red-team generation create separate costs. Reviewed September 1, 2026.
A fair buyer test
Take 100 production-like cases and 30 human-labeled failures. Compare deterministic checks, two model judges, and blinded human review. Add prompt injection, data exposure, harmful action, tool abuse, and policy cases. Measure judge agreement, false passes, false blocks, reproducibility, runtime, cost, secret handling, and remediation time.
Final verdict
Promptfoo earns a shortlist for engineering and security teams that want evals in source control and CI. Begin with a small human-labeled suite, favor deterministic assertions where possible, and treat red-team results as a triage input rather than a compliance certificate.
This is a research-based assessment, not a claim of hands-on product testing. Product, pricing, privacy, security, licensing, and usage claims were checked against the first-party sources below on September 1, 2026. Verify current terms and run the proposed test with approved data before adoption.
Sources and verification
Product details and claims were checked against the following primary sources.
Frequently asked questions
Is Promptfoo free?
Yes. The core CLI is open source; model calls, infrastructure, hosted collaboration, and enterprise services can add cost.
Can Promptfoo test for prompt injection?
Yes. Its red-team framework includes prompt-injection and other vulnerability strategies, which should be supplemented with application-specific attacks.
Does Promptfoo work in CI?
Yes. Tests can run in CI and gate changes using configured assertions and thresholds.
Can Promptfoo prove an AI app is secure?
No tool can provide that guarantee. Its results cover the configured targets, attacks, models, and test period and require human triage and remediation.
Continue exploring
A useful next step

Ragas Review 2026: RAG and Agent Evaluation, Cost, and Fit
A research-based Ragas review covering capabilities, pricing, privacy, limitations, and a fair buyer test.
Ragas helps teams replace informal AI vibe checks with datasets, experiments, custom metrics, and model-assisted evaluation, but metric validity, judge alignment, token cost, and human labels remain essential.
Read guide

Braintrust AI Review 2026: Evals, Observability, Pricing, and Fit
A research-based Braintrust review covering capabilities, pricing, privacy, limitations, and a fair buyer test.
Braintrust connects production traces, datasets, experiments, scorers, prompts, and human review, but judge validity, sensitive logs, retention, score volume, and release-gate design require calibration.
Read guide

Zep Review 2026: Agent Memory, Pricing, Security, and Fit
A research-based Zep review covering capabilities, pricing, privacy, limitations, and a fair buyer test.
Zep turns conversations and business events into time-aware agent memory, but extraction quality, stale facts, deletion, credit usage, and the deployment trust boundary need controlled evaluation.
Read guide

Portkey AI Review 2026: Gateway, Pricing, Guardrails, and Fit
A research-based Portkey review covering capabilities, pricing, privacy, limitations, and a fair buyer test.
Portkey centralizes model access, fallbacks, caching, guardrails, keys, budgets, and traces, but a gateway becomes a critical data and availability boundary that needs failure testing.
Read guide
The five-minute weekly AI briefing
One useful change, workflow, and decision—already filtered.
Stay current without tracking every launch. Built for lean teams weighing budget, privacy, and implementation effort.
Recommended tool
Use Promptfoo if this workflow fits your team
Open-source local workflow
Tools mentioned in this article
Promptfoo
Open-source evaluation and security testing for prompts, models, RAG systems, and agents
Promptfoo brings repeatable evals, model comparisons, assertions, red teaming, and CI gates to AI development, but test-set quality, judge calibration, sensitive traces, and remediation ownership determine its value.
Langfuse
Open-source tracing, evaluation, prompt management, and metrics for LLM applications
Langfuse unifies traces, costs, prompts, datasets, and evaluation with cloud and self-hosted options, but telemetry sensitivity, retention, operational load, and fast-rising plan costs demand a scoped pilot.
Helicone
An open-source AI gateway with request monitoring, cost tracking, caching, fallbacks, prompts, and evaluations
Helicone combines multi-provider routing and observability behind a familiar API, but proxy trust, logged payloads, retention, usage-based costs, and gateway dependency need careful architecture review.
Ragas
An open-source framework for systematic evaluation of RAG, prompts, workflows, and agents
Ragas helps teams replace informal AI vibe checks with datasets, experiments, custom metrics, and model-assisted evaluation, but metric validity, judge alignment, token cost, and human labels remain essential.