Promptfoo brings repeatable evals, model comparisons, assertions, red teaming, and CI gates to AI development, but test-set quality, judge calibration, sensitive traces, and remediation ownership determine its value.
Direct verdict
Promptfoo earns a shortlist for engineering and security teams that want evals in source control and CI. Begin with a small human-labeled suite, favor deterministic assertions where possible, and treat red-team results as a triage input rather than a compliance certificate.
What to verify
Take 100 production-like cases and 30 human-labeled failures. Compare deterministic checks, two model judges, and blinded human review. Add prompt injection, data exposure, harmful action, tool abuse, and policy cases. Measure judge agreement, false passes, false blocks, reproducibility, runtime, cost, secret handling, and remediation time.