Giskard generates adversarial and quality scenarios for AI agents, but generated attacks and LLM judges need human review and cannot establish complete security.
Direct verdict
Giskard earns a shortlist for Python teams that want a practical bridge from exploratory AI red teaming to regression tests. Use generated scans for discovery, promote confirmed failures into deterministic checks, control provider data paths, and keep manual security review around the agent's permissions and infrastructure.
What to verify
Create a threat model and 40 human-authored cases covering indirect injection, data exfiltration, excessive agency, harmful output, hallucination, and multi-turn escalation. Run generated scans with fixed seeds, compare judge decisions with two reviewers, replay confirmed failures in CI, and measure coverage, disagreement, reproducibility, latency, and model cost.