ReviewUpdated 2026-09-03

Giskard Review 2026: AI Agent Red Teaming, Scans, and Fit

A research-based Giskard review covering capabilities, pricing, privacy, limitations, and a fair buyer test.

By DiscoverAI Editorial TeamReviewed by DiscoverAI Editorial Review2 min readBuild, Design & GovernHow we evaluate
Paper-cut shield protecting an AI conversation from adversarial arrows and unsafe outcomes
Original DiscoverAI editorial illustration. Generated red-team tests are most useful when humans verify failures and preserve them as repeatable regression checks.

Bottom line

Giskard generates adversarial and quality scenarios for AI agents, but generated attacks and LLM judges need human review and cannot establish complete security.

Editorial accountability

Who checked this guide

Meet the editorial team →
Evaluation type
Hands-on evaluation
Last materially checked
Evidence
4 listed sources

Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.

Editorial freshness

Checked this month

Pricing and material product claims were checked September 3, 2026.

Review evidence

What this guidance is based on

Editorial basis
Current first-party product, pricing, documentation, privacy, security, and license material
Review type
Research-based product assessment
Material review date
September 3, 2026
Buyer test
Controlled workflow test with evidence, cost, permission, privacy, and ownership checks

Important limits

  • DiscoverAI did not complete the proposed long-term paid deployment for this research-based review.
  • Features, prices, limits, security controls, privacy terms, licensing, and provider data paths can change; verify the linked first-party pages before purchase.
In this guide
  1. Short answer
  2. Best for
  3. Look elsewhere if
  4. What Giskard verifiably does
  5. Important limitations
  6. Pricing snapshot
  7. A fair buyer test
  8. Final verdict

Short answer

Giskard is worth testing when a team needs repeatable prompt-injection, harmful-content, hallucination, and policy tests for an AI agent. It can generate tailored scenarios and turn failures into regression checks. A clean scan is not an exhaustive security assessment, and its own documentation warns that model judges can be wrong.

Best for

  • AI agent red teaming
  • RAG quality regression
  • Python teams adding AI security checks

Look elsewhere if

  • Treating one scan as certification
  • Sensitive outputs sent to unapproved judges
  • Teams without a threat model

What Giskard verifiably does

Official documentation describes vulnerability scans, quality scans against a knowledge base, custom checks, generated test suites, multi-turn attacks, configurable scenario budgets and seeds, CI replay, JUnit export, multiple model providers, and optional adapters for third-party scanners. Giskard Hub adds managed probes and continuous testing.

Important limitations

The generator and judge send agent descriptions and replies to the configured model provider. Weak generators produce shallow attacks; weak judges create false passes or failures. Scenario caps trade coverage for time and cost. Generated tests do not replace architecture review, authorization tests, conventional application security, or manual red teaming.

Pricing snapshot

Giskard Scan and Checks are available as open-source Python tooling. Teams pay the configured model provider for attack generation, target calls, and grading, plus CI and engineering costs. Giskard Hub provides managed continuous red teaming under sales-led commercial terms. Reviewed September 3, 2026.

A fair buyer test

Create a threat model and 40 human-authored cases covering indirect injection, data exfiltration, excessive agency, harmful output, hallucination, and multi-turn escalation. Run generated scans with fixed seeds, compare judge decisions with two reviewers, replay confirmed failures in CI, and measure coverage, disagreement, reproducibility, latency, and model cost.

Final verdict

Giskard earns a shortlist for Python teams that want a practical bridge from exploratory AI red teaming to regression tests. Use generated scans for discovery, promote confirmed failures into deterministic checks, control provider data paths, and keep manual security review around the agent's permissions and infrastructure.

This is a research-based assessment, not a claim of hands-on product testing. Product, pricing, privacy, security, licensing, and usage claims were checked against the first-party sources below on September 3, 2026. Verify current terms and run the proposed test with approved data before adoption.

Reusable trial worksheet

Test Giskard before you commit

Turn this review’s buyer test into evidence. Your entries autosave only in this browser and are never added to shared shortlist links.

0/7 checks complete
  1. Confirm the tool meets every must-have workflow and stakeholder requirement.

    Review starting point: AI agent red teaming; RAG quality regression; Python teams adding AI security checks

  2. Run the same representative work you would use in production; do not score a polished demo.

    Review starting point: Create a threat model and 40 human-authored cases covering indirect injection, data exfiltration, excessive agency, harmful output, hallucination, and multi-turn escalation. Run generated scans with fixed seeds, compare judge decisions with two reviewers, replay confirmed failures in CI, and measure coverage, disagreement, reproducibility, latency, and model cost.

  3. Calculate the effective cost per accepted result, including usage, review, corrections, and required add-ons.

    Review starting point: Giskard Scan and Checks are available as open-source Python tooling. Teams pay the configured model provider for attack generation, target calls, and grading, plus CI and engineering costs. Giskard Hub provides managed continuous red teaming under sales-led commercial terms. Reviewed September 3, 2026.

  4. Define an acceptance threshold, test known answers and edge cases, and record every correction.

    Review starting point: Editorial quality signals: features 4.1/5; AI quality 4.0/5. Validate these signals in your own work.

  5. Verify what data enters the product, who can access it, how long it is retained, and whether it trains models.

    Review starting point: Use approved low-risk data first. Check roles, consent, deletion, subprocessors, model-training settings, and the contract—not only the marketing page.

  6. Test the real handoffs, permissions, failure states, and export path your team depends on.

    Review starting point: Python, OpenAI, Anthropic, Google, LiteLLM, CI systems

  7. Record training, governance, reliability, accessibility, ownership, and change-management risks before rollout.

    Review starting point: Model calls add cost and data paths; Coverage depends on scenario quality; Managed Hub pricing is sales-led

Open Decision Workspace

Loading saved worksheet… · private to this device or your optional account

Community evidence

How verified users put Giskard to work

Structured, editor-moderated experience—not star ratings. This complements our independent review and never changes its score.

No approved community evidence yet. Be the first verified user to contribute.

Sources and verification

Product details and claims were checked against the following primary sources.

Frequently asked questions

Is Giskard free?

Giskard's open-source scanning and checks tooling is free, while the model calls, CI runtime, engineering work, and managed Hub offering cost separately.

What does Giskard scan for?

Its current documentation describes vulnerability scenarios such as prompt injection and harmful behavior, plus quality scans that look for unsupported answers against a supplied knowledge base.

Can Giskard run in CI?

Yes. Generated suites can be saved, replayed on pull requests, and exported in JUnit XML according to the official documentation.

Does passing a Giskard scan prove an agent is secure?

No. The vendor explicitly says a clean run is not exhaustive and judge decisions require review.

Found this useful?

Get the next one in your inbox.

One five-minute briefing a week: a meaningful change, a practical workflow, and a clearer tool decision—already filtered for lean teams.

Free · one email a week · unsubscribe any time

Recommended tool

Use Giskard if this workflow fits your team

Open-source scan and check workflow

Tools mentioned in this article

Giskard

Open-source scans and managed continuous testing for AI agents

4.0

Giskard generates adversarial and quality scenarios for AI agents, but generated attacks and LLM judges need human review and cannot establish complete security.

FreemiumCodeResearch

Promptfoo

Open-source evaluation and security testing for prompts, models, RAG systems, and agents

4.0

Promptfoo brings repeatable evals, model comparisons, assertions, red teaming, and CI gates to AI development, but test-set quality, judge calibration, sensitive traces, and remediation ownership determine its value.

FreemiumCodeResearch

Ragas

An open-source framework for systematic evaluation of RAG, prompts, workflows, and agents

4.0

Ragas helps teams replace informal AI vibe checks with datasets, experiments, custom metrics, and model-assisted evaluation, but metric validity, judge alignment, token cost, and human labels remain essential.

FreemiumCodeResearch

Arize Phoenix

Open-source tracing and evaluation for LLM, RAG, and agent applications

4.0

Phoenix gives teams OpenTelemetry-based traces, evaluations, experiments, datasets, and prompt tooling in a self-hostable project, but telemetry volume, sensitive content, evaluator validity, and operations remain buyer-owned.

FreeCodeResearch

Read next

More on Build, Design & Govern