ReviewUpdated 2026-09-01

Zep Review 2026: Agent Memory, Pricing, Security, and Fit

A research-based Zep review covering capabilities, pricing, privacy, limitations, and a fair buyer test.

By DiscoverAI Editorial TeamReviewed by DiscoverAI Editorial Review2 min readBuild, Design & GovernHow we evaluate
Paper-cut illustration of conversation fragments becoming a time-aware knowledge graph
Original DiscoverAI editorial illustration. Agent memory earns trust when changing facts, corrections, access boundaries, and deletion are tested together.

Bottom line

Zep turns conversations and business events into time-aware agent memory, but extraction quality, stale facts, deletion, credit usage, and the deployment trust boundary need controlled evaluation.

Editorial accountability

Who checked this guide

Meet the editorial team →
Evaluation type
Hands-on evaluation
Last materially checked
Evidence
4 listed sources

Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.

Review evidence

What this guidance is based on

Editorial basis
Current first-party product, pricing, documentation, privacy, security, and license material
Review type
Research-based product assessment
Material review date
September 1, 2026
Buyer test
Controlled workflow test with evidence, correction, cost, permission, privacy, and ownership checks

Important limits

  • DiscoverAI did not complete the proposed long-term paid deployment for this research-based review.
  • Features, prices, limits, security controls, privacy terms, licensing, and provider data paths can change; verify the linked first-party pages before purchase.
In this guide
  1. Short answer
  2. Best for
  3. Look elsewhere if
  4. What Zep verifiably does
  5. Important limitations
  6. Pricing snapshot
  7. A fair buyer test
  8. Final verdict

Short answer

Zep is worth testing when an agent must remember changing people, organizations, preferences, and events across sessions. Its temporal graph is more purposeful than stuffing every old message into a vector store. The decision hinges on whether extracted relationships stay accurate, corrections supersede old facts, and deletion reaches every derived memory.

Best for

  • Agents needing cross-session memory
  • Teams modeling changing relationships
  • Enterprises comparing cloud and private deployment

Look elsewhere if

  • High-stakes memory without correction workflows
  • Teams unable to govern behavioral data
  • Simple short-history chat

What Zep verifiably does

Official materials describe Graphiti-powered temporal context graphs, episodic ingestion, entity and relationship extraction, hybrid retrieval, custom entity types, Memory MCP, observations, webhooks, analytics, cloud, customer-key encryption, and bring-your-own-cloud deployment.

Important limitations

Extraction can merge identities, invent relationships, or preserve obsolete facts. Larger episodes consume more credits. Memory contains sensitive behavioral history, so retention, deletion, access separation, model subprocessors, and backup behavior require review.

Pricing snapshot

Zep includes 10,000 monthly prototype credits. Flex is $125 per month with 50,000 credits, and Flex Plus is $375 with 200,000. Extra credits, rate limits, projects, logs, Memory MCP seats, and features vary. Episode size determines ingestion credits; Enterprise is negotiated. Reviewed September 1, 2026.

A fair buyer test

Create 100 synthetic users with changing employers, preferences, aliases, contradictions, deletions, and access boundaries. Measure extraction precision, temporal ordering, corrections, forbidden cross-user retrieval, deletion, ingestion latency, retrieval usefulness, credits per accepted memory, and correction effort.

Final verdict

Zep earns a shortlist for teams that need durable, time-aware agent context and can validate memory as a governed data system. Start with synthetic identities, define correction and deletion service levels, and forecast credits from real episode sizes.

This is a research-based assessment, not a claim of hands-on product testing. Product, pricing, privacy, security, licensing, and usage claims were checked against the first-party sources below on September 1, 2026. Verify current terms and run the proposed test with approved data before adoption.

Sources and verification

Product details and claims were checked against the following primary sources.

Frequently asked questions

Is Zep free?

Zep offers a limited prototype plan with 10,000 monthly credits; production self-serve and enterprise plans are paid.

How does Zep pricing work?

Credits are based mainly on Episode size. Test representative payloads because larger Episodes consume multiple credits.

How is Zep different from a vector database?

Zep builds a temporal graph of entities, relationships, and episodes, then combines graph and semantic retrieval.

Can Zep run privately?

Zep advertises bring-your-own-cloud enterprise deployment and customer-managed encryption-key options; verify architecture and contract terms.

Continue exploring

A useful next step

View topic →
Paper-cut illustration of AI outputs crossing test lanes with calibrated gauges
ReviewContent & Search

Braintrust AI Review 2026: Evals, Observability, Pricing, and Fit

A research-based Braintrust review covering capabilities, pricing, privacy, limitations, and a fair buyer test.

Braintrust connects production traces, datasets, experiments, scorers, prompts, and human review, but judge validity, sensitive logs, retention, score volume, and release-gate design require calibration.

Read guide

Paper-cut illustration of an agent trace with a failure under inspection
ReviewContent & Search

Arize Phoenix Review 2026: LLM Tracing, Evals, Pricing, and Fit

A research-based Arize Phoenix review covering capabilities, pricing, privacy, limitations, and a fair buyer test.

Phoenix gives teams OpenTelemetry-based traces, evaluations, experiments, datasets, and prompt tooling in a self-hostable project, but telemetry volume, sensitive content, evaluator validity, and operations remain buyer-owned.

Read guide

Paper-cut illustration of adversarial prompt fragments meeting an AI shield and a governed continuous-testing pipeline
ReviewContent & Search

Promptfoo Review 2026: LLM Testing, Red Teaming, Pricing, and Fit

A research-based Promptfoo review covering capabilities, pricing, privacy, limitations, and a fair buyer test.

Promptfoo brings repeatable evals, model comparisons, assertions, red teaming, and CI gates to AI development, but test-set quality, judge calibration, sensitive traces, and remediation ownership determine its value.

Read guide

Paper-cut illustration of varied AI outputs passing through calibrated metric lenses into a comparison notebook
ReviewBuild, Design & Govern

Ragas Review 2026: RAG and Agent Evaluation, Cost, and Fit

A research-based Ragas review covering capabilities, pricing, privacy, limitations, and a fair buyer test.

Ragas helps teams replace informal AI vibe checks with datasets, experiments, custom metrics, and model-assisted evaluation, but metric validity, judge alignment, token cost, and human labels remain essential.

Read guide

The five-minute weekly AI briefing

One useful change, workflow, and decision—already filtered.

Stay current without tracking every launch. Built for lean teams weighing budget, privacy, and implementation effort.

Recommended tool

Use Zep if this workflow fits your team

Temporal memory model

Tools mentioned in this article

Zep

Temporal knowledge-graph memory infrastructure for production AI agents

4.0

Zep turns conversations and business events into time-aware agent memory, but extraction quality, stale facts, deletion, credit usage, and the deployment trust boundary need controlled evaluation.

FreemiumCodeResearch

Mem0

Memory infrastructure that helps AI agents retain and retrieve user context across sessions

4.0

Mem0 gives developers managed and open-source memory layers for AI agents, but retrieval quality, deletion, sensitive-data handling, training terms, and add-versus-retrieve economics need production testing.

FreemiumCodeAutomation

Langfuse

Open-source tracing, evaluation, prompt management, and metrics for LLM applications

4.0

Langfuse unifies traces, costs, prompts, datasets, and evaluation with cloud and self-hosted options, but telemetry sensitivity, retention, operational load, and fast-rising plan costs demand a scoped pilot.

FreemiumCodeAnalytics

Ragas

An open-source framework for systematic evaluation of RAG, prompts, workflows, and agents

4.0

Ragas helps teams replace informal AI vibe checks with datasets, experiments, custom metrics, and model-assisted evaluation, but metric validity, judge alignment, token cost, and human labels remain essential.

FreemiumCodeResearch