Technology · Agentic → Evaluations · Open Core · wiki:deep
DeepEval is an open-source LLM evaluation framework (Confident AI) shaped like Pytest for LLM apps: metrics such as G-Eval, task completion, answer relevancy, and hallucination, using LLM-as-a-judge and NLP models that can run locally. It evaluates end-to-end LLM apps, full agent trajectories, and individual steps (LLM calls, tools, retrieval, handoffs).
Related DeepTeam covers red-teaming. Confident AI is the commercial layer for iteration compare, shared reports, and production monitoring — open-core pattern.
Weekly prompt/tool changes need unit/integration tests that score quality, not screenshots. DeepEval is the research default when you want metrics as code in CI for agents and RAG. Pair with Promptfoo when the bottleneck is declarative provider matrices or security probes; pair with TruLens when you need OTEL-native step traces plus judges.
Scores must attach to cases; the eval store is not the company ledger.
Related: abm-agent-behavior-mining · agent-evaluation · hal-the-holistic-agent-leaderboard · topics/32-eval-harnesses
Scroll inside the canvas to pan
Catalog backlinks — what points here (wiki “what links here”).
Claims below are backed by science sources on disk.
Framework / pytest-style eval
confident-ai/deepeval · MODERATE
“DeepEval is a simple-to-use, open-source LLM evaluation framework, for evaluating large-language model systems.”
Local / quickstart path
DeepEval 5-min Quickstart · MODERATE
“DeepEval incorporates the latest research to run evals via metrics such as G-Eval, task completion, answer relevancy, hallucination, etc., which uses LLM-as-a-judge and other NLP models that run locally on your machine.”
Local metric runs
Local vs cloud · MODERATE
“It is similar to Pytest but specialized for unit testing LLM apps.”
Features and peers linked from the catalog map — not a second product surface.
No DM vendor crosswalk edges yet.
Primary repo github.com/confident-ai/deepeval · Open Core
technologies/deepeval/README.md
9 tags · 16 out · 17 in · 3 artifacts · 1 gaps · 32 corpus docs
map_edge · 9
tech_features · 1
tech_quote · 8
tech_readme · 1
tech_science_source · 3
tech_section · 10