Technology · Agentic → Evaluations · Open Source · wiki:deep
TruLens is an open-source stack that finds where an agent fails and where cost can drop without losing quality. It is OpenTelemetry-native: instrument apps with decorators, score steps with self-explaining LLM judges, compare versions, and export traces to any OTLP backend. Evaluations run inline as traces land or in batch over datasets.
Core concepts include feedback/metrics, the RAG triad, and honest/harmless/helpful evals. Provider and framework extras install as separate packages (trulens-providers-*, trulens-apps-langchain, etc.). MIT licensed (TruEra / Snowflake ecosystem).
Eval without step-level traces cannot distinguish a bad retrieval from a bad generator or a bad tool call. TruLens is the research pick when Evaluations must attach scores to OTEL spans and compare app versions on quality × latency × cost — complementary to Promptfoo/DeepEval CI matrices and to process mining (pm4py) on XES.
pip install trulens (+ provider/app extras). @instrument spans (retrieval, generation, tools, MCP, …). RunConfig). Related: abm-agent-behavior-mining · agent-evaluation · deepeval · topics/32-eval-harnesses
Scroll inside the canvas to pan
Catalog backlinks — what points here (wiki “what links here”).
Claims below are backed by science sources on disk.
Product / tracing + feedback
truera/trulens · MODERATE
“The results… show that our proposed metrics are much closer aligned with the human judgements than the predictions from the two baselines.”
Docs / core concepts
TruLens documentation · MODERATE
“For faithfulness and context relevance, the two annotators agreed in around 95% of cases. For answer relevance, they agreed in around 90% of the cases.”
Peer contrast
RAGAS contrast · MODERATE
“Ragas is available at https://github.com/explodinggradients/ragas.”
Features and peers linked from the catalog map — not a second product surface.
No DM vendor crosswalk edges yet.
Primary repo github.com/truera/trulens · Open Source
technologies/trulens/README.md
9 tags · 14 out · 15 in · 3 artifacts · 1 gaps · 30 corpus docs
map_edge · 7
tech_features · 1
tech_quote · 8
tech_readme · 1
tech_science_source · 3
tech_section · 10