TruLens
What is TruLens?
TruLens is an open-source AI evaluation tool for teams that measures agent quality by tracing execution flow and scoring retrieved context, tool calls, plans, and outputs. Its Evaluate, Iterate, and Test workflow helps teams compare versions and catch regressions with groundedness, context relevance, coherence, and answer relevance. It works through the Python SDK or OpenTelemetry traces and is used by teams at Equinix, Snowflake, and KBC Group.
Last verifiedHow we evaluate
At a glance
- TruLens is best for AI teams who need trace-level evaluation of agent behavior across versions.
- Yes — TruLens can be used via the Python SDK or by ingesting OpenTelemetry traces.
What it is and who maintains it
TruLens instruments an AI agent or RAG pipeline with a Python decorator, records every span (latency, inputs, outputs, tokens, cost) as OpenTelemetry data, and scores steps with LLM-as-judge metrics. It was created by TruEra, which open-sourced it in 2021; Snowflake acquired TruEra in May 2024 and now stewards the project (https://www.snowflake.com/en/blog/snowflake-acquires-truera-to-bring-llm-ml-observability-to-data-cloud/). It remains MIT-licensed and free — there is no separate paid tier of TruLens itself.
Tracing architecture: OpenTelemetry-native
Every instrumented call becomes a structured OTEL span, so traces are portable to any OTLP-compatible backend (Jaeger, Grafana Tempo, Datadog) rather than being locked in a proprietary format. This is a genuine differentiator among LLM-eval libraries, most of which use a custom trace schema. TruLens 2.9.0 (2026-07-23) added dual-emission of standard gen_ai.* OTel GenAI semantic-convention attributes, and 2.10.0 (2026-07-28) corrected span-kind values to match that spec, improving interoperability with OTel-native tooling (https://github.com/truera/trulens/releases/tag/trulens-2.10.0).
Judge quality: one independent benchmark, one vendor-authored one
AIMultiple, an independent research firm, benchmarked five context-relevance evaluators (TruLens, WandB Weave, Ragas, DeepEval, UpTrain) across 1,010 adversarial QA samples in March 2026. TruLens led on ranking quality — NDCG@5 of 0.932, Spearman ρ of 0.750, and the highest discrimination ratio (4.2:1) between a correct passage and a near-identical entity-swapped fake — while WandB Weave had the highest raw Top-1 accuracy due to its coarser binary scoring (https://aimultiple.com/rag-evaluation-tools). Separately, TruLens's own site cites a 95%-error-coverage figure on the TRAIL/GAIA agent benchmark from 'Agent GPA', a framework described in an arXiv paper authored by the TruEra/Snowflake research team, including TruEra co-founder Anupam Datta (https://arxiv.org/abs/2510.08847) — that number is the vendor's own research, not third-party corroboration, though it is publicly posted and citable.
Adoption evidence and how it's sourced
TruLens's GitHub repo maintains an ADOPTERS.md file that separates 'published references' (companies with a public blog post or press quote, e.g. Walmart Global Tech's RAG-triad case study, Thomson Reuters, phData, the SOA Research Institute) from 'community-reported usage' sourced from named engineers' own GitHub issues (Cisco, J.P. Morgan Chase, VMware by Broadcom, Hitachi Digital Services, and others) — and lists three logos it removed for lack of evidence, including Snowflake itself, explicitly because Snowflake maintains the project and isn't an independent adopter (https://github.com/truera/trulens/blob/main/ADOPTERS.md). That level of self-imposed sourcing discipline on a vendor's own logo wall is unusual and worth taking at face value; most of the individual claims (Equinix's '2 weeks to 2 hours' iteration-time reduction, KBC Group's early-adopter quote) trace to a March 2024 press release rather than an audited study, so they remain attributed vendor/customer claims, not verified metrics.
Release pace and maintenance
The repo is actively maintained: five tagged releases between 2026-07-23 and 2026-08-20 (2.9.0 through 2.13.1), roughly weekly, each shipping features from a rotating set of contributors — 2.9.0 and 2.10.0 each list 8-14 first-time contributors credited with shipping the release's headline features (https://github.com/truera/trulens/releases). GitHub shows 3,526 stars, 333 forks, and 59 open issues as of 2026-08-28, with the last push on 2026-08-27 (https://api.github.com/repos/truera/trulens). Current PyPI release is 2.13.1, published 2026-08-20 (https://pypi.org/pypi/trulens/json).
Cost and what you're actually installing
TruLens itself is free and MIT-licensed; pip install trulens costs nothing (https://pypi.org/project/trulens/). What you pay for is infrastructure you already run (a place to store traces) plus the LLM API calls the judges make when scoring your traces — TruLens 2.11.0 (2026-08-04) added a configure_online_eval() sampling/throttle/cost-budget mechanism specifically because always-on judging can be expensive on high-traffic apps (https://github.com/truera/trulens/releases/tag/trulens-2.11.0). Snowflake also offers a Cortex AI Observability feature inside its own platform; whether that is a separately-priced managed version of TruLens or a distinct product, and what it costs, is not established from public sources checked in this run.
Where it doesn't fit
TruLens is a library you instrument yourself and either self-host the dashboard for or query programmatically — there is no vendor-hosted TruLens SaaS with a signup page, so teams wanting a zero-ops managed eval platform will need to look elsewhere or go through Snowflake's own product. Its PyPI package still carries a 'Development Status :: 3 - Alpha' trove classifier despite five years of history and version 2.13.1 (https://pypi.org/pypi/trulens/json) — worth noting as a calibration point even though the release cadence and adopter evidence suggest the classifier is stale rather than a live warning.
Frequently asked questions
What is TruLens?
TruLens is an open-source AI evaluation tool for teams that measures agent quality by tracing execution flow and scoring retrieved context, tool calls, plans, and outputs. Its Evaluate, Iterate, and Test workflow helps teams compare versions and catch regressions with groundedness, context relevance, coherence, and answer relevance. It works through the Python SDK or OpenTelemetry traces and is used by teams at Equinix, Snowflake, and KBC Group.
What is TruLens used for? Who is it for?
TruLens is used for Evaluate, Iterate, and Test. It's built for AI engineers, ML teams, and Platform teams.
Does TruLens have an API and what does it integrate with?
TruLens can be used via the Python SDK or by ingesting OpenTelemetry traces. It integrates with OpenTelemetry.
