RAGAS
What is RAGAS?
Ragas is an open-source LLM evaluation library for AI teams that need repeatable experiments across prompts, RAG systems, workflows, and agents. It combines Experiments-first approach, Ragas Metrics, dataset management, result tracking, and Test Data Generation, and integrates with LangChain, LlamaIndex, Haystack, LangGraph, Amazon Bedrock, Arize, LangSmith, Google Gemini, and OCI Gen AI. Public references include Atomicwork, Pinecone, Weaviate, Qdrant, Deepset Haystack, Mixedbread.ai, LangChain, and OpenAI.
Last verifiedHow we evaluate
At a glance
- Ragas is best for AI teams who need systematic evaluation loops for prompts, RAG systems, and agents.
- Yes — The page links to API documentation and technical references for the Ragas library.
What it does well
Ragas' metric library is unusually comprehensive for an open-source eval tool: RAG-specific metrics (faithfulness, context precision/recall, context entity recall, noise sensitivity, response relevancy, multimodal faithfulness/relevance), agent and tool-use metrics (topic adherence, tool-call accuracy, tool-call F1, agent goal accuracy), NLP comparison metrics (factual correctness, semantic similarity, BLEU/ROUGE/CHRF, exact match), SQL-specific metrics (execution-based Datacompy score, SQL semantic equivalence), and general-purpose LLM-graded metrics (aspect critic, rubric scoring). Few open-source eval libraries cover this much ground under one API. It also ships synthetic test-data generation (knowledge-graph-based multi-hop query synthesis) so teams without a labeled eval set can bootstrap one. Source: https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/
A currently-open bug that breaks fresh installs
As of this check, installing the latest PyPI release (ragas 0.4.3, published 2026-01-13) and running from ragas import evaluate on a modern LangChain stack raises ModuleNotFoundError: No module named 'langchain_community.chat_models.vertexai' — the package unconditionally imports a class that LangChain removed. This was first reported 2026-05-24 (GitHub issue #2741) and affects all users regardless of which LLM provider they use, per a later fix PR's own description. Three separate community pull requests (#2739, #2746, #2749, and a further attempt #2862 opened 2026-07-19) have proposed fixes; none is merged as of 2026-08-26. The documented workaround is to pin langchain-community<0.4 (or downgrade the whole LangChain family to <1.0). One commenter on the thread wrote 'Looking at the activity of this project, it looks like its dead,' and another said they were switching to a different eval tool after hitting this on a fresh install. Sources: https://github.com/vibrantlabsai/ragas/issues/2741 , https://github.com/vibrantlabsai/ragas/pull/2862
Maintenance pace
The GitHub repo (renamed from explodinggradients/ragas to vibrantlabsai/ragas; old URLs redirect) has 15,485 stars and 1,656 forks and shows real historical activity — regular releases through v0.4.3 in January 2026. But the most recent commit to the main branch is from 2026-02-24, and the most recent tagged release is v0.4.3 from 2026-01-13 — over seven months old at the time of this check, despite 577 open issues and multiple outstanding fix PRs for the import-crash bug above. Sources: https://api.github.com/repos/vibrantlabsai/ragas (pushed_at, open_issues_count), https://github.com/vibrantlabsai/ragas/releases
Company behind it
Ragas is maintained by Vibrant Labs, a Y Combinator W24 company (team size 6, San Francisco) founded by Jithin James and Shahul ES. Vibrant Labs' current YC profile describes the company's focus as 'autonomously scaling environments' for post-training AI agents — not evaluation tooling — suggesting Ragas may no longer be the founders' primary product focus, which is consistent with the slowed release cadence and unmerged critical fix above. Source: https://www.ycombinator.com/companies/vibrant-labs
Hosted dashboard is gone, no pricing page exists
The library and its documentation are entirely free and open source (Apache 2.0) — there is no paid tier to evaluate. Ragas' docs and README historically pointed to a hosted dashboard at app.ragas.io for viewing evaluation reports; that subdomain returns DNS_PROBE_FINISHED_NXDOMAIN and has been reported broken since at least February 2026 with no vendor response in the tracking issue. www.ragas.io/pricing returns a 404. There is no commercial product to price — everything runs self-hosted. Sources: https://github.com/vibrantlabsai/ragas/issues/2584 , https://www.ragas.io/pricing (404)
Support channels
Community support runs through GitHub issues (577 open at time of check) and a Discord server; there's a weekly 'office hours' booking link (cal.com/team/vibrantlabs/office-hours) that remains live. No SLA, status page, or security/trust page is published — reasonable for a free OSS library, but worth knowing if you need any of those for a compliance review. Sources: https://docs.ragas.io/en/stable/community/ , https://cal.com/team/vibrantlabs/office-hours
Frequently asked questions
What is RAGAS?
Ragas is an open-source LLM evaluation library for AI teams that need repeatable experiments across prompts, RAG systems, workflows, and agents. It combines Ragas Metrics, dataset management, result tracking, and test data generation, and integrates with LangChain, LlamaIndex, Haystack, and Amazon Bedrock. Public references include Atomicwork, Pinecone, Weaviate, Qdrant, LangChain, and OpenAI.
What is RAGAS used for? Who is it for?
RAGAS is used for Experiments-first approach, Ragas Metrics, and Easy to integrate. It's built for ML engineers, AI product teams, and RAG developers.
Does RAGAS have an API and what does it integrate with?
The page links to API documentation and technical references for the Ragas library.
Editor's read
Check whether your evaluation workflow depends on frameworks beyond the listed integrations. The docs show LangChain, LlamaIndex, Haystack, LangGraph, Arize, LangSmith, Amazon Bedrock, Google Gemini, OCI Gen AI, and Arize Phoenix, so confirm your stack is covered before standardizing on it.
