Skip to main content
Favicon of UpTrain

UpTrain

What is UpTrain?

UpTrain is an LLM evaluation and improvement platform for teams shipping AI apps that scores outputs, compares prompt or model changes, and catches regressions before production. It combines Diverse evaluations, Faster and Systematic Experimentation, Automated Regression Testing, and Root Cause Analysis, with Google sign-in and an open-source core that supports self-hosting. The platform says it has evaluated more than 1,000,000 responses and handles 100, 10k, or million rows without failures.

Last verifiedHow we evaluate

Screenshot of UpTrain website

At a glance

Best for
UpTrain is best for teams shipping LLM apps who need reliable evaluation, testing, and monitoring before production.

What it actually is

UpTrain ships two things under one brand: an open-source Python package (pip install uptrain, latest release 0.7.1 from May 2024) that runs LLM-graded and classical NLP checks — response completeness, factual accuracy, context relevance, hallucination, jailbreak/prompt-injection detection, and more — locally against your own LLM API key, and a self-hostable web dashboard (run via Docker, bash run_uptrain.sh) for browsing results and root-cause analysis. The marketing site talks about an "enterprise-grade" managed/hosted version, but there is no pricing page (uptrain.ai/pricing 404s) and no plan tiers anywhere — the only calls to action are "Book Demo" and a Google-login dashboard.

Where it's genuinely good

The breadth of pre-built evals is real and documented in the repo's README with working code samples — response quality, context/retrieval quality, language quality, conversation-level checks, custom-guideline evals, and safety checks (jailbreak, prompt injection, system-prompt leak) all ship as ready-made classes you can call in a few lines (EvalLLM(...).evaluate(...)). The library architecture is genuinely local-first: evaluation logic runs on your machine and the only network calls are to whichever LLM provider you configure as the judge, which is a clean privacy story for the library itself (not the hosted dashboard). It supports OpenAI, Anthropic, Mistral, Azure and Anyscale-hosted open models as the judge LLM, and integrates with LlamaIndex, Langfuse, Helicone, Qdrant, FAISS and Chroma.

The dashboard has four unpatched RCE/auth CVEs

GitHub Security Lab published CVE-2025-27621 (hardcoded default API key combined with an open CORS policy, letting any website make authenticated cross-origin requests) and CVE-2025-27770/27771/27772 (remote code execution via unsanitized eval() calls on the checks and metadata parameters at the /create_project, /add_prompts and /new_run endpoints) on 2026-08-17–19. All four affect every released version up to and including 0.7.1 — i.e. the current latest release — and as of this check none has a merged fix: three pull requests from an outside contributor (#754, #755, #756, opened 2026-08-10–11) remain open and unmerged against the uptrain-ai/uptrain main branch. GitHub's disclosure notes state the RCE was first reported privately on 2024-09-05 with no maintainer response, and that Slack messages from 2025-03-10 suggested the project was no longer maintained.

The company behind it has moved on

UpTrain was built by Sourabh Agrawal and Shikha Mohanty through Y Combinator's Winter 2023 batch, first launched as an ML observability tool (Feb 2023) and relaunched as an LLMOps platform in early 2024. That same YC company profile (id 27966) is today listed under the name CombineHealth with the one-liner "Automating Healthcare Revenue Cycle Management with AI Workforce" — a healthcare RCM startup, still marked "Active," with UpTrain's GitHub repo listed as a legacy artifact on the page rather than the company's current product. The GitHub repository's last code push was 2024-08-18; the last PyPI release (0.7.1) shipped 2024-05-14. Recent repository activity is limited to third-party dependency-license fixes and unmerged security patches from outside contributors, not the original team.

Docs and support channels

Documentation at docs.uptrain.ai (Mintlify-hosted) is still live and reasonably complete. The company blog (blog.uptrain.ai, which resolves to a uptrainai.wordpress.com site) now returns "This site is currently private" for every post, including the one linked from the current homepage banner. The hosted "Try Evals Playground" demo (demo.uptrain.ai) did not load in two attempts on 2026-08-27. The Slack community invite link on the homepage still resolves.

Frequently asked questions

What is UpTrain?

UpTrain is an LLM evaluation and improvement platform for teams shipping AI apps that scores outputs, compares prompt or model changes, and catches regressions before production. It combines Diverse evaluations, Faster and Systematic Experimentation, Automated Regression Testing, and Root Cause Analysis, with Google sign-in and an open-source core that supports self-hosting. The platform says it has evaluated more than 1,000,000 responses and handles 100, 10k, or million rows without failures.

What is UpTrain used for? Who is it for?

UpTrain is used for Diverse evaluations, Faster and Systematic Experimentation, and Automated Regression Testing. It's built for LLM developers, Product managers, and AI platform teams.

Does UpTrain have an API and what does it integrate with?

UpTrain doesn't publish a public API. It integrates with Google.

Share:

Sponsored
Favicon

 

  
 

Explore other Agent Tools & Integrations

Favicon

 

  
  
Favicon

 

  
  
Favicon