Skip to main content
Favicon of Inspect AI

Inspect AI

What is Inspect AI?

Inspect AI is an open-source evaluation framework for AI evaluation engineers that benchmarks models and agents through a Python API, CLI, and Inspect View. It combines Inspect View, Tool calling, Structured Output, Batch Mode, and Model Concurrency for repeatable runs, and supports external agents like Claude Code, Codex CLI, and Gemini CLI. Developed by the UK AI Security Institute and Meridian Labs, it includes over 200 pre-built evaluations.

Last verifiedHow we evaluate

Screenshot of Inspect AI website

At a glance

Best for
Inspect AI is best for AI eval engineers who need a flexible framework for benchmarking models and agents.

What it actually is

Inspect is a pip-installable Python library (pip install inspect-ai) plus a CLI (inspect eval, inspect view) for defining evaluations as three composable pieces — a dataset, a solver (how the model produces an answer, from a single generate() call to a full ReAct agent loop), and a scorer (how the answer is graded). It was built and open-sourced by the UK AI Security Institute (then AI Safety Institute) in May 2024, and development is now shared with Meridian Labs, a 501(c)(3) nonprofit the AISI team spun out to steward the project alongside sibling tools (Inspect Scout for transcript analysis, Inspect Petri for alignment auditing).

What it's good at

It's a real, load-bearing piece of infrastructure, not a hobby project. The GitHub repo (UKGovernmentBEIS/inspect_ai) has 2,651 stars, 681 forks, and a companion repo (inspect_evals) with 648 stars holds 200+ pre-built benchmark implementations spanning coding (SWE-bench, HumanEval, LiveCodeBench), cybersecurity (Cybench, CyberGym, GDM Dangerous Capabilities), safety/scheming (AgentHarm, StrongREJECT, WMDP, Agentic Misalignment), science, math, and multimodal tasks — that count is corroborated directly on the vendor's own evals listing page. It has built-in support for 24+ model providers (OpenAI, Anthropic, Google, Mistral, AWS Bedrock/SageMaker, Azure, Groq, Together, local inference via vLLM/Ollama/SGLang, and more), and a genuine third-party ecosystem has grown around it: METR built and open-sourced 'Hawk', a Kubernetes-based platform for running Inspect evals at scale in AWS; the Vector Institute, Arcadia Impact, and independent contributors maintain sandbox backends (Podman, Modal, Daytona, Proxmox) and analysis plugins; and IBM Research cited Inspect AI as one of three eval harnesses (alongside HELM and lm-eval-harness) it built automated translators for in its EveryEvalEver benchmarking-standardization project. Commit activity is current as of this check (last push within the last day) and the CLI ships a VS Code extension with about 11,000 installs.

Maturity and stability

PyPI still classifies the package as 'Development Status :: 4 - Beta', and the project is on a 0.3.x version scheme (current release 0.3.260 as of this check) with no 1.0 yet, despite roughly 113 million cumulative PyPI downloads and multi-year production use by AISI itself. Inspect ships frequent point releases (hundreds of 0.3.x tags), which is normal for an actively developed framework but means buyers building long-lived internal tooling on top of it should pin versions and expect occasional breaking changes between releases rather than treating it as a stable public API.

Security disclosure and scope

The repo publishes a formal SECURITY.md (contact: [email protected]) with a stated 3-business-day acknowledgment window and a 30-day target for fixes. It's explicit about scope: sandbox escapes, secrets leaking into logs, and log-viewer vulnerabilities are in scope; running untrusted model-generated code inside a sandbox, or using the (explicitly unsandboxed) 'local' sandbox mode, is called out as expected behavior, not a vulnerability — worth knowing if your eval design assumes stronger isolation than Docker provides by default. No published security advisories exist on the repo as of this check.

Governance and who's behind it

Ownership is split between a UK government body (UKGovernmentBEIS, i.e. AI Security Institute, under the Department for Science, Innovation and Technology) and Meridian Labs, a US 501(c)(3) nonprofit. This is unusual for a widely-used dev tool: it means the roadmap answers to a state safety institute's priorities (the evals catalogue leans heavily toward dangerous-capability, cyber, and alignment benchmarks) rather than a commercial vendor's, and there is no pricing or support tier to buy — everything is free and community-supported via GitHub issues and a public Slack.

Frequently asked questions

What is Inspect AI?

Inspect AI is an open-source evaluation framework for AI evaluation engineers that benchmarks models and agents through a Python API, CLI, and Inspect View. It combines Inspect View, Tool calling, Structured Output, Batch Mode, and Model Concurrency for repeatable runs, and supports external agents like Claude Code, Codex CLI, and Gemini CLI. Developed by the UK AI Security Institute and Meridian Labs, it includes over 200 pre-built evaluations.

What is Inspect AI used for? Who is it for?

Inspect AI is used for Inspect View, VS Code Extension, and Tool calling. It's built for AI evaluation engineers, Research teams, and Platform engineers.

Does Inspect AI have an API and what does it integrate with?

Inspect AI doesn't publish a public API.

Share:

Sponsored
Favicon

 

  
 

Explore other Agent Tools & Integrations

Favicon

 

  
  
Favicon

 

  
  
Favicon