Skip to main content

Security AI Agents: What They Monitor, and What They Cannot See

AI security agents split into four jobs: runtime gateways, red-teaming, SOC detection, governance. Each has a published blind spot. Pick by job, not by category.

AgentsIndex's profile

Written by AgentsIndex

Editorial team••5 min read

The category splits into four different jobs, not one

"AI security agent" gets used for products that do genuinely different things, and the differences matter more than any feature-comparison table suggests. Before picking one, know which job you actually have:

  • Runtime gateways sit inline on traffic between your app and the model, inspecting prompts, responses, and tool calls as they pass through — Lakera Guard, NeuralTrust's TrustGate, Lasso Security.
  • Pre-deployment red-teaming throws a library of known attacks at an agent before it ships, to find what breaks — Promptfoo, Straiker's Ascend AI, Lakera's red-teaming product.
  • SOC-style detection and response reads telemetry from tools you already run (SIEM, EDR, identity) and reasons about what's already happened — Simbian, Straiker Defend AI.
  • Governance and discovery platforms build an inventory of every agent, model, and application running across the company, then check it against policy — Holistic AI, Credo AI, Vijil, NeuralTrust's TrustLens.

A buyer who wants "AI agent security" and picks a red-teaming tool when they needed a runtime gateway — or a governance dashboard when they needed a SOC tool — will get a product that works exactly as advertised and still leaves them exposed, because it was never built to watch the thing they're worried about.

What each type actually sees, and where the edge is

Runtime gateways only see traffic that passes through them. Wire the gateway into one integration path and an agent that calls tools through a different route — a direct shell command, an MCP connection the gateway isn't proxying — is invisible to it by construction. Vendor-reported accuracy numbers for this category are all self-tested, on each vendor's own traffic and attack set, not an independent benchmark: Lakera cites a 0.01% false-positive rate on production traffic; Straiker's homepage cites 98.1% threat detection accuracy; Lasso cites 98.6%. None of these numbers are directly comparable to each other — they measure different things against different test sets — and none of them is zero. A gateway inspecting everything that passes through it still misses some fraction of what passes through it, by the vendors' own figures.

Pre-deployment red-teaming tests against a catalogued library of attacks before an agent ships. It's the right tool for catching what's already known to work — prompt injection patterns, jailbreak templates, tool-chain manipulation — but by definition it can't catch an attack technique invented after the test ran, and it doesn't continue watching once the agent is in production and its configuration drifts.

SOC-style detection and response is the category most honest about its own ceiling, and that honesty is worth taking at face value. Simbian's own published research benchmarked 27 frontier models — Claude Opus, GPT-5.x, Grok, DeepSeek, Kimi, and others — on MITRE ATT&CK coverage and found the best of them covers 45.1% of tactics, clearing the 50% bar on only 7 of 13. Simbian's pitch is built directly on this gap: "point Claude or GPT at your SIEM, it will not defend you, and a better model will not change that." Whether or not Simbian's own product closes that gap, the underlying finding — that no current model defends comprehensively, because none was trained to — applies to every product in this category that relies on an LLM doing the reasoning, including the ones selling against it.

Governance and discovery platforms are only as complete as their integrations. Holistic AI connects to AWS, Azure, GitHub, and Databricks (plus "20+ integrations") to build its agent inventory; an agent running somewhere not connected — a laptop, an unlisted SaaS tool, a contractor's own account — doesn't appear in the inventory. That's the shadow-AI problem these platforms are sold to solve, and it's also the boundary of what they can solve: visibility into the connected estate, not omniscience over the whole one. Holistic AI's own marketing cites an IBM figure that 63% of organizations have no AI governance policy at all — a plausible finding about the state of the market, worth treating as a vendor-sourced claim rather than an independent one.

The one tool built the opposite way

HOL Guard takes a genuinely different design stance from every product above: it runs locally, on the machine, intercepting the specific actions a coding agent is about to take — deleting a directory, reading a credentials file, pushing to a protected branch — and asking before it executes them. Nothing about this requires a cloud account; local interception is free forever and works offline, which the other vendors in this piece can't say, because their entire value proposition depends on being between the agent and the network. The tradeoff is the mirror image of the governance platforms: HOL Guard sees everything that happens on the one machine it's installed on and nothing across a fleet unless you pay for Cloud sync ($4.99/mo Solo, $15/mo Pro, $30/seat/mo Team), which adds cross-device history, alerts, and shared policy — still not the kind of enterprise-wide discovery Holistic AI or Credo AI are built for.

What free actually buys you

Promptfoo is the strongest free option in this list by a clear margin: MIT-licensed, 24,192 GitHub stars, pushed to in the last two months, and its Community tier is free forever with 10,000 red-team probes a month and full local evaluation — enough for most individual teams to run meaningful pre-deployment testing without talking to sales. Guardrails AI (Apache-2.0, 7,473 stars, pushed today) and NVIDIA's NeMo Guardrails (Apache-2.0, 7,228 stars, pushed yesterday) are both actively maintained, open-source frameworks you wire validation policy into directly — they're libraries, not hosted dashboards, so there's no proxy layer watching anything you haven't coded a check for.

LLM Guard is the cautionary case in the open-source set. It's still MIT-licensed with 3,213 stars, and it still shows up in searches for prompt-injection and PII-screening libraries — but Protect AI archived the repository on GitHub, with the last push on 2026-07-08 and no commits since. It will keep working exactly as it does today; nothing in it will get fixed or extended. Don't start a new project on it.

The listing to flag before you book a demo

If you're evaluating CalypsoAI, stop: its website now redirects entirely to F5's AI Guardrails and AI Security Platform product pages, under F5's own branding, navigation, and footer. F5 has folded CalypsoAI into its product line; CalypsoAI no longer exists as an independent platform or pricing entity. Anything you read about CalypsoAI's standalone pricing or roadmap is stale by definition — what you'd actually be buying is an F5 enterprise contract.

Which one to pick

  • You're worried about a coding agent on a developer's laptop (destructive shell commands, leaked credentials): HOL Guard. Free, local, works offline, nothing to wire up.
  • You need pre-deployment testing in CI before an agent ships: Promptfoo. Free tier covers meaningful volume, MIT-licensed, actively maintained.
  • You're running a customer-facing LLM app and need inline prompt-injection and PII filtering: Lakera Guard or a comparable runtime gateway — understand that you're buying detection with a published, non-zero miss rate, not a guarantee.
  • You already have a SOC and want AI-specific triage on top of your SIEM/EDR: Simbian or Straiker Defend AI — and go in knowing that no LLM-based detector currently covers the MITRE ATT&CK matrix comprehensively, by the vendors' own published research.
  • You need to find every agent running across the company and prove governance to an auditor: Holistic AI or Credo AI — and budget time to connect every system that matters, because unconnected systems are invisible to the inventory.

No single product in this list does more than one of these jobs well. The buyers who get burned are the ones who bought a gateway expecting governance, or a governance platform expecting runtime protection.

Share: