Capsule Security Launches ‘AI Circuit Breaker’ to Stop Rogue Agents
Capsule Security has launched an AI-powered ‘circuit breaker’ designed to proactively prevent damage from autonomous AI agents. This new solution utilizes specialized Small Language Models (SLMs) to evaluate an agent’s intended action in real-time, offering an independent control layer to stop potentially harmful actions before they execute. The system boasts high accuracy and minimal latency, addressing a critical security gap created by increasingly sophisticated AI agents.
Capsule Security has released its ‘AI Circuit Breaker,’ a new runtime security layer designed to mitigate the risks posed by autonomous AI agents. Founded in 2025 by Naor Paz and Lidan Hazout, the company recognized a growing security vulnerability stemming from the ability of AI agents to operate independently and potentially cause harm. Their solution aims to address this by providing real-time evaluation and intervention.
The core of the AI Circuit Breaker is Capsule’s own specialized AI, trained using NVIDIA Nemotron 3 Ultra, incorporating real agent traces, human review, and adversarial examples. This training process focused on identifying the boundary between authorized and rogue agent behavior.
Two models were developed, achieving a 96.9% detection accuracy compared to 86% for the strongest third-party model. Crucially, the models can make decisions in as little as 71 milliseconds, minimizing latency and allowing the system to operate seamlessly within an agent’s workflow. Capsule reduced the infrastructure for its larger model, decreasing memory requirements by almost 50%.
The AI Circuit Breaker evaluates an agent’s intended action immediately prior to execution, enabling organizations to allow, flag, or block it in real-time. This creates an independent control layer for agents accessing sensitive data, writing code, operating infrastructure, and interacting with other systems – a critical feature for organizations deploying agentic workflows.
Capsule claims 98% efficiency when tested against StepShield, an independent academic benchmark for rogue agent behavior detection. The company emphasizes the shift towards specialized Small Language Models (SLMs) as the key to safely scaling trusted agentic workflows across enterprises, moving beyond general-purpose models to achieve better performance and efficiency.