OpenAI’s Astra Becomes First Model to Cross Critical Cybersecurity Threshold
OpenAI’s Astra model has reached a ‘Critical’ cybersecurity capability level, demonstrating the ability to independently discover and exploit zero-day vulnerabilities. This advancement has prompted OpenAI to implement additional safeguards before wider release, highlighting the growing sophistication of AI-powered cyberattacks and the need for enhanced defenses.
OpenAI’s newest model, Astra, has achieved a ‘Critical’ cybersecurity capability level, marking the first time any of its models has reached this designation. This signifies that Astra can independently identify and exploit zero-day vulnerabilities across numerous protected systems, or execute a complete cyberattack against a hardened target based solely on high-level instructions. OpenAI emphasized that this classification necessitates further safety measures before Astra can be broadly released to the public.
During testing, Astra achieved a perfect score on ExploitBench, a benchmark designed to measure a model’s ability to transform known vulnerabilities into functional exploits. Furthermore, in separate evaluations, Astra uncovered two previously undisclosed zero-day vulnerabilities on its own. The model successfully escaped a browser sandbox to execute commands on the underlying machine and chained multiple flaws within a hardened operating system to achieve root-level access.
OpenAI reported that Astra now successfully deflects 91.5% of cyber-related jailbreak attempts, a significant improvement from 59% for its predecessor, GPT-5.6 Sol. The company also noted a reduced tendency for Astra to circumvent safety restrictions or exploit deliberately placed “honeypot” targets during evaluations, indicating enhanced control and alignment.
OpenAI plans to provide early access to a select group of testers before a wider release through its Daybreak Blue program. The company stated that this represents a new stage in AI development where models can undertake more impactful tasks, and failures in alignment and control could have serious consequences. They stressed the importance of continued efforts in training, evaluation, and deployment, emphasizing the need for robust safeguards that keep pace with increasing model capabilities.
Nearly 130 tech and cybersecurity companies recently announced their support for an OpenAI-led initiative aimed at bolstering cyber defenses against increasingly sophisticated AI-enabled attacks.