news.mlab.sh
Back to the feed
threat-intel

Anthropic Releases Claude Fable 5, Its Most Powerful AI Yet, With Cyber Safeguards

High
Summary

Anthropic has released Claude Fable 5, its most powerful AI model, alongside a specialized version called Claude Mythos 5 designed for cybersecurity professionals. Fable 5 incorporates cyber-focused classifiers to mitigate potential misuse and exploit development, while Mythos 5 retains full cyber capabilities. This split approach reflects concerns about the potential for a general-purpose AI model to be used for malicious cyber activities, particularly due to its ability to identify and exploit vulnerabilities. Anthropic’s testing demonstrated the effectiveness of the safeguards, though some false positives were observed.

Anthropic’s release of Claude Fable 5 marks a significant step in AI development, offering a powerful model alongside a specialized version tailored for cybersecurity applications. The core strategy involves a bifurcated approach: Fable 5, accessible to the public, utilizes classifiers to manage potentially dangerous requests, routing them to a less capable model (Claude Opus 4.8) for processing. Conversely, Claude Mythos 5 remains restricted to vetted cyber defenders and critical infrastructure operators, retaining the full suite of cyber capabilities. This decision stems from concerns highlighted during the Project Glasswing testing of Claude Mythos Preview, where the model autonomously identified and exploited zero-day vulnerabilities across major operating systems and web browsers, including a 27-year-old flaw in OpenBSD. Anthropic’s red team emphasized that these capabilities emerged as a side effect of the model’s general improvements, posing a significant risk if unleashed without safeguards. The company’s testing, involving over 1,000 hours of a bug bounty program, confirmed the effectiveness of the classifiers in preventing harmful requests related to cyberattack planning and exploit development. While some false positives were identified, Anthropic plans to refine the safeguards to minimize disruption.

Read the full article at The Hacker News