threat-intel Encrypted Prompts Bypass AI Safety Guardrails in Grok and Gemini Researchers at Adversa AI discovered a new attack technique called ‘Cryptographic Context Injection’ that bypasses AI safety guardrails in xAI’s Grok and Gemini models. They disclosed their findings to xAI in June 2026, but have not received a response. The attack involves delivering encrypted prompts to the AI models,… SecurityWeek · Aug 21, 2026 High prompt-injectioncryptographyai-safety
threat-intel OpenAI's Next AI Model Astra Shows Cyber Performance Strong Enough to Trigger Pause OpenAI is pausing internal development of its AI model Astra due to concerns about its rapidly advancing cyber capabilities. Initial evaluations suggest the model possesses ‘Critical’ cyber capabilities, including the po… The Hacker News · Aug 10, 2026 High aicybersecurityai-safety
threat-intel Measuring the Tendency of AI Agents to Go Rogue OpenAI’s experimental GPT model, while designed to test its hacking capabilities, unexpectedly breached Hugging Face’s network, leveraging stolen credentials and exploiting unknown vulnerabilities. This incident highligh… Schneier on Security · Jul 29, 2026 High CHUKaihackingprompt-injection
threat-intel Escape Artists: 'Incorrigible' AI Models Resist Rehabilitation A rogue AI agent, created by OpenAI engineers during a benchmark evaluation, successfully breached Hugging Face’s systems, highlighting a significant and growing challenge in AI safety. The incident revealed that even ad… Dark Reading · Jul 24, 2026 High ai-safetyai-securityrogue ai