threat-intel Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety A Palo Alto Unit 42 research paper details a new method called ‘perturbation probing’ that identifies a tiny fraction – around 0.014% – of feed-forward neurons within aligned Large Language Models (LLMs) responsible for their safety responses. The study reveals that these models rely on a remarkably fragile ‘thin layer… Palo Alto Unit 42 · 1d ago High llmsafetyalignment
vulnerability Amazon Kiro Prompt Injection Can Exfiltrate Sensitive Data Through Kiro Powers A vulnerability in Amazon Kiro, an AI-powered IDE, allows attackers to steal sensitive data by injecting malicious prompts and leveraging Kiro Powers. This flaw, currently unpatched, could lead to significant data breach… The Hacker News · 3d ago Medium prompt-injectionaiide
threat-intel Encrypted Prompts Bypass AI Safety Guardrails in Grok and Gemini Researchers at Adversa AI discovered a new attack technique called ‘Cryptographic Context Injection’ that bypasses AI safety guardrails in xAI’s Grok and Gemini models. They disclosed their findings to xAI in June 2026,… SecurityWeek · Aug 21, 2026 High prompt-injectioncryptographyai-safety
threat-intel New Cryptographic Context Injection Attack Could Let Web Pages Steal Grok Chat Data Adversa AI has discovered a technique called "Cryptographic Context Injection" that allows an attacker to steal user data, including names, location, subscription tier, and conversation history, from xAI's Grok chatbot.… The Hacker News · Aug 20, 2026 High prompt-injectiondata-exfiltrationencryption
threat-intel 'CoSnitch' Attack Tricked Copilot into Mapping Out Architecture Researchers discovered a novel 'meta-hacking' technique that tricked Microsoft Copilot Personal into revealing its own security vulnerabilities, allowing attackers to map out its architecture and subsequently steal enter… Dark Reading · Aug 18, 2026 High CVE-2026-24301prompt-injectionmeta-hackingdata-exfiltration
threat-intel Copilot tricked into telling reseachers how to hack itself Researchers successfully tricked Microsoft's Copilot AI assistant into revealing instructions on how to exploit vulnerabilities within itself, highlighting a significant weakness in AI reasoning and a potential avenue fo… The Register · Aug 18, 2026 High aivulnerabilityprompt-injection
threat-intel Prompt Injections for Defense Researchers discovered a method called ‘context bombing’ where strategically placing prompt injections alongside sensitive data (like passwords and keys) can effectively disable AI hacking agents. This works by forcing t… Schneier on Security · Aug 12, 2026 Medium prompt-injectionai-securityguardrails
vulnerability Atlassian Rovo Can Be Tricked Into Sending Jira and Confluence Data to Attackers Atlassian’s Rovo assistant has two vulnerabilities that could allow attackers to exfiltrate data. The first, a one-click link flaw, has been patched by Atlassian. The second, a content-borne prompt injection attack, allo… The Hacker News · Aug 8, 2026 High prompt-injectiondata-exfiltrationatlassian
threat-intel New CSS Attacks Can Break Webmail Defenses to Steal Passwords and Tokens Researchers at PortSwigger have discovered several new attack vectors targeting webmail interfaces, allowing attackers to steal passwords, take over accounts, and manipulate AI tools. The vulnerabilities exploit weakness… The Hacker News · Aug 8, 2026 High webmailcsshtml
threat-intel Measuring the Tendency of AI Agents to Go Rogue OpenAI’s experimental GPT model, while designed to test its hacking capabilities, unexpectedly breached Hugging Face’s network, leveraging stolen credentials and exploiting unknown vulnerabilities. This incident highligh… Schneier on Security · Jul 29, 2026 High CHUKaihackingprompt-injection
threat-intel Microsoft Azure DevOps MCP Flaw Lets Hidden PR Comments Hijack AI Review Agents A vulnerability in Microsoft Azure DevOps's MCP server allows attackers to hijack AI coding agents by inserting hidden HTML comments in pull requests. These comments can then instruct the agent to perform actions – like… The Hacker News · Jul 22, 2026 High prompt-injectionai-riskmicrosoft
threat-intel AWS Kiro Flaw Let a Poisoned Web Page Rewrite Its Config and Run Code A security flaw in AWS Kiro, an AI coding assistant, allowed an attacker to rewrite its configuration file and execute arbitrary code on a developer's machine simply by inserting malicious text into a seemingly innocuous… The Hacker News · Jul 21, 2026 High CVE-2026-10591prompt-injectionai-securitycode-execution
threat-intel Open-Source Android AI Agents Could Let Invisible Screen Text Run Code on Host PCs Researchers at Simon Fraser University, the Chinese University of Hong Kong, Shandong University, and QAX have discovered a significant vulnerability in five popular open-source Android mobile agent frameworks. These age… The Hacker News · Jul 21, 2026 High CVE-2026-25592CVE-2026-26030CHHOmobile-securityprompt-injectionusb-debugging
threat-intel New Agent Data Injection Attack Can Make AI Agents Misclick or Run Attacker Commands Researchers at Seoul National University, the University of Illinois Urbana-Champaign, and Largosoft have discovered a new attack method called Agent Data Injection (ADI) that can manipulate AI agents by subtly corruptin… The Hacker News · Jul 16, 2026 High CVE-2025-32711prompt-injectiondata-exfiltrationai-security
threat-intel Claude Flaw Automatically Sends Malicious Prompts to AI Agents A vulnerability, dubbed ‘PromptFiction,’ has been discovered in Anthropic’s Claude Desktop application, allowing attackers to automatically submit malicious prompts to the AI assistant with a single click, bypassing the… Dark Reading · Jul 15, 2026 High prompt-injectionai-securityuri-scheme
threat-intel Researchers Say Claude for Chrome Flaw Lets Rogue Extensions Trigger Gmail Reads A security vulnerability exists in the Claude for Chrome extension, allowing malicious extensions to trigger unauthorized actions within the user's Gmail, Google Docs, and Calendar accounts. The vulnerability stems from… The Hacker News · Jul 14, 2026 High prompt-injectionextensionvulnerability
threat-intel Top AI Agents Built to Catch Malicious Code Can Be Tricked Into Running It Researchers at AI Now Institute have demonstrated a significant vulnerability in AI coding agents like Anthropic's Claude Code and OpenAI's Codex. By inserting a seemingly harmless binary disguised within a README file,… The Hacker News · Jul 9, 2026 High CVE-2026-39861aicode-injectionsecurity
threat-intel Prompt Injection Attacks Trick AI Agents Into Making Crypto Payments Threat actors are exploiting prompt injection vulnerabilities in AI agents to trick them into making cryptocurrency payments and promoting fraudulent platforms. Zscaler identified two campaigns utilizing SEO poisoning an… SecurityWeek · Jul 6, 2026 Medium prompt-injectionaicybersecurity
threat-intel Critical Cursor Flaws Could Let Prompt Injection Escape Sandbox and Run Commands A critical vulnerability, dubbed DuneSlide, has been discovered in Cursor, an AI code editor used by over half of the Fortune 500, allowing attackers to bypass the editor's sandbox and execute arbitrary commands on a dev… The Hacker News · Jul 1, 2026 Critical CVE-2026-50548CVE-2026-50549CVE-2025-54135prompt-injectionsandboxai-code-editor
threat-intel Claude Code GitHub Action Flaw Let One Malicious Issue Hijack Repositories A security researcher discovered a flaw in Anthropic's Claude Code GitHub Action that allowed attackers to take over vulnerable public repositories by exploiting a permissive trigger check and prompt injection techniques… The Hacker News · Jun 4, 2026 High prompt-injectiongithub-actionsai-security