threat-intel Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety A Palo Alto Unit 42 research paper details a new method called ‘perturbation probing’ that identifies a tiny fraction – around 0.014% – of feed-forward neurons within aligned Large Language Models (LLMs) responsible for their safety responses. The study reveals that these models rely on a remarkably fragile ‘thin layer… Palo Alto Unit 42 · 1d ago High llmsafetyalignment
threat-intel Senators press TikTok over withholding of safety features for some users U.S. senators are investigating TikTok over allegations that the company intentionally disabled safety features for a subset of users, including a 16-year-old who died after viewing harmful content. The company conducted… The Record · Aug 20, 2026 High safetyalgorithmchild-safety
threat-intel Anthropic’s Opus 5 Nears Mythos 5 on Finding Bugs, but Falls Short on Exploits Anthropic’s Opus 5 AI model excels at finding software vulnerabilities, nearly matching Mythos 5’s capabilities, but it cannot automatically generate exploits. Anthropic deliberately limits Opus 5’s exploit generation ab… SecurityWeek · Jul 27, 2026 Medium aivulnerabilitycybersecurity
threat-intel Why AI Needs a “Genie Coefficient” This article introduces the concept of the ‘Genie Coefficient’ – a metric to measure the gap between a user’s request and an AI’s actual action, reflecting the tendency of AI agents to go beyond the explicit instructions… Schneier on Security · Jul 24, 2026 High aialignmentreward hacking
threat-intel Exclusive: Meet AIVEX, a New Triage Model Built to Reduce Supply Chain Threat and Risk This article discusses the limitations of current vulnerability triage methods (SBOMs, VEX, and CVSS) in addressing supply chain attacks, particularly in the context of increasingly complex AI-driven systems. The author,… SecurityWeek · Jun 24, 2026 High supply chainaiautonomous robots
threat-intel Patch Now: Critical Flaw in OT Robot OS Gives Attackers Control A critical command injection vulnerability (CVE-2026-8153) was discovered in the operating system of Universal Robots’ PolyScope 5 collaborative robots. This flaw allows unauthenticated attackers to gain remote access an… Dark Reading · May 20, 2026 Critical CVE-2026-8153DEcommand injectionotrobotics