news.mlab.sh
21 results
threat-intel

Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety

A Palo Alto Unit 42 research paper details a new method called ‘perturbation probing’ that identifies a tiny fraction – around 0.014% – of feed-forward neurons within aligned Large Language Models (LLMs) responsible for their safety responses. The study reveals that these models rely on a remarkably fragile ‘thin layer…

Palo Alto Unit 42 · 1d ago High
threat-intel

Prompt Injections for Defense

Researchers discovered a method called ‘context bombing’ where strategically placing prompt injections alongside sensitive data (like passwords and keys) can effectively disable AI hacking agents. This works by forcing t…

Schneier on Security · Aug 12, 2026 Medium