threat-intel
Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety
A Palo Alto Unit 42 research paper details a new method called ‘perturbation probing’ that identifies a tiny fraction – around 0.014% – of feed-forward neurons within aligned Large Language Models (LLMs) responsible for…
High

