threat-intel Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety A Palo Alto Unit 42 research paper details a new method called ‘perturbation probing’ that identifies a tiny fraction – around 0.014% – of feed-forward neurons within aligned Large Language Models (LLMs) responsible for their safety responses. The study reveals that these models rely on a remarkably fragile ‘thin layer… Palo Alto Unit 42 · 2d ago High llmsafetyalignment
threat-intel Phantom Squatting: AI-Hallucinated Domains as a Software Supply Chain Vector Palo Alto Unit 42 researchers have identified a new supply chain threat: "phantom squatting," where large language models (LLMs) hallucinate web domains that adversaries can then register to intercept traffic generated b… Palo Alto Unit 42 · Jul 1, 2026 High USllmaisupply chain