“Sorry, I can’t help with that”: How your guardrails might become the attacker’s best friend
This week’s Threat Source newsletter focuses on the potential pitfalls of relying too heavily on AI guardrails in security operations. Cisco Talos argues that overly restrictive AI systems can actually hinder investigations by slowing them down or outright blocking them, effectively creating a ‘safety penalty’ for defenders. The article emphasizes the importance of operational sovereignty – maintaining control over security policies and technical controls – and advocates for organizations to customize guardrails to match their specific threat models, rather than accepting inflexible solutions from third-party providers. Several other security headlines highlight ongoing threats, including a maturing banking Trojan, a disruption of West African crime infrastructure, and a new malware targeting car head units.
Welcome to this week’s edition of the Threat Source newsletter. Hello, everyone. Long time reader, first time writer here at the Threat Source newsletter! I wanted to start out by introducing myself. My colleague and friend Mick Baccio set the bar pretty high last week, so I was planning to tell you all about myself, including:
- How I did my first real IR under the influence of The Cuckoo’s Egg while an undergraduate (and failed)
- My pre-bug bounty flirtation with vulnerability research, including an arbitrary file overwrite in biff(1) and how I once hacked MIT’s website
- My first ever hands-on experience with a computer, the display demo Commodore 64 at the Montgomery Ward
Unfortunately, my editor says we don’t have the “space” for that, the MIT thing might open me up to “liability,” and it’s not the kind of “professional image” we strive for here at Talos. (I'm watching. Always watching. -Amy)
So instead, I’ll just play it safe and say that I’ve been in the security field for a little over 30 years now, mostly concentrating on the defensive side (Go, Team Blue!). I’ve helped set up SOCs, run threat hunting teams, and even published a few things you might have heard of.
Speaking of things I’ve published, I’ve written before about the Attacker’s Dilemma. The idea that defenders have inherent advantages over attackers runs contrary to what most of us have heard throughout our careers. An attacker must evade monitoring and technical controls at every step of their attack lifecycle, because the defender only needs to notice once in order to respond and prevent them from achieving their goal. This is one of the most important advantages of any security team has, but we are currently witnessing a self-imposed erosion of this advantage through the rise of poorly-designed AI guardrails.
I’m not opposed to guardrails, but we have to carefully consider what we’re guarding against and where we deploy them. As I explored in a recent piece on The Safety Penalty, by allowing third-party AI providers to implement and control safety filters and the policies behind them, we may in fact be helping the attacker. If agentic SOC process experience refusals, it can slow or even halt investigations. Of course, these should get flagged for human intervention, but that takes time and may give the attacker breathing room in which to complete their mission.
It may turn out that the where of the guardrails is even more important than the what. Operational sovereignty relies on having control of our own limits. Any vision of an agentic SOC must allow the security teams to customize the guardrails to match their own threat model. They should also have the flexibility to temporarily remove specific safeguards under authorized circumstances, something you won’t get with guardrails from a frontier provider. These controls belong inside your organization’s agentic harness where you can set the policies and technical controls to allow you to analyze threats while ensuring your agents stay within their lanes.
Ultimately, operational sovereignty means engaging with the reality of the threat landscape, ensuring that the adversary can’t derail the defender’s investigation and response processes, either accidentally or intentionally. We need to move toward a model where each organization can choose the guardrails that work for them, rather than having inflexible guardrails chosen for them.
**Top security headlines of the week**
- **ToxicPanda banking trojan matures into enterprise threat:** ToxicPanda 2.0 expands substantially on its predecessor, adding 167 remote commands and broadening its targeting from 16 financial institutions to 349 banking, e-wallet, and cryptocurrency applications. (Dark Reading)
- **Interpol's Jackal IV disrupts West African crime infrastructure:** Law enforcement from 22 countries across six continents worked together to arrest 58 suspects and identify 263 more. The first two Jackal operations in 2022 and 2023 led to approximately 200 arrests in total and millions of dollars more in seized assets. (Dark Reading)
- **First malware built specifically for car head units fuels botnet:** Researchers have found what appears to be the first malware specifically designed for car head units, with links to the notorious BadBox botnet, on an Android-powered aftermarket infotainment system made by Chinese company DoFun, which is widely used in China and other APAC countries. (SecurityWeek)
- **A Tale of Two SOCs: Insights From Two Red Team Assessments:** A CISA red team fully compromised two critical infrastructure organizations at the domain level and reached sensitive business systems and cloud resources. Organization A failed to detect or contain the activity. Organization B rapidly identified initial compromise attempts, isolated affected systems, and forced the red team into an assume breach model. (CISA)
- **NovaCookies campaigns abuse genuine Docusign notifications to steal M365 sessions:** The $320/month service is a subscription-based phishing platform that facilitates real-time M365 session theft. The kit has been used to target hundreds of organizations across multiple sectors in the U.S., the U.K., Canada, Germany, and more. (The Hacker News)
**Can’t get enough Talos?**
- **JavaScript obfuscation: From party trick to phishing kit:** We've spent a lot of time pulling apart suspicious JavaScript from phishing kits, malware packages, compromised sites, and more. Learn the basics of what obfuscation is, why a researcher would try to reverse it, and several ways to approach the problem.
- **The safety penalty: Reclaiming operational sovereignty in the age of AI:** As frontier AI models become increasingly restrictive, security teams are facing a “safety penalty” that hampers real-time incident response. Discover how organizations can move toward operational sovereignty to ensure their defensive AI keeps pace with unconstrained adversaries.
- **Back-to-school cybersecurity: Protecting education networks from ransomware and threats:** As the new academic year begins, school districts face a surge in cybersecurity threats, from phishing attacks and ransomware to student experimentation with network devices. In this episode, Amy sits down with Cisco Talos expert Pierre Cadieux to discuss practical strategies for IT practitioners.
**Upcoming events where you can find Talos**
- International European Cyber Threat Intelligence Conference (IECTIC) (Sept. 9) Kassel, Germany
- Secure Iowa (Sept. 9) Altoona, IA
- .conf26 (Sept. 14 – 17) Denver, CO
- LABSCon (Sept. 16 – 19) Scottsdale, AZ
- VB (Oct. 14 – 16) Seville, Spain
- CAMLIS (Oct. 21 – 23) Arlington, VA
**Most prevalent malware files from Talos telemetry over the past week**
- SHA256: 9f1f11a708d393e0a4109ae189bc64f1f3e312653dcf317a2bd406f18ffcc507
- MD5: 2915b3f8b703eb744fc54c81f4a9c67f
- Talos Rep: https://talosintelligence.com/talos_file_reputation?s=9f1f11a708d393e0a4109ae189bc64f1f3e312653dcf317a2bd406f18ffcc507
- Example Filename: VID001.exe
- Detection Name: W32.9F1F11A708-100.SBX.TG**
