news.mlab.sh
Back to the feed
threat-intel

'Yellow Teams' Are Defining the Future of AI Security

High
Summary

A growing trend of ‘yellow teams’ – engineering groups building both attack and defense tools – is emerging as a crucial response to the increasing threat of AI-powered cyberattacks. These teams are using advanced AI models like Mythos (Anthropic) and GPT-5.5 (OpenAI) to proactively identify vulnerabilities and develop defenses. Organizations are increasingly collaborating between ‘red’ and ‘blue’ teams, with yellow teams providing the technical expertise to effectively wield these AI models. This collaborative approach is vital for organizations to stay ahead of the rapidly evolving AI threat landscape and integrate AI security into their software development lifecycle.

A growing trend of ‘yellow teams’ – engineering groups building both attack and defense tools – is emerging as a crucial response to the increasing threat of AI-powered cyberattacks. These teams are using advanced AI models like Mythos (Anthropic) and GPT-5.5 (OpenAI) to proactively identify vulnerabilities and develop defenses. Organizations are increasingly collaborating between ‘red’ and ‘blue’ teams, with yellow teams providing the technical expertise to effectively wield these AI models.

In April, Anthropic invited more than 50 organizations to participate in its Project Glasswing initiative to preview Claude Mythos, and shortly after, OpenAI followed suit with its Daybreak program offering access to GPT-5.5. Since then, red teams – penetration testers – have been using these AI models to exploit their own systems, and rival blue teams have tried to detect and defend against those exploits. The core of this movement is the ‘yellow team’ – a group of engineers building both attack and defense frameworks.

“Yellow team is a build team,” explains Sam Curry, CISO at Zscaler. “They’ll say: ‘Hey, Red, what tools could you use to do better attacks?’ As if they were the engineering department behind a major attacker. They’ll also turn to Blue and go: ‘What would you like on defense that isn’t supplied to you by your vendor community?’”

Organizations are developing ‘harnesses’ – software cocoons – to control and focus the AI models’ capabilities. Cisco’s Foundry Security Spec and Microsoft’s MDASH are examples of these harnesses, which restrict the AI’s actions and define its permissions. Netskope, for instance, developed a tradition of daily scrums between red and yellow teams to discuss how they were using the models and documenting their findings.

“If the blue team doesn’t take advantage of AI for analysis, signals and events are going to overwhelm them significantly,” argues Levi Bolourie, CISO at Zscaler. “And I think because the volume is increasing, blue and yellow are going to have to become way more integrated.”

Beyond simply identifying vulnerabilities, yellow teams are integrating AI findings directly into the software development lifecycle (SDLC). Cisco’s Foundry Security Spec is paired with its CodeGuard framework to create a closed-loop system where attack tools feed vulnerabilities back to defense, and vice versa. Netskope, for example, has seen a shift where red and yellow team members are cross-educating each other, with the yellow team members contributing to patching processes.

“We know what’s coming up — in six months, we know that these models are going to be available, and people are going to use them to attack our systems,” says James Robinson, CISO at Netskope. “CISOs need to start working with development teams to take an engineering-forward approach and create some new processes. Definitely get them involved.”

Read the full article at Dark Reading