Nuclear-Sabotage Malware Benchmark Trips Up Most Frontier AI Models
SentinelOne has developed a new benchmark to assess how well frontier AI models can conduct thorough investigations of complex malware, using Fast16 – a malware linked to Iran’s nuclear program – as the test case. The benchmark revealed that while some models, like GPT-5.6 Sol, could complete a multi-stage investigation, others struggled to correct their findings and acknowledge new evidence, highlighting the continued need for human oversight in AI-driven security analysis.
SentinelOne has created a novel benchmark designed to evaluate the investigative capabilities of advanced AI models, specifically focusing on their ability to analyze complex malware. The benchmark’s foundation is built upon the investigation of Fast16, a 2005 Windows malware that has been linked to Iran’s nuclear weapons development program.
SentinelLabs, SentinelOne’s research arm, utilized Fast16 as a test case, assessing models like OpenAI’s GPT-5.5 and GPT-5.6 Sol, Z.ai’s GLM-5.2, and Anthropic’s Opus 4.x. The benchmark tracks whether a model can sustain a trustworthy investigation across eight escalating stages, with new evidence consistently contradicting its initial conclusions.
GPT-5.6 Sol demonstrated the most robust performance, successfully completing all eight stages across three separate runs with varying reasoning efforts. GPT-5.5, GLM-5.2, and Opus 4.7 and 4.8 produced localized analysis but ultimately stalled, failing to correct their findings and acknowledge new evidence. GPT-5.5 never progressed beyond the initial stage, while the Opus models frequently declared their work complete before addressing underlying defects.
SentinelLabs attributes this discrepancy not to the models’ technical skill or insight, but to what they term ‘project-scale recovery’ – a model’s capacity to withdraw disproven conclusions, trace all downstream dependencies, fix the root cause, and carry that correction through the entire investigation.
The researchers emphasized that human oversight remains crucial, even with GPT-5.6 Sol’s successful performance, noting that the model still made significant technical errors, accepted weak quality controls, and prematurely claimed readiness. “Senior reverse engineers remain essential,” the researchers explained. “Even the best current runs made semantic errors, accepted weak quality controls, and claimed readiness prematurely. We assess the best current use as supervised investigative agency, with human analysts defining objectives, exposing blind spots, and retaining final publication authority.”