VYPR
researchPublished Jul 23, 2026· 1 source

Frontier AI Models Struggle with Complex Malware Investigations, SentinelOne Benchmark Reveals

A new benchmark by SentinelOne demonstrates that even advanced AI models falter when tasked with sustained, complex malware investigations, highlighting the need for human oversight.

SentinelOne has developed what it claims is the first long-horizon reverse-engineering benchmark for frontier AI models, using its own investigation into the recently documented Fast16 malware as a test case. The Fast16 malware, detailed by SentinelLabs in April, is a 2005 Windows-based program designed to interfere with LS-DYNA, engineering software that appears to have been used by Iran as part of its nuclear weapons development program. Similar to the notorious Stuxnet, which it predates, Fast16 may have been developed by the United States and used to sabotage Iran’s nuclear program.

SentinelLabs researchers put leading AI models to the test, evaluating OpenAI’s GPT-5.5 and latest GPT-5.6 Sol model, Z.ai’s GLM-5.2, and Anthropic’s Opus 4.x. The benchmark was designed not to score models on isolated tasks, but rather to track whether a model could sustain a trustworthy investigation across eight escalating stages, particularly when new evidence repeatedly contradicted its own earlier conclusions. This approach aims to simulate the dynamic and often contradictory nature of real-world incident response.

Out of the tested models, only GPT-5.6 Sol managed to complete all eight stages of the benchmark, and this was achieved across three separate runs with different reasoning-effort settings. In contrast, GPT-5.5, GLM-5.2, and the Opus models (4.7 and 4.8) demonstrated solid local analysis capabilities but ultimately stalled during the investigation. GPT-5.5, for instance, never progressed beyond the initial stage, while the Opus models frequently declared the investigation complete prematurely, before critical defects were fully resolved.

SentinelLabs attributes the performance gap not to a lack of technical skill or insight among the AI models, but to what they describe as 'project-scale recovery.' This refers to a model's ability to retract a disproven conclusion, trace all downstream dependencies, identify and fix the root cause of the error, and then propagate that correction throughout the remainder of the investigation. This is a more complex cognitive task than simply patching an immediate error or providing a static analysis.

Despite the advanced capabilities shown by GPT-5.6 Sol, the SentinelLabs researchers concluded that human oversight remains absolutely essential. Even the strongest AI runs made significant technical mistakes, accepted weak quality controls, and claimed readiness prematurely. The researchers emphasized that the best current use of these AI models is as a supervised investigative agency, where human analysts define objectives, identify blind spots, and retain final authority over any published findings.

The findings underscore the current limitations of AI in complex cybersecurity tasks that require sustained critical thinking, adaptation to new information, and the ability to manage large-scale investigative projects. While AI shows promise in assisting with cybersecurity, its current state necessitates a collaborative approach with human experts to ensure thoroughness and accuracy in incident response and malware analysis.

This research highlights a critical area for AI development in cybersecurity: the ability to perform long-horizon, adaptive investigations. As AI models become more integrated into security operations, understanding their limitations in complex, evolving scenarios like malware analysis is crucial for effective deployment and risk management.

Synthesized by Vypr AI