Mid-Tier AI Models Emerge as Potent Hacking Tools, Outpacing Frontier Models in Cost-Effectiveness
Researchers warn that increasingly capable 'middle-class' AI models, offering a strategic advantage through lower costs, are becoming a significant threat in offensive cybersecurity.

While much attention has focused on the advanced hacking capabilities of frontier AI models, a new report from XBOW highlights a growing concern: the 'middle class' of AI, including models like GLM-5.2, Grok 4.5, and Opus 4.7, is rapidly improving and becoming a more practical and cost-effective tool for cyber adversaries.
These mid-tier models, both proprietary and open-source, have crossed a critical threshold, enabling them to perform complex agentic tasks that were challenging just six months ago. Their key advantage lies in their affordability. Unlike expensive frontier models, the lower cost of these mid-tier AIs allows users to run them repeatedly, iterating on tasks to achieve better results and potentially leapfrog the performance of more powerful, but cost-prohibitive, models.
"It’s not even that the open-source variants or…not quite frontline competitors are catching up to frontier models as such," explained Albert Ziegler, head of AI at XBOW. "It’s that they are crossing a certain threshold, which means that suddenly they are providing net value at a cheaper price."
GPT 5.5, now considered a near-frontier model, demonstrated remarkable progress in autonomous web application testing. XBOW's benchmarks show a significant reduction in vulnerability miss rates compared to previous iterations, particularly in black-box scenarios where attackers lack source code access. This improvement is crucial, as it mirrors real-world attack conditions where source code is typically unavailable.
Furthermore, GPT 5.5's performance without source code access surpassed that of GPT 5 when it *did* have access. This indicates a shift in how AI models are achieving success in offensive workflows, prioritizing live interaction with the target system over static code analysis. This capability is vital for attackers seeking to exploit vulnerabilities in running applications.
While frontier models like Mythos and GPT 5.6 offer superior performance on individual tasks, their exponentially higher token costs make them inaccessible for many. Anthropic's research, for instance, showed that while coordinating agent swarms could find significantly more vulnerabilities, the token expenditure was immense. This cost barrier makes the more economical mid-tier models a more attractive option for widespread use.
The implications for cybersecurity are profound. As these more affordable AI tools become more capable, they lower the barrier to entry for sophisticated cyberattacks. Malicious actors, who are less concerned with collateral damage and budget constraints than legitimate organizations, stand to benefit significantly from this trend.
Ziegler noted that recent incidents involving frontier models escaping sandboxes underscore the upper-tier capabilities of LLMs. However, he emphasized that the cybersecurity industry, like most others, favors tools that are both effective and affordable. The increasing prowess and cost-effectiveness of mid-tier AI models suggest they will play a pivotal role in shaping the future threat landscape, potentially democratizing advanced hacking techniques on an unprecedented scale.