Anthropic's Claude Haiku 5.5 Enhances Cybersecurity Capabilities While Improving Resistance to Misuse
Anthropic's latest budget-friendly AI model, Claude Haiku 5.5, demonstrates significant improvements in identifying vulnerabilities and generating exploits, alongside enhanced safeguards against misuse and prompt injection attacks.

Anthropic has released Claude Haiku 5.5, an updated version of its AI model designed for speed-sensitive and repetitive tasks. This new iteration boasts enhanced capabilities in cybersecurity, including a greater proficiency in identifying software vulnerabilities and generating exploit code compared to its predecessor, Haiku 4.5. While its offensive security skills have improved, they remain less advanced than Anthropic's more sophisticated models like Sonnet and Opus.
To gauge its offensive potential, Anthropic tested Haiku 5.5 with its cybersecurity safeguards disabled. In a specific test involving known flaws within Chrome's V8 engine, the model successfully achieved arbitrary code execution in four out of 410 test runs. Further evaluations of multi-stage cyber operations showed a pre-release version of Haiku 5.5 completing 3.3% of challenges, a figure significantly lower than the 46.1% achieved by Sonnet 5.5 and 67.6% by Opus 5.5, indicating a deliberate limitation on its offensive capabilities.
The company has implemented stricter cybersecurity safeguards in Haiku 5.5 compared to Haiku 4.5. However, these safeguards are intentionally less restrictive than those on its more advanced models. This configuration allows for a broader range of defensive security tasks while still blocking activities more likely to be associated with malicious actors, such as penetration testing. Anthropic also offers a Cyber Verification Program for qualified security professionals seeking reduced restrictions.
Anthropic conducted extensive safety evaluations to assess Haiku 5.5's response to harmful requests, sensitive topics, and scenarios where a user attempts to steer the model toward malicious outcomes. Across various categories including weapons, extremist activity, and child safety, Haiku 5.5 achieved a 99.71% harmless response rate on explicitly harmful requests when using a near-final version of the Claude.ai production system prompt. Its rate of incorrectly refusing harmless requests also decreased significantly from 3.05% in Haiku 4.5 to 0.82% in Haiku 5.5.
In longer conversational tests, Haiku 5.5 showed improved performance in handling influence operations and ad surveillance scenarios. However, its performance declined in weapons-related scenarios when tested via API without a system prompt. Anthropic acknowledges that these evaluations excluded additional production-level protections like real-time monitoring and advises API developers to implement their own safeguards, particularly for sensitive conversations involving self-harm or eating disorders.
The model also demonstrates increased resistance to misuse and prompt injection attacks. In Claude Code tests, Haiku 5.5 refused 84.3% of malicious requests, a substantial increase from Haiku 4.5's 66.6%. These malicious requests included attempts to generate malware or support DDoS attacks. Similarly, in computer-use tests, Haiku 5.5 refused approximately 82.6% of harmful requests, up from 58.9% in its predecessor. This improved refusal rate also surpassed that of Sonnet 5.5 and Opus 5.5 in this specific evaluation.
Anthropic highlights Haiku 5.5 as its most resistant Haiku model to prompt injection, a technique where malicious instructions are hidden within AI inputs. Its resistance in coding and computer-use environments against adaptive attackers largely matched Anthropic's frontier models. However, it was less resistant on a separate Gray Swan benchmark, with remaining vulnerabilities primarily in graphical computer use scenarios. These adaptive attack evaluations were conducted both with and without prompt-injection probes enabled.
Beyond its security enhancements, Haiku 5.5 is also more cost-effective than its predecessor, with significant price reductions for shorter requests. Anthropic has also reduced the price of cache reads for Sonnet 5.5, making AI agent tasks approximately 20% cheaper. The model is now available across major cloud platforms including AWS, Google Cloud, and Microsoft Azure, with developers able to access it via the Claude Platform.