VYPR
researchPublished Jul 29, 2026· 1 source

AI Models Demonstrate Advanced Cryptanalysis Capabilities, Discovering Novel Attacks

A new benchmark, CryptanalysisBench, reveals that advanced AI models can perform sophisticated cryptanalysis, even discovering previously unknown vulnerabilities in cryptographic schemes.

A novel benchmark, dubbed "CryptanalysisBench," is emerging as a critical tool for evaluating the capabilities of large language models (LLMs) in the complex field of mathematical cryptanalysis. This benchmark aims to assess how effectively LLMs can discover new mathematical attacks against a range of historical cryptographic algorithms. Cryptanalysis, the art and science of breaking codes, sits at the nexus of mathematical reasoning and cybersecurity – two domains where LLMs have shown remarkable progress.

The research paper introducing CryptanalysisBench highlights its dual purpose: serving as a rigorous testbed for the advanced reasoning abilities of frontier AI models, and as a high-stakes domain where the security primitives under study are fundamental to digital infrastructure. The benchmark is structured into three tiers. The first tier comprises cryptographic primitives for which practical attacks are already known. The second tier includes primitives with no known practical breaks, tested both at their full strength and in scaled-down variants. The third tier presents a challenge set of production-level primitives that are at the current frontier of cryptanalytic research.

Results from testing five leading LLMs – Anthropic's Claude Opus and Sonnet, Google's Mythos, OpenAI's GPT-5.5, and the open-weight GLM-5.2 – have been striking. These models successfully broke between 65% and 86% of the Tier 1 schemes. For Tier 2 schemes at full strength, they managed to break between 6 and 12 tasks. Across all scaled-down variants in Tier 2, the models achieved success rates ranging from 24% to 61%, demonstrating a broad capability to find weaknesses.

Perhaps more significantly, these AI models have gone beyond merely replicating known cryptanalytic results. They have independently produced novel attacks, including a key-recovery attack that exploits a specific design flaw in the SpoC Authenticated Encryption with Associated Data (AEAD) scheme. Furthermore, an error was identified in KINDI's published CCA-security proof, a finding that, to the best of the researchers' knowledge, had not been previously reported. These discoveries underscore the potential for AI to contribute meaningfully to the field of cryptanalysis.

Anthropic itself utilized the benchmark to test its Mythos Preview model, leading to the discovery of new vulnerabilities in the Hawk post-quantum signature scheme and a reduced-round version of the AES block cipher. These findings, which had eluded human experts for years, suggest that AI is rapidly approaching, and in some cases potentially exceeding, the published state-of-the-art in cryptanalysis.

The researchers are releasing CryptanalysisBench as an open tool. Its purpose is to help track the evolving landscape of AI-driven cryptanalysis and to serve as a platform for stress-testing candidate cryptographic schemes before they are deployed. The early successes of LLMs in this domain indicate that AI is poised to become a significant factor in cybersecurity, necessitating a proactive approach to understanding and mitigating the risks associated with AI-powered cryptanalytic capabilities.

As AI models become more sophisticated, their ability to uncover complex mathematical weaknesses in cryptographic systems will likely grow. This trend poses both opportunities for improving security through AI-assisted analysis and significant challenges for maintaining the confidentiality and integrity of digital communications. The cybersecurity community must adapt to this new reality, integrating AI into defensive strategies while also developing countermeasures against AI-driven offensive capabilities.

Synthesized by Vypr AI