Attackers Exploit LLM Safety Classifiers by Decomposing Harmful Requests
CrowdStrike researchers detail a novel attack method where adversaries break down malicious prompts into seemingly benign subtasks to bypass AI safety classifiers, enabling the generation of harmful content.

Modern large language models (LLMs) are protected by sophisticated safety classifiers designed to prevent the generation of harmful or malicious content. These classifiers, themselves AI models, act as gatekeepers, evaluating user requests in real time before they reach the main LLM. While highly effective against direct attacks, these systems have a structural blind spot: they analyze requests individually, not as part of a larger sequence. This limitation is being exploited by attackers who can decompose a harmful objective into a series of individually benign subtasks.
The technique, detailed by CrowdStrike's Cyber Superintelligence Lab, involves a three-step process: decomposition of the offensive goal, benign reframing of each subtask to appear as a legitimate request, and recomposition of the results by an unclassified, smaller LLM. This bypass method has been independently validated across nine out of ten offensive security categories, demonstrating its broad applicability. The research highlights that while direct bypass attempts against robust classifiers fail, this indirect approach circumvents the safety mechanisms by exploiting the sequential nature of interactions.
CrowdStrike's findings echo research published by Microsoft, titled "Capability Laundering," which independently discovered similar bypass techniques. Both teams found that per-request filtering is insufficient when harmful tasks are broken down into seemingly innocuous queries. While Microsoft's research focused on specific benchmark evaluations, CrowdStrike's work provides a more comprehensive mapping of the classifier's boundary surface across approximately 515 technique classes, identifying specific reframing strategies and demonstrating the full pipeline across multiple MITRE ATT&CK® categories.
The core of the attack lies in the "benign reframing" step. Adversaries can frame their requests as legitimate software engineering tasks, such as writing code for game modifications, detection engineering rules, or legitimate software development. For instance, a request to generate code for process injection, a common malicious technique, can be reframed as a request to write Sigma rules for detecting that very process injection. The safety classifier, seeing a request for defensive rule generation, would likely approve it, unaware that the underlying intent is to extract exploitation mechanics.
Once the benign subtasks are approved and their outputs collected, an unclassified, smaller LLM can then reassemble these pieces. This recomposition step occurs outside the observation boundary of the safety classifier, allowing the attacker to construct functional malicious code or exploit details. CrowdStrike demonstrated this with a proof-of-concept that generated a complete C program for Windows remote process injection by combining outputs from three benignly reframed requests.
Beyond process injection, the technique is effective for developing exploits for known vulnerabilities (CVEs). By asking the LLM to generate detection rules for a specific CVE, attackers can extract detailed exploitation mechanics. This information can then be used by a local, unclassified model to synthesize working proof-of-concept exploits. This method effectively weaponizes the LLM's knowledge base for offensive purposes, bypassing the intended safety guardrails.
The implications of this bypass technique are significant for the burgeoning field of AI security. As LLMs become more integrated into development workflows and security operations, understanding and mitigating these emergent vulnerabilities is crucial. The ability for adversaries to systematically circumvent safety classifiers poses a direct threat, potentially enabling the creation and dissemination of more sophisticated malware, exploit code, and malicious content.
Defenders must adapt their strategies to account for these advanced adversarial tactics. This includes developing methods to detect sequential request patterns indicative of decomposition attacks, enhancing LLM safety architectures to consider conversational context rather than just individual prompts, and fostering ongoing research into the evolving threat landscape of AI-powered attacks. The parallel discovery by independent research teams underscores the structural nature of this vulnerability and the urgent need for robust defenses.