Researchers Bypass LLM Token Limits Using Overlapping Fragments
A novel technique allows attackers to exfiltrate larger data tokens from large language models by exploiting overlapping fragment vulnerabilities, potentially bypassing security controls.

Security researchers have unveiled a new method for circumventing the token limits imposed on large language models (LLMs), enabling the exfiltration of significantly larger data payloads than previously thought possible. This technique, detailed by PortSwigger Research, exploits a vulnerability related to how LLMs process overlapping fragments of input data.
Large language models typically operate with strict token limits to manage computational resources and prevent abuse. These limits define the maximum amount of text a model can process or generate in a single interaction. However, the newly discovered attack vector bypasses these restrictions by carefully crafting input that causes the model to misinterpret or re-process data segments.
The core of the exploit lies in the manipulation of "overlapping fragments." When an LLM processes input, it often breaks it down into smaller pieces or "fragments." By strategically designing these fragments to overlap in specific ways, attackers can trick the model into treating a larger amount of data as a single, valid token, thereby exceeding the intended limit.
This method allows for the exfiltration of "larger tokens" than previously demonstrated in similar research. While the exact size of the exfiltrated data can vary depending on the specific LLM architecture and implementation, the implications are significant for data security and privacy. Sensitive information that should remain within the model's processing boundaries could potentially be extracted.
The research highlights a critical, yet often overlooked, class of vulnerabilities in the rapidly evolving field of artificial intelligence. As LLMs become more integrated into various applications and services, understanding and mitigating such bypass techniques is paramount to ensuring their secure deployment.
While the specific LLMs affected and the full scope of the vulnerability are still being investigated, this discovery underscores the need for robust security testing and validation of AI systems. Developers and security professionals must consider these novel attack vectors when designing, implementing, and securing AI-powered applications.
Further research is expected to explore the precise mechanisms behind this exploit and to develop effective countermeasures. The findings serve as a stark reminder that the security landscape for AI is dynamic and requires continuous vigilance and innovation to stay ahead of emerging threats.