Vast AI Relay Network Facilitates Chinese Access to Frontier US Models
An estimated 80,000 AI relay servers are being used by individuals in China to mask their access to advanced large language models hosted in the US, likely for cloning purposes.

A massive network of over 80,000 intermediary servers, known as LLM relay servers, is enabling users in China to access frontier AI models hosted in the United States while obscuring their true identities and locations. This sophisticated operation, detailed in a recent report by Team Cymru, is believed to be a systematic effort to circumvent access restrictions and facilitate the cloning of advanced AI technologies.
The scale of this relay network suggests a coordinated campaign to exfiltrate sensitive AI capabilities. Researchers identified these relay servers acting as transfer stations, pooling credentials for multiple AI accounts and routing user requests through them to services from major providers like Anthropic, OpenAI, Google, and xAI. This architecture fundamentally breaks the assumption that the account making a request belongs to the party consuming the AI's output, undermining controls such as account attribution, usage metering, and abuse detection.
Team Cymru's analysis initially identified over 10,000 such transfer stations, a figure later revised to more than 80,000 relays. A significant portion of the traffic originated from over 4,000 IP addresses in China and Hong Kong connecting to relay servers hosted by US virtual private server providers. Over an eight-day period, these addresses transmitted approximately 14 terabytes of data to the relay stations and received more than 7 terabytes in return.
The observed traffic patterns, particularly a high upload-to-download ratio when interacting with Anthropic's API, are consistent with large-scale model distillation attempts. This technique involves using the outputs of a frontier model as training data to develop a less capable but significantly cheaper version that mimics its behavior. Scott Fisher, a Team Cymru researcher, noted that Chinese users are not only violating access blocks but also potentially stealing outputs to clone AI models.
Further investigation revealed that two open-source software packages, Claude Relay Service (CRS 1.x) and its successor sub2api, both published on GitHub by developer Wei-Shaw, are being used to power many of these relay servers. The newer sub2api software offers features like user management, per-user billing, and prompt auditing, indicating a robust infrastructure for managing access and potentially commercializing the illicitly obtained AI capabilities.
The widespread interest in this software is evident from its numerous forks on GitHub and a large subscriber base on its associated Telegram channel. While these figures don't confirm the exact number of active malicious users, they highlight the significant demand and adoption of tools designed to facilitate unauthorized access to and exfiltration of AI model outputs.
This discovery comes shortly after US government agencies issued a warning about China's systematic efforts to extract frontier AI capabilities through distillation. The findings underscore the growing threat posed by nation-state actors and sophisticated criminal organizations seeking to acquire advanced AI technology through illicit means, potentially impacting national security and technological competitiveness.