Critical LMCache Flaw Allows Unauthenticated Remote Code Execution
A critical vulnerability in LMCache's multiprocess mode, affecting LLM acceleration servers like vLLM, enables unauthenticated attackers to execute arbitrary code remotely.

A critical vulnerability has been discovered in LMCache, an open-source software designed to accelerate large language model (LLM) servers such as vLLM. The flaw, identified in LMCache's multiprocess mode, allows unauthenticated attackers to execute arbitrary code remotely on the cache server. As of the disclosure, no patched version of the software is available, leaving affected systems exposed.
The vulnerability resides within LMCache's multiprocess mode, where the cache operates as a standalone server. LLM workers communicate with this server via the ZeroMQ messaging library. By sending specially crafted network messages, an attacker can trigger the execution of commands with the privileges of the LMCache process. This poses a significant risk, especially since JFrog's security research team noted that on official container images, this process often runs as root.
Exploitation is contingent on the server being configured to listen on a routable network address rather than the default localhost. While this default setting limits accessibility to the local machine, operators may configure it to listen on broader network interfaces, particularly in multi-node deployments where a shared cache is necessary. LMCache's own example Kubernetes deployment, for instance, starts the server listening on all network interfaces, increasing its attack surface.
JFrog disclosed the vulnerability on October 7, assigning it a severity score of 9.8 out of 10, placing it in the critical range. The flaw, tracked as CVE-2026-105192, impacts LMCache versions from 0.3.9 up to the latest stable release, 0.5.5, and is also present in release candidates and the development branch. The absence of a patched version means that organizations relying on these versions are vulnerable.
The core of the exploit lies in the ZeroMQ socket used by the multiprocess server. This socket lacks authentication, and a specific type of message is unpacked using Python's pickle module. pickle is capable of serializing and deserializing arbitrary Python objects, including executable code. The server unpacks this pickle data while still processing the message's arguments, before any validation checks are performed, allowing a malicious message to execute the sender's code.
While LMCache has not yet published a formal security advisory, JFrog recommends that operators avoid assigning the multiprocess server a routable address. Keeping the port restricted to the local machine or a trusted cluster network is advised. Implementing a firewall can reduce the risk by limiting access, but it does not eliminate it entirely, as any host that can establish a connection remains susceptible to code execution.
In addition to CVE-2026-105192, JFrog's advisory also mentions other security reports filed by a GitHub user on October 6, alleging unauthenticated access to cached data and network services. While these reports lack CVE assignments and official confirmation, they highlight broader security concerns within LMCache. Separately, a related denial-of-service vulnerability (CVE-2026-105756) in vLLM, which could crash the engine when using the LMCache multiprocess connector, has already been fixed in version 0.30.0.
The underlying issue of handing data from an unauthenticated network socket to pickle for deserialization is a recurring theme in AI inference frameworks. Researchers previously identified similar flaws in other projects in November 2025, collectively termed "ShadowMQ." The extent to which LMCache's code shares a common origin with those affected projects remains to be determined, but the pattern underscores a critical security blind spot in the rapid development of AI infrastructure.