Vulnerabilities
LMCache flaw runs code from one network message, no fix yet
CVE-2026-105192 (CVSS 9.8) lets anyone who can reach LMCache's multiprocess port run code, as root in official images. No patch exists. Workarounds and checks.
LMCache is an open source KV cache layer for vLLM, the open source large language model (LLM) inference server: it stores the attention state of prompts the model has already processed, so repeated context does not have to be computed again. JFrog Security Research published CVE-2026-105192 on October 7, rated CVSS 9.8. In LMCache's multiprocess mode, a single unauthenticated message to the cache server's ZeroMQ port, 5555 by default, runs code as the user of the LMCache process. On the project's official container images that user is root.
There is no patched release. Every version from 0.3.9 onwards is affected, including 0.5.5 (the latest on PyPI), the 0.5.6 release candidates up to 0.5.6rc3 and the development branch as of October 7, and LMCache has not published its own advisory. If you run LMCache as a separate server so that several vLLM instances or nodes can share a cache, and that server listens on anything other than localhost, treat it as exposed today and restrict who can reach it.
Who is exposed
LMCache can run in two ways. Inside the vLLM process, through the in-process connector, it opens no network port and this flaw does not apply. In multiprocess mode, started with lmcache server, it runs as a standalone daemon that vLLM connects to over ZeroMQ, a message-passing library. The LMCache documentation describes this mode as the way to share one cache across several vLLM pods on a node and across instances.
The transport binds to localhost by default, and a default single-host install cannot be reached from other machines. JFrog's 9.8 score applies when an operator passes a routable address with --host, which is what multi-node deployments do so that peers can connect. In that setup, any host that can open a TCP connection to the port can run code.
How it works
The multiprocess server listens on a ZeroMQ ROUTER socket, a socket type that accepts messages from many clients and routes replies back to each one. An attacker talks to it with an ordinary DEALER client socket, and JFrog found no authentication of any kind on it: no CURVE encryption, no ZAP (ZeroMQ's authentication protocol), no password and no per-message signature.
Requests are encoded with msgpack, a compact binary format that lets an application register extension types for its own objects. LMCache registers extension code 1 for DeviceIPCWrapper, the object that carries GPU memory handles between processes. When the server meets that code, ext_hook in lmcache/v1/multiprocess/custom_types.py hands the bytes to DeviceIPCWrapper.Deserialize in lmcache/v1/platform/base/ipc_wrapper.py, which calls Python's pickle.loads. Pickle is Python's native object format, and loading a pickle can call any function the data names, so an attacker who controls the bytes controls what runs.
The order of operations is what makes this a one-message attack. The decode happens while the server is still unpacking the arguments of a REGISTER_KV_CACHE request, before the request handler runs and before any check on message type or caller. The attacker's command has already executed by the time the server notices that the argument is the wrong type and logs an error.
What attackers are doing
As of October 8, no public source reports exploitation of CVE-2026-105192, and CISA has not listed it in its Known Exploited Vulnerabilities (KEV) catalog. The advisory explains the vulnerable code path down to the file and function, and building a payload from that is simple for anyone familiar with pickle, so the lack of reports should not slow down the workaround.
LMCache had other security reports the same week. A GitHub user opened six reports on October 6 alleging unauthenticated access to cached data across tenants and to other LMCache network services. Separate CVEs have also been published for LMCache, among them CVE-2026-107204, a separate code execution flaw in a /run_script endpoint that affects versions up to 0.5.5, and CVE-2026-107206 for the multiprocess HTTP server, which listens on 0.0.0.0:8080 by default. The maintainers have not confirmed the GitHub reports, and none of these issues has a fix either.
What to do
1. Find every multiprocess server
On each GPU host, run ss -ltnp 'sport = :5555' and look at the local address. 127.0.0.1:5555 is the safe default; 0.0.0.0:5555 or a node IP means the port can be reached from the network. Check the server's command line (ps -ef | grep 'lmcache server') and your container specs for --host, and for --port in case someone moved the port. In Kubernetes, look for Services, hostNetwork: true pods and hostPort entries that publish the LMCache port.
2. Workarounds until a fix ships
JFrog's interim advice is to stop setting --host to a routable address, and to keep the port on localhost or on a trusted cluster network. If your vLLM instances and the LMCache server share a node, bind to localhost. If you only need offload on a single node and do not need a shared server, the in-process connector avoids the port entirely.
Where cross-node sharing is a requirement, allow port 5555 only from the specific inference hosts that use it, with a host firewall rule or a Kubernetes NetworkPolicy, and block it everywhere else. JFrog notes this lowers exposure without removing it: any host that is allowed to connect can still run code. Two further steps of our own limit the damage. Run the LMCache container as a non-root user (runAsNonRoot: true with an explicit runAsUser), because the official images default to root. And apply the same restriction to port 8080, given the separate HTTP server report.
3. Check whether you were hit
JFrog does not publish indicators, but the exploit leaves one trace in normal operation: the server logs a type error for REGISTER_KV_CACHE after the payload runs, because the handler expected a DeviceIPCWrapper and received something else. Search LMCache logs, or kubectl logs for the LMCache pods, for that error. Legitimate vLLM clients send valid wrappers, so any such error deserves a closer look.
Also review connections to port 5555 from addresses that are not your inference hosts, child processes of the LMCache process such as shells, curl or wget, and new outbound connections from GPU nodes. If anything looks wrong, rebuild the container or host from a clean image and rotate what the process could read: Hugging Face and model registry tokens, cloud and object storage credentials used for the cache's storage tier, and any Kubernetes service account token mounted in the pod.
4. Watch for the fix
JFrog's recommended code fix is to replace pickle behind extension code 1 with a safe format, authenticate the transport with CURVE or a per-message HMAC (a keyed hash), and refuse to bind a routable address without authentication. Follow the LMCache GitHub repository and PyPI releases, and upgrade once a release explicitly lists CVE-2026-105192. Until then, keep the workaround in place.
The wider lesson
Inference stacks are being assembled from research code that assumed a trusted cluster: internal ZeroMQ or HTTP ports with no authentication, pickle on the wire and root containers. LMDeploy (CVE-2026-76850) and ktransformers (CVE-2026-63767) had the same pickle over ZeroMQ pattern earlier this year. When an AI platform team adds a component like LMCache, list every port it opens and what authenticates each one, and put that in the deployment review before the first multi-node rollout.
Our AI-native systems reviews map the services an inference stack exposes and test them the way an attacker would, and our cloud and Kubernetes security work checks which pod ports are reachable across the cluster. Open the chat and Yaali, our AI agent, will pass your question to the engineer who would do the work.
Sources: JFrog Security Research advisory, The Hacker News, Strix CVE entry, The Hacker Wire, LMCache multiprocess mode docs, LMCache issue 5511, The Hacker Wire on CVE-2026-107204, The Hacker Wire on CVE-2026-107206.