JFrog Security Research disclosed CVE-2026-105192 on October 7 — a CVSS 9.8 unauthenticated remote code execution flaw in LMCache, the distributed KV-cache layer used alongside vLLM at Google Cloud GKE Inference, CoreWeave, NVIDIA Dynamo, and IBM’s LLM serving stack. No patch exists. A single ZeroMQ message to the default port 5555 executes arbitrary code as root on affected systems. If you run LMCache in a multi-node or Kubernetes deployment, check your firewall rules before reading further.
How the Attack Works
LMCache’s multiprocess mode runs as a standalone server that vLLM worker replicas connect to for KV-cache sharing. To coordinate this, it opens a ZeroMQ ROUTER socket on port 5555 with no authentication, no encryption, and no message validation. When a message arrives with msgpack extension code 1, the server calls pickle.loads() directly on the incoming data — before checking what the message actually contains.
Python’s pickle module can carry arbitrary executable code. Deserializing attacker-controlled pickle data is therefore equivalent to running attacker-supplied Python. According to JFrog researcher Yuval Moravchick, who discovered the flaw: “The server unpacks messages before validating their type.” A public proof-of-concept appeared on GitHub the same day as disclosure. Official LMCache container images run as root, so exploitation means complete system compromise — not merely container escape, but root on the host.
This is not a novel attack class. Pickle deserialization on untrusted network input has been a documented critical risk since 2011. That it appears in 2026 production AI infrastructure is the part that warrants attention.
Related: Pwn2Own 2026: AI Infrastructure Falls, LiteLLM Hit Twice
Who Is Actually at Risk
Single-node deployments are not remotely exploitable. By default, LMCache’s ZMQ socket binds to 127.0.0.1 — an attacker would need local access to exploit it. However, the danger is concentrated in multi-node and Kubernetes deployments, where operators pass --host <routable_address> to expose the LMCache server across a cluster. That flag is explicitly recommended in official deployment guides for the vLLM production-stack.
LMCache is not a research project. It graduated to production in January 2026, joined the PyTorch Foundation in October 2025, and now powers inference at Google Cloud GKE Inference, CoreWeave (Cohere’s production infra), NVIDIA Dynamo, and IBM’s open-source LLM serving stack. The project has 9.2K+ GitHub stars. Anyone running the official vLLM production-stack uses LMCache by default. Consequently, if your team deployed that stack on Kubernetes and followed the documented networking guidance, your LMCache port may be reachable inside your cluster namespace — and potentially beyond it, depending on your NetworkPolicy configuration.
What to Do Right Now
There is no patch. All versions from 0.3.9 through 0.5.5 are affected, including recent release candidates. LMCache has not published a security advisory or patch timeline as of October 8. The mitigations are operational, not code-based.
Three actions to take immediately. First, block port 5555 at the network level for any traffic from untrusted sources — firewall rules or Kubernetes NetworkPolicy. Second, audit whether your LMCache deployment uses --host with a routable address; if so, switch to 127.0.0.1 until a patch is available. Third, confirm that your NetworkPolicy restricts LMCache pod access to only the vLLM worker pods that need it. JFrog’s full advisory contains additional hardening recommendations.
There is also the question of what comes next. A GitHub user reported six additional LMCache security issues on October 6 — the day before this CVE went public. None have CVE assignments or maintainer confirmation yet. The security picture may be incomplete.
The Broader Pattern in AI Infrastructure
This CVE landed during the same week as Pwn2Own Ireland’s first-ever AI Infrastructure category, where LiteLLM was cracked twice and OpenAI Codex was compromised. Those were competition findings in a controlled environment. CVE-2026-105192 is an unpatched production vulnerability in software that Google Cloud deploys for customers today. The common thread is that LLM inference infrastructure is scaling faster than the security practices around it.
Teams using managed cloud inference may be partially insulated by their provider’s network controls. Teams running self-hosted vLLM stacks — the majority of high-volume enterprise deployments — are responsible for that exposure themselves. The AI serving stack has become critical infrastructure. It is being audited accordingly, whether teams are ready for that or not.
Key Takeaways
- CVE-2026-105192 is a CVSS 9.8 unpatched RCE in LMCache (versions 0.3.9–0.5.5) exploitable via pickle deserialization on an unauthenticated ZeroMQ socket on port 5555
- Only multi-node and Kubernetes deployments using a routable
--hostbinding are remotely exploitable; single-node localhost setups are not - No patch exists as of October 8 — block port 5555, avoid routable
--hostbindings, and audit Kubernetes NetworkPolicy immediately - Six additional LMCache security reports were filed the day before this CVE; the full security surface may not yet be known













