Illustration of a padlock over a server rack, symbolizing a critical cybersecurity vulnerability in LMCache affecting LLM servers.
Uncategorized

LMCache Under Siege: Critical Unpatched Flaw Exposes LLM Servers to Remote Code Execution

Share
Share
Pinterest Hidden

A severe, unpatched vulnerability in LMCache, an open-source tool designed to accelerate large language model (LLM) servers like vLLM, poses a critical threat, enabling unauthenticated attackers to execute arbitrary code on affected cache servers. With no fix currently available, this flaw, identified as CVE-2026-105192, demands immediate attention from operators of AI infrastructure.

The Heart of the Vulnerability: Unauthenticated RCE

The core of this critical flaw lies within LMCache’s multiprocess mode. In this configuration, LMCache operates as a standalone server, communicating with LLM worker processes via the ZeroMQ messaging library. Researchers at JFrog discovered that a specially crafted network message sent to this server can trigger remote code execution (RCE) without any authentication whatsoever. The malicious code runs with the privileges of the LMCache process, which, alarmingly, often runs as root on official container images.

How the Exploit Works

The vulnerability exploits the server’s handling of specific ZeroMQ messages. One particular message type is deserialized using Python’s pickle format. Pickle, notorious for its ability to embed and execute arbitrary code during the decoding process, is processed by the LMCache server early in the message parsing sequence. This means a malicious actor can inject their code before any type-checking or authentication mechanisms are engaged, effectively running their commands on the target system.

Exposure and Risk: When Default Isn’t Enough

While LMCache’s multiprocess server defaults to listening only on the local machine (localhost), making it unreachable from external hosts, the risk escalates significantly when operators configure it to listen on a routable address. This is a common practice for multi-node deployments that require shared cache access across machines. JFrog highlights that LMCache’s own example Kubernetes deployment configures the server to listen on every network interface, directly exposing it to potential attacks. A single vLLM process running LMCache internally, however, does not open this vulnerable port.

JFrog, who disclosed the flaw on October 7, assigned it a critical severity score of 9.8 out of 10, specifically for servers exposed via a routable address. The vulnerability impacts LMCache versions from 0.3.9 (October 2025) through 0.5.5, including 0.5.6 release candidates and the development branch. The absence of a patched version leaves organizations vulnerable.

Mitigation Strategies in the Absence of a Patch

Given the lack of a current fix, JFrog strongly advises LMCache operators to avoid assigning a routable address to the multiprocess server. Instead, they recommend keeping its port confined to the local machine or within a trusted cluster network. While a firewall can reduce the attack surface by limiting who can reach the port, it does not eliminate the risk entirely, as any host capable of establishing a connection can still execute code remotely.

LMCache has yet to publish its own security advisory regarding this flaw, and JFrog’s advisory does not offer methods to detect if a server has already been compromised.

Broader Context: Related Reports and Prior Flaws

Adding to the concerns, a GitHub user surfaced six additional LMCache security reports just a day before CVE-2026-105192 went public. These reports allege unauthenticated access to cached data and several network services capable of command execution without login. While these claims lack CVE assignments, maintainer confirmation, and fixes, they underscore a potential pattern of security oversights within the project. Notably, one report highlighted a default admin HTTP server that listened on all network interfaces in version 0.5.5, a setting that has since been restricted to localhost in 0.5.6 release candidates.

Separately, a related denial-of-service (DoS) flaw (CVE-2026-105756) in vLLM, rated 6.5, has already been addressed in versions 0.30.0 and later. This bug could crash the engine via a malformed cache_salt value in deployments using the LMCache multiprocess connector, though it did not permit code execution.

The fundamental error of passing unauthenticated network data to pickle echoes similar vulnerabilities discovered in other AI inference frameworks in November 2025, collectively dubbed “ShadowMQ” flaws. While a direct link between LMCache’s codebase and these projects remains unconfirmed, the recurring pattern highlights a critical security blind spot in the broader AI ecosystem.

Stay informed on the latest cybersecurity threats impacting AI infrastructure by following us on Google News, Twitter, and LinkedIn.


For more details, visit our website.

Source: Link

Share

Leave a comment

Leave a Reply

Your email address will not be published. Required fields are marked *