CVE-2026-76850: LMDeploy: RCE via unauth pickle deserialization
CRITICAL PoC AVAILABLE CISA: ATTENDLMDeploy's disaggregated-serving mode deserializes peer-to-peer messages with Python's pickle before any type validation occurs, and the two HTTP endpoints that establish that peer connection ship with no authentication unless an operator explicitly sets api_keys — which is not the default. A remote, unauthenticated attacker can point a vulnerable engine at a ZMQ endpoint they control and achieve arbitrary code execution in the inference engine process, a textbook CVSS 9.8 with no user interaction required. There is no CISA KEV listing, no EPSS score, and no public exploit or Nuclei template yet, so this hasn't been weaponized in the wild as of this writing — but the technique (pickle deserialization RCE) is well-documented and trivial to operationalize once a target is fingerprinted, so the absence of a scanner is a lead-time window, not a reason to deprioritize. Only deployments running LMDeploy's disaggregated (prefill/decode-split) serving mode are exposed, since the vulnerable receive loop only starts once a migration backend connection is accepted. Patch to v0.16.0 immediately, or as a compensating control set api_keys and restrict network access to the /distserve/p2p_initialize and /distserve/p2p_connect endpoints and the ZMQ ports to a trusted internal network only.
What is the risk?
Critical. The vulnerability combines the three conditions that define worst-case exploitability: network-reachable attack vector, no authentication required by default, and no user interaction — all confirmed by the CVSS 3.1 vector (AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H). The isinstance() type check that was meant to constrain deserialized objects executes only after pickle.loads() has already run, which is a textbook bypass of a security control placed after the dangerous operation instead of before it. Mitigating factors are narrow: exploitation requires the target to have disaggregated (prefill/decode-split) serving enabled, which is an opt-in deployment topology rather than the default single-process mode, and no confirmed public exploit or scanner template exists yet. That said, pickle-based RCE is one of the most well-understood exploitation primitives in the Python ecosystem, so time-to-weaponization is likely short once attackers fingerprint LMDeploy disaggregated deployments at scale (e.g., via Shodan/Censys on exposed inference APIs).
How does the attack unfold?
What systems are affected?
| Package | Ecosystem | Vulnerable Range | Patched |
|---|---|---|---|
| LMDeploy | pip | — | No patch |
Do you use LMDeploy? You're affected.
How severe is it?
What is the attack surface?
What should I do?
1 step-
1) Patch to LMDeploy v0.16.0 or later, which addresses the unsafe pickle path (see fix commit f05b4ad and the v0.16.0 release notes). 2) Until patched, set the
api_keysparameter when starting the API server — it is None/disabled by default — to require authentication on the /distserve/p2p_initialize and /distserve/p2p_connect endpoints. 3) Network-isolate disaggregated-serving deployments: restrict the API server's distserve endpoints and the engine's ZMQ ports to a private, trusted network segment (VPC-internal only, no public ingress) — these endpoints have no business being internet-facing. 4) Detection: monitor for engine processes making outbound ZMQ PULL connections to endpoints outside your known engine fleet, and alert on unexpected child processes or network connections spawned from the LMDeploy engine process (a strong RCE indicator). 5) If disaggregated serving is not in active use, no action is required beyond patching on your normal upgrade cycle, since the vulnerable code path never starts.
What does CISA's SSVC say?
Source: CISA Vulnrichment (SSVC v2.0). Decision based on the CISA Coordinator decision tree.
How is it classified?
Which compliance frameworks are affected?
This CVE is relevant to:
Frequently Asked Questions
What is CVE-2026-76850?
LMDeploy's disaggregated-serving mode deserializes peer-to-peer messages with Python's pickle before any type validation occurs, and the two HTTP endpoints that establish that peer connection ship with no authentication unless an operator explicitly sets api_keys — which is not the default. A remote, unauthenticated attacker can point a vulnerable engine at a ZMQ endpoint they control and achieve arbitrary code execution in the inference engine process, a textbook CVSS 9.8 with no user interaction required. There is no CISA KEV listing, no EPSS score, and no public exploit or Nuclei template yet, so this hasn't been weaponized in the wild as of this writing — but the technique (pickle deserialization RCE) is well-documented and trivial to operationalize once a target is fingerprinted, so the absence of a scanner is a lead-time window, not a reason to deprioritize. Only deployments running LMDeploy's disaggregated (prefill/decode-split) serving mode are exposed, since the vulnerable receive loop only starts once a migration backend connection is accepted. Patch to v0.16.0 immediately, or as a compensating control set api_keys and restrict network access to the /distserve/p2p_initialize and /distserve/p2p_connect endpoints and the ZMQ ports to a trusted internal network only.
Is CVE-2026-76850 actively exploited?
Proof-of-concept exploit code is publicly available for CVE-2026-76850, increasing the risk of exploitation.
How to fix CVE-2026-76850?
1) Patch to LMDeploy v0.16.0 or later, which addresses the unsafe pickle path (see fix commit f05b4ad and the v0.16.0 release notes). 2) Until patched, set the `api_keys` parameter when starting the API server — it is None/disabled by default — to require authentication on the /distserve/p2p_initialize and /distserve/p2p_connect endpoints. 3) Network-isolate disaggregated-serving deployments: restrict the API server's distserve endpoints and the engine's ZMQ ports to a private, trusted network segment (VPC-internal only, no public ingress) — these endpoints have no business being internet-facing. 4) Detection: monitor for engine processes making outbound ZMQ PULL connections to endpoints outside your known engine fleet, and alert on unexpected child processes or network connections spawned from the LMDeploy engine process (a strong RCE indicator). 5) If disaggregated serving is not in active use, no action is required beyond patching on your normal upgrade cycle, since the vulnerable code path never starts.
What systems are affected by CVE-2026-76850?
This vulnerability affects the following AI/ML architecture patterns: model serving, distributed/disaggregated LLM inference, LLM inference pipelines.
What is the CVSS score for CVE-2026-76850?
CVE-2026-76850 has a CVSS v3.1 base score of 9.8 (CRITICAL). The EPSS exploitation probability is 1.27%.
What is the AI security impact?
Affected AI Architectures
MITRE ATLAS Techniques
AML.T0044 Full AI Model Access AML.T0049 Exploit Public-Facing Application AML.T0050 Command and Scripting Interpreter Compliance Controls Affected
What are the technical details?
Original Advisory
LMDeploy deserializes disaggregated-serving peer messages with pickle. The handle_zmq_recv coroutine in lmdeploy/pytorch/disagg/conn/engine_conn.py reads peer-to-peer cache-free requests with recv_pyobj(), which deserializes the received bytes with pickle.loads(), and the isinstance check against DistServeCacheFreeRequest runs only after deserialization has already completed. The peer that supplies those bytes is caller-controlled: p2p_connect passes remote_engine_endpoint_info.zmq_address from the request body to connect() on the ZMQ PULL socket, and the POST /distserve/p2p_initialize and /distserve/p2p_connect endpoints in lmdeploy/serve/openai/api_server.py apply no authentication unless the server is started with api_keys, which defaults to None. A remote attacker can direct an engine to pull from a ZMQ endpoint under their control and execute arbitrary code in the engine process. Deployments that do not enable disaggregated serving are not affected, because the receive loop is only started once the migration backend accepts the connection.
Exploitation Scenario
An attacker scans for internet- or network-exposed LMDeploy API servers running disaggregated serving (identifiable via the /distserve/* endpoints responding). Finding one without api_keys configured, the attacker stands up a rogue ZMQ PULL-compatible endpoint and crafts a malicious pickle payload using a standard __reduce__-based gadget that executes an OS command (e.g., a reverse shell) upon deserialization. The attacker then sends a POST request to /distserve/p2p_initialize followed by /distserve/p2p_connect, supplying their attacker-controlled zmq_address in remote_engine_endpoint_info within the request body. The victim engine's p2p_connect logic dutifully connects its ZMQ PULL socket to the attacker's endpoint and starts the handle_zmq_recv coroutine, which calls recv_pyobj() and pickle.loads() on whatever bytes the attacker sends — executing the attacker's payload before the isinstance(DistServeCacheFreeRequest) check ever runs. The attacker now has a shell inside the inference engine process, with access to the loaded model, GPU resources, and any credentials or lateral network paths available to that host.
Weaknesses (CWE)
CWE-502 — Deserialization of Untrusted Data: The product deserializes untrusted data without sufficiently ensuring that the resulting data will be valid.
- [Architecture and Design, Implementation] If available, use the signing/sealing features of the programming language to assure that deserialized data has not been tainted. For example, a hash-based message authentication code (HMAC) could be used to ensure that data has not been modified.
- [Implementation] When deserializing data, populate a new object rather than just deserializing. The result is that the data flows through safe input validation and that the functions are safe.
Source: MITRE CWE corpus.
CVSS Vector
CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H References
- github.com/InternLM/lmdeploy
- github.com/InternLM/lmdeploy/blob/v0.15.0/lmdeploy/pytorch/disagg/conn/engine_conn.py
- github.com/InternLM/lmdeploy/blob/v0.15.0/lmdeploy/pytorch/disagg/conn/engine_conn.py
- github.com/InternLM/lmdeploy/commit/f05b4ad8bf2e2d84101a1d63b3c44fadd99223b2
- github.com/InternLM/lmdeploy/issues/4804
- github.com/InternLM/lmdeploy/releases/tag/v0.16.0
- vulncheck.com/advisories/lmdeploy-remote-code-execution-via-unsafe-pickle-deserialization-in-the-disaggregated-serving-peer-connector
Timeline
Related Vulnerabilities
CVE-2025-66455 9.8 Analysis pending
Same package: lmdeploy CVE-2025-59953 9.8 Analysis pending
Same package: lmdeploy CVE-2025-67729 8.8 lmdeploy: Deserialization enables RCE
Same package: lmdeploy CVE-2026-33625 8.8 Analysis pending
Same package: lmdeploy CVE-2026-63764 8.6 lmdeploy: SSRF via redirect bypasses private IP guard
Same package: lmdeploy