vLLM's prompt-embeds ingestion path is vulnerable to a race condition: when multiple prompt_embeds parts are submitted concurrently to POST /v1/chat/completions, the process-global state used by torch.sparse.check_sparse_tensor_invariants can be manipulated by one request while another is validated, letting a malformed sparse tensor slip past the safety check added for CVE-2025-62164 and reach tensor.to_dense unchecked. For CISOs, the practical concern today is availability rather than remote code execution — an attacker who can send concurrent requests to a vLLM endpoint with enable_prompt_embeds turned on could trigger crashes or undefined behavior in the inference worker, and the flaw only affects deployments that explicitly expose prompt-embeddings ingestion, which narrows but does not eliminate exposure. EPSS sits at a low 0.0025 with no CISA KEV listing, no public PoC, and no Nuclei template, and CISA's own SSVC decision is TRACK — this is not an urgent, actively-exploited bug, but it does bypass a previously shipped security control, which is the kind of regression worth closing quickly. Patch to vLLM 0.26.0, and until upgraded, disable enable_prompt_embeds or restrict access to trusted callers only, since that flag is the sole precondition for exploitation. Monitor inference worker crash/restart logs for anomalies correlated with bursts of concurrent chat-completions requests carrying prompt_embeds payloads.
What is the risk?
Low-to-moderate exploitability due to the timing precision required to win the race, combined with the precondition that enable_prompt_embeds must be explicitly enabled — a non-default configuration limited to deployments accepting client-supplied prompt embeddings instead of raw text. EPSS (0.0025) and CISA SSVC (TRACK) both indicate low near-term exploitation likelihood, and there is no public exploit code or scanner template. However, severity if triggered is meaningful: the bug bypasses a previously shipped security fix (CVE-2025-62164), and successful exploitation reaches native tensor deserialization code (tensor.to_dense on an invalid sparse tensor), which can produce worker crashes and is the type of memory-unsafe path that historically escalates beyond simple DoS. Overall risk: MEDIUM for exposed deployments with prompt_embeds enabled, LOW for the majority of vLLM deployments running default text-only chat completions.
How does the attack unfold?
What systems are affected?
| Package | Ecosystem | Vulnerable Range | Patched |
|---|---|---|---|
| vLLM | pip | >= 0.21.0, < 0.26.0 | 0.26.0 |
Do you use vLLM? You're affected.
How severe is it?
What should I do?
1 step-
Upgrade to vLLM 0.26.0 or later, which fixes the race. If immediate upgrade isn't possible, disable enable_prompt_embeds unless prompt-embeds ingestion is a hard product requirement — this closes the only entry point for the bug. Where prompt_embeds must stay enabled, consider rate-limiting or serializing concurrent requests to a given worker, or isolating prompt_embeds-handling workers from general inference traffic to reduce the chance of a successful race. For detection, monitor vLLM worker process restarts/crashes and correlate with concurrent chat-completions traffic carrying multimodal prompt_embeds parts; review the fix commit (793cf79) and GHSA-pr7f-p5mw-fc87 for technical detail before deciding on interim mitigations.
What does CISA's SSVC say?
Source: CISA Vulnrichment (SSVC v2.0). Decision based on the CISA Coordinator decision tree.
How is it classified?
Which compliance frameworks are affected?
This CVE is relevant to:
Frequently Asked Questions
What is CVE-2026-73557?
vLLM's prompt-embeds ingestion path is vulnerable to a race condition: when multiple prompt_embeds parts are submitted concurrently to POST /v1/chat/completions, the process-global state used by torch.sparse.check_sparse_tensor_invariants can be manipulated by one request while another is validated, letting a malformed sparse tensor slip past the safety check added for CVE-2025-62164 and reach tensor.to_dense unchecked. For CISOs, the practical concern today is availability rather than remote code execution — an attacker who can send concurrent requests to a vLLM endpoint with enable_prompt_embeds turned on could trigger crashes or undefined behavior in the inference worker, and the flaw only affects deployments that explicitly expose prompt-embeddings ingestion, which narrows but does not eliminate exposure. EPSS sits at a low 0.0025 with no CISA KEV listing, no public PoC, and no Nuclei template, and CISA's own SSVC decision is TRACK — this is not an urgent, actively-exploited bug, but it does bypass a previously shipped security control, which is the kind of regression worth closing quickly. Patch to vLLM 0.26.0, and until upgraded, disable enable_prompt_embeds or restrict access to trusted callers only, since that flag is the sole precondition for exploitation. Monitor inference worker crash/restart logs for anomalies correlated with bursts of concurrent chat-completions requests carrying prompt_embeds payloads.
Is CVE-2026-73557 actively exploited?
No confirmed active exploitation of CVE-2026-73557 has been reported, but organizations should still patch proactively.
How to fix CVE-2026-73557?
Upgrade to vLLM 0.26.0 or later, which fixes the race. If immediate upgrade isn't possible, disable enable_prompt_embeds unless prompt-embeds ingestion is a hard product requirement — this closes the only entry point for the bug. Where prompt_embeds must stay enabled, consider rate-limiting or serializing concurrent requests to a given worker, or isolating prompt_embeds-handling workers from general inference traffic to reduce the chance of a successful race. For detection, monitor vLLM worker process restarts/crashes and correlate with concurrent chat-completions traffic carrying multimodal prompt_embeds parts; review the fix commit (793cf79) and GHSA-pr7f-p5mw-fc87 for technical detail before deciding on interim mitigations.
What systems are affected by CVE-2026-73557?
This vulnerability affects the following AI/ML architecture patterns: model serving, LLM inference APIs, multimodal inference pipelines.
What is the CVSS score for CVE-2026-73557?
No CVSS score has been assigned yet.
What is the AI security impact?
Affected AI Architectures
MITRE ATLAS Techniques
AML.T0029 Denial of AI Service AML.T0040 AI Model Inference API Access AML.T0049 Exploit Public-Facing Application Compliance Controls Affected
What are the technical details?
Original Advisory
vLLM is an inference and serving engine for large language models. From 0.20.2rc0 until 0.26.0, safe_load_prompt_embeds in vllm/renderers/embed_utils.py uses torch.sparse.check_sparse_tensor_invariants, whose process-global save, enable, and restore state can be raced by concurrent prompt_embeds parts submitted to POST /v1/chat/completions through AsyncMultiModalItemTracker.resolve_items, asyncio.gather, and the default executor, allowing an invalid sparse tensor to reach tensor.to_dense despite the CVE-2025-62164 guard when enable_prompt_embeds is enabled. This issue is fixed in version 0.26.0.
Exploitation Scenario
An attacker with API access to a vLLM deployment running with enable_prompt_embeds enabled fires multiple concurrent POST /v1/chat/completions requests, each carrying a prompt_embeds part — one well-formed sparse tensor and another crafted to violate sparse tensor invariants, the same class of payload CVE-2025-62164 was meant to block. Because AsyncMultiModalItemTracker.resolve_items processes these parts concurrently via asyncio.gather and the default executor, the save/enable/restore sequence around torch.sparse.check_sparse_tensor_invariants can interleave: the invariant check meant to gate the malicious tensor gets disabled or restored at the wrong moment by the other concurrent request, letting the invalid tensor reach tensor.to_dense unchecked. Depending on how the invalid sparse tensor is structured, this crashes the inference worker or triggers undefined behavior in native tensor code, disrupting the shared inference service for every tenant on that worker.
Weaknesses (CWE)
CWE-362 Concurrent Execution using Shared Resource with Improper Synchronization ('Race Condition')
Primary
CWE-362 Concurrent Execution using Shared Resource with Improper Synchronization ('Race Condition')
Primary
CWE-362 Concurrent Execution using Shared Resource with Improper Synchronization ('Race Condition') CWE-362 — Concurrent Execution using Shared Resource with Improper Synchronization ('Race Condition'): The product contains a concurrent code sequence that requires temporary, exclusive access to a shared resource, but a timing window exists in which the shared resource can be modified by another code sequence operating concurrently.
- [Architecture and Design] In languages that support it, use synchronization primitives. Only wrap these around critical code to minimize the impact on performance.
- [Architecture and Design] Use thread-safe capabilities such as the data access abstraction in Spring.
Source: MITRE CWE corpus.
References
- github.com/advisories/GHSA-pr7f-p5mw-fc87
- nvd.nist.gov/vuln/detail/CVE-2026-73557
- github.com/vllm-project/vllm/commit/793cf79c89d4049124e756915468ac30318f2e50
- github.com/vllm-project/vllm/pull/48583
- github.com/vllm-project/vllm/releases/tag/v0.26.0
- github.com/vllm-project/vllm/security/advisories/GHSA-pr7f-p5mw-fc87
Timeline
Related Vulnerabilities
CVE-2026-61732 10.0 Analysis pending
Same package: vllm CVE-2024-11041 9.8 vllm: RCE via unsafe pickle deserialization in MessageQueue
Same package: vllm CVE-2026-25960 9.8 vllm: SSRF allows internal network access
Same package: vllm CVE-2025-47277 9.8 vLLM: RCE via exposed TCPStore in distributed inference
Same package: vllm CVE-2024-9053 9.8 vllm: RCE via unsafe pickle deserialization in RPC server
Same package: vllm