CVE-2026-73557: vLLM: race condition bypasses prompt-embeds tensor guard

GHSA-pr7f-p5mw-fc87 MEDIUM
Published August 13, 2026
CISO Take

vLLM's prompt-embeds ingestion path is vulnerable to a race condition: when multiple prompt_embeds parts are submitted concurrently to POST /v1/chat/completions, the process-global state used by torch.sparse.check_sparse_tensor_invariants can be manipulated by one request while another is validated, letting a malformed sparse tensor slip past the safety check added for CVE-2025-62164 and reach tensor.to_dense unchecked. For CISOs, the practical concern today is availability rather than remote code execution — an attacker who can send concurrent requests to a vLLM endpoint with enable_prompt_embeds turned on could trigger crashes or undefined behavior in the inference worker, and the flaw only affects deployments that explicitly expose prompt-embeddings ingestion, which narrows but does not eliminate exposure. EPSS sits at a low 0.0025 with no CISA KEV listing, no public PoC, and no Nuclei template, and CISA's own SSVC decision is TRACK — this is not an urgent, actively-exploited bug, but it does bypass a previously shipped security control, which is the kind of regression worth closing quickly. Patch to vLLM 0.26.0, and until upgraded, disable enable_prompt_embeds or restrict access to trusted callers only, since that flag is the sole precondition for exploitation. Monitor inference worker crash/restart logs for anomalies correlated with bursts of concurrent chat-completions requests carrying prompt_embeds payloads.

Sources: NVD GitHub Advisory EPSS ATLAS CISA SSVC

What is the risk?

Low-to-moderate exploitability due to the timing precision required to win the race, combined with the precondition that enable_prompt_embeds must be explicitly enabled — a non-default configuration limited to deployments accepting client-supplied prompt embeddings instead of raw text. EPSS (0.0025) and CISA SSVC (TRACK) both indicate low near-term exploitation likelihood, and there is no public exploit code or scanner template. However, severity if triggered is meaningful: the bug bypasses a previously shipped security fix (CVE-2025-62164), and successful exploitation reaches native tensor deserialization code (tensor.to_dense on an invalid sparse tensor), which can produce worker crashes and is the type of memory-unsafe path that historically escalates beyond simple DoS. Overall risk: MEDIUM for exposed deployments with prompt_embeds enabled, LOW for the majority of vLLM deployments running default text-only chat completions.

How does the attack unfold?

Initial Access
Attacker with API access sends multiple concurrent POST /v1/chat/completions requests carrying prompt_embeds parts to a vLLM deployment with enable_prompt_embeds turned on.
AML.T0049
Guard Bypass via Race
Concurrent processing through AsyncMultiModalItemTracker.resolve_items and asyncio.gather races the global save/enable/restore state of torch.sparse.check_sparse_tensor_invariants, letting a malformed sparse tensor skip the CVE-2025-62164 validation.
AML.T0040
Impact
The invalid sparse tensor reaches tensor.to_dense unchecked, crashing the inference worker or causing undefined behavior and disrupting service for all users on that worker.
AML.T0029

What systems are affected?

Package Ecosystem Vulnerable Range Patched
vLLM pip >= 0.21.0, < 0.26.0 0.26.0
92.7K 95 dependents Pushed 5d ago 26% patched ~47d to patch Full package profile →

Do you use vLLM? You're affected.

How severe is it?

CVSS 3.1
N/A
EPSS
0.4%
chance of exploitation in 30 days
Higher than 32% of all CVEs
Exploitation Status
No known exploitation
Sophistication
Advanced

What should I do?

1 step
  1. Upgrade to vLLM 0.26.0 or later, which fixes the race. If immediate upgrade isn't possible, disable enable_prompt_embeds unless prompt-embeds ingestion is a hard product requirement — this closes the only entry point for the bug. Where prompt_embeds must stay enabled, consider rate-limiting or serializing concurrent requests to a given worker, or isolating prompt_embeds-handling workers from general inference traffic to reduce the chance of a successful race. For detection, monitor vLLM worker process restarts/crashes and correlate with concurrent chat-completions traffic carrying multimodal prompt_embeds parts; review the fix commit (793cf79) and GHSA-pr7f-p5mw-fc87 for technical detail before deciding on interim mitigations.

What does CISA's SSVC say?

Decision Track
Exploitation none
Automatable No
Technical Impact partial

Source: CISA Vulnrichment (SSVC v2.0). Decision based on the CISA Coordinator decision tree.

How is it classified?

Which compliance frameworks are affected?

This CVE is relevant to:

EU AI Act
Article 15 - Accuracy, Robustness and Cybersecurity
ISO 42001
Annex A.6.2 - AI system operation and monitoring
NIST AI RMF
MEASURE 2.7 - AI system security and resilience is evaluated and documented
OWASP LLM Top 10
LLM10:2025 - Unbounded Consumption

Frequently Asked Questions

What is CVE-2026-73557?

vLLM's prompt-embeds ingestion path is vulnerable to a race condition: when multiple prompt_embeds parts are submitted concurrently to POST /v1/chat/completions, the process-global state used by torch.sparse.check_sparse_tensor_invariants can be manipulated by one request while another is validated, letting a malformed sparse tensor slip past the safety check added for CVE-2025-62164 and reach tensor.to_dense unchecked. For CISOs, the practical concern today is availability rather than remote code execution — an attacker who can send concurrent requests to a vLLM endpoint with enable_prompt_embeds turned on could trigger crashes or undefined behavior in the inference worker, and the flaw only affects deployments that explicitly expose prompt-embeddings ingestion, which narrows but does not eliminate exposure. EPSS sits at a low 0.0025 with no CISA KEV listing, no public PoC, and no Nuclei template, and CISA's own SSVC decision is TRACK — this is not an urgent, actively-exploited bug, but it does bypass a previously shipped security control, which is the kind of regression worth closing quickly. Patch to vLLM 0.26.0, and until upgraded, disable enable_prompt_embeds or restrict access to trusted callers only, since that flag is the sole precondition for exploitation. Monitor inference worker crash/restart logs for anomalies correlated with bursts of concurrent chat-completions requests carrying prompt_embeds payloads.

Is CVE-2026-73557 actively exploited?

No confirmed active exploitation of CVE-2026-73557 has been reported, but organizations should still patch proactively.

How to fix CVE-2026-73557?

Upgrade to vLLM 0.26.0 or later, which fixes the race. If immediate upgrade isn't possible, disable enable_prompt_embeds unless prompt-embeds ingestion is a hard product requirement — this closes the only entry point for the bug. Where prompt_embeds must stay enabled, consider rate-limiting or serializing concurrent requests to a given worker, or isolating prompt_embeds-handling workers from general inference traffic to reduce the chance of a successful race. For detection, monitor vLLM worker process restarts/crashes and correlate with concurrent chat-completions traffic carrying multimodal prompt_embeds parts; review the fix commit (793cf79) and GHSA-pr7f-p5mw-fc87 for technical detail before deciding on interim mitigations.

What systems are affected by CVE-2026-73557?

This vulnerability affects the following AI/ML architecture patterns: model serving, LLM inference APIs, multimodal inference pipelines.

What is the CVSS score for CVE-2026-73557?

No CVSS score has been assigned yet.

What is the AI security impact?

Affected AI Architectures

model servingLLM inference APIsmultimodal inference pipelines

MITRE ATLAS Techniques

AML.T0029 Denial of AI Service
AML.T0040 AI Model Inference API Access
AML.T0049 Exploit Public-Facing Application

Compliance Controls Affected

EU AI Act: Article 15
ISO 42001: Annex A.6.2
NIST AI RMF: MEASURE 2.7
OWASP LLM Top 10: LLM10:2025

What are the technical details?

Original Advisory

vLLM is an inference and serving engine for large language models. From 0.20.2rc0 until 0.26.0, safe_load_prompt_embeds in vllm/renderers/embed_utils.py uses torch.sparse.check_sparse_tensor_invariants, whose process-global save, enable, and restore state can be raced by concurrent prompt_embeds parts submitted to POST /v1/chat/completions through AsyncMultiModalItemTracker.resolve_items, asyncio.gather, and the default executor, allowing an invalid sparse tensor to reach tensor.to_dense despite the CVE-2025-62164 guard when enable_prompt_embeds is enabled. This issue is fixed in version 0.26.0.

Exploitation Scenario

An attacker with API access to a vLLM deployment running with enable_prompt_embeds enabled fires multiple concurrent POST /v1/chat/completions requests, each carrying a prompt_embeds part — one well-formed sparse tensor and another crafted to violate sparse tensor invariants, the same class of payload CVE-2025-62164 was meant to block. Because AsyncMultiModalItemTracker.resolve_items processes these parts concurrently via asyncio.gather and the default executor, the save/enable/restore sequence around torch.sparse.check_sparse_tensor_invariants can interleave: the invariant check meant to gate the malicious tensor gets disabled or restored at the wrong moment by the other concurrent request, letting the invalid tensor reach tensor.to_dense unchecked. Depending on how the invalid sparse tensor is structured, this crashes the inference worker or triggers undefined behavior in native tensor code, disrupting the shared inference service for every tenant on that worker.

Weaknesses (CWE)

CWE-362 — Concurrent Execution using Shared Resource with Improper Synchronization ('Race Condition'): The product contains a concurrent code sequence that requires temporary, exclusive access to a shared resource, but a timing window exists in which the shared resource can be modified by another code sequence operating concurrently.

  • [Architecture and Design] In languages that support it, use synchronization primitives. Only wrap these around critical code to minimize the impact on performance.
  • [Architecture and Design] Use thread-safe capabilities such as the data access abstraction in Spring.

Source: MITRE CWE corpus.

Timeline

Published
August 13, 2026
Last Modified
September 4, 2026
First Seen
August 13, 2026

Related Vulnerabilities