CVE-2026-73558: vLLM: integer overflow leaks other users' output
GHSA-7m6h-x95x-82q5 MEDIUM CISA: TRACK*A bug in vLLM's fused activation CUDA kernel (act_and_mul_kernel) lets an integer overflow in batch-index arithmetic cause one user's inference request to read memory belonging to another user processed in the same GPU batch, exposing partial or complete copies of a stranger's prompt or model output. This matters for any organization running a shared, multi-tenant vLLM inference endpoint — internal LLM gateways, RAG backends, or customer-facing AI APIs — where a request from one tenant could unintentionally leak another tenant's data purely by co-existing in the same batch. Exploitability is limited today: no public PoC, no Nuclei template, not in CISA KEV, and CVSS rates attack complexity High with required user interaction, placing it in the top 82% of EPSS scores but still a low absolute probability (0.26%). The fix is to upgrade to vLLM 0.27.0 immediately on any multi-tenant deployment; until patched, avoid co-batching untrusted requests from different trust boundaries (e.g., disable request batching across tenants or isolate tenants to separate vLLM instances) and monitor for anomalous or duplicated content appearing across otherwise unrelated inference responses.
What is the risk?
Medium severity (CVSS 5.3) reflecting a confidentiality-only impact (C:H, I:N, A:N) with no integrity or availability loss. Attack complexity is High and requires user interaction, meaning an adversary cannot trivially trigger the leak on demand — they need their request to land in the same inference batch as a victim's, which depends on timing, load, and batching configuration. There is no evidence of active exploitation (not in CISA KEV, SSVC decision is TRACK*), no public exploit code, and no scanner template, keeping near-term opportunistic risk low. However, the underlying flaw is a memory-safety bug in a widely deployed open-source inference engine, and any organization serving multiple tenants, customers, or business units from the same vLLM instance carries meaningful confidentiality exposure until patched.
How does the attack unfold?
What systems are affected?
| Package | Ecosystem | Vulnerable Range | Patched |
|---|---|---|---|
| vLLM | pip | < 0.27.0 | 0.27.0 |
Do you use vLLM? You're affected.
How severe is it?
What is the attack surface?
What should I do?
1 step-
Upgrade vLLM to 0.27.0 or later, which contains the fix (see commit 451227c and PR #49660). Until patched, do not run vLLM as a shared multi-tenant inference service for workloads with different confidentiality requirements — isolate tenants to separate vLLM processes/GPUs or disable dynamic batching across trust boundaries if your deployment supports it. Detection is difficult at the application layer since the leak is silent and occurs in GPU memory; as a compensating control, review inference logs for anomalies such as response content that doesn't match the corresponding prompt, or implement output validation/canary tokens per tenant to catch cross-contamination. Track the GitHub Security Advisory GHSA-7m6h-x95x-82q5 for any follow-up guidance.
What does CISA's SSVC say?
Source: CISA Vulnrichment (SSVC v2.0). Decision based on the CISA Coordinator decision tree.
How is it classified?
Which compliance frameworks are affected?
This CVE is relevant to:
Frequently Asked Questions
What is CVE-2026-73558?
A bug in vLLM's fused activation CUDA kernel (act_and_mul_kernel) lets an integer overflow in batch-index arithmetic cause one user's inference request to read memory belonging to another user processed in the same GPU batch, exposing partial or complete copies of a stranger's prompt or model output. This matters for any organization running a shared, multi-tenant vLLM inference endpoint — internal LLM gateways, RAG backends, or customer-facing AI APIs — where a request from one tenant could unintentionally leak another tenant's data purely by co-existing in the same batch. Exploitability is limited today: no public PoC, no Nuclei template, not in CISA KEV, and CVSS rates attack complexity High with required user interaction, placing it in the top 82% of EPSS scores but still a low absolute probability (0.26%). The fix is to upgrade to vLLM 0.27.0 immediately on any multi-tenant deployment; until patched, avoid co-batching untrusted requests from different trust boundaries (e.g., disable request batching across tenants or isolate tenants to separate vLLM instances) and monitor for anomalous or duplicated content appearing across otherwise unrelated inference responses.
Is CVE-2026-73558 actively exploited?
No confirmed active exploitation of CVE-2026-73558 has been reported, but organizations should still patch proactively.
How to fix CVE-2026-73558?
Upgrade vLLM to 0.27.0 or later, which contains the fix (see commit 451227c and PR #49660). Until patched, do not run vLLM as a shared multi-tenant inference service for workloads with different confidentiality requirements — isolate tenants to separate vLLM processes/GPUs or disable dynamic batching across trust boundaries if your deployment supports it. Detection is difficult at the application layer since the leak is silent and occurs in GPU memory; as a compensating control, review inference logs for anomalies such as response content that doesn't match the corresponding prompt, or implement output validation/canary tokens per tenant to catch cross-contamination. Track the GitHub Security Advisory GHSA-7m6h-x95x-82q5 for any follow-up guidance.
What systems are affected by CVE-2026-73558?
This vulnerability affects the following AI/ML architecture patterns: model serving, shared/multi-tenant inference endpoints, RAG pipelines.
What is the CVSS score for CVE-2026-73558?
CVE-2026-73558 has a CVSS v3.1 base score of 5.3 (MEDIUM). The EPSS exploitation probability is 0.41%.
What is the AI security impact?
Affected AI Architectures
MITRE ATLAS Techniques
AML.T0024 Exfiltration via AI Inference API AML.T0040 AI Model Inference API Access AML.T0057 LLM Data Leakage Compliance Controls Affected
What are the technical details?
Original Advisory
vLLM is an inference and serving engine for large language models. Prior to 0.27.0, an integer overflow in blockIdx.x * 2 * d in activation_kernels.cu can cause act_and_mul_kernel to consume another batched user's input, allowing a request processed in the same inference batch to receive a partial or complete copy of another user's inference result. This issue is fixed in version 0.27.0.
Exploitation Scenario
An attacker with legitimate but low-privilege access to a shared vLLM-backed inference endpoint (e.g., a SaaS product offering LLM completions to many customers on shared GPU capacity) sends a stream of requests timed to increase the likelihood of being scheduled into the same inference batch as a target victim's request — for example, flooding the endpoint during a period when the victim is known or suspected to be active. Due to the integer overflow in blockIdx.x * 2 * d, the attacker's request occasionally reads memory belonging to the victim's batched request instead of (or in addition to) its own, and the attacker receives a partial or complete copy of the victim's prompt or generated output in their own API response. Repeated over many requests, this could allow harvesting fragments of other tenants' proprietary prompts, confidential data included in RAG context, or generated outputs meant to remain private.
Weaknesses (CWE)
CWE-190 Integer Overflow or Wraparound
Primary
CWE-190 Integer Overflow or Wraparound
Primary
CWE-190 Integer Overflow or Wraparound CWE-190 — Integer Overflow or Wraparound: The product performs a calculation that can produce an integer overflow or wraparound when the logic assumes that the resulting value will always be larger than the original value. This occurs when an integer value is incremented to a value that is too large to store in the associated representation. When this occurs, the value may become a very small or negative number.
- [Requirements] Ensure that all protocols are strictly defined, such that all out-of-bounds behavior can be identified simply, and require strict conformance to the protocol.
- [Requirements] Use a language that does not allow this weakness to occur or provides constructs that make this weakness easier to avoid. If possible, choose a language or compiler that performs automatic bounds checking.
Source: MITRE CWE corpus.
CVSS Vector
CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:H/I:N/A:N References
- github.com/advisories/GHSA-7m6h-x95x-82q5
- nvd.nist.gov/vuln/detail/CVE-2026-73558
- github.com/vllm-project/vllm/commit/451227cb3ff07989698fed982c2d3e4300257924
- github.com/vllm-project/vllm/issues/42860
- github.com/vllm-project/vllm/pull/49660
- github.com/vllm-project/vllm/releases/tag/v0.27.0
- github.com/vllm-project/vllm/security/advisories/GHSA-7m6h-x95x-82q5
Timeline
Related Vulnerabilities
CVE-2026-61732 10.0 Analysis pending
Same package: vllm CVE-2024-11041 9.8 vllm: RCE via unsafe pickle deserialization in MessageQueue
Same package: vllm CVE-2026-25960 9.8 vllm: SSRF allows internal network access
Same package: vllm CVE-2025-47277 9.8 vLLM: RCE via exposed TCPStore in distributed inference
Same package: vllm CVE-2024-9053 9.8 vllm: RCE via unsafe pickle deserialization in RPC server
Same package: vllm