CVE-2026-73559: vLLM: unbounded prompt batch triggers single-request DoS

GHSA-87x5-vmc3-756j MEDIUM
Published August 13, 2026
CISO Take

vLLM's OpenAI-compatible completions endpoint accepts a prompt field that can be a list of unbounded length, and the serving layer allocates a dedicated engine generator plus response slot for every element in that list — so a single authenticated API call can exhaust CPU, memory, async scheduling capacity, and response buffering on the inference server. This matters because vLLM underpins a large share of self-hosted LLM serving (128 downstream dependents tracked here) and exploitation needs nothing more than a valid low-privilege API key and network access (AC:L, PR:L, no user interaction) — the attack is a single oversized request, not a sustained flood. There's no public exploit or scanner template circulating yet, it isn't in CISA KEV, and EPSS sits at just 0.39% (68th percentile), so this looks like opportunistic-DoS risk rather than active exploitation — but multi-tenant deployments (shared inference gateways, RAG backends, agent frameworks routing through a common vLLM instance) turn one abusive caller into an outage for every other consumer. Patch to vLLM 0.26.0, and until then cap request body/array size and rate-limit per-API-key at the reverse proxy in front of the OpenAI-compatible endpoint.

Sources: NVD GitHub Advisory EPSS ATLAS

What is the risk?

Medium severity (CVSS 6.5, availability-only impact) reflects a real but bounded risk: the vulnerability is trivial to trigger technically (AC:L, no crafted payload beyond an oversized array) but requires authenticated access (PR:L), which limits exposure mostly to insider abuse, compromised API keys, or misconfigured multi-tenant gateways rather than fully anonymous internet-wide attackers. No active exploitation (not in CISA KEV), no public PoC/exploit, and a low EPSS score (0.39%) all point to low near-term exploitation likelihood. The real risk driver is architectural: any vLLM deployment that exposes the completions API to more than one trust boundary (SaaS inference providers, internal platform teams serving multiple product teams) faces disproportionate blast radius from a single bad actor or buggy client.

How does the attack unfold?

Initial Access
Attacker obtains or already holds a valid, low-privilege API key for a vLLM-served /v1/completions endpoint.
AML.T0040
Malicious Request Crafting
Attacker builds a single completions request with the prompt field set to an unbounded list[str] or list[list[int]] containing thousands of elements.
AML.T0034.001
Resource Exhaustion
vLLM's serving layer allocates one engine generator and response slot per prompt element, rapidly consuming CPU, memory, scheduling capacity, and response buffers.
AML.T0034
Service Impact
The vLLM instance degrades or becomes unresponsive, denying inference service to all concurrent users sharing that deployment.
AML.T0029

What systems are affected?

Package Ecosystem Vulnerable Range Patched
vLLM pip >= 0.19.0, < 0.26.0 0.26.0
92.7K 95 dependents Pushed 5d ago 26% patched ~47d to patch Full package profile →

Do you use vLLM? You're affected.

How severe is it?

CVSS 3.1
6.5 / 10
EPSS
0.6%
chance of exploitation in 30 days
Higher than 46% of all CVEs
Exploitation Status
No known exploitation
Sophistication
Trivial

What is the attack surface?

AV AC PR UI S C I A
AV Network
AC Low
PR Low
UI None
S Unchanged
C None
I None
A High

What should I do?

1 step
  1. 1) Upgrade vLLM to 0.26.0 or later, which fixes the unbounded prompt handling. 2) Until patched, enforce a maximum array length / request body size at the API gateway or reverse proxy in front of vLLM (e.g., reject completions requests with prompt lists above a sane threshold, such as 32-64 items). 3) Apply per-API-key rate limiting and concurrent-request caps so a single client cannot monopolize engine slots. 4) For multi-tenant deployments, run vLLM instances per-tenant or with resource quotas/cgroup limits so one abusive client cannot starve others. 5) Detection: monitor for anomalous spikes in request payload size, sudden increases in engine queue depth/memory usage correlated with a single API key, and CPU/memory saturation alerts on vLLM hosts without a corresponding increase in request count.

How is it classified?

Which compliance frameworks are affected?

This CVE is relevant to:

EU AI Act
Article 15 - Accuracy, robustness and cybersecurity
ISO 42001
A.6.2.4 - AI system operation and monitoring
NIST AI RMF
MANAGE-2.3 - Mechanisms for AI system availability and resilience
OWASP LLM Top 10
LLM10:2025 - Unbounded Consumption

Frequently Asked Questions

What is CVE-2026-73559?

vLLM's OpenAI-compatible completions endpoint accepts a prompt field that can be a list of unbounded length, and the serving layer allocates a dedicated engine generator plus response slot for every element in that list — so a single authenticated API call can exhaust CPU, memory, async scheduling capacity, and response buffering on the inference server. This matters because vLLM underpins a large share of self-hosted LLM serving (128 downstream dependents tracked here) and exploitation needs nothing more than a valid low-privilege API key and network access (AC:L, PR:L, no user interaction) — the attack is a single oversized request, not a sustained flood. There's no public exploit or scanner template circulating yet, it isn't in CISA KEV, and EPSS sits at just 0.39% (68th percentile), so this looks like opportunistic-DoS risk rather than active exploitation — but multi-tenant deployments (shared inference gateways, RAG backends, agent frameworks routing through a common vLLM instance) turn one abusive caller into an outage for every other consumer. Patch to vLLM 0.26.0, and until then cap request body/array size and rate-limit per-API-key at the reverse proxy in front of the OpenAI-compatible endpoint.

Is CVE-2026-73559 actively exploited?

No confirmed active exploitation of CVE-2026-73559 has been reported, but organizations should still patch proactively.

How to fix CVE-2026-73559?

1) Upgrade vLLM to 0.26.0 or later, which fixes the unbounded prompt handling. 2) Until patched, enforce a maximum array length / request body size at the API gateway or reverse proxy in front of vLLM (e.g., reject completions requests with prompt lists above a sane threshold, such as 32-64 items). 3) Apply per-API-key rate limiting and concurrent-request caps so a single client cannot monopolize engine slots. 4) For multi-tenant deployments, run vLLM instances per-tenant or with resource quotas/cgroup limits so one abusive client cannot starve others. 5) Detection: monitor for anomalous spikes in request payload size, sudden increases in engine queue depth/memory usage correlated with a single API key, and CPU/memory saturation alerts on vLLM hosts without a corresponding increase in request count.

What systems are affected by CVE-2026-73559?

This vulnerability affects the following AI/ML architecture patterns: model serving, LLM inference APIs, RAG pipelines, agent frameworks, multi-tenant inference gateways.

What is the CVSS score for CVE-2026-73559?

CVE-2026-73559 has a CVSS v3.1 base score of 6.5 (MEDIUM). The EPSS exploitation probability is 0.58%.

What is the AI security impact?

Affected AI Architectures

model servingLLM inference APIsRAG pipelinesagent frameworksmulti-tenant inference gateways

MITRE ATLAS Techniques

AML.T0029 Denial of AI Service
AML.T0034 Cost Harvesting
AML.T0034.001 Resource-Intensive Queries
AML.T0040 AI Model Inference API Access

Compliance Controls Affected

EU AI Act: Article 15
ISO 42001: A.6.2.4
NIST AI RMF: MANAGE-2.3
OWASP LLM Top 10: LLM10:2025

What are the technical details?

Original Advisory

vLLM is an inference and serving engine for large language models. From 0.19.0 until 0.26.0, the /v1/completions CompletionRequest.prompt field in vllm/entrypoints/openai/completion/protocol.py accepts an unbounded list[str] or list[list[int]], prompt_to_seq() in vllm/renderers/inputs/preprocess.py and OnlineRenderer.preprocess_completion() in vllm/renderers/online_renderer.py expand every element, and vllm/entrypoints/openai/completion/serving.py creates one engine generator and response slot per prompt, allowing an authenticated API client to exhaust CPU, memory, async scheduling capacity, engine request slots, and response buffering with one request. This issue is fixed in version 0.26.0.

Exploitation Scenario

An attacker with a valid but low-privilege API key to a vLLM-served completions endpoint sends a single POST to /v1/completions with the prompt field set to a list containing thousands of strings (or list[list[int]] token sequences). vllm/renderers preprocesses and expands every element, and the serving layer spins up one engine generator and response slot per prompt — instantly consuming available CPU, memory, and scheduling capacity meant for all concurrent users. If the deployment is multi-tenant (a shared internal LLM gateway, RAG backend serving multiple applications, or a SaaS inference provider), this single request degrades or crashes inference for every other tenant, effectively a low-effort denial-of-service against the whole platform with one HTTP call and no elevated privileges.

Weaknesses (CWE)

CWE-400 — Uncontrolled Resource Consumption: The product does not properly control the allocation and maintenance of a limited resource.

  • [Architecture and Design] Design throttling mechanisms into the system architecture. The best protection is to limit the amount of resources that an unauthorized user can cause to be expended. A strong authentication and access control model will help prevent such attacks from occurring in the first place. The login application should be protected against DoS attacks as much as possible. Limiting the database access, perhaps by caching result sets, can help minimize the resources expended. To further limit the potential for a DoS attack, consider tracking the rate of requests received from users and blocking requests that exceed a defined rate threshold.
  • [Architecture and Design] Mitigation of resource exhaustion attacks requires that the target system either: The first of these solutions is an issue in itself though, since it may allow attackers to prevent the use of the system by a particular valid user. If the attacker impersonates the valid user, they may be able to prevent the user from accessing the server in question. The second solution is simply difficult to effectively institute -- and even when properly done, it does not provide a full solution. It simply makes the attack require more resources on the part of the attacker. recognizes the attack and denies that user further access for a given amount of time, or uniformly throttles all requests in order to make it more difficult to consume resources more quickly than they can again be freed.

Source: MITRE CWE corpus.

CVSS Vector

CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H

Timeline

Published
August 13, 2026
Last Modified
August 14, 2026
First Seen
August 13, 2026

Related Vulnerabilities