vLLM's OpenAI-compatible completions endpoint accepts a prompt field that can be a list of unbounded length, and the serving layer allocates a dedicated engine generator plus response slot for every element in that list — so a single authenticated API call can exhaust CPU, memory, async scheduling capacity, and response buffering on the inference server. This matters because vLLM underpins a large share of self-hosted LLM serving (128 downstream dependents tracked here) and exploitation needs nothing more than a valid low-privilege API key and network access (AC:L, PR:L, no user interaction) — the attack is a single oversized request, not a sustained flood. There's no public exploit or scanner template circulating yet, it isn't in CISA KEV, and EPSS sits at just 0.39% (68th percentile), so this looks like opportunistic-DoS risk rather than active exploitation — but multi-tenant deployments (shared inference gateways, RAG backends, agent frameworks routing through a common vLLM instance) turn one abusive caller into an outage for every other consumer. Patch to vLLM 0.26.0, and until then cap request body/array size and rate-limit per-API-key at the reverse proxy in front of the OpenAI-compatible endpoint.
What is the risk?
Medium severity (CVSS 6.5, availability-only impact) reflects a real but bounded risk: the vulnerability is trivial to trigger technically (AC:L, no crafted payload beyond an oversized array) but requires authenticated access (PR:L), which limits exposure mostly to insider abuse, compromised API keys, or misconfigured multi-tenant gateways rather than fully anonymous internet-wide attackers. No active exploitation (not in CISA KEV), no public PoC/exploit, and a low EPSS score (0.39%) all point to low near-term exploitation likelihood. The real risk driver is architectural: any vLLM deployment that exposes the completions API to more than one trust boundary (SaaS inference providers, internal platform teams serving multiple product teams) faces disproportionate blast radius from a single bad actor or buggy client.
How does the attack unfold?
What systems are affected?
| Package | Ecosystem | Vulnerable Range | Patched |
|---|---|---|---|
| vLLM | pip | >= 0.19.0, < 0.26.0 | 0.26.0 |
Do you use vLLM? You're affected.
How severe is it?
What is the attack surface?
What should I do?
1 step-
1) Upgrade vLLM to 0.26.0 or later, which fixes the unbounded prompt handling. 2) Until patched, enforce a maximum array length / request body size at the API gateway or reverse proxy in front of vLLM (e.g., reject completions requests with prompt lists above a sane threshold, such as 32-64 items). 3) Apply per-API-key rate limiting and concurrent-request caps so a single client cannot monopolize engine slots. 4) For multi-tenant deployments, run vLLM instances per-tenant or with resource quotas/cgroup limits so one abusive client cannot starve others. 5) Detection: monitor for anomalous spikes in request payload size, sudden increases in engine queue depth/memory usage correlated with a single API key, and CPU/memory saturation alerts on vLLM hosts without a corresponding increase in request count.
How is it classified?
Which compliance frameworks are affected?
This CVE is relevant to:
Frequently Asked Questions
What is CVE-2026-73559?
vLLM's OpenAI-compatible completions endpoint accepts a prompt field that can be a list of unbounded length, and the serving layer allocates a dedicated engine generator plus response slot for every element in that list — so a single authenticated API call can exhaust CPU, memory, async scheduling capacity, and response buffering on the inference server. This matters because vLLM underpins a large share of self-hosted LLM serving (128 downstream dependents tracked here) and exploitation needs nothing more than a valid low-privilege API key and network access (AC:L, PR:L, no user interaction) — the attack is a single oversized request, not a sustained flood. There's no public exploit or scanner template circulating yet, it isn't in CISA KEV, and EPSS sits at just 0.39% (68th percentile), so this looks like opportunistic-DoS risk rather than active exploitation — but multi-tenant deployments (shared inference gateways, RAG backends, agent frameworks routing through a common vLLM instance) turn one abusive caller into an outage for every other consumer. Patch to vLLM 0.26.0, and until then cap request body/array size and rate-limit per-API-key at the reverse proxy in front of the OpenAI-compatible endpoint.
Is CVE-2026-73559 actively exploited?
No confirmed active exploitation of CVE-2026-73559 has been reported, but organizations should still patch proactively.
How to fix CVE-2026-73559?
1) Upgrade vLLM to 0.26.0 or later, which fixes the unbounded prompt handling. 2) Until patched, enforce a maximum array length / request body size at the API gateway or reverse proxy in front of vLLM (e.g., reject completions requests with prompt lists above a sane threshold, such as 32-64 items). 3) Apply per-API-key rate limiting and concurrent-request caps so a single client cannot monopolize engine slots. 4) For multi-tenant deployments, run vLLM instances per-tenant or with resource quotas/cgroup limits so one abusive client cannot starve others. 5) Detection: monitor for anomalous spikes in request payload size, sudden increases in engine queue depth/memory usage correlated with a single API key, and CPU/memory saturation alerts on vLLM hosts without a corresponding increase in request count.
What systems are affected by CVE-2026-73559?
This vulnerability affects the following AI/ML architecture patterns: model serving, LLM inference APIs, RAG pipelines, agent frameworks, multi-tenant inference gateways.
What is the CVSS score for CVE-2026-73559?
CVE-2026-73559 has a CVSS v3.1 base score of 6.5 (MEDIUM). The EPSS exploitation probability is 0.58%.
What is the AI security impact?
Affected AI Architectures
MITRE ATLAS Techniques
AML.T0029 Denial of AI Service AML.T0034 Cost Harvesting AML.T0034.001 Resource-Intensive Queries AML.T0040 AI Model Inference API Access Compliance Controls Affected
What are the technical details?
Original Advisory
vLLM is an inference and serving engine for large language models. From 0.19.0 until 0.26.0, the /v1/completions CompletionRequest.prompt field in vllm/entrypoints/openai/completion/protocol.py accepts an unbounded list[str] or list[list[int]], prompt_to_seq() in vllm/renderers/inputs/preprocess.py and OnlineRenderer.preprocess_completion() in vllm/renderers/online_renderer.py expand every element, and vllm/entrypoints/openai/completion/serving.py creates one engine generator and response slot per prompt, allowing an authenticated API client to exhaust CPU, memory, async scheduling capacity, engine request slots, and response buffering with one request. This issue is fixed in version 0.26.0.
Exploitation Scenario
An attacker with a valid but low-privilege API key to a vLLM-served completions endpoint sends a single POST to /v1/completions with the prompt field set to a list containing thousands of strings (or list[list[int]] token sequences). vllm/renderers preprocesses and expands every element, and the serving layer spins up one engine generator and response slot per prompt — instantly consuming available CPU, memory, and scheduling capacity meant for all concurrent users. If the deployment is multi-tenant (a shared internal LLM gateway, RAG backend serving multiple applications, or a SaaS inference provider), this single request degrades or crashes inference for every other tenant, effectively a low-effort denial-of-service against the whole platform with one HTTP call and no elevated privileges.
Weaknesses (CWE)
CWE-400 Uncontrolled Resource Consumption
Primary
CWE-400 Uncontrolled Resource Consumption
Primary
CWE-400 Uncontrolled Resource Consumption CWE-400 — Uncontrolled Resource Consumption: The product does not properly control the allocation and maintenance of a limited resource.
- [Architecture and Design] Design throttling mechanisms into the system architecture. The best protection is to limit the amount of resources that an unauthorized user can cause to be expended. A strong authentication and access control model will help prevent such attacks from occurring in the first place. The login application should be protected against DoS attacks as much as possible. Limiting the database access, perhaps by caching result sets, can help minimize the resources expended. To further limit the potential for a DoS attack, consider tracking the rate of requests received from users and blocking requests that exceed a defined rate threshold.
- [Architecture and Design] Mitigation of resource exhaustion attacks requires that the target system either: The first of these solutions is an issue in itself though, since it may allow attackers to prevent the use of the system by a particular valid user. If the attacker impersonates the valid user, they may be able to prevent the user from accessing the server in question. The second solution is simply difficult to effectively institute -- and even when properly done, it does not provide a full solution. It simply makes the attack require more resources on the part of the attacker. recognizes the attack and denies that user further access for a given amount of time, or uniformly throttles all requests in order to make it more difficult to consume resources more quickly than they can again be freed.
Source: MITRE CWE corpus.
CVSS Vector
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H References
- github.com/advisories/GHSA-87x5-vmc3-756j
- nvd.nist.gov/vuln/detail/CVE-2026-73559
- github.com/vllm-project/vllm/commit/675f4295cdfe0d870471c2b51bfeca3a68a9569e
- github.com/vllm-project/vllm/pull/47845
- github.com/vllm-project/vllm/releases/tag/v0.26.0
- github.com/vllm-project/vllm/security/advisories/GHSA-87x5-vmc3-756j
Timeline
Related Vulnerabilities
CVE-2026-61732 10.0 Analysis pending
Same package: vllm CVE-2024-11041 9.8 vllm: RCE via unsafe pickle deserialization in MessageQueue
Same package: vllm CVE-2025-47277 9.8 vLLM: RCE via exposed TCPStore in distributed inference
Same package: vllm CVE-2026-25960 9.8 vllm: SSRF allows internal network access
Same package: vllm CVE-2024-9053 9.8 vllm: RCE via unsafe pickle deserialization in RPC server
Same package: vllm