CVE-2026-71486: vLLM: unbounded /derender inputs cause resource DoS

GHSA-8737-qx52-hjff MEDIUM
Published August 17, 2026
CISO Take

vLLM's /v1/completions/derender and /v1/chat/completions/derender endpoints process caller-supplied response objects — including token IDs, logprobs, and expert-routing data — through the tokenizer and OnlineDerenderer before any size or length limits are checked, letting an authenticated client force excessive CPU and memory consumption and bloated responses. This is a low-severity, low-complexity availability issue (CVSS 4.3, AV:N/AC:L/PR:L) rather than a data-exposure or code-execution bug, but any organization running multi-tenant or shared vLLM inference endpoints should treat it as a real service-disruption risk: a single misbehaving or malicious API key can degrade throughput and latency for every other user on that deployment. There is no evidence of active exploitation, no CISA KEV listing, and no public PoC or scanner template, so this is a routine patch-cycle item rather than an emergency. Upgrade to vLLM 0.26.0 or later, and in the interim consider rate-limiting or payload-size caps at the reverse proxy in front of any exposed derender endpoints, since these paths bypass the engine's normal max_model_len/max_tokens/max_num_seqs guardrails.

Sources: NVD GitHub Advisory ATLAS

What is the risk?

Low-to-moderate risk in practice: exploitability is trivial for any authenticated API client (AC:L, PR:L, no user interaction), but impact is confined to availability (A:L) with no confidentiality or integrity loss, and CVSS lands at 4.3 (medium). The requirement for API authentication meaningfully narrows the exposure compared to unauthenticated DoS bugs — this is primarily a threat from malicious or compromised legitimate users, or from noisy-neighbor tenants in shared inference deployments, rather than from opportunistic internet scanning. No EPSS data, KEV listing, or public exploit/scanner exists yet, so near-term mass exploitation is unlikely; the main residual risk is targeted abuse against organizations running multi-tenant vLLM services.

How does the attack unfold?

Authenticated Access
Attacker obtains or already holds valid API credentials to a vLLM deployment exposing the OpenAI-compatible completions/chat API.
AML.T0040
Crafted Oversized Payload
Attacker submits a POST to /v1/chat/completions/derender or /v1/completions/derender with an inflated GenerateResponse object (excessive token_ids, logprobs, routed_experts).
AML.T0034.001
Unthrottled Resource Consumption
OnlineDerenderer and tokenizer.decode process the payload before max_model_len/max_tokens/max_num_seqs limits are enforced, spiking CPU and memory usage.
AML.T0034
Service Degradation
Sustained or repeated oversized requests degrade throughput and latency for all tenants sharing the vLLM inference service.
AML.T0029

What systems are affected?

Package Ecosystem Vulnerable Range Patched
vLLM pip < 0.26.0 0.26.0
92.7K 95 dependents Pushed 5d ago 26% patched ~47d to patch Full package profile →

Do you use vLLM? You're affected.

How severe is it?

CVSS 3.1
4.3 / 10
EPSS
0.5%
chance of exploitation in 30 days
Higher than 38% of all CVEs
Exploitation Status
No known exploitation
Sophistication
Moderate

What is the attack surface?

AV AC PR UI S C I A
AV Network
AC Low
PR Low
UI None
S Unchanged
C None
I None
A Low

What should I do?

1 step
  1. Upgrade vLLM to 0.26.0 or later, which enforces size/length limits before derender processing begins. If immediate upgrade isn't possible, restrict or disable the /v1/completions/derender and /v1/chat/completions/derender endpoints at the API gateway or reverse proxy for clients that don't need them, and enforce request body size caps and per-key rate limiting in front of the vLLM server. Monitor for anomalous CPU/memory spikes or unusually large request payloads correlated with specific API keys as a detection signal, and review access logs for repeated derender calls from a single authenticated client.

What does CISA's SSVC say?

Decision Track
Exploitation none
Automatable No
Technical Impact partial

Source: CISA Vulnrichment (SSVC v2.0). Decision based on the CISA Coordinator decision tree.

How is it classified?

Which compliance frameworks are affected?

This CVE is relevant to:

ISO 42001
8.2 - Operational planning and control of AI system resources
NIST AI RMF
MEASURE 2.7 - AI system security and resilience is evaluated
OWASP LLM Top 10
LLM10:2025 - Unbounded Consumption

Frequently Asked Questions

What is CVE-2026-71486?

vLLM's /v1/completions/derender and /v1/chat/completions/derender endpoints process caller-supplied response objects — including token IDs, logprobs, and expert-routing data — through the tokenizer and OnlineDerenderer before any size or length limits are checked, letting an authenticated client force excessive CPU and memory consumption and bloated responses. This is a low-severity, low-complexity availability issue (CVSS 4.3, AV:N/AC:L/PR:L) rather than a data-exposure or code-execution bug, but any organization running multi-tenant or shared vLLM inference endpoints should treat it as a real service-disruption risk: a single misbehaving or malicious API key can degrade throughput and latency for every other user on that deployment. There is no evidence of active exploitation, no CISA KEV listing, and no public PoC or scanner template, so this is a routine patch-cycle item rather than an emergency. Upgrade to vLLM 0.26.0 or later, and in the interim consider rate-limiting or payload-size caps at the reverse proxy in front of any exposed derender endpoints, since these paths bypass the engine's normal max_model_len/max_tokens/max_num_seqs guardrails.

Is CVE-2026-71486 actively exploited?

No confirmed active exploitation of CVE-2026-71486 has been reported, but organizations should still patch proactively.

How to fix CVE-2026-71486?

Upgrade vLLM to 0.26.0 or later, which enforces size/length limits before derender processing begins. If immediate upgrade isn't possible, restrict or disable the /v1/completions/derender and /v1/chat/completions/derender endpoints at the API gateway or reverse proxy for clients that don't need them, and enforce request body size caps and per-key rate limiting in front of the vLLM server. Monitor for anomalous CPU/memory spikes or unusually large request payloads correlated with specific API keys as a detection signal, and review access logs for repeated derender calls from a single authenticated client.

What systems are affected by CVE-2026-71486?

This vulnerability affects the following AI/ML architecture patterns: model serving, multi-tenant LLM inference platforms, MoE inference deployments.

What is the CVSS score for CVE-2026-71486?

CVE-2026-71486 has a CVSS v3.1 base score of 4.3 (MEDIUM). The EPSS exploitation probability is 0.47%.

What is the AI security impact?

Affected AI Architectures

model servingmulti-tenant LLM inference platformsMoE inference deployments

MITRE ATLAS Techniques

AML.T0029 Denial of AI Service
AML.T0034 Cost Harvesting
AML.T0034.001 Resource-Intensive Queries

Compliance Controls Affected

ISO 42001: 8.2
NIST AI RMF: MEASURE 2.7
OWASP LLM Top 10: LLM10:2025

What are the technical details?

Original Advisory

vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the /v1/completions/derender and /v1/chat/completions/derender endpoints accept caller-supplied GenerateResponse objects whose generate_responses, choices, token_ids, prompt_logprobs, logprobs.content, top_logprobs, and routed_experts structures are processed by OnlineDerenderer and tokenizer.decode before max_model_len, max_tokens, max_num_seqs, or response-size limits are enforced, allowing an authenticated API client to consume excessive CPU and memory and produce oversized responses. This issue is fixed in version 0.26.0.

Exploitation Scenario

An attacker who holds valid (even low-privilege) API credentials to a shared vLLM deployment crafts a POST to /v1/chat/completions/derender containing an artificially bloated GenerateResponse object — thousands of token_ids, deeply nested logprobs.content/top_logprobs structures, and inflated routed_experts arrays — sized well beyond what max_model_len or max_num_seqs would normally allow. Because these limits aren't enforced until after tokenizer.decode and OnlineDerenderer have already processed the payload, the request consumes disproportionate CPU and memory and returns an oversized response, degrading service for other tenants on the same inference cluster. Repeating this with concurrent requests from the same or multiple API keys amplifies the effect into a sustained denial-of-service against the shared model-serving endpoint.

Weaknesses (CWE)

CWE-400 — Uncontrolled Resource Consumption: The product does not properly control the allocation and maintenance of a limited resource.

  • [Architecture and Design] Design throttling mechanisms into the system architecture. The best protection is to limit the amount of resources that an unauthorized user can cause to be expended. A strong authentication and access control model will help prevent such attacks from occurring in the first place. The login application should be protected against DoS attacks as much as possible. Limiting the database access, perhaps by caching result sets, can help minimize the resources expended. To further limit the potential for a DoS attack, consider tracking the rate of requests received from users and blocking requests that exceed a defined rate threshold.
  • [Architecture and Design] Mitigation of resource exhaustion attacks requires that the target system either: The first of these solutions is an issue in itself though, since it may allow attackers to prevent the use of the system by a particular valid user. If the attacker impersonates the valid user, they may be able to prevent the user from accessing the server in question. The second solution is simply difficult to effectively institute -- and even when properly done, it does not provide a full solution. It simply makes the attack require more resources on the part of the attacker. recognizes the attack and denies that user further access for a given amount of time, or uniformly throttles all requests in order to make it more difficult to consume resources more quickly than they can again be freed.

Source: MITRE CWE corpus.

CVSS Vector

CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:L

Timeline

Published
August 17, 2026
Last Modified
September 4, 2026
First Seen
August 17, 2026

Related Vulnerabilities