vLLM's /v1/completions/derender and /v1/chat/completions/derender endpoints process caller-supplied response objects — including token IDs, logprobs, and expert-routing data — through the tokenizer and OnlineDerenderer before any size or length limits are checked, letting an authenticated client force excessive CPU and memory consumption and bloated responses. This is a low-severity, low-complexity availability issue (CVSS 4.3, AV:N/AC:L/PR:L) rather than a data-exposure or code-execution bug, but any organization running multi-tenant or shared vLLM inference endpoints should treat it as a real service-disruption risk: a single misbehaving or malicious API key can degrade throughput and latency for every other user on that deployment. There is no evidence of active exploitation, no CISA KEV listing, and no public PoC or scanner template, so this is a routine patch-cycle item rather than an emergency. Upgrade to vLLM 0.26.0 or later, and in the interim consider rate-limiting or payload-size caps at the reverse proxy in front of any exposed derender endpoints, since these paths bypass the engine's normal max_model_len/max_tokens/max_num_seqs guardrails.
What is the risk?
Low-to-moderate risk in practice: exploitability is trivial for any authenticated API client (AC:L, PR:L, no user interaction), but impact is confined to availability (A:L) with no confidentiality or integrity loss, and CVSS lands at 4.3 (medium). The requirement for API authentication meaningfully narrows the exposure compared to unauthenticated DoS bugs — this is primarily a threat from malicious or compromised legitimate users, or from noisy-neighbor tenants in shared inference deployments, rather than from opportunistic internet scanning. No EPSS data, KEV listing, or public exploit/scanner exists yet, so near-term mass exploitation is unlikely; the main residual risk is targeted abuse against organizations running multi-tenant vLLM services.
How does the attack unfold?
What systems are affected?
| Package | Ecosystem | Vulnerable Range | Patched |
|---|---|---|---|
| vLLM | pip | < 0.26.0 | 0.26.0 |
Do you use vLLM? You're affected.
How severe is it?
What is the attack surface?
What should I do?
1 step-
Upgrade vLLM to 0.26.0 or later, which enforces size/length limits before derender processing begins. If immediate upgrade isn't possible, restrict or disable the /v1/completions/derender and /v1/chat/completions/derender endpoints at the API gateway or reverse proxy for clients that don't need them, and enforce request body size caps and per-key rate limiting in front of the vLLM server. Monitor for anomalous CPU/memory spikes or unusually large request payloads correlated with specific API keys as a detection signal, and review access logs for repeated derender calls from a single authenticated client.
What does CISA's SSVC say?
Source: CISA Vulnrichment (SSVC v2.0). Decision based on the CISA Coordinator decision tree.
How is it classified?
Which compliance frameworks are affected?
This CVE is relevant to:
Frequently Asked Questions
What is CVE-2026-71486?
vLLM's /v1/completions/derender and /v1/chat/completions/derender endpoints process caller-supplied response objects — including token IDs, logprobs, and expert-routing data — through the tokenizer and OnlineDerenderer before any size or length limits are checked, letting an authenticated client force excessive CPU and memory consumption and bloated responses. This is a low-severity, low-complexity availability issue (CVSS 4.3, AV:N/AC:L/PR:L) rather than a data-exposure or code-execution bug, but any organization running multi-tenant or shared vLLM inference endpoints should treat it as a real service-disruption risk: a single misbehaving or malicious API key can degrade throughput and latency for every other user on that deployment. There is no evidence of active exploitation, no CISA KEV listing, and no public PoC or scanner template, so this is a routine patch-cycle item rather than an emergency. Upgrade to vLLM 0.26.0 or later, and in the interim consider rate-limiting or payload-size caps at the reverse proxy in front of any exposed derender endpoints, since these paths bypass the engine's normal max_model_len/max_tokens/max_num_seqs guardrails.
Is CVE-2026-71486 actively exploited?
No confirmed active exploitation of CVE-2026-71486 has been reported, but organizations should still patch proactively.
How to fix CVE-2026-71486?
Upgrade vLLM to 0.26.0 or later, which enforces size/length limits before derender processing begins. If immediate upgrade isn't possible, restrict or disable the /v1/completions/derender and /v1/chat/completions/derender endpoints at the API gateway or reverse proxy for clients that don't need them, and enforce request body size caps and per-key rate limiting in front of the vLLM server. Monitor for anomalous CPU/memory spikes or unusually large request payloads correlated with specific API keys as a detection signal, and review access logs for repeated derender calls from a single authenticated client.
What systems are affected by CVE-2026-71486?
This vulnerability affects the following AI/ML architecture patterns: model serving, multi-tenant LLM inference platforms, MoE inference deployments.
What is the CVSS score for CVE-2026-71486?
CVE-2026-71486 has a CVSS v3.1 base score of 4.3 (MEDIUM). The EPSS exploitation probability is 0.47%.
What is the AI security impact?
Affected AI Architectures
MITRE ATLAS Techniques
AML.T0029 Denial of AI Service AML.T0034 Cost Harvesting AML.T0034.001 Resource-Intensive Queries Compliance Controls Affected
What are the technical details?
Original Advisory
vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the /v1/completions/derender and /v1/chat/completions/derender endpoints accept caller-supplied GenerateResponse objects whose generate_responses, choices, token_ids, prompt_logprobs, logprobs.content, top_logprobs, and routed_experts structures are processed by OnlineDerenderer and tokenizer.decode before max_model_len, max_tokens, max_num_seqs, or response-size limits are enforced, allowing an authenticated API client to consume excessive CPU and memory and produce oversized responses. This issue is fixed in version 0.26.0.
Exploitation Scenario
An attacker who holds valid (even low-privilege) API credentials to a shared vLLM deployment crafts a POST to /v1/chat/completions/derender containing an artificially bloated GenerateResponse object — thousands of token_ids, deeply nested logprobs.content/top_logprobs structures, and inflated routed_experts arrays — sized well beyond what max_model_len or max_num_seqs would normally allow. Because these limits aren't enforced until after tokenizer.decode and OnlineDerenderer have already processed the payload, the request consumes disproportionate CPU and memory and returns an oversized response, degrading service for other tenants on the same inference cluster. Repeating this with concurrent requests from the same or multiple API keys amplifies the effect into a sustained denial-of-service against the shared model-serving endpoint.
Weaknesses (CWE)
CWE-400 Uncontrolled Resource Consumption
Primary
CWE-400 Uncontrolled Resource Consumption
Primary
CWE-770 Allocation of Resources Without Limits or Throttling
Primary
CWE-770 Allocation of Resources Without Limits or Throttling
Primary
CWE-400 Uncontrolled Resource Consumption CWE-770 Allocation of Resources Without Limits or Throttling CWE-400 — Uncontrolled Resource Consumption: The product does not properly control the allocation and maintenance of a limited resource.
- [Architecture and Design] Design throttling mechanisms into the system architecture. The best protection is to limit the amount of resources that an unauthorized user can cause to be expended. A strong authentication and access control model will help prevent such attacks from occurring in the first place. The login application should be protected against DoS attacks as much as possible. Limiting the database access, perhaps by caching result sets, can help minimize the resources expended. To further limit the potential for a DoS attack, consider tracking the rate of requests received from users and blocking requests that exceed a defined rate threshold.
- [Architecture and Design] Mitigation of resource exhaustion attacks requires that the target system either: The first of these solutions is an issue in itself though, since it may allow attackers to prevent the use of the system by a particular valid user. If the attacker impersonates the valid user, they may be able to prevent the user from accessing the server in question. The second solution is simply difficult to effectively institute -- and even when properly done, it does not provide a full solution. It simply makes the attack require more resources on the part of the attacker. recognizes the attack and denies that user further access for a given amount of time, or uniformly throttles all requests in order to make it more difficult to consume resources more quickly than they can again be freed.
Source: MITRE CWE corpus.
CVSS Vector
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:L References
- github.com/advisories/GHSA-8737-qx52-hjff
- nvd.nist.gov/vuln/detail/CVE-2026-71486
- github.com/vllm-project/vllm/commit/8e61b646e2d157f9b93451fa048f9c8530c8a67b
- github.com/vllm-project/vllm/pull/47260
- github.com/vllm-project/vllm/releases/tag/v0.26.0
- github.com/vllm-project/vllm/security/advisories/GHSA-8737-qx52-hjff
Timeline
Related Vulnerabilities
CVE-2026-61732 10.0 Analysis pending
Same package: vllm CVE-2024-11041 9.8 vllm: RCE via unsafe pickle deserialization in MessageQueue
Same package: vllm CVE-2025-47277 9.8 vLLM: RCE via exposed TCPStore in distributed inference
Same package: vllm CVE-2026-25960 9.8 vllm: SSRF allows internal network access
Same package: vllm CVE-2024-9053 9.8 vllm: RCE via unsafe pickle deserialization in RPC server
Same package: vllm