CVE-2026-73556: vLLM: ReDoS in structured output stalls engine
GHSA-48jh-3gj7-fg8v MEDIUM CISA: TRACK*An unauthenticated attacker can send a single /v1/completions request carrying a malicious regular expression through the structured_outputs.regex parameter, and because vLLM's lm-format-enforcer backend never validates or bounds regex compilation time, that request pins a CPU core and stalls the structured-output engine path for every user sharing that inference server. This matters for any team running vLLM as a shared multi-tenant inference endpoint — one low-effort request degrades service for all downstream consumers of that node, which is a direct availability hit against production LLM serving infrastructure. The EPSS score (0.00315, top 76th percentile) is low and there's no KEV listing, public exploit, or Nuclei template yet, so this looks opportunistic rather than actively weaponized today, but the vulnerability class (catastrophic backtracking / ReDoS, CWE-400/CWE-1333) is trivial to reproduce once attackers notice the unauthenticated endpoint accepts arbitrary regex. Patch to vLLM 0.26.0, and until upgraded, restrict or rate-limit access to the lm-format-enforcer structured-output backend (or disable it in favor of the default outlines/xgrammar backend) on any internet-reachable deployment.
What is the risk?
Medium severity (CVSS 5.3) reflecting a pure availability impact — no confidentiality or integrity loss. Attack complexity is low and no authentication or user interaction is required, which lowers the bar for exploitation to 'anyone who can reach the API'. The mitigating factor is that this only affects deployments explicitly configured to use the lm-format-enforcer structured-output backend (not the default), which narrows real-world exposure. EPSS remains low (0.3%) and CISA SSVC rates it TRACK_STAR (track, no urgent action), consistent with a DoS-only bug with no known exploitation.
How does the attack unfold?
What systems are affected?
| Package | Ecosystem | Vulnerable Range | Patched |
|---|---|---|---|
| vLLM | pip | < 0.26.0 | 0.26.0 |
Do you use vLLM? You're affected.
How severe is it?
What is the attack surface?
What should I do?
1 step-
1) Upgrade to vLLM 0.26.0 or later, which adds compile_regex_with_timeout/validation to the lm-format-enforcer request path. 2) Until patched, avoid exposing the lm-format-enforcer structured-output backend on internet-facing endpoints, or switch to the default outlines/xgrammar backend, which is not affected. 3) Add request-level rate limiting and regex-length/complexity caps at the API gateway in front of vLLM. 4) Monitor for CPU saturation or worker-thread stalls correlated with /v1/completions requests carrying structured_outputs.regex payloads — this is the detection signature for exploitation attempts.
What does CISA's SSVC say?
Source: CISA Vulnrichment (SSVC v2.0). Decision based on the CISA Coordinator decision tree.
How is it classified?
Which compliance frameworks are affected?
This CVE is relevant to:
Frequently Asked Questions
What is CVE-2026-73556?
An unauthenticated attacker can send a single /v1/completions request carrying a malicious regular expression through the structured_outputs.regex parameter, and because vLLM's lm-format-enforcer backend never validates or bounds regex compilation time, that request pins a CPU core and stalls the structured-output engine path for every user sharing that inference server. This matters for any team running vLLM as a shared multi-tenant inference endpoint — one low-effort request degrades service for all downstream consumers of that node, which is a direct availability hit against production LLM serving infrastructure. The EPSS score (0.00315, top 76th percentile) is low and there's no KEV listing, public exploit, or Nuclei template yet, so this looks opportunistic rather than actively weaponized today, but the vulnerability class (catastrophic backtracking / ReDoS, CWE-400/CWE-1333) is trivial to reproduce once attackers notice the unauthenticated endpoint accepts arbitrary regex. Patch to vLLM 0.26.0, and until upgraded, restrict or rate-limit access to the lm-format-enforcer structured-output backend (or disable it in favor of the default outlines/xgrammar backend) on any internet-reachable deployment.
Is CVE-2026-73556 actively exploited?
No confirmed active exploitation of CVE-2026-73556 has been reported, but organizations should still patch proactively.
How to fix CVE-2026-73556?
1) Upgrade to vLLM 0.26.0 or later, which adds compile_regex_with_timeout/validation to the lm-format-enforcer request path. 2) Until patched, avoid exposing the lm-format-enforcer structured-output backend on internet-facing endpoints, or switch to the default outlines/xgrammar backend, which is not affected. 3) Add request-level rate limiting and regex-length/complexity caps at the API gateway in front of vLLM. 4) Monitor for CPU saturation or worker-thread stalls correlated with /v1/completions requests carrying structured_outputs.regex payloads — this is the detection signature for exploitation attempts.
What systems are affected by CVE-2026-73556?
This vulnerability affects the following AI/ML architecture patterns: model serving, structured/constrained generation pipelines, multi-tenant inference endpoints.
What is the CVSS score for CVE-2026-73556?
CVE-2026-73556 has a CVSS v3.1 base score of 5.3 (MEDIUM). The EPSS exploitation probability is 0.52%.
What is the AI security impact?
Affected AI Architectures
MITRE ATLAS Techniques
AML.T0029 Denial of AI Service AML.T0034.001 Resource-Intensive Queries AML.T0049 Exploit Public-Facing Application Compliance Controls Affected
What are the technical details?
Original Advisory
vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the structured_outputs.regex parameter in vllm/v1/structured_output/backend_lm_format_enforcer.py is passed to lmformatenforcer.RegexParser without compile_regex_with_timeout or validation in validate_structured_output_request_lm_format_enforcer, allowing an unauthenticated /v1/completions request against the lm-format-enforcer backend to consume a CPU core and stall the structured-output engine path with a catastrophic regular expression. This issue is fixed in version 0.26.0.
Exploitation Scenario
An adversary discovers a publicly reachable vLLM inference endpoint (e.g., via Shodan or a known company API domain) running a version prior to 0.26.0 with the lm-format-enforcer backend enabled. Without needing any credentials, they submit a single /v1/completions request with a structured_outputs.regex value engineered for catastrophic backtracking (e.g., nested quantifiers like (a+)+$). The regex compilation pegs a CPU core and the structured-output engine path stalls, degrading response times or availability for all other tenants/users of that inference server — a low-cost, repeatable denial-of-service the attacker can trigger at will simply by re-sending the request.
Weaknesses (CWE)
CWE-1333 Inefficient Regular Expression Complexity
Primary
CWE-1333 Inefficient Regular Expression Complexity
Primary
CWE-400 Uncontrolled Resource Consumption
Primary
CWE-400 Uncontrolled Resource Consumption
Primary
CWE-1333 Inefficient Regular Expression Complexity CWE-400 Uncontrolled Resource Consumption CWE-1333 — Inefficient Regular Expression Complexity: The product uses a regular expression with a worst-case computational complexity that is inefficient and possibly exponential.
- [Architecture and Design] Use regular expressions that do not support backtracking, e.g. by removing nested quantifiers.
- [System Configuration] Set backtracking limits in the configuration of the regular expression implementation, such as PHP's pcre.backtrack_limit. Also consider limits on execution time for the process.
Source: MITRE CWE corpus.
CVSS Vector
CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:L References
- github.com/advisories/GHSA-48jh-3gj7-fg8v
- nvd.nist.gov/vuln/detail/CVE-2026-73556
- github.com/vllm-project/vllm/commit/c9a788eedc412acceaa5112e0d44624b49841577
- github.com/vllm-project/vllm/pull/47595
- github.com/vllm-project/vllm/releases/tag/v0.26.0
- github.com/vllm-project/vllm/security/advisories/GHSA-48jh-3gj7-fg8v
Timeline
Related Vulnerabilities
CVE-2026-61732 10.0 Analysis pending
Same package: vllm CVE-2026-25960 9.8 vllm: SSRF allows internal network access
Same package: vllm CVE-2025-47277 9.8 vLLM: RCE via exposed TCPStore in distributed inference
Same package: vllm CVE-2024-11041 9.8 vllm: RCE via unsafe pickle deserialization in MessageQueue
Same package: vllm CVE-2024-9053 9.8 vllm: RCE via unsafe pickle deserialization in RPC server
Same package: vllm