A silent integer truncation bug in vLLM's GGUF dequantize CUDA kernels causes output tensors to be only partially initialized, leaving the remainder populated with stale GPU memory from prior operations — which in multi-tenant inference deployments may contain other users' tensor data. An attacker needs only to publish a GGUF model file with tensor dimensions whose product exceeds INT_MAX to trigger this silently: no error is thrown, no warning logged, and contaminated tensors pass undetected through all downstream computation. With 130 downstream dependents and EPSS placing this vulnerability in the top 87th percentile for exploitation likelihood, any team running shared vLLM inference infrastructure or loading externally-sourced GGUF models faces genuine cross-tenant data leakage exposure. Apply the upstream fix from PR #44971 (commit f219788f) immediately and validate all GGUF model files for oversized tensor dimensions before loading in production.
What is the risk?
Medium CVE severity but elevated contextual risk for multi-tenant AI inference environments. The attack surface is realistic: exploitation requires only publishing or social-engineering the loading of a malicious GGUF model file — a plausible threat given how widely models are sourced from public hubs. The vulnerability is entirely passive and silent once triggered: the dequantize kernel truncates its work silently, leaving uninitialized GPU buffers to propagate through model computation. Single-tenant deployments with trusted model sources face low risk; shared inference platforms and model-as-a-service providers running GGUF-quantized models face genuine confidentiality exposure. No patched release version is currently published — the fix exists only as a commit on the upstream repo, meaning users must manually apply or cherry-pick the patch.
How does the attack unfold?
What systems are affected?
| Package | Ecosystem | Vulnerable Range | Patched |
|---|---|---|---|
| vLLM | pip | >= 0.5.5, < 0.24.0 | 0.24.0 |
Do you use vLLM? You're affected.
How severe is it?
What is the attack surface?
What should I do?
5 steps-
Apply upstream fix from PR #44971 (commit f219788f91952827132fa4fdf916427cd20d225e): changes the int k parameter to int64_t in to_cuda_ggml_t and all derived dequantize functions.
-
Until patched, implement a pre-load validation layer that rejects any GGUF model file containing weight tensor dimensions whose product exceeds INT_MAX (2,147,483,647).
-
Audit all GGUF models currently loaded in production for weight matrices with shapes such as [65536, 65536] or any m×n > 2.1B configuration.
-
For multi-tenant deployments, consider process-level isolation for GGUF model loading with dedicated GPU memory allocations pending the patch.
-
Enforce model provenance controls — only load GGUF files from cryptographically attested, internal or verified sources; disable loading from arbitrary public model hubs in production.
What does CISA's SSVC say?
Source: CISA Vulnrichment (SSVC v2.0). Decision based on the CISA Coordinator decision tree.
How is it classified?
Which compliance frameworks are affected?
This CVE is relevant to:
Frequently Asked Questions
What is CVE-2026-53923?
A silent integer truncation bug in vLLM's GGUF dequantize CUDA kernels causes output tensors to be only partially initialized, leaving the remainder populated with stale GPU memory from prior operations — which in multi-tenant inference deployments may contain other users' tensor data. An attacker needs only to publish a GGUF model file with tensor dimensions whose product exceeds INT_MAX to trigger this silently: no error is thrown, no warning logged, and contaminated tensors pass undetected through all downstream computation. With 130 downstream dependents and EPSS placing this vulnerability in the top 87th percentile for exploitation likelihood, any team running shared vLLM inference infrastructure or loading externally-sourced GGUF models faces genuine cross-tenant data leakage exposure. Apply the upstream fix from PR #44971 (commit f219788f) immediately and validate all GGUF model files for oversized tensor dimensions before loading in production.
Is CVE-2026-53923 actively exploited?
No confirmed active exploitation of CVE-2026-53923 has been reported, but organizations should still patch proactively.
How to fix CVE-2026-53923?
1. Apply upstream fix from PR #44971 (commit f219788f91952827132fa4fdf916427cd20d225e): changes the int k parameter to int64_t in to_cuda_ggml_t and all derived dequantize functions. 2. Until patched, implement a pre-load validation layer that rejects any GGUF model file containing weight tensor dimensions whose product exceeds INT_MAX (2,147,483,647). 3. Audit all GGUF models currently loaded in production for weight matrices with shapes such as [65536, 65536] or any m×n > 2.1B configuration. 4. For multi-tenant deployments, consider process-level isolation for GGUF model loading with dedicated GPU memory allocations pending the patch. 5. Enforce model provenance controls — only load GGUF files from cryptographically attested, internal or verified sources; disable loading from arbitrary public model hubs in production.
What systems are affected by CVE-2026-53923?
This vulnerability affects the following AI/ML architecture patterns: multi-tenant LLM inference, model serving, GGUF model pipelines, quantized model deployment, shared GPU inference clusters.
What is the CVSS score for CVE-2026-53923?
CVE-2026-53923 has a CVSS v3.1 base score of 7.5 (HIGH). The EPSS exploitation probability is 0.28%.
What is the AI security impact?
Affected AI Architectures
MITRE ATLAS Techniques
AML.T0010.003 Model AML.T0011.000 Unsafe AI Artifacts AML.T0025 Exfiltration via Cyber Means AML.T0049 Exploit Public-Facing Application Compliance Controls Affected
What are the technical details?
Original Advisory
vLLM is an inference and serving engine for large language models (LLMs). From 0.5.5 until 0.23.1rc0, integer truncation of tensor dimensions in vLLM's GGUF dequantize kernels (csrc/quantization/gguf/gguf_kernel.cu) causes partial tensor processing. The output tensor is allocated at full size via torch::empty (uninitialized memory), but the dequantize CUDA kernel processes only a truncated number of elements. The unfilled portion of the output tensor retains whatever was previously in GPU memory. In multi-tenant inference deployments, this residual GPU memory may contain tensor data from other users' inference requests, constituting information disclosure. This vulnerability is fixed in 0.23.1rc0.
Exploitation Scenario
An adversary targeting a shared vLLM inference platform publishes a GGUF model to HuggingFace with a weight tensor shaped [65536, 65536], yielding m*n = 4,294,967,296 — exceeding INT_MAX by roughly one full INT_MAX. A victim organization loads this model onto their multi-tenant GPU cluster serving multiple enterprise clients. During GGUF weight dequantization at model load time, the int64 product is silently cast to int32, producing a near-zero truncated k. The CUDA kernel fills essentially none of the ~16GB output buffer, leaving it populated entirely with prior GPU allocations from active user sessions. The contaminated weight tensor is then used in all subsequent inference matrix multiplications. A co-located adversary submitting concurrent inference requests may craft queries that amplify or surface fragments of this residual data, potentially recovering portions of other users' prompt text, token embeddings, or intermediate activations without any authentication bypass or elevated privilege.
Weaknesses (CWE)
CWE-200 Exposure of Sensitive Information to an Unauthorized Actor
Primary
CWE-200 Exposure of Sensitive Information to an Unauthorized Actor
Primary
CWE-681 Incorrect Conversion between Numeric Types
Primary
CWE-681 Incorrect Conversion between Numeric Types
Primary
CWE-200 Exposure of Sensitive Information to an Unauthorized Actor CWE-681 Incorrect Conversion between Numeric Types CWE-200 — Exposure of Sensitive Information to an Unauthorized Actor: The product exposes sensitive information to an actor that is not explicitly authorized to have access to that information.
- [Architecture and Design] Compartmentalize the system to have "safe" areas where trust boundaries can be unambiguously drawn. Do not allow sensitive data to go outside of the trust boundary and always be careful when interfacing with a compartment outside of the safe area. Ensure that appropriate compartmentalization is built into the system design, and the compartmentalization allows for and reinforces privilege separation functionality. Architects and designers should rely on the principle of least privilege to decide the appropriate time to use privileges and the time to drop privileges.
Source: MITRE CWE corpus.
CVSS Vector
CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N References
- github.com/advisories/GHSA-5jv2-g5wq-cmr4
- github.com/pypa/advisory-database/tree/main/vulns/vllm/PYSEC-2026-3403.yaml
- github.com/vllm-project/vllm
- github.com/vllm-project/vllm/commit/f219788f91952827132fa4fdf916427cd20d225e
- github.com/vllm-project/vllm/pull/44971
- github.com/vllm-project/vllm/security/advisories/GHSA-5jv2-g5wq-cmr4
- nvd.nist.gov/vuln/detail/CVE-2026-53923
- pypi.org/project/vllm
Timeline
Related Vulnerabilities
CVE-2024-9053 9.8 vllm: RCE via unsafe pickle deserialization in RPC server
Same package: vllm CVE-2024-11041 9.8 vllm: RCE via unsafe pickle deserialization in MessageQueue
Same package: vllm CVE-2026-25960 9.8 vllm: SSRF allows internal network access
Same package: vllm CVE-2025-47277 9.8 vLLM: RCE via exposed TCPStore in distributed inference
Same package: vllm CVE-2025-32444 9.8 vLLM: RCE via pickle deserialization on ZeroMQ
Same package: vllm