CVE-2026-53923: vLLM: integer truncation leaks GPU memory cross-tenant

GHSA-5jv2-g5wq-cmr4 HIGH
Published June 17, 2026
CISO Take

A silent integer truncation bug in vLLM's GGUF dequantize CUDA kernels causes output tensors to be only partially initialized, leaving the remainder populated with stale GPU memory from prior operations — which in multi-tenant inference deployments may contain other users' tensor data. An attacker needs only to publish a GGUF model file with tensor dimensions whose product exceeds INT_MAX to trigger this silently: no error is thrown, no warning logged, and contaminated tensors pass undetected through all downstream computation. With 130 downstream dependents and EPSS placing this vulnerability in the top 87th percentile for exploitation likelihood, any team running shared vLLM inference infrastructure or loading externally-sourced GGUF models faces genuine cross-tenant data leakage exposure. Apply the upstream fix from PR #44971 (commit f219788f) immediately and validate all GGUF model files for oversized tensor dimensions before loading in production.

Sources: NVD EPSS GitHub Advisory ATLAS

What is the risk?

Medium CVE severity but elevated contextual risk for multi-tenant AI inference environments. The attack surface is realistic: exploitation requires only publishing or social-engineering the loading of a malicious GGUF model file — a plausible threat given how widely models are sourced from public hubs. The vulnerability is entirely passive and silent once triggered: the dequantize kernel truncates its work silently, leaving uninitialized GPU buffers to propagate through model computation. Single-tenant deployments with trusted model sources face low risk; shared inference platforms and model-as-a-service providers running GGUF-quantized models face genuine confidentiality exposure. No patched release version is currently published — the fix exists only as a commit on the upstream repo, meaning users must manually apply or cherry-pick the patch.

How does the attack unfold?

Supply Chain Staging
Attacker crafts a GGUF model file with weight tensor dimensions whose product exceeds INT_MAX (e.g., shape [65536, 65536] = 4.29B elements) and publishes it to a public model hub such as HuggingFace.
AML.T0010.003
Model Loading
Victim's vLLM multi-tenant inference server loads the malicious GGUF model, triggering the vulnerable dequantize CUDA kernel path (ggml_dequantize or ggml_mul_mat_*) during weight initialization.
AML.T0011.000
Silent Integer Truncation
The 64-bit tensor size (m*n) is silently cast to int32, truncating to a near-zero value; the CUDA kernel fills only that fraction of the torch::empty output buffer while the remainder retains raw GPU memory from prior operations.
AML.T0049
Cross-Tenant Data Leakage
The contaminated output tensor — containing residual GPU memory from concurrent users' inference sessions — propagates silently through all downstream model computation, enabling potential recovery of other tenants' prompt content or activation data.
AML.T0025

What systems are affected?

Package Ecosystem Vulnerable Range Patched
vLLM pip >= 0.5.5, < 0.24.0 0.24.0
88.6K 130 dependents Pushed today 26% patched ~51d to patch Full package profile →

Do you use vLLM? You're affected.

How severe is it?

CVSS 3.1
7.5 / 10
EPSS
0.3%
chance of exploitation in 30 days
Higher than 20% of all CVEs
Exploitation Status
No known exploitation
Sophistication
Moderate

What is the attack surface?

AV AC PR UI S C I A
AV Network
AC Low
PR None
UI None
S Unchanged
C High
I None
A None

What should I do?

5 steps
  1. Apply upstream fix from PR #44971 (commit f219788f91952827132fa4fdf916427cd20d225e): changes the int k parameter to int64_t in to_cuda_ggml_t and all derived dequantize functions.

  2. Until patched, implement a pre-load validation layer that rejects any GGUF model file containing weight tensor dimensions whose product exceeds INT_MAX (2,147,483,647).

  3. Audit all GGUF models currently loaded in production for weight matrices with shapes such as [65536, 65536] or any m×n > 2.1B configuration.

  4. For multi-tenant deployments, consider process-level isolation for GGUF model loading with dedicated GPU memory allocations pending the patch.

  5. Enforce model provenance controls — only load GGUF files from cryptographically attested, internal or verified sources; disable loading from arbitrary public model hubs in production.

What does CISA's SSVC say?

Decision Track
Exploitation none
Automatable No
Technical Impact partial

Source: CISA Vulnrichment (SSVC v2.0). Decision based on the CISA Coordinator decision tree.

How is it classified?

Which compliance frameworks are affected?

This CVE is relevant to:

EU AI Act
Article 15 - Accuracy, robustness and cybersecurity
ISO 42001
A.10.1 - Supply chain management A.6.2 - AI system design and development
NIST AI RMF
MANAGE 2.2 - Mechanisms are in place and applied to sustain the value of deployed AI systems
OWASP LLM Top 10
LLM05:2025 - Supply Chain Vulnerabilities LLM06:2025 - Sensitive Information Disclosure

Frequently Asked Questions

What is CVE-2026-53923?

A silent integer truncation bug in vLLM's GGUF dequantize CUDA kernels causes output tensors to be only partially initialized, leaving the remainder populated with stale GPU memory from prior operations — which in multi-tenant inference deployments may contain other users' tensor data. An attacker needs only to publish a GGUF model file with tensor dimensions whose product exceeds INT_MAX to trigger this silently: no error is thrown, no warning logged, and contaminated tensors pass undetected through all downstream computation. With 130 downstream dependents and EPSS placing this vulnerability in the top 87th percentile for exploitation likelihood, any team running shared vLLM inference infrastructure or loading externally-sourced GGUF models faces genuine cross-tenant data leakage exposure. Apply the upstream fix from PR #44971 (commit f219788f) immediately and validate all GGUF model files for oversized tensor dimensions before loading in production.

Is CVE-2026-53923 actively exploited?

No confirmed active exploitation of CVE-2026-53923 has been reported, but organizations should still patch proactively.

How to fix CVE-2026-53923?

1. Apply upstream fix from PR #44971 (commit f219788f91952827132fa4fdf916427cd20d225e): changes the int k parameter to int64_t in to_cuda_ggml_t and all derived dequantize functions. 2. Until patched, implement a pre-load validation layer that rejects any GGUF model file containing weight tensor dimensions whose product exceeds INT_MAX (2,147,483,647). 3. Audit all GGUF models currently loaded in production for weight matrices with shapes such as [65536, 65536] or any m×n > 2.1B configuration. 4. For multi-tenant deployments, consider process-level isolation for GGUF model loading with dedicated GPU memory allocations pending the patch. 5. Enforce model provenance controls — only load GGUF files from cryptographically attested, internal or verified sources; disable loading from arbitrary public model hubs in production.

What systems are affected by CVE-2026-53923?

This vulnerability affects the following AI/ML architecture patterns: multi-tenant LLM inference, model serving, GGUF model pipelines, quantized model deployment, shared GPU inference clusters.

What is the CVSS score for CVE-2026-53923?

CVE-2026-53923 has a CVSS v3.1 base score of 7.5 (HIGH). The EPSS exploitation probability is 0.28%.

What is the AI security impact?

Affected AI Architectures

multi-tenant LLM inferencemodel servingGGUF model pipelinesquantized model deploymentshared GPU inference clusters

MITRE ATLAS Techniques

AML.T0010.003 Model
AML.T0011.000 Unsafe AI Artifacts
AML.T0025 Exfiltration via Cyber Means
AML.T0049 Exploit Public-Facing Application

Compliance Controls Affected

EU AI Act: Article 15
ISO 42001: A.10.1, A.6.2
NIST AI RMF: MANAGE 2.2
OWASP LLM Top 10: LLM05:2025, LLM06:2025

What are the technical details?

Original Advisory

vLLM is an inference and serving engine for large language models (LLMs). From 0.5.5 until 0.23.1rc0, integer truncation of tensor dimensions in vLLM's GGUF dequantize kernels (csrc/quantization/gguf/gguf_kernel.cu) causes partial tensor processing. The output tensor is allocated at full size via torch::empty (uninitialized memory), but the dequantize CUDA kernel processes only a truncated number of elements. The unfilled portion of the output tensor retains whatever was previously in GPU memory. In multi-tenant inference deployments, this residual GPU memory may contain tensor data from other users' inference requests, constituting information disclosure. This vulnerability is fixed in 0.23.1rc0.

Exploitation Scenario

An adversary targeting a shared vLLM inference platform publishes a GGUF model to HuggingFace with a weight tensor shaped [65536, 65536], yielding m*n = 4,294,967,296 — exceeding INT_MAX by roughly one full INT_MAX. A victim organization loads this model onto their multi-tenant GPU cluster serving multiple enterprise clients. During GGUF weight dequantization at model load time, the int64 product is silently cast to int32, producing a near-zero truncated k. The CUDA kernel fills essentially none of the ~16GB output buffer, leaving it populated entirely with prior GPU allocations from active user sessions. The contaminated weight tensor is then used in all subsequent inference matrix multiplications. A co-located adversary submitting concurrent inference requests may craft queries that amplify or surface fragments of this residual data, potentially recovering portions of other users' prompt text, token embeddings, or intermediate activations without any authentication bypass or elevated privilege.

Weaknesses (CWE)

CWE-200 — Exposure of Sensitive Information to an Unauthorized Actor: The product exposes sensitive information to an actor that is not explicitly authorized to have access to that information.

  • [Architecture and Design] Compartmentalize the system to have "safe" areas where trust boundaries can be unambiguously drawn. Do not allow sensitive data to go outside of the trust boundary and always be careful when interfacing with a compartment outside of the safe area. Ensure that appropriate compartmentalization is built into the system design, and the compartmentalization allows for and reinforces privilege separation functionality. Architects and designers should rely on the principle of least privilege to decide the appropriate time to use privileges and the time to drop privileges.

Source: MITRE CWE corpus.

CVSS Vector

CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N

Timeline

Published
June 17, 2026
Last Modified
July 17, 2026
First Seen
June 17, 2026

Related Vulnerabilities