AI Component

Training Data

Training data is both the model's most valuable input and its most underprotected one. Three problem classes dominate. First, poisoning: an attacker who can influence a public dataset, a web crawl, or a fine-tuning corpus can plant backdoors or biases that survive into the deployed model — BadNets-style attacks on image classifiers, trigger-phrase attacks on LLMs, and reward-hacking on RLHF datasets. Second, memorization and leakage: models can regurgitate verbatim training data, exposing PII and copyrighted content; this has driven the active New York Times v. OpenAI litigation and is a recurring GDPR concern. Third, provenance: when training data origins are unclear, downstream users inherit legal and security risk they can't assess. EU AI Act Article 10 (Data Governance) and ISO 42001 Annex A treat training-data quality as a controlled asset. Defenses: data lineage tracking, deduplication, PII scrubbing before training, and adversarial training against known trigger families.

228
Total CVEs
12
Pages
Page 11 of 12
Current
Severity CVE CVSS
HIGH CVE-2026-54058 -
HIGH CVE-2026-27775 8.8
MEDIUM CVE-2026-58435 5.4
HIGH CVE-2026-64832 8.8
MEDIUM CVE-2026-65010 6.6
MEDIUM CVE-2026-66007 6.5
HIGH CVE-2021-47816 8.8
CRITICAL CVE-2026-68771 9.8
HIGH CVE-2026-71281 8.8
HIGH CVE-2026-18947 8.5
MEDIUM CVE-2026-18942 5.5
CRITICAL CVE-2026-18948 9.9
HIGH CVE-2026-18941 7.7
UNKNOWN CVE-2026-20728 -
UNKNOWN CVE-2021-33627 -
UNKNOWN CVE-2022-24069 -
MEDIUM CVE-2026-28707 -
HIGH CVE-2026-75111 7.5
MEDIUM CVE-2026-69146 6.5
HIGH GHSA-wg9g-w2j2-8pgr 7.8

Page 11 of 12