Revealing Interpretable Failure Modes of VLMs
Isha Chaudhary, Vedaant V Jain, Kavya Sachdeva +2 more
Vision-Language Models (VLMs) are increasingly used in safety-critical applications because of their broad reasoning capabilities and ability to...
AI Threat Alert indexes 3,371+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 121–140 of 471 papers
Clear filtersIsha Chaudhary, Vedaant V Jain, Kavya Sachdeva +2 more
Vision-Language Models (VLMs) are increasingly used in safety-critical applications because of their broad reasoning capabilities and ability to...
Sina Mavali, David Pape, Jonathan Evertz +5 more
Terminal agents are increasingly capable of executing complex, long-horizon tasks autonomously from a single user prompt. To do so, they must...
Xinyi Zeng, Xue Yang, Jingyuan Zhang +5 more
Multimodal large language models (MLLMs) are gaining increasing attention. Due to the heterogeneity of their input features, they face significant...
John T. Halloran
Large language model (LLM) alignment algorithms typically consist of post-training over preference pairs. While such algorithms are widely used to...
Christian Moya, Alex Semendinger, Guang Lin +1 more
Preference learning methods such as Direct Preference Optimization (DPO) are known to induce reliance on spurious correlations, leading to sycophancy...
Fatima Z. Abacha, Sin G. Teo, Yuanxiang Wu +2 more
Federated Learning remains highly susceptible to backdoor attacks--malicious clients inject targeted behaviours into the global model. Existing...
Nikita Kezins, Urbas Ekka, Pascal Berrang +1 more
Guardrail Classifiers defend production language models against harmful behavior, but although results seem promising in testing, they provide no...
Giordano De Marzo, Alessandro Bellina, Claudio Castellano +2 more
Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly...
Krishak Aneja, Manas Mittal, Anmol Goel +2 more
Fine-tuning Large Language Models (LLMs) on benign narrow data can sometimes induce broad harmful behaviors, a vulnerability termed emergent...
Tianyuan Zhang, Peng Yue, Zihao Peng +8 more
Multimodal large language models (MLLMs) are increasingly integrated into autonomous driving (AD) systems; however, they remain vulnerable to diverse...
Wenxin Tang, Xiang Zhang, Junliang Liu +11 more
Automated vulnerability detection is a fundamental task in software security, yet existing learning-based methods still struggle to capture the...
Aleksandr Bowkis, Marie Davidsen Buhl, Jacob Pfau +1 more
A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as...
Leo Linqian Gan, Jeffery Wu, Longyuan Ge +6 more
Autonomous LLM agents face a critical security risk known as workflow hijacking, where attackers subtly alter tool and skill invocations. Existing...
Guoxin Lu, Letian Sha, Qing Wang +4 more
The safety alignment of Large Language Models (LLMs) remains vulnerable to Harmful Fine-tuning (HFT). While existing defenses impose constraints on...
Siyuan Li, Aodu Wulianghai, Xi Lin +6 more
The increasing prevalence of Large Language Models (LLMs) in content creation has made distinguishing human-written textual content from...
Xinjie Shen, Rongzhe Wei, Peizhi Niu +6 more
Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful...
Fabrice Harel-Canada, Amit Sahai
LLM watermarks must be detectable without compromising text quality, yet most existing schemes bias the next-token distribution and pay for detection...
Marco Arazzi, Vignesh Kumar Kembu, Antonino Nocera +2 more
The open-source ecosystem has accelerated the democratization of Large Language Models (LLMs) through the public distribution of specialized Low-Rank...
Hanum Ko, Sangheum Yeon, Jong Hwan Ko +1 more
As DRAM scales in density and adopts 3D integration, raw fault rates increase and multi-bit errors are no longer rare. Such errors can severely...
Zhenning Yang, Yuhan Chen, Patrick Tser Jern Kon +5 more
To unleash the full potential of AI for Science, we must untether the agents from a purely digital environment. The agent's ability to control and...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,371+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial