Multi-Agent AI Safety as an Institutional Design Problem
Abdullah X
AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent...
AI Threat Alert indexes 3,371+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 1–20 of 471 papers
Clear filtersAbdullah X
AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent...
Juncheng Dong, Ding Tong, Ishan Gupta +1 more
Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preference. As...
Joseph Bingham
Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment....
Praphul Chandra, Sujit Gujar, Ganesh Ghalme
We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mechanism is built on the principle...
Abdulkadir Külçe, Alihan Esen, Cağla Fikir +4 more
This paper presents ECHO (Enhanced Care \& Health Observer), a locally-deployable conversational health assistant for long-term chronic care...
Buzhao Liu, Xinhang Ma, Yevgeniy Vorobeychik
The remarkable instruction-following ability of modern LLMs has enabled their practical use as the minds of agents that can autonomously complete...
Alina Klerings, Jannik Brinkmann, Heiner Stuckenschmidt +1 more
Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep...
Ethen Santana, Gabriel Gyaase, Hao Zheng
Hardware designs, like software, are susceptible to bugs that can introduce security vulnerabilities and create opportunities for malicious...
Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi +6 more
Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated...
Ben Falchuk, Himanshu Garg, Euthimios Panagos +1 more
Just like software and hardware, business processes are susceptible to vulnerabilities that can lead to product quality issues, delays, and increased...
Alexander Meulemans, Maciej Wołczyk, Marissa A. Weis +11 more
As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles...
Nizhang Li, Zonghao Ying, Xiangfan Wu +7 more
External skills extend the capabilities of large language model agents, but also introduce an execution-time attack surface: a skill that appears...
Yiming Chen, Kemou Li, Haiwei Wu +1 more
Backdoor attacks in multimodal contrastive learning (MCL) have garnered growing attention in recent years, as many downstream tasks critically depend...
Francis Heylighen
AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate humanity, or...
Sungju Yun, Sijune Hwang, Yeonjoon Lee +2 more
Detecting source code vulnerabilities is increasingly difficult as modern security flaws are rooted in complex causal dependencies between execution...
Kaustuv Mukherji, Jaikrishna Manojkumar Patil, Colton Payne +4 more
Large language models are increasingly used to reason about software vulnerabilities, but their outputs can silently violate domain knowledge,...
Pingyu Wu, Lingyao Zhu, Weiming Zhang +1 more
Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks:...
Xiangyu Yin, Tora Bodin, Rohan Menon +1 more
Inference time defences against vision language model jailbreaks often subtract a calibrated direction from the residual stream at a chosen decoder...
Yongjian Guo, Wanlun Ma, Lingyu Shen +2 more
Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers...
Rui Yang, Michael Fu, Kla Tantithamthavorn +2 more
Software engineering teams now deploy AI coding agents (Cursor, Claude Code, GitHub Copilot) as first-class productivity tools, installing...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,371+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial