Multi-Agent AI Safety as an Institutional Design Problem
Abdullah X
AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent...
AI Threat Alert indexes 3,371+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 1–10 of 10 papers
Clear filtersAbdullah X
AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent...
Juncheng Dong, Ding Tong, Ishan Gupta +1 more
Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preference. As...
Joseph Bingham
Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment....
Praphul Chandra, Sujit Gujar, Ganesh Ghalme
We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mechanism is built on the principle...
Abdulkadir Külçe, Alihan Esen, Cağla Fikir +4 more
This paper presents ECHO (Enhanced Care \& Health Observer), a locally-deployable conversational health assistant for long-term chronic care...
Buzhao Liu, Xinhang Ma, Yevgeniy Vorobeychik
The remarkable instruction-following ability of modern LLMs has enabled their practical use as the minds of agents that can autonomously complete...
Alina Klerings, Jannik Brinkmann, Heiner Stuckenschmidt +1 more
Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep...
Ethen Santana, Gabriel Gyaase, Hao Zheng
Hardware designs, like software, are susceptible to bugs that can introduce security vulnerabilities and create opportunities for malicious...
Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi +6 more
Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated...
Ben Falchuk, Himanshu Garg, Euthimios Panagos +1 more
Just like software and hardware, business processes are susceptible to vulnerabilities that can lead to product quality issues, delays, and increased...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,371+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial