LACUNA: Safe Agents as Recursive Program Holes
Yaoyu Zhao, Yichen Xu, Oliver Bračevac +3 more
LLM agents increasingly act by writing code, yet a split persists between the runtime that drives the agent and the code the model writes. The...
AI Threat Alert indexes 3,771+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 1041–1060 of 3,771 papers
Yaoyu Zhao, Yichen Xu, Oliver Bračevac +3 more
LLM agents increasingly act by writing code, yet a split persists between the runtime that drives the agent and the code the model writes. The...
Jaydip Sen
Artificial Intelligence has achieved remarkable success across diverse application domains. However, its vulnerability to adversarial attacks poses...
Matteo Gioele Collu, Riccardo Conte, Alberto Giaretta +4 more
In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained...
Avidan Shah, Jannik Brinkmann, Rico Angell
As LLMs gain stronger reasoning capabilities, their extended chain-of-thought introduces new degrees of complexity for defending against adversarial...
Taojie Zhu, Wentao Zhao, Rui Sun +7 more
Evaluating whether large language model (LLM) agents can profit in capital markets is increasingly framed as end-to-end trading: place an agent in a...
Chao Ding, Mouxiao Bian, Tianbin Li +12 more
Large language models(LLMs) increasingly match expert performance on licensing examinations, yet routine clinical use remains limited because...
Chenxi Wang, Ruiyang Huang, Jiayan Sun +2 more
Latent-based multi-agent systems replace parts of explicit inter-agent communication with hidden representations, offering a new direction for...
Yongxiang Li, Moxin Li, Zhixin Ma +4 more
Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into...
Víctor Mayoral-Vilches
We present CAI Dataset, a fourteen-month corpus of cybersecurity LLM trajectories collected through the open-source CAI agent framework, built in...
Yujie Ma, Jialin Rong, Chenxi Yang +4 more
Large Language Models(LLMs) have been actively integrated into modern software systems as critical components. LLM-in-the-loop vulnerabilities, where...
Ruoqi Guo, Yi Liu, Gelei Deng +7 more
Mobile graphical user interface (GUI) agents driven by vision-language models (VLMs) perceive the screen as rendered pixels and choose actions from...
Junjie Mu, Qiongxiu Li
Federated Retrieval-Augmented Generation (FedRAG) is attractive for privacy-sensitive applications because raw data remain local. As a result,...
Jiachen Qian
Retrieval-Augmented Generation (RAG) mitigates LLM hallucinations but introduces a critical vulnerability: corpus integrity. We present...
Yu Yin, Shuai Wang, Bevan Koopman +1 more
Recent generative engine optimisation (GEO) research has shown that prompt-injection attacks can push a target product to the top of an LLM's...
Yongwoo Kim, Sojung An, Yunjin Park +8 more
Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and...
Yuan Tian, Bing Hu, Fang Wu +3 more
Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly...
Xuesi Hu, Peng Wang, Jinpeng Miao +7 more
Recently, large language models (LLMs) have achieved superior performance in static financial reasoning and simple dynamic trading tasks. However,...
Xiang Fang, Wanlong Fang
Large Language Models (LLMs) are increasingly vulnerable to adversarial prompts that exploit semantic ambiguities to bypass safety mechanisms,...
Qiyuan Wang, Yao Li, Raymond K. W. Wong
Despite recent progress in backdoor attacks, existing methods remain susceptible to post-training defenses that erase the backdoor through...
Akshaj Murhekar, Abhijit Mishra
Recent advances in large language models have accelerated open-vocabulary EEG-to-imagined-text decoding, where non-invasive neural activity recorded...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,771+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial