Forensic Trajectory Signatures for Agent Memory Poisoning Detection
Jun Wen Leong
We discover a behavioral invariant in LLM agents under persistent memory poisoning: in architectures where routing information is retrieved through...
AI Threat Alert indexes 3,771+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 421–440 of 888 papers
Clear filtersJun Wen Leong
We discover a behavioral invariant in LLM agents under persistent memory poisoning: in architectures where routing information is retrieved through...
Dvir Alsheich, Adar Peleg, Ben Hagag +3 more
The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where specialized agents collaborate to...
Chen Frydman, Aviram Zilberman, Rubin Krief +6 more
Mitigating an observed adversary in an enterprise network typically takes weeks of expert work: an analyst derives a mitigation tailored to that...
Buğra Alperen Uluırmak, Rifat Kurban
LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while...
Nicole Geumheon Liu, Haeun Jang, Yonghyun Jun +1 more
Large Language Models (LLMs) often fail to maintain instruction hierarchies (IH) when processing multi-source inputs with varying role-level...
Yikai Hua, Peter West
Large Vision-Language Models (VLMs) are increasingly deployed as content moderation tools, yet they remain vulnerable to jailbreak attacks in which...
Yingjie Wang, Yi Dong, Edmund Lau +3 more
Rare events govern the safety profile of modern AI systems, yet their probabilities are extremely difficult to estimate: direct Monte Carlo requires...
Alex Kwon
LLM agents carry conclusions across steps and sessions in compressed memory, and memory products (e.g., mem0, LangMem) rewrite conversation into...
Ryan Fetterman
LLMs fine-tuned for security classification are usually evaluated on held-out examples from the same distribution as their training data. We show...
Padmaraj Madatha
LLM coding harnesses grant agents broad file and shell access, yet the configuration layer that steers them -- rules files, agent definitions,...
Nada Lahjouji, Ashwin Gerard Colaco
Large language model agents increasingly query databases, search document collections, call external APIs, remember past interactions, and act on a...
Manar Alsaid, Chimdumebi Nebolisa, Faris Abbas
Security misconfigurations in Terraform Infrastructure-as-Code are a growing risk in cloud deployments, and large language models are increasingly...
A. S. Ushakov, Yu. N. Berdinsk
We propose a complete architectural blueprint for safe artificial general intelligence based on a closed reentry loop (D <-> I cycle). In contrast to...
Praneet Suresh, Jack Stanley, Sonia Joseph +2 more
Pre-trained transformers have demonstrated remarkable generalization abilities, at times extending beyond the scope of their training data. Yet,...
Poojitha Thota, Yun Lei, Santhosh Thangaraj +2 more
Large language models (LLMs) are increasingly deployed in interactive applications, yet they remain vulnerable to adversarial interactions that...
William Aiken, Paula Branco, Guy-Vincent Jourdan +1 more
Noise-based backdoor attacks on diffusion models typically rely on input-time trigger injection, untargeted activation, and out-of-distribution...
Jintao Huang, Fengqing Jiang, Radha Poovendran +1 more
We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulnerability...
Poojitha Thota, Shirin Nilizadeh
Training-time data poisoning during fine-tuning poses a significant threat to large language models (LLMs) deployed for abstractive text...
Nasrin Malekzadeh Goradel, Niccolo Pancino, Yaser Gholizade Atani +3 more
Several theoretical works have tried to explain the adversarial vulnerability of deep neural networks through properties of high-dimensional...
Hyejun Jeong, Dzung Pham, Amir Houmansadr +1 more
To better assist users with completing challenging tasks, AI agents mediate communications, access data, and interact with different APIs. Many...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,771+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial