aiAuthZ: Off-Host, Identity-Bound Authorization for AI Agents
Sai Varun Kodathala
AI agents issue tool calls on the basis of text they cannot verify, so any party who controls part of the context can forge the appearance of...
AI Threat Alert indexes 3,397+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 301–320 of 1,755 papers
Clear filtersSai Varun Kodathala
AI agents issue tool calls on the basis of text they cannot verify, so any party who controls part of the context can forge the appearance of...
Paul K. Mandal, Pavan Reddy, Tristan Malatynski
Model-specific adversarial attacks have been extensively studied. We study a different failure mode: naturally occurring statistical signals in...
Kristina Nikolić, Egor Zverev, Javier Rando +3 more
Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data....
Konstantin Berlin
Detection models running in adversarial environments face a malicious distribution that drifts rapidly while the benign distribution stays...
Yechao Zhang, Shiqian Zhao, Jiawen Zhang +5 more
Persistent personal agents combine long-term memory with access to users' external environments, enabling personalized foreground assistance and...
Woohyuk Choi, Juhee Kim, Taehyun Kang +3 more
AI agents act on behalf of user prompts, consuming external data and taking actions based on the agent context. Prior research on AI agent security...
Neeraj Karamchandani, Piyush Nagasubramaniam, Sencun Zhu +1 more
Persistent memory has enabled large language model (LLM) agents to store factual knowledge, prior decisions, reasoning histories, tool usage...
Mouhamed Amine Bouchiha, Gregory Blanc
Cyber Threat Intelligence (CTI) reports are predominantly unstructured, heterogeneous, and noisy, which limits their direct usability for automated...
Mouhamed Amine Bouchiha, Gregory Blanc
Cyber Threat Intelligence (CTI) reports are predominantly unstructured, heterogeneous, and noisy, which limits their direct usability for automated...
Yuanmin Xie, Xiangfan Wu, Wenhao Wu +6 more
Algorithmic Complexity Vulnerabilities (ACVs) arise when adversarial inputs trigger worst-case execution behavior, causing severe performance...
Samira Hajizadeh
Safety alignment in large language models is typically evaluated against direct, imperative harmful requests. We show that this alignment is highly...
Om Solanki, Lopamudra Praharaj, Deepti Gupta +1 more
This paper presents an adversarial security study of the Policy-Aware LLM Retrieval-Augmented Generation (PA-LLM-RAG) framework for Internet of...
Bogdan Banu
The transition of Large Language Models (LLMs) from passive generators to autonomous agents has introduced significant challenges in reliability,...
Amit LeVi, Elad David, Max Fomin
Interpretability methods aim to reveal the features represented inside large language models (LLMs). Many existing methods begin with labeled...
Jingfeng Wu, Yiyuan He, Minxian Xu +7 more
Online large language model (LLM) serving has become the backbone of modern AI applications, powering diverse downstream services through shared...
Josh Hills, Ida Caspary, Asa Cooper Stickland
As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence...
Mona Schirmer, Metod Jazbec, Alexander Timans +3 more
Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when...
Bohan Liu, Wenqian Ye, Guangzhi Xiong +3 more
Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language...
Zimo Ji, Congying Xu, Zongjie Li +4 more
LLM coding agents increasingly rely on third-party agent skills from public marketplaces, which execute with the agent's privileges and create a...
Seren Yenikent, Jack Vinijtrongjit, Katherine Ng
Mental health disorders affect nearly one billion people globally, yet 75% of individuals in low- and middle-income countries receive no treatment...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,397+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial