OCELOT: Inference-Leakage Budgets for Privacy-Preserving LLM Agents
Jin Xie, Songze Li
Large language model (LLM) agents increasingly act on a user's behalf -- reading personal files, calling tools, transacting with external services --...
AI Threat Alert indexes 3,371+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 101–120 of 586 papers
Clear filtersJin Xie, Songze Li
Large language model (LLM) agents increasingly act on a user's behalf -- reading personal files, calling tools, transacting with external services --...
Pinran Gao, Lingxiang Wang, Ying Zhang +1 more
The rapid integration of large language models (LLMs) into mobile applications has introduced a new class of credential security risk: leaked...
Malikeh Ehghaghi, Boglárka Ecsedi, Marsha Chechik +1 more
Adversarial robustness evaluations of large language models (LLMs) typically report attack success rate (ASR) under fixed query budgets, implicitly...
Naihao Deng, Yilun Zhu, Naichen Shi +2 more
Warning: This paper contains several toxic and offensive statements. Modern large language models (LLMs) are typically aligned through large-scale...
Peiyang Li, Songping Wang, Yi Huang +9 more
Autonomous AI agents have driven the transition from conversation to task execution, shifting security failures from textual deception to system...
Saeid Jamshidi, Amin Nikanjam, Arghavan Moradi Dakhel +2 more
Large Language Models (LLMs) in multi-turn interactions maintain evolving context rather than generating isolated responses, making them vulnerable...
Carlos S. Sepúlveda, Gonzalo A. Ruz
Neural combinatorial optimization (NCO) trains autoregressive policies to solve routing problems. The standard training algorithm, REINFORCE with a...
Shixiong Jiang, Taozheng Zhu, Fanxin Kong
Offline safe reinforcement learning (Safe RL) enables policy learning without online interactions, making it suitable for safety-critical systems...
Yinan Wang
AI Scientist agents are often evaluated as if capability were mainly a function of model quality, prompting, or reasoning scaffolds. We test a...
Ethan Rathbun, Ahmed Agha, Saaduddin Mahmud +3 more
World models have recently seen a rapid growth in both their popularity and capability as more data efficient tools for generating robot training...
Srravya Chandhiramowuli, Ding Wang, Alex Taylor
We present an ethnographic study of an alternative approach to data work, developed by a civic-tech initiative that builds datasets for training and...
Bartłomiej Marek, Lorenzo Rossi, Vincent Hanke +4 more
Recent work has applied differential privacy (DP) to adapt large language models (LLMs) for sensitive applications, offering theoretical guarantees....
Ziqian Zhong, Ivgeni Segal, Ivan Bercovich +3 more
Agent benchmarks score submissions with outcome verifiers that are typically hand-written and brittle, leaving them open to reward hacking. We audit...
Victor De Marez, Luna De Bruyne, Walter Daelemans
Factual sycophancy occurs when a language model abandons a correct, verifiable answer under social pressure. Because a flip occurs only when pressure...
Vincent Koc, Patrick Erichsen, Jacob Tomlinson +3 more
Agent skills extend AI agents with reusable instructions, tools, scripts, references, and workflows, establishing a security boundary distinct from...
Abdelrahman Abouelenein, Marwan Torki
It is crucial for modern on-device AI systems that rely on retrieval-augmented inference to release and share datastores without compromising...
Mark Vero, Fabian Kaczmarczyck, Ivan Petrov +6 more
Honeypots are decoy systems mimicking real system components designed to defend against cyber attacks. Recently, LLMs increasingly serve as...
Víctor Mayoral-Vilches
We present CAI Dataset, a fourteen-month corpus of cybersecurity LLM trajectories collected through the open-source CAI agent framework, built in...
Yujie Ma, Jialin Rong, Chenxi Yang +4 more
Large Language Models(LLMs) have been actively integrated into modern software systems as critical components. LLM-in-the-loop vulnerabilities, where...
Yongwoo Kim, Sojung An, Yunjin Park +8 more
Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,371+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial