Human-on-the-Bridge: Scalable Evaluation for AI Agents
Fouad Bousetouane
AI agents must be evaluated as behavioral systems, not as isolated response generators. They reason across turns, call tools, preserve context,...
AI Threat Alert indexes 3,771+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 501–520 of 888 papers
Clear filtersFouad Bousetouane
AI agents must be evaluated as behavioral systems, not as isolated response generators. They reason across turns, call tools, preserve context,...
Chen Chen, Xiang Gao, Xianshun Wang +6 more
Split learning provides a practical paradigm for resource-constrained users to train Large Language Models (LLMs) by offloading computation-intensive...
Ziniu Liu, Aiping Li
When a person's records appear in k independent data silos, each protected by (epsilon, delta)-differential privacy, standard composition yields a...
Rowdy Chotkan, Bulat Nasrulin, Johan Pouwelse +1 more
Distributed systems handle adversarial nodes through redundancy, which imposes a significant performance overhead. In blockchain systems, Byzantine...
Hankyul Baek, Jaewon Noh, Sang Seo +5 more
AI agents are increasingly being adopted in enterprise and personal settings with access to emails, databases, documents, and other tools where they...
Sipeng Xie, Qianhong Wu, Hengrun Lu +4 more
Agents increasingly access large language models (LLMs) through API routers. A router terminates the client's transport-layer security session and...
Hao-Hsuan Chen
Paper A defines a time-consistent actuarial runtime that prices each side-effect-bearing action against a contractually fixed safe default and gates...
Liuyang Yao, Zhouyu Li, Junguang He +1 more
AI systems are increasingly deployed for credit assessment and investment advisory in global financial markets, yet the integrity of their inference...
Jiahao Zhang, Xiuyu Li, Suhang Wang
As Large Language Model (LLM) APIs become ubiquitous, users increasingly rely on black-box fingerprinting to verify that providers are serving the...
Ismail Hossain, Sai Puppala, Md Jahangir Alam +2 more
Open-source LLM agent ecosystems are growing rapidly, yet the security of community-contributed skills - modular tool definitions that extend agent...
Liran Tal, Johannes Kloos, Arsenii Rudich +2 more
We ran 300 repeated vulnerability-finding scans to measure how repeatable agentic large language model (LLM) security review is on the same...
Achraf Hsain, Sultan Almuhammadi
Shielded reinforcement learning is typically presented as a runtime safety mechanism that compiles temporal-logic specifications into automata...
Jiaqi Luo, Jiarun Dai, Zhile Chen +8 more
Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical red lines...
Xu Yang, Zhizhou Sha, Junbo Li +10 more
As AI-generated reviews move from experimental tools into peer-review infrastructure, most robustness concerns have focused on explicit attacks such...
Haowei Qian
As LLM agents proliferate in prediction markets and collective decision-making, they risk a cognitive monoculture: agents built on shared foundation...
Junfeng Guo Heng Huang
While real-world applications of reinforcement learning (RL) are becoming increasingly popular, the security of RL systems deserve more attention and...
Md Jafrin Hossain, Mohammad Arif Hossain, Weiqi Liu +1 more
Agentic large language model systems that autonomously invoke tools, maintain persistent memory, and execute multi-step plans are increasingly...
Andy Wang, Parv Mahajan, David Demitri Africa +3 more
Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model...
Timothy McAllister, Sina Abdidizaji, Ivan Garibay +1 more
As LLM-based multi-agent systems (MAS) are deployed in the wild, the resilience of their collaboration structures against adversarial compromise...
Tarun Sharma
Retrieval-augmented generation (RAG) agents increasingly run with persistent memory that accumulates across user sessions. This creates a new attack...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,771+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial