Human-on-the-Bridge: Scalable Evaluation for AI Agents
Fouad Bousetouane
AI agents must be evaluated as behavioral systems, not as isolated response generators. They reason across turns, call tools, preserve context,...
AI Threat Alert indexes 3,771+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 881–900 of 3,771 papers
Fouad Bousetouane
AI agents must be evaluated as behavioral systems, not as isolated response generators. They reason across turns, call tools, preserve context,...
Yimeng Chen, Zhe Ren, Firas Laakom +3 more
Large language model (LLM)-based search agents synthesize open-web content into actionable recommendations on behalf of users, creating a risk that...
Chen Chen, Xiang Gao, Xianshun Wang +6 more
Split learning provides a practical paradigm for resource-constrained users to train Large Language Models (LLMs) by offloading computation-intensive...
Ziniu Liu, Aiping Li
When a person's records appear in k independent data silos, each protected by (epsilon, delta)-differential privacy, standard composition yields a...
Qi Wang, Chengcheng Wan, Weijia He +4 more
Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. However, their safety remains a critical concern...
Rowdy Chotkan, Bulat Nasrulin, Johan Pouwelse +1 more
Distributed systems handle adversarial nodes through redundancy, which imposes a significant performance overhead. In blockchain systems, Byzantine...
Bojie Li
A personalized AI agent needs a user memory: a persistent model of who the user is, built across many conversations and consulted on each new one....
Xuanyu Yin, Yilin Jiang, Jun Zhou +3 more
As large language models (LLMs) are increasingly deployed in user-facing systems, black-box jailbreak defense has become an important practical...
Hankyul Baek, Jaewon Noh, Sang Seo +5 more
AI agents are increasingly being adopted in enterprise and personal settings with access to emails, databases, documents, and other tools where they...
Sipeng Xie, Qianhong Wu, Hengrun Lu +4 more
Agents increasingly access large language models (LLMs) through API routers. A router terminates the client's transport-layer security session and...
Hao-Hsuan Chen
Paper A defines a time-consistent actuarial runtime that prices each side-effect-bearing action against a contractually fixed safe default and gates...
David Huang, Jaewon Chang, Avidan Shah +2 more
The Rapid Response (RR) framework, deployed in production systems, including Anthropic's ASL-3 safeguards, continuously improves jailbreak-detection...
Liuyang Yao, Zhouyu Li, Junguang He +1 more
AI systems are increasingly deployed for credit assessment and investment advisory in global financial markets, yet the integrity of their inference...
Jiahao Zhang, Xiuyu Li, Suhang Wang
As Large Language Model (LLM) APIs become ubiquitous, users increasingly rely on black-box fingerprinting to verify that providers are serving the...
Ismail Hossain, Sai Puppala, Md Jahangir Alam +2 more
Open-source LLM agent ecosystems are growing rapidly, yet the security of community-contributed skills - modular tool definitions that extend agent...
Yuyang Dai, Yushun Dong
Large language models deployed as commercial APIs are vulnerable to model extraction attacks, while existing defenses either act too late or degrade...
Aman Anifer, Vignesh Kumar Kembu, Vishnu M +4 more
Large Language Models (LLMs) constitute pivotal components within the AI-dominated information technology ecosystem. To mitigate risks associated...
Liran Tal, Johannes Kloos, Arsenii Rudich +2 more
We ran 300 repeated vulnerability-finding scans to measure how repeatable agentic large language model (LLM) security review is on the same...
Achraf Hsain, Sultan Almuhammadi
Shielded reinforcement learning is typically presented as a runtime safety mechanism that compiles temporal-logic specifications into automata...
Zihao Wang, Yiming Li, Yutong Wu +8 more
Web agents driven by large language models (LLMs) are increasingly deployed in real-world environments, where they operate over untrusted web content...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,771+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial