CoT-Guard: Small Models for Strong Monitoring
Nirav Diwan, Han Wang, Berkcan Kapusuzoglu +6 more
Monitoring the chain-of-thought (CoT) of reasoning models is a promising approach for detecting covert misbehavior (i.e., hidden objectives) in code...
AI Threat Alert indexes 3,771+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 661–680 of 891 papers
Clear filtersNirav Diwan, Han Wang, Berkcan Kapusuzoglu +6 more
Monitoring the chain-of-thought (CoT) of reasoning models is a promising approach for detecting covert misbehavior (i.e., hidden objectives) in code...
Shravan Doda
Final-token safety probes monitor a single hidden state after prompt prefill, but jailbreak prompts can contain probe-visible unsafe evidence...
Hao Wang, Hanchen Li, Qiuyang Mang +3 more
Agent benchmarks have become the de facto measure of frontier AI competence, guiding model selection, investment, and deployment. However, reward...
Sam Herring, Jake Naviasky, Karan Malhotra
Language models are instruction-tuned to refuse harmful requests, but the mechanisms underlying this behavior remain poorly understood. Popular...
Sae Furukawa, Alina Oprea
Supervised Finetuning (SFT) has become one of the primary methods for adapting a large language model (LLM) with extensive pre-trained knowledge to...
Sina Mavali, David Pape, Jonathan Evertz +5 more
Terminal agents are increasingly capable of executing complex, long-horizon tasks autonomously from a single user prompt. To do so, they must...
Olga Sorokoletova, Francesco Giarrusso, Giacomo De Luca +6 more
Large language models are increasingly deployed in safety-critical applications, where their ability to resist harmful instructions is essential....
Yuhao Wu, Tung-Ling Li, Hongliang Liu
Agent skills extend LLM agents with privileged third-party capabilities such as filesystem access, credentials, network calls, and shell execution....
Xinyi Zeng, Xue Yang, Jingyuan Zhang +5 more
Multimodal large language models (MLLMs) are gaining increasing attention. Due to the heterogeneity of their input features, they face significant...
Pranshav Gajjar, Vijay K Shah
This position paper argues that to achieve Level 5 autonomous 6G networks, the next generation of Artificial Intelligence in Radio Access Networks...
Fanxiao Li, Jiaying Wu, Tingchao Fu +3 more
Multi-agent systems (MAS) powered by large language models (LLMs) increasingly adopt planner--executor architectures, where planners convert prompts...
Khondaker Tasnia Hoque, Toukir Ahammed
Flaky tests, which exhibit non-deterministic pass/fail behavior for the same version of code, pose significant challenges to reliable regression...
John T. Halloran
Large language model (LLM) alignment algorithms typically consist of post-training over preference pairs. While such algorithms are widely used to...
Christian Moya, Alex Semendinger, Guang Lin +1 more
Preference learning methods such as Direct Preference Optimization (DPO) are known to induce reliance on spurious correlations, leading to sycophancy...
Fatima Z. Abacha, Sin G. Teo, Yuanxiang Wu +2 more
Federated Learning remains highly susceptible to backdoor attacks--malicious clients inject targeted behaviours into the global model. Existing...
Youssef Zaazou, Mark Thomas
Vision-language models (VLMs), such as CLIP and SigLIP 2, are widely used for image classification, yet their vision encoders remain vulnerable to...
Joel Rorseth, Parke Godfrey, Lukasz Golab +2 more
This paper demonstrates RUBEN, an interactive tool for discovering minimal rules to explain the outputs of retrieval-augmented large language models...
Pedro Conde, Henrique Branquinho, Valerio Mazzone +3 more
AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will...
Saba Pourhanifeh, AbdulAziz AbdulGhaffar, Ashraf Matrawy
Large Language Models(LLMs) are increasingly explored for cybersecurity applications such as vulnerability detection. In the domain of threat...
Johann Knechtel, Ozgur Sinanoglu, Ramesh Karri
The integration of Large Language Models (LLMs) into Electronic Design Automation (EDA) and hardware security is rapidly reshaping the semiconductor...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,771+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial