Approval Integrity and Recovery in LLM Answer Publication
Faruk Alpay, Taylan Alpay
Publication integrity in LLM systems requires binding approved content to its current authorization context. We examine exact-content binding,...
AI Threat Alert indexes 3,771+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 21–40 of 523 papers
Clear filtersFaruk Alpay, Taylan Alpay
Publication integrity in LLM systems requires binding approved content to its current authorization context. We examine exact-content binding,...
Mark Russinovich, Blake Bullwinkel, Giorgio Severi +2 more
Language model safety is typically evaluated one interaction at a time. We show that a weaker, unaligned model can split a harmful task into...
Xue Tan, Changhui Wang, Sanrui Yang +6 more
Modern large language model (LLM) agents often construct prompts by aggregating retrieved passages, user reviews, and documents from multiple...
Xue Tan, Xuandi Zeng, Yu Shao +5 more
Modern Vision-Language Retrieval-Augmented Generation (VLRAG) systems augment Large Vision-Language Models (LVLMs) with retrieved visual and textual...
Juanen Li, Peng Qian, Guanyan Li +6 more
Smart contracts, serving as the cornerstone of decentralized applications, autonomously manage trillion-dollar digital assets, making them attractive...
Bokang Zeng, Zheng Gao, Xiaoyu Li +2 more
Watermarking the final patch produced by a coding agent provides provenance evidence for the submitted artifact, but does not authenticate the...
Zhantong Xue, Pingchuan Ma, Zhaoyu Wang +3 more
Zero-knowledge (ZK) proof systems for neural-network inference compile the model into a system of arithmetic constraints. Many of these constraints...
Ryan Lum, Yongfeng Zhang
AI systems become more useful when they can adapt to the people using them, but in multi-agent systems, useful context learned by one agent often...
Jing Guan, Yachao Yang, Zhaoliang Liu +3 more
Large language models remain fragile against malicious fine-tuning, motivating training-time defenses against harmful persona drift. Preventative...
Arnab Chattopadhayay, Debdipta Halder
Large language model agents produce fluent action sequences across a wide range of tasks, yet they fail in characteristic ways once the environment...
Rafael M. Mamede, Pedro C. Neto, Ana F. Sequeira
Deepfake detectors remain vulnerable to transfer-based black-box attacks, in which adversarial examples are generated on a source surrogate model and...
Iliano Fasolino
Retrieval-augmented generation (RAG) grounds a language model in retrieved documents, which reduces hallucination but creates a new attack surface:...
Hyun Gu Kang, Daniil Gurgurov, Tanja Baeumel +2 more
This study examines the compositionality of steering vectors for language and behavioral control in large language models. Focusing on language,...
Han Jin
We introduce HoneyRoute, an inference-serving layer that detects whether an incoming request is malicious and, if so, routes it to a dedicated...
Xinhong Xie, Piyush Nagasubramaniam, Neeraj Karamchandani +1 more
Large language model (LLM) agents are increasingly employed for offensive cybersecurity tasks such as automated vulnerability discovery,...
Xiao Ma, Hong Shen, Hui Tian +2 more
This paper proposes a robust decentralized personalized federated learning method R-DPFL, that enables clients to reduce the impact of Byzantine...
Branka Stojanović, Andreas Flatscher, Michael Somma
Anomaly-based intrusion detection systems in industrial control systems (ICS) and operational technology (OT) environments are increasingly required...
Syed Ghazanfar Abbas, Dongyan Xu
Large language model (LLM) security has largely focused on role-playing jailbreaks, with less attention to what happens when a user asks an LLM to...
Jiarui Li, Jiahao Chen, Chunyi Zhou +5 more
Reusable agent skills extend large language model (LLM) agents with task procedures, tool-use guidance, and output constraints. Yet these skills also...
Gang-Hyun Park, Ju-Hyeong Lee, Hee-Youl Kwak +1 more
Multi-bit watermarking for large language models (LLMs) enables content source tracing by embedding user-identifiable messages into generated text....
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,771+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial