Latent-space Attacks for Refusal Evasion in Language Models
Giorgio Piras, Raffaele Mura, Fabio Brau +4 more
Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal...
AI Threat Alert indexes 3,771+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 1121–1140 of 3,771 papers
Giorgio Piras, Raffaele Mura, Fabio Brau +4 more
Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal...
Shahnewaz Karim Sakib, Swati Kar, Anindya Bijoy Das
Large Language Models (LLMs) are widely deployed in diverse real-world settings, yet remain vulnerable to jailbreaking, where prompt-based attacks...
Heajun An, Qi Zhang, Vedanth Achanta +1 more
Large language models (LLMs) are increasingly embedded in adolescent digital environments, mediating information seeking, advice, and emotionally...
Roland Pihlakas, Jan Llenzl Dagohoy
Large language models (LLMs) are increasingly deployed as autonomous agents that make sequences of decisions over extended interactions in...
Pengyu Sun, Qishu Jin, Enhao Huang +4 more
Model Context Protocol (MCP) has emerged as a standard interface for connecting LLM agents to external tools. Because MCP servers expose privileged...
Abdullah Al Nomaan Nafi, Fnu Suya, Swarup Bhunia +1 more
Jailbreak attacks expose a persistent gap between the intended safety behavior of aligned large language models and their behavior under adversarial...
Samuele Pasini, Jinhan Kim, Paolo Tonella
Modern DNNs are repeatedly fine-tuned to incorporate new data and functionality. This evolutionary workflow introduces a security risk when updated...
Scott Freitas, Amir Gharib
Defending against today's increasingly sophisticated cyberattacks requires security analysts to continuously translate evolving attacker tradecraft...
Leitao Yuan, Qinghua Mao, Daizong Liu +5 more
Multimodal large language models (MLLMs) remain vulnerable to transfer-based targeted attacks, where perturbations optimized on open-source surrogate...
Laura Jiang, Reza Ryan, Qian Li +1 more
Single-turn safety evaluation is a poor proxy for real fraud defense, where attackers escalate across multiple rounds. This paper evaluates fraud...
Amit Roth, Ankur Samanta, Matan Halevy +2 more
Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking,...
Jiachen Ma, Jiawen Zhang, Xiangtian Li +3 more
While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that...
Yifei Wang, Tianlin Li, Xiaohan Zhang +3 more
Inference optimization is a vital technique for deploying LLMs at scale. Compilation is the most widely adopted optimization technique for LLMs....
Yong Jin Chun, Iftekhar Ahmed
Large Language Models (LLMs) have enabled collaborative Multi-Agent (MA) systems, where interacting agents improve performance through diverse...
Richard J. Young, Gregory D. Moody
The evaluation of large language model refusal on malicious-coding tasks now spans at least thirteen publicly released prompt corpora (AdvBench, the...
Chengcai Gao, Zhihong Sun, Xiaochuan Shi +2 more
The growing adoption of Retrieval-Augmented Generation (RAG) has led to a rise in adversarial attacks. Existing defenses, relying on semantic...
Mohammed Alshaalan, Miguel R. D. Rodrigues
Optimization-based adversarial suffixes can jailbreak aligned large language models (LLMs) while remaining fluent, weakening static and windowed...
Florian A. D. Burnat, Brittany I. Davidson
Multi-tenant retrieval-augmented generation (RAG) services advertise per-account differential privacy as the operative leakage boundary: each...
Isaac David, Arthur Gervais
Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents?...
Yasmine Hayder
Knowledge Graphs (KGs) are a powerful representation of linked data, offering flexibility, semantic richness, and support for knowledge enrichment...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,771+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial