PAST2HARM: A Simple Adaptive Past Tense Attack for Jailbreaking Multimodal AI
Snehasis Mukhopadhyay
Jailbreak attacks on multimodal AI systems remain underexplored, even though unsafe image generation can have more severe consequences than unsafe...
AI Threat Alert indexes 3,406+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 701–720 of 1,756 papers
Clear filtersSnehasis Mukhopadhyay
Jailbreak attacks on multimodal AI systems remain underexplored, even though unsafe image generation can have more severe consequences than unsafe...
Tamerlan Aghayev, Maxime Elkael, Michele Polese +11 more
Cellular research and development (R&D) is throttled by six structural processes that each consume months of manual engineering work per iteration:...
Dongyoon Hahm, Dylan Hadfield-Menell, Kimin Lee
Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work,...
Yifan Jiang, Ruoxi Ning, Sheng Yao +1 more
Visual inputs are often assumed to improve language understanding in multimodal models. We examine this assumption by asking whether vision-language...
Syed Huma Shah
Modern retrieval-augmented generation(RAG) deployments increasingly rely on caching to reduce token cost and time-to-first-token(TTFT). Prefix-level...
Md Hafizur Rahman, Zafaryab Haider, Tanzim Mahfuz +1 more
Multi-agent LLM systems decompose workflows across agents, tools, shared context, memory, and decision gates. This modularity improves...
Xuan Luo, Yue Wang, Geng Tu +2 more
In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal...
Akindoyin Akinrele, Shreyank N Gowda
Prompt injection poses a critical threat to the safe deployment of large language models, yet existing detection approaches are typically evaluated...
Yujie Lin, Kaidi Jia, Jiayao Ma +2 more
Vision-language models (VLMs) may memorize undesirable information from training data, motivating growing interest in machine unlearning. In this...
Wajdi Zaghouani, Kholoud K. Aldous, Isra Fejzullaj
Safety evaluation of Large Language Models (LLMs) has largely focused on high-resource languages, leaving low-resource languages critically...
Anas H. Alzahrani
Background: Large language models are typically evaluated as models, benchmarks, or short conversational episodes. Less is known about what happens...
Yunbo Long, Haolang Zhao, Lukas Beckenbauer +2 more
Post-trained LLMs are often optimized to align responses with human preferences, making them safe, polite, and conversationally appropriate. In...
Zhe Yu, Wenpeng Xing, Gaolei Li +4 more
Retrieval-augmented generation (RAG) increasingly underpins high-stakes applications, yet remains vulnerable to Confundo-style poisoning where...
Zedian Shao, Charles Fleming, Teodora Baluta
Large language models (LLMs) are often fine-tuned on uncurated text datasets that adversaries can poison. Existing poisoning attacks primarily rely...
Haodong Zhao, Tianyi Xu, Tianhang Zhao +2 more
Fine-tuning Large Language Models with untrusted data exposes models to backdoor attacks, where poisoned samples cause targeted misbehavior. Existing...
Hwiwon Lee, Jiawei Liu, Dongjun Kim +3 more
Large language models (LLMs) now support automated software security tasks, including vulnerability discovery and proof-of-concept (PoC) generation....
Kevin Kuo, Chhavi Yadav, Virginia Smith
Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an...
Peiran Wang, Ying Li, Yuan Tian
LLM-based agents are increasingly deployed in high-stakes scenarios such as email management, financial transactions, and code execution, where they...
Ujwal Kumar, Arth Singh, Hershraj Niranjani +5 more
Frontier LLM agents engage in blackmail, sabotage, and document leaks under goal conflicts in agentic settings, exposing limitations of alignment...
Arian Komaei Koma, Seyed Amir Kasaei, AmirMahdi Sadeghzadeh +1 more
Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white- and black-box attacks have...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,406+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial