Jailbreaks on Vision Language Model via Multimodal Reasoning
Aarush Noheria, Yuguang Yao
Vision-language models (VLMs) have become central to tasks such as visual question answering, image captioning, and text-to-image generation....
AI Threat Alert indexes 3,795+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 2401–2420 of 3,795 papers
Aarush Noheria, Yuguang Yao
Vision-language models (VLMs) have become central to tasks such as visual question answering, image captioning, and text-to-image generation....
Chanwoo Park, Chanwoo Kim
Evasion attacks pose significant threats to AI systems, exploiting vulnerabilities in machine learning models to bypass detection mechanisms. The...
Yavuz Bakman, Duygu Nur Yaldiz, Salman Avestimehr +1 more
Large Language Models (LLMs) are rarely static and are frequently updated in practice. A growing body of alignment research has shown that models...
Javier Carnerero-Cano, Luis Muñoz-González, Phillippa Spencer +1 more
Regression models are widely used in industrial processes, engineering, and in natural and physical sciences, yet their robustness to poisoning has...
Amirhossein Taherpour, Xiaodong Wang
Federated learning (FL) enables collaborative model training while preserving data privacy, yet both centralized and decentralized approaches face...
Pedro H. Barcha Correia, Ryan W. Achjian, Diego E. G. Caetano de Oliveira +5 more
The rapid advancement and widespread adoption of generative artificial intelligence (GenAI) and large language models (LLMs) has been accompanied by...
Gloria Felicia, Michael Eniolade, Jinfeng He +4 more
Existing agent safety benchmarks report binary accuracy, conflating early intervention with post-mortem analysis. A detector that flags a violation...
Xiaogeng Liu, Xinyan Wang, Yechao Zhang +5 more
Large reasoning models (LRMs) extend large language models with explicit multi-step reasoning traces, but this capability introduces a new class of...
Xiaoyu Xu, Minxin Du, Kun Fang +6 more
Large language models (LLMs) demonstrate impressive capabilities across diverse tasks but raise concerns about privacy, copyright, and harmful...
Mingyang Liao, Yichen Wan, shuchen wu +6 more
LLM-based role-playing has rapidly improved in fidelity, yet stronger adherence to persona constraints commonly increases vulnerability to jailbreak...
Ningyuan He, Ronghong Huang, Qianqian Tang +3 more
In-context learning (ICL) has become a powerful, data-efficient paradigm for text classification using large language models. However, its robustness...
Wenhui Zhang, Huiyu Xu, Zhibo Wang +4 more
Recent advancements in multi-model AI systems have leveraged LLM routers to reduce computational cost while maintaining response quality by assigning...
Devanshu Sahoo, Manish Prasad, Vasudev Majhi +5 more
The rapid integration of Large Language Models (LLMs) into educational assessment rests on the unverified assumption that instruction following...
Xiang Zheng, Yutao Wu, Hanxun Huang +5 more
Autonomous code agents built on large language models are reshaping software and AI development through tool use, long-horizon reasoning, and...
Alvi Md Ishmam, Najibul Haque Sarker, Zaber Ibn Abdul Hakim +1 more
Multimodal Large Language Models (MLLMs) have achieved remarkable performance across vision-language tasks. Recent advancements allow these models to...
Arther Tian, Alex Ding, Frank Chen +2 more
Decentralized large language model inference networks require lightweight mechanisms to reward high quality outputs under heterogeneous latency and...
Jarrod Barnes
As large language models (LLMs) improve, so do their offensive applications: frontier agents now generate working exploits for under $50 in compute...
Onkar Shelar, Travis Desell
Evolutionary prompt search is a practical black-box approach for red teaming large language models (LLMs), but existing methods often collapse onto a...
Mingqiao Mo, Yunlong Tan, Hao Zhang +2 more
Large language models (LLMs) have achieved remarkable progress in code generation, yet their potential for software protection remains largely...
Xingwei Lin, Wenhao Lin, Sicong Cao +4 more
Multi-turn jailbreak attacks have emerged as a critical threat to Large Language Models (LLMs), bypassing safety mechanisms by progressively...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,795+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial