Curriculum Learning for Safety Alignment
Sandeep Kumar, Virginia Smith, Chhavi Yadav
Direct Preference Optimisation (DPO) is widely used for safety alignment in large language models. However, prior work shows it is brittle and...
AI Threat Alert indexes 3,406+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 721–740 of 1,756 papers
Clear filtersSandeep Kumar, Virginia Smith, Chhavi Yadav
Direct Preference Optimisation (DPO) is widely used for safety alignment in large language models. However, prior work shows it is brittle and...
Mohammed N. Swileh, Shengli Zhang, Kai Lei
Software-Defined Networking (SDN) provides flexible and programmable network management; however, its centralized control architecture remains highly...
Faruk Alpay, Taylan Alpay
LLM agents process trusted instructions, retrieved records, and tool observations through a common generative channel. This conflates data flow with...
Sofiat Abioye, Ufaq Khan, Shazad Ashraf +6 more
Urgent suspected colorectal cancer (CRC) referrals create operational bottlenecks because semi-structured clinical documents often require manual...
David Košťál, Martin Jureček
We present a dataset of adversarial malware samples derived from the public RawMal-TF collection of real-world malware binaries. Using a suite of...
Xiao Liu, Jiaxiang Liu, Boci Peng +6 more
Vision Language Models adapt well to downstream tasks but are highly vulnerable to adversarial perturbations that disrupt cross-modal semantic...
Jianwei Tai
Vision-Language-Action (VLA) models are increasingly deployed on real robots, where each predicted action is executed and each failure carries a...
Jianwei Tai
Vision-Language-Action (VLA) models are increasingly deployed on real robots, where each predicted action is executed and each failure carries a...
Yue Liu, Yanjie Zhao, Yunbo Lyu +3 more
Agentic AI coding assistants can edit files, run commands, and access the internet on behalf of developers. However, their reliance on unvetted...
Yutong Cheng, Changze Li, Raihan Sultan Pasha Basuki +3 more
Extracting MITRE ATT&CK techniques from cyber threat intelligence (CTI) reports is an open-set, multi-label problem requiring both high recall (not...
Miel Verkerken, Laurens D'hooge, Bruno Volckaert +2 more
Network Intrusion Detection Systems (NIDS) are now increasingly leveraging Machine Learning (ML) techniques to detect malicious network activities....
Chenxin Mao, Shangyu Liu, Zhenzhe Zheng +3 more
Retrieval-Augmented Generation (RAG) empowers LLMs with external knowledge, making cross-institutional domain-specific knowledge base integration a...
Lingyao Li, Junjie Xiong, Changjia Zhu +5 more
Large language models (LLMs) are increasingly used in academic peer review, yet their reliability, alignment with human judgment, and robustness to...
Bingyu Yan, Xiaoming Zhang, Jinyu Hou +4 more
While Large Language Model-based Multi-Agent Systems (LLM-MAS) demonstrate remarkable capabilities in solving complex tasks by orchestrating...
Aditya Sridhar
Concept Bottleneck Models (CBMs) have emerged as a cornerstone approach for interpretable machine learning, providing human-understandable...
Dongpeng Zhang, Ke Ma, Yangbangyan Jiang +4 more
Adversarial images pose a severe security threat to multimodal large language models through prompt injection. Existing defenses largely lack a...
Wenjuan Li, Yitao Liu, Runze Chen +1 more
Background: Fine-tuning is central to adapting pre-trained Large Language Models (LLMs) to downstream tasks, but its reliance on training data,...
Esra Yeniaras
Quantum machine learning (QML) is moving from research prototypes to deployed cloud services. As QML enters regulated industries, the integrity of...
Haobo Zhang, Xutao Mao, Guangyuan Dong +5 more
Memory-backed agents need provenance that can survive leaked or migrated snapshots, where logs, visible outputs, and trusted metadata may be absent....
William Guanting Li, Alsharif Abuadbba, Kristen Moore +1 more
Penetration testing is essential to securing modern web infrastructures, yet traditional manual methods struggle to keep pace with their scale and...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,406+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial