Detecting Trojaned DNNs via Spectral Regression Analysis
Samuele Pasini, Jinhan Kim, Paolo Tonella
Modern DNNs are repeatedly fine-tuned to incorporate new data and functionality. This evolutionary workflow introduces a security risk when updated...
AI Threat Alert indexes 3,397+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 221–240 of 495 papers
Clear filtersSamuele Pasini, Jinhan Kim, Paolo Tonella
Modern DNNs are repeatedly fine-tuned to incorporate new data and functionality. This evolutionary workflow introduces a security risk when updated...
Laura Jiang, Reza Ryan, Qian Li +1 more
Single-turn safety evaluation is a poor proxy for real fraud defense, where attackers escalate across multiple rounds. This paper evaluates fraud...
Amit Roth, Ankur Samanta, Matan Halevy +2 more
Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking,...
Paul Wang, Jade Garcia-Bourrée, Anne-Marie Kermarrec +1 more
As jailbreaks, adversarially crafted inputs that bypass safety constraints, continue to be discovered in Large Language Models, practitioners...
Doguhuan Yeke, Yanming Zhou, Leo Y. Lin +3 more
Recent advances in Vision-Language Models (VLMs) facilitate a new class of embodied AI systems, where these models are integrated into physical...
Quanxing Xu, Yuhao Tian, Ling Zhou +4 more
Visual Question Answering (VQA), as the representative multimodal task, serves as a key benchmark for evaluating the reasoning capabilities of...
Daniel Yiming Cao, Chengzhong Wang, Sheng-Yen Chou +3 more
Masked diffusion language models (MDLMs) are emerging as a compelling new paradigm for text generation, but their training-time security remains...
Yubin Qu, Ying Zhang, Yanjun Zhang +4 more
Coding agents now run autonomously with shell, file, and network privileges. When a user issues a benign request, the agent sometimes does more than...
Jonathan Diller, David Barnes, Rebekah Bogdanoff +14 more
As autonomous systems grow more advanced, objective metrics to evaluate their ethical and legal compliance are critical for informing end users of...
Lei Zhao, Abhay Bhaskar, Edgar Dobriban
AI agents such as OpenClaw are increasingly deployed in local workflows with access to external tools. This creates indirect prompt-injection (IPI)...
Xiao Yang, Ronghao Fu, Zhiwen Lin +10 more
Remote sensing vision-language models commonly rely on pretrained visual encoders to convert images into semantic features before language-model...
Qiqi Liu, Thorsten Holz, Shilin Ye +1 more
Multi-agent systems extend large language models (LLMs) by decomposing tasks among specialized agents, but their distributed decision process creates...
Simiao Liu, Fang Liu, Li Zhang +2 more
Large language model (LLM) agents are increasingly used for automated vulnerability repair (AVR), where repository-level reasoning enables them to...
Simiao Liu, Li Zhang, Fang Liu +3 more
Modern software ecosystems face a rapidly growing number of disclosed vulnerabilities, increasing the need for automated repair techniques that can...
Zhen Xu, Zihao Wang, Yuhua Sun +1 more
Side-channel attacks exploit unintended information leakage from system behavior and continue to pose serious privacy risks in modern platforms....
Nils Loose, Joseph Bienhüls, Kristoffer Hempel +2 more
Automated detection of vulnerability-fixing commits (VFCs) is critical for timely security patch deployment, as advisory databases lag patch releases...
Toluwani Aremu, Nils Lukas, Jie Zhang
Watermarking is widely proposed for provenance, attribution, and safety monitoring in generative models, yet is typically evaluated only under...
Joana Pasquali, Ramiro N. Barros, Arthur S. Bianchessi +7 more
LoRA is widely adopted for continual fine-tuning of Large Language Models due to its parameter efficiency, modularity across tasks, and compatibility...
Hao Wang, Hanchen Li, Qiuyang Mang +3 more
Agent benchmarks have become the de facto measure of frontier AI competence, guiding model selection, investment, and deployment. However, reward...
Rezarta Islamaj, Robert Leaman, Joey Chan +13 more
Evaluating large language models (LLMs) in the biomedical domain requires benchmarks that can distinguish reasoning from pattern matching and remain...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,397+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial