Exposing the Illusion of Erasure in Knowledge Editing for LLMs
Advik Raj Basani, Anshuman Chhabra
Knowledge Editing (KE) has emerged as a frontier for updating specific facts in LLMs without costly retraining, but its reliability and underlying...
AI Threat Alert indexes 3,771+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 461–480 of 888 papers
Clear filtersAdvik Raj Basani, Anshuman Chhabra
Knowledge Editing (KE) has emerged as a frontier for updating specific facts in LLMs without costly retraining, but its reliability and underlying...
Zewen Liu
Large Language Model (LLM) agents increasingly rely on memory systems to maintain long-term coherence. Recent work shows that agent memories degrade...
Jaehyuk Jang, Minseok Seo. Seungju Cho, Kangwook Ko +1 more
Vision-language models (VLMs) achieve strong zero-shot recognition, but they remain highly vulnerable to adversarial perturbations. Recent test-time...
Junquan Deng, Zhiyu Fan, Ruijie Meng
Recent advances in large language models (LLMs) have enabled vibe coding, an emerging software development paradigm in which users create...
Junquan Deng, Zhiyu Fan, Ruijie Meng
Recent advances in large language models (LLMs) have enabled vibe coding, an emerging software development paradigm in which users create...
Ruixiao Lin, Xinhao Deng, Qingming Li +12 more
Self-evolving LLM agent systems, which autonomously update their model parameters, memory, tools, and architectures, introduce a qualitatively new...
Aymen Bouferroum, Ildi Alla, Valeria Loscri +2 more
Radio frequency jamming poses a critical threat to the availability of wireless Industrial Internet of Things (IIoT) networks. Existing detection and...
Parth Bramhecha, Smit Deshmukh, Sairaj Bodhale +2 more
As Large Language Models (LLMs) achieve widespread integration across diverse linguistic landscapes, ensuring their safety and alignment with...
Xingwei Zhong, Varun Sharma, Kar Wai Fok +1 more
Vision language models (VLMs) employ both visual and textual modalities to enable advanced vision-language inference. However, incorporating visual...
Osama Wehbi, Sarhad Arisdakessian, Omar Abdel Wahab +2 more
Federated Learning (FL) enables collaborative model training without sharing raw data, making it a promising paradigm for privacy-sensitive...
Shivam Ratnakar, Kartikeya Vats
Modern Large Language Models (LLMs) rely on extensive safety alignment, yet the mechanistic basis of refusal remains opaque. In this work, we...
Weidi Luo, Qiming Zhang, Yihao Quan +5 more
Coding agents based on large language models (LLMs) demonstrate remarkable autonomous capabilities, but they also introduce significant safety and...
Sihui Dai, Mann Patel
Prior work has shown that in-context demonstrations can jailbreak language models, but it remains unclear how models interpret different types of...
Shihao Ji, HongXi Li, Zihui Song +1 more
Scaling end-to-end autonomous driving to complex, open-world environments requires perceptual models that generalize to anomalous scenarios and...
Kaihsun Yang, Min-Yan Tsai, Chia-Mu Yu
Model quantization is widely adopted to reduce memory usage and inference cost when deploying deep neural networks on resource-constrained devices....
Serge Sharoff
The increasing prominence of Large Language Models (LLMs) in public discourse presents both opportunities and challenges for democratic deliberation....
Prashanti Nilayam, Kiran Kumar Ramanna, Prashil Tumbade +1 more
Heterogeneous LLM debate is motivated by the promise that diverse peers correct one another, but the same exchange that carries correction also...
Licheng Pan, Haocheng Yang, Haoxuan Li +7 more
Reinforcement learning from human feedback (RLHF) aligns large language models by training reward models on preference data and optimizing policies...
Haotian Xu, Zeyang Zhang, Linbao Li +3 more
Speculative inference accelerates large language model (LLM) decoding but provides no inherent safety guarantees. Existing safety defenses are...
Zunchen Huang, Songgaojun Deng
Formal tools such as SAT and SMT solvers are increasingly embedded in language model reasoning pipelines when a safety or security critical question...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,771+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial