Cheap, Fallible Cognition and the Political Economy of Expertise
Christophe Kolb, Jim Caron
The question of whether artificial intelligence will "destroy jobs" is too coarse to guide economic analysis or institutional design. A job is not an...
AI Threat Alert indexes 3,771+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 381–400 of 3,771 papers
Christophe Kolb, Jim Caron
The question of whether artificial intelligence will "destroy jobs" is too coarse to guide economic analysis or institutional design. A job is not an...
Yan Wen, Zhenyi Wang, Heng Huang
Graph Neural Networks (GNNs) serve as the backbone for high-stakes applications in Machine-Learning-as-a-Service (MLaaS). Still, their black-box...
Soumya Mazumdar, Vineet Kumar Rakesh, Tapas Samanta
Robust comparisons of federated aggregation methods require joint consideration of predictive performance, threat definitions, metric semantics, and...
Md. Nahid Hasan, Mohammad Arif Hossain
Anyone can upload a fine-tuned large language model (LLM) to a public repository and claim it is safe. A backdoored model behaves normally on...
Gabriel Huang, Abhay Puri, Léo Boisvert +4 more
Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met...
Huafeng Chen, Yueming Lyu, Ziyuan Chen +4 more
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in storing and recalling rich person-related knowledge, raising...
Abigail Oppong, P Sam Sahil, Tadesse Destaw Belay +12 more
Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings....
Clemens Vetter, David Kaczér, Lucie Flek +1 more
Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A...
Xinrui Lin, Sha Zhang, Shumin Wang +3 more
Task-driven 3D affordance grounding aims to localize the functional region in a cluttered 3D scene that enables an action specified by a...
Tao Lin, Gaojie Jin, Zongxin Liu +2 more
Existing vision-language model (VLM) backdoors are usually treated as static vulnerabilities: one-to-one and N-to-N attacks bind one or more triggers...
Lukasz Olejnik, Wenchao Dong, Jonas R. Kunst +4 more
We introduce IO Factory, an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes. The...
Zixing Chen, Xingyuan Liu, Jie Zhu +6 more
Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit...
Ze Yu, Hongwei Zhen, Chao Shen +1 more
The rapid growth of large language model (LLM) services is accelerating the expansion of AI data centers (AIDCs), intensifying concerns over power...
Xinzhe Huang, Biwu Yao, Kedong Xiu +4 more
Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically...
Albus W. Ng, Yi Han, Jusheng Zhang +1 more
The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is...
Rui Zhang, Wenbo Jiang, Hongwei Li +5 more
Expert parallelism (EP) is a common strategy for serving large Mixture-of-Experts (MoE) models across multiple GPUs by distributing experts among...
Chuqiao Lin, Shivaji Sondhi, Xiao-Liang Qi
Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding...
Md Jafrin Hossain, Mohammad Arif Hossain, Nirwan Ansari
Large Language Models (LLMs) have undergone a shift from stateless conversational interfaces to autonomous agents capable of multi-step planning,...
Caoyuan Ma, Wenpu Liu, Weichu Xie +12 more
Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from...
Berkay Ozcam, Irem Onen, Mehmet Fatih Amasyali +1 more
The rapid advancement of Large Language Models (LLMs) has facilitated their ubiquitous integration into various domains, leading to widespread...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,771+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial