Cheap, Fallible Cognition and the Political Economy of Expertise
Christophe Kolb, Jim Caron
The question of whether artificial intelligence will "destroy jobs" is too coarse to guide economic analysis or institutional design. A job is not an...
AI Threat Alert indexes 3,771+ peer-reviewed and preprint papers on AI/ML security — covering adversarial attacks, model defenses, red-teaming benchmarks, surveys, and security tooling. Papers are sourced from arXiv, classified by type and by relevance to real-world threats, and cross-referenced with the CVEs and incidents they relate to.
Showing 221–240 of 888 papers
Clear filtersChristophe Kolb, Jim Caron
The question of whether artificial intelligence will "destroy jobs" is too coarse to guide economic analysis or institutional design. A job is not an...
Yan Wen, Zhenyi Wang, Heng Huang
Graph Neural Networks (GNNs) serve as the backbone for high-stakes applications in Machine-Learning-as-a-Service (MLaaS). Still, their black-box...
Md. Nahid Hasan, Mohammad Arif Hossain
Anyone can upload a fine-tuned large language model (LLM) to a public repository and claim it is safe. A backdoored model behaves normally on...
Abigail Oppong, P Sam Sahil, Tadesse Destaw Belay +12 more
Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings....
Clemens Vetter, David Kaczér, Lucie Flek +1 more
Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A...
Tao Lin, Gaojie Jin, Zongxin Liu +2 more
Existing vision-language model (VLM) backdoors are usually treated as static vulnerabilities: one-to-one and N-to-N attacks bind one or more triggers...
Lukasz Olejnik, Wenchao Dong, Jonas R. Kunst +4 more
We introduce IO Factory, an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes. The...
Xinzhe Huang, Biwu Yao, Kedong Xiu +4 more
Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically...
Rui Zhang, Wenbo Jiang, Hongwei Li +5 more
Expert parallelism (EP) is a common strategy for serving large Mixture-of-Experts (MoE) models across multiple GPUs by distributing experts among...
Chuqiao Lin, Shivaji Sondhi, Xiao-Liang Qi
Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding...
Md Jafrin Hossain, Mohammad Arif Hossain, Nirwan Ansari
Large Language Models (LLMs) have undergone a shift from stateless conversational interfaces to autonomous agents capable of multi-step planning,...
Caoyuan Ma, Wenpu Liu, Weichu Xie +12 more
Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from...
Alexander Panfilov, David Schmotz, Ilia Shumailov +5 more
Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and...
Puyu Zeng, Simeng Qin, Jingzhi Li +3 more
Agent skills are emerging as an important attack surface in LLM-based agent systems. Through an empirical study of existing skill scanners, we find...
Hongli Shen, Shaopeng Fu, Qinbo Zhang +2 more
Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent...
Hanlin Jiang, Puyi Wang, Jiandong Jin +10 more
Real-world cyberattacks often require sustained progress across multiple hosts and network segments, making multi-hop cyber ranges essential...
Hongwei Yao, Yiming Liu, Meihui Chen +6 more
Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define...
Tejasvi C. Addagada
The key-value (KV) cache is the primary throughput optimization in modern large language model (LLM) inference, enabling prefix reuse across...
Sam Siavoshian, Omar Ramadan, Amir K. Saeed +3 more
Temporal decisions in language-model systems often depend on both symbolic task state and elapsed wall-clock time, such as cache expiration, job...
Lier Jin, Lan Hu, Binqi Shen +2 more
Large language models (LLMs) are increasingly integrated into healthcare, education, public services, and everyday decision making. They should...
AI security research studies how AI and machine-learning systems can be attacked and defended — covering adversarial examples, prompt injection, model poisoning, training-data extraction, and the mitigations against them. AI Threat Alert curates this research from academic sources so security teams can track the threats behind emerging AI risks.
AI Threat Alert indexes 3,771+ papers on AI/ML security, classified across attack, defense, benchmark, survey, and tool categories and updated continuously.
Papers are sourced from arXiv, then classified by type and by relevance to real-world AI/ML threats, and cross-referenced with the CVEs and incidents they relate to.
Coverage spans adversarial attacks, model and system defenses, red-teaming benchmarks, literature surveys, and security tooling for LLMs, ML libraries, AI agents, and inference pipelines.
Every paper is filtered for AI security relevance and linked to the vulnerabilities, vendors, and incidents it relates to, so the research connects directly to operational threat intelligence.
Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.
Start 14-Day Free Trial