340 results in 91ms
Paper 2603.01019v1

BadRSSD: Backdoor Attacks on Regularized Self-Supervised Diffusion Models

backdoor attack targeting the representation layer of self-supervised diffusion models. Specifically, it hijacks the semantic representations of poisoned samples with triggers in Principal Component Analysis (PCA) space toward those

high relevance attack
Paper 2607.20759v1

IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

security risks from both the LLM backbone, where adversarial prompts, poisoned training data, and backdoor triggers can cause models to emit insecure or attacker-chosen code, and their agentic architecture

medium relevance benchmark
Paper 2601.05504v2

Memory Poisoning Attack and Defense on Memory Based LLM-Agents

Large language model agents equipped with persistent memory are vulnerable to memory poisoning attacks, where adversaries inject malicious instructions through query only interactions that corrupt the agents long term memory

high relevance attack
Paper 2509.21761v2

Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models

Fine-tuned Large Language Models (LLMs) are vulnerable to backdoor attacks through data poisoning, yet the internal mechanisms governing these attacks remain a black box. Previous research on interpretability

medium relevance attack
Paper 2607.25502v1

Anti-Backdoor Coreset Selection via Cumulative Entropy

time defenses against neural backdoors isolate a benign subset from poisoned training data, to learn a backdoor-free model from it. In this paper, we formulate this defense strategy

medium relevance attack
Paper 2607.06484v1

Assessing the Operational Impact of Poisoning Attacks over Augmented 3D Point Cloud Public Datasets for Connected and Autonomous Vehicles

system apply over the learned models. Its impact over data augmentation models is unclear. While data augmentation reduces the likelihood of poisoning attack success, some valid questions remain. Is data

high relevance benchmark
Paper 2604.04289v1

Poisoned Identifiers Survive LLM Deobfuscation: A Case Study on Claude Opus 4.6

poisoned identifier names in the string table survive into the model's reconstructed code, even when the model demonstrably understands the correct semantics? Using Claude Opus 4.6 across 192 inference

low relevance other
Paper 2604.07403v1

RefineRAG: Word-Level Poisoning Attacks via Retriever-Guided Text Refinement

Retrieval-Augmented Generation (RAG) significantly enhances Large Language Models (LLMs), but simultaneously exposes a critical vulnerability to knowledge poisoning attacks. Existing attack methods like PoisonedRAG remain detectable due to coarse

high relevance attack
Paper 2605.26754v1

Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control

Confundo-style poisoning where adversarially optimized documents manipulate generated outputs. Existing defenses assume that detecting poisoned evidence prevents harm. We show this assumption is incorrect: models exhibit a monitoring-control

medium relevance attack
Paper 2602.04899v1

Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoning

data-level defences are insufficient for stopping sophisticated data poisoning attacks. We suggest that future work should focus on model audits and white-box security methods

medium relevance attack
Paper 2604.07536v1

TRUSTDESC: Preventing Tool Poisoning in LLM Applications via Trusted Description Generation

injection attack surface: tool poisoning attacks (TPAs). Attackers manipulate tool descriptions by embedding malicious instructions (explicit TPAs) or misleading claims (implicit TPAs) to influence model behavior and tool selection. Existing

medium relevance tool
Paper 2602.02629v1

Trustworthy Blockchain-based Federated Learning for Electronic Health Records: Securing Participant Identity with Decentralized Identifiers and Verifiable Credentials

patient data. Despite its potential, FL remains vulnerable to poisoning and Sybil attacks, in which malicious participants corrupt the global model or infiltrate the network using fake identities. While recent

medium relevance benchmark
Paper 2602.19547v1

CIBER: A Comprehensive Benchmark for Security Evaluation of Code Interpreter Agents

four major types of adversarial attacks: Direct/Indirect Prompt Injection, Memory Poisoning, and Prompt-based Backdoor. We evaluate six foundation models across two representative code interpreter agents (OpenInterpreter and OpenCodeInterpreter), incorporating

medium relevance benchmark
Paper 2607.19292v1

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

legitimacy laundering in retrieval, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional human oversight, synthetic evidence pollution, and model collapse. We conclude with design and governance recommendations

medium relevance tool
Paper 2602.07200v1

BadSNN: Backdoor Attacks on Spiking Neural Networks via Adversarial Spiking Neuron

converts input data into spikes following the Leaky Integrate-and-Fire (LIF) neuron model. This model includes several important hyperparameters, such as the membrane potential threshold and membrane time constant

high relevance attack
Paper 2607.14493v1

Context Contamination in LLM Analysis of Network Security Logs: Poison with Passive Prompt Injection and Mitigation Evaluation

Large Language Models are increasingly deployed in Security Operations Centers

high relevance benchmark
Paper 2605.19262v1

Backdooring Masked Diffusion Language Models

training-time security remains largely unexplored. Existing backdoor attacks on Gaussian diffusion models or autoregressive language models do not directly apply to MDLMs because MDLMs rely on discrete state corruption

medium relevance benchmark
Paper 2603.25164v1

PIDP-Attack: Combining Prompt Injection with Database Poisoning Attacks on Retrieval-Augmented Generation Systems

Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of applications. However, their practical deployment is often hindered by issues such as outdated knowledge and the tendency

high relevance attack
Paper 2602.01942v1

Human Society-Inspired Approaches to Agentic AI Security: The 4C Framework

software components. Although recent work has strengthened defenses against model and pipeline level vulnerabilities such as prompt injection, data poisoning, and tool misuse, these system centric approaches may fail

medium relevance tool
Paper 2510.09710v2

SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAG

Retrieval-augmented generation (RAG) systems enhance large language models (LLMs) with external knowledge but are vulnerable to corpus poisoning and contamination attacks, which can compromise output integrity. Existing defenses often

medium relevance benchmark
Previous Page 7 of 17 Next