Paper 2607.17550v1

(A)iSpy: Parasitic Trojans for Machine Learning Infrastructure

iSpy module easily evades standard malware scanners, while the associated poisoned inputs and resulting compromised models bypass typical inspection tools. We demonstrate the practicality of this threat with an implementation

medium relevance attack
Paper 2604.05809v1

Stealthy and Adjustable Text-Guided Backdoor Attacks on Multimodal Pretrained Models

significantly improving stealthiness and practicality. Furthermore, we introduce visual adversarial perturbations on poisoned samples to modulate the model's learning of textual triggers, enabling a controllable and adjustable TGB attack

high relevance attack
Paper 2605.28112v1

A Wolf in Sheep's Clothing: Targeted Routing Hijacking in Federated RAG

including missing evidence, poisoning, incorrect answers, and hallucinations. In a high-stakes MedQA-USMLE case study, we further show that poisoned retrieved evidence can mislead models across scales, leading

medium relevance attack
Paper 2512.16962v1

MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval

Large Language Model (LLM) agents increasingly rely on long-term memory and Retrieval-Augmented Generation (RAG) to persist experiences and refine future performance. While this experience learning capability enhances agentic

medium relevance benchmark
Paper 2603.27986v1

FedFG: Privacy-Preserving and Robust Federated Learning via Flow-Matching Generation

data; on the other hand, they may compromise clients to launch poisoning attacks that corrupt the global model. To balance accuracy and security, we propose FedFG, a robust FL framework

medium relevance attack
Paper 2511.14715v2

FLARE: Adaptive Multi-Dimensional Reputation for Robust Client Reliability in Federated Learning

learning (FL) enables collaborative model training while preserving data privacy. However, it remains vulnerable to malicious clients who compromise model integrity through Byzantine attacks, data poisoning, or adaptive adversarial behaviors

medium relevance benchmark
Paper 2605.01834v1

Repurposing and Evaluating the (In)Feasibility of Dataset Poisoning enabled Watermarking for Contrastive Learning

third-party or internet data is common. Recent studies show CL models are vulnerable to data-poisoning backdoor attacks, but their generalization and robustness are underexplored. We systematically evaluate existing

medium relevance benchmark
Paper 2510.03636v1

From Theory to Practice: Evaluating Data Poisoning Attacks and Defenses in In-Context Learning on Social Media Health Discourse

This study explored how in-context learning (ICL) in large language models can be disrupted by data poisoning attacks in the setting of public health sentiment analysis. Using tweets

high relevance attack
Paper 2512.12921v1

Cisco Integrated AI Security and Safety Framework Report

threats now span content safety failures (e.g., harmful or deceptive outputs), model and data integrity compromise (e.g., poisoning, supply-chain tampering), runtime manipulations (e.g., prompt injection, tool and agent misuse

medium relevance tool
Paper 2607.28075v1

Temporal Poisoning: Clean-Label Backdoors via Event Redistribution in SNNs

inspection are blind by construction, while feature-space methods detect the poison only in selected settings. Our model-free detector, based on per-step event mass, detects the evaluated temporal

medium relevance attack
Paper 2512.10998v1

SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models

Backdoor attacks create significant security threats to language models by

high relevance attack
Paper 2601.22308v2

Stealthy Poisoning Attacks Bypass Defenses in Regression Settings

natural and physical sciences, yet their robustness to poisoning has received less attention. When it has, studies often assume unrealistic threat models and are thus less useful in practice

high relevance attack
Paper 2601.04266v1

State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space

Vision-Language-Action (VLA) models are widely deployed in safety

high relevance attack
Paper 2603.04859v1

Osmosis Distillation: Model Hijacking with the Fewest Samples

generated by dataset distillation methods, where an adversary can perform a model hijacking attack with only a few poisoned samples in the synthetic dataset. To reveal this threat, we propose

medium relevance benchmark
Paper 2605.01782v1

Needle-in-RAG: Prompt-Conditioned Character-Level Traceback of Poisoned Spans in Retrieved Evidence

evidence, but it also opens a data-layer attack surface: poisoned corpus entries can steer outputs without changing model parameters. Existing defenses and traceback methods are largely passage-level, which

medium relevance benchmark
Paper 2604.21477v1

MCP Pitfall Lab: Exposing Developer Pitfalls in MCP Tool Server Security under Multi-Vector Attacks

email, document, crypto) with six server variants (baseline and hardened) and model three attack families: tool-metadata poisoning, puppet servers, and multimodal image-to-tool chains, in a unified, trace

high relevance tool
Paper 2603.12989v1

Test-Time Attention Purification for Backdoored Large Vision Language Models

defenses across diverse datasets and backdoor attack types, while preserving the model's utility on both clean and poisoned samples

medium relevance benchmark
Paper 2509.26032v2

Stealthy Yet Effective: Distribution-Preserving Backdoor Attacks on Graph Classification

semantic deviation caused by label flipping, both of which make poisoned graphs easily detectable by anomaly detection models. To address this, we propose DPSBA, a clean-label backdoor framework that

high relevance attack
Paper 2603.18034v1

Semantic Chameleon: Corpus-Dependent Poisoning Attacks and Defenses in RAG Systems

documents are preferentially retrieved at inference time, enabling targeted manipulation of model outputs. We study gradient-guided corpus poisoning attacks against modern RAG pipelines and evaluate retrieval-layer defenses that

high relevance attack
Paper 2602.11213v1

Transferable Backdoor Attacks for Code Models via Sharpness-Aware Adversarial Perturbation

software development but remain vulnerable to backdoor attacks via poisoned training data. Existing backdoor attacks on code models face a fundamental trade-off between transferability and stealthiness. Static trigger-based

high relevance attack
Previous Page 6 of 18 Next