Paper 2607.28226v1

Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation

tools. We show that familiar attack families: poisoning, backdoors, adversarial examples, sensor spoofing, prompt injection, trajectory manipulation, and supply-chain attacks take on distinct meanings when they corrupt world states

medium relevance benchmark

@dynatrace-oss/dynatrace-mcp-server's create_dynatrace_notebook missing the human

CVSS 3.7 @dynatrace-oss/dynatrace-mcp-server View details
Paper 2607.25255v1

SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems

global risk context before irreversible actions are committed. Evaluated on four benchmarks spanning prompt injection, jailbreak-based unsafe tool use, risky code execution, and harmful web-agent behavior, SafeFlow reduces

medium relevance benchmark
Paper 2607.24625v1

Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents

Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors. While dynamic Information Flow Control (IFC) provides structural security guarantees, traditional taint

medium relevance benchmark
Paper 2607.23586v1

Are You Still the Agent I Authorized? Earned Authority under a Fixed Ceiling for Evolving Agents

creates a distinct authorization problem. Tool-enabled agents can turn model errors and prompt injections into consequential external actions; when evolution occurs under a live grant, the subject exercising that

medium relevance attack
Paper 2607.21151v1

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

visual understanding activates a weaker refusal tendency than textual understanding. Furthermore, we introduce a prompt injection intervention method that reduces attack success rates by an average of 48.24 percentage points

medium relevance defense
Paper 2607.20090v1

Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results

Both trained policies reduce forbidden-content adoption and produce shorter, more focused responses, yet prompt-injection following does not improve. The paired gains are modest and fail to survive Holm

medium relevance benchmark
Paper 2607.19837v1

Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents

used, and which agent weaknesses they exploit to give adversaries leverage in indirect prompt injection attacks. We instantiate these insights in Know Your Agent (KYA), a framework that automates black

medium relevance benchmark
Paper 2607.19292v1

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

identify under-recognized risk patterns, including overreliance, uncertainty and legitimacy laundering in retrieval, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional human oversight, synthetic evidence pollution, and model collapse

medium relevance tool
Paper 2607.18485v1

Trusted Credentials, Untrusted Behavior: Benchmarking LLM-Agent Security in High-Performance Computing

hijacked authorized agent problem. Existing agent-security studies explain relevant mechanisms, such as indirect prompt injection and tool misuse, but generally evaluate them in web, enterprise, or personal-assistant settings

medium relevance benchmark
Paper 2607.18063v1

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security

based agents process external content, exposing them to prompt injection and multi-turn manipulation. Most safety benchmarks evaluate defenders against fixed attack pools collected before evaluation, single-turn or multi

medium relevance benchmark
Paper 2607.17117v1

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models

features concentrate topic-level information in a persistent state. Moreover, as shown in a prompt-injection monitoring case study, slow features preserve detection signals and remain causally effective over long

medium relevance attack
Paper 2607.13987v1

Agent Skill Security: Threat Models, Attacks, Defenses, and Evaluation

packaged, shared, and reused across diverse applications. However, existing security research primarily focuses on prompt injection and runtime execution, leaving security risks throughout the broader skill lifecycle largely unexplored

high relevance benchmark
Paper 2607.13718v1

How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement

agents gain prevalance, users are increasingly exposed to the risks such systems entail. Prompt injection attacks, as well as hallucination, can cause agents to leak private information to third parties

medium relevance survey
Paper 2607.13683v1

Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity

much by the harness around its model as by the frozen model itself: its prompts, injected knowledge, runtime control, and configuration. In deployment the harness is often the only lever

low relevance other
Paper 2607.12406v1

Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions

types, applications, and benchmarks. This makes it hard to explain why failures such as prompt injection, tool misuse, and memory poisoning often share the same structural cause, and how they

medium relevance survey
Paper 2607.10712v1

Distributed Denial of Science: How Indirect Data Poisoning of AI Systems Can Industrialize Scientific Fraud

only 6.0%. The attack requires no topic-specific trigger-words, agent access, indirect prompt injection, or fabricated papers, only the open data ecosystem and misleading metadata. To mitigate the attacks

medium relevance tool
Paper 2607.08282v1

Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language Models

proprietary code leakage prevention, and extensible components designed for future security enhancements such as prompt injection evasion. The framework's layered architecture enables deployment across heterogeneous environments, allowing organizations

medium relevance defense
Paper 2607.07529v1

FedMark-FM: Auditable, Risk-Adjusted Data Markets for Federated Foundation-Model Adaptation

retrieval, held-out generator-backed RAG, and trained PEFT/LoRA tracks. Under a held-out prompt-injection poisoner, FedMark-FM improves downstream accuracy by 7.5-8.1 points over volume, leave

medium relevance attack
Paper 2607.05277v1

Untrusted Content Masking for Web Agents with Security Guarantees

Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data. In text-based environments such as tool-use APIs, this separation

medium relevance attack
Previous Page 15 of 26 Next