AI Security Research

2,583+ academic papers on AI security, attacks, and defenses

Total
2,583
Attack
994
Benchmark
740
Defense
355
Tool
275
Survey
146

Showing 381–400 of 928 papers

Clear filters
Attack MEDIUM

Training Agents to Self-Report Misbehavior

Bruce W. Lee, Chen Yueh-Han, Tomek Korbak

Frontier AI agents may pursue hidden goals while concealing their pursuit from oversight. Alignment training aims to prevent such behavior by...

2 months ago cs.LG cs.AI PDF

Track AI security vulnerabilities in real time

Get breaking CVE alerts, compliance reports (ISO 42001, EU AI Act), and CISO risk assessments for your AI/ML stack.

Start 14-Day Free Trial