Defense MEDIUM relevance

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

Caoyuan Ma Wenpu Liu Weichu Xie Tian Gu Shilei Zhao Lingxi Min Shuai Dong Yuqi Xu Ji Zhao Ziyue Wang Wenzheng Chang Taiqiang Wu Yongfu Zhu Wenqi Shao Yinqiang Zheng
Published
August 11, 2026
Updated
August 11, 2026

Abstract

Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propose SafeCap, a reinforcement-learning framework that aligns LVLMs through learned self-captioning. SafeCap trains a policy model to first generate a safety-relevant image caption and then produce a final answer; the caption is further optimized by whether it enables a frozen LLM to reach a safety-aligned decision. This caption-mediated objective encourages the policy to expose visual cues relevant to safe response generation rather than relying solely on direct refusal supervision. Across five multimodal safety benchmarks and six vision-utility benchmarks, SafeCap substantially improves aggregate safety performance under its intended DirectCap protocol, with gains of 3.7-19.0 points in safety average across four model settings while maintaining comparable or improved vision utility. Under controlled comparisons on matched backbones and data, SafeCap outperforms safety SFT, DPO, and SafeGRPO, demonstrating the effectiveness of caption-mediated reinforcement learning for multimodal safety alignment.

Metadata

Comment
16pages, 4 figures. Preprint. Project page: https://safe-vlm.github.io/SafeCap/ ; code: https://github.com/Safe-VLM/SafeCap

Pro Analysis

Full threat analysis, ATLAS technique mapping, compliance impact assessment (ISO 42001, EU AI Act), and actionable recommendations are available with a Pro subscription.

Threat Deep-Dive
ATLAS Mapping
Compliance Reports
Actionable Recommendations
Start 14-Day Free Trial