Attack MEDIUM relevance

Through Human Eyes and Machine Eyes: Understanding View Mismatch in Video See-Through Extended Reality

Yanming Xiu
Published
September 24, 2026
Updated
September 24, 2026

Abstract

Video see-through extended reality (VST XR) systems commonly use headset screenshots or captured frames as proxies for the user's first-person visual context. However, the system-captured view and the user's effective visible field do not necessarily coincide: a screenshot records a rectangular machine-readable frame, whereas the user's effective visible region can be more constrained and non-rectangular. This paper studies this human-system view mismatch in VST XR. We formalize the relationship between the system-captured region and the human-visible region by defining their co-visible, system-only, and human-only regions. \rev{We then conduct a pilot-level boundary measurement on Meta Quest 3, revealing a clear mismatch between the rectangular screenshot frame and the approximate human-visible boundary. Building on this model, we analyze how view mismatch can affect screenshot-based XR sensing and downstream vision-language model tasks. Through four representative case studies, we illustrate potential risks and failure modes including prompt injection, privacy leakage, human-invisible information bias, and missing human-visible information. Our results show that view mismatch is not only a geometric artifact, but can also introduce security, privacy, and reliability concerns for AI-integrated VST XR systems.

Pro Analysis

Full threat analysis, ATLAS technique mapping, compliance impact assessment (ISO 42001, EU AI Act), and actionable recommendations are available with a Pro subscription.

Threat Deep-Dive
ATLAS Mapping
Compliance Reports
Actionable Recommendations
Start 14-Day Free Trial