Through Human Eyes and Machine Eyes: Understanding View Mismatch in Video See-Through Extended Reality

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the mismatch between the rectangular field of view (FoV) captured by system screenshots and the actual non-rectangular visible region perceived by users in video see-through extended reality (VST XR). By formalizing this human-machine FoV discrepancy through geometric modeling and boundary measurement, we systematically evaluate its impact on vision-language model (VLM) tasks. This work is the first to reveal that such a mismatch can introduce critical security vulnerabilities, including prompt injection and privacy leakage. Our findings demonstrate that the FoV mismatch extends beyond a mere geometric artifact; it constitutes a substantial threat to the security and reliability of AI-integrated XR systems. Consequently, this research establishes a novel security analysis perspective for the field, highlighting the urgent need to account for perceptual discrepancies when deploying VLMs within immersive environments.
📝 Abstract
Video see-through extended reality (VST XR) systems commonly use headset screenshots or captured frames as proxies for the user's first-person visual context. However, the system-captured view and the user's effective visible field do not necessarily coincide: a screenshot records a rectangular machine-readable frame, whereas the user's effective visible region can be more constrained and non-rectangular. This paper studies this human-system view mismatch in VST XR. We formalize the relationship between the system-captured region and the human-visible region by defining their co-visible, system-only, and human-only regions. \rev{We then conduct a pilot-level boundary measurement on Meta Quest 3, revealing a clear mismatch between the rectangular screenshot frame and the approximate human-visible boundary. Building on this model, we analyze how view mismatch can affect screenshot-based XR sensing and downstream vision-language model tasks. Through four representative case studies, we illustrate potential risks and failure modes including prompt injection, privacy leakage, human-invisible information bias, and missing human-visible information. Our results show that view mismatch is not only a geometric artifact, but can also introduce security, privacy, and reliability concerns for AI-integrated VST XR systems.
Problem

Research questions and friction points this paper is trying to address.

Video See-Through XR
View Mismatch
Vision-Language Models
Security and Privacy
Extended Reality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video See-Through Extended Reality
View Mismatch
Vision-Language Models
Prompt Injection
Privacy Leakage
🔎 Similar Papers
No similar papers found.