🤖 AI Summary
This study addresses the limitation of vision-language models (VLMs) in perspective-taking reasoning, where they are often constrained by default camera viewpoints and struggle to interpret spatial relations from specified entity perspectives. To overcome this, we propose LeRF, a novel framework introducing a learnable reference frame mechanism. LeRF localizes reference objects, predicts their origins and axes, and renders visual cues to facilitate reasoning. By integrating selective tool invocation with a lightweight renderer, it achieves end-to-end viewpoint-dependent reasoning without requiring external perception or 3D reconstruction. The framework trains VLMs for coordinate frame estimation through a combination of supervised fine-tuning and reinforcement learning. Extensive evaluations across multiple benchmarks demonstrate that LeRF significantly outperforms existing methods, effectively enhancing both reference frame localization accuracy and orientation estimation capabilities.
📝 Abstract
Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as that of another entity or an imagined observer. Despite the increasing spatial reasoning capabilities of Vision-Language Models (VLMs), they still struggle with perspective taking, often defaulting to the camera viewpoint when a query requires reasoning from a different perspective. We introduce Learning Reference Coordinate Frames for Perspective Taking (LeRF), a framework that trains VLMs to construct and use explicit reference frames for viewpoint-dependent reasoning. Given an image and a query, LeRF decides whether a coordinate frame is necessary. If so, it grounds the reference entity and predicts the frame's origin and entity-centered reference frame. A lightweight renderer overlays the frame onto the image, enabling subsequent reasoning over these visual cues without external perception models or explicit 3D reconstruction. To learn this process, we first perform supervised fine-tuning to teach selective tool invocation and reference coordinate frame prediction, followed by reinforcement learning on spatial VQA pairs to improve frame-guided reasoning. Across diverse perspective-taking benchmarks, LeRF consistently improves over its backbone and achieves strong performance against existing open-source methods. Further evaluations also show improved reference-frame grounding and orientation estimation, supporting the effectiveness of learned reference frames for viewpoint-dependent reasoning.