Rendering-Free Lookahead for Question-Guided Active Vision

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of anticipating unobserved viewpoints to select optimal camera motions in active vision by proposing a Rendered Foresight Learning (RFL) strategy. RFL pioneers the shift of rendering-dependent visual foresight from online inference to an offline training phase. By integrating 3D Gaussian Splatting with Vision-Language Models, it employs two-stage knowledge distillation to transfer privileged rendering capabilities into a student model, thereby enabling render-free, real-time prediction of future answerability and action value estimation. Evaluated on the E3VS-Bench benchmark, RFL improves the average score by 43% over baselines, substantially enhancing performance on viewpoint-dependent visual question answering tasks.
📝 Abstract
Active robot vision requires controlling the camera to reveal task-relevant information that is hidden from the current viewpoint. For example, determining what is inside a box may require raising the camera and looking down into it. For viewpoint-dependent question answering, the challenge is to select camera motions that expose the visual evidence needed to answer the question. Although vision-language models (VLMs) can interpret observed images, selecting such motions requires anticipating the usefulness of unseen views. We quantify this usefulness as answerability, a VLM's estimate that a view suffices to answer the question, and present Rendering-Free Lookahead (RFL), a viewpoint-selection policy that ranks candidate camera motions by predicted future answerability. RFL transfers visual lookahead from deployment to offline training. At training, a privileged teacher renders candidate future views in 3D Gaussian Splatting (3DGS) scenes and uses a frozen VLM to compute one- and two-step answerability targets. Through two-stage distillation, a student learns to predict these action values from the question, recent visual observations, and a candidate camera motion. At deployment, RFL uses these predicted values to select camera motions without rendering future views. On 377 E3VS-Bench test episodes in unseen environments, RFL improves the mean judge score by 43\% over a direct-action baseline using the same VLM. These results support learning camera-control policies from privileged visual lookahead for viewpoint-dependent question answering.
Problem

Research questions and friction points this paper is trying to address.

Active robot vision
Viewpoint-dependent question answering
Camera motion selection
Answerability prediction
Vision-language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rendering-Free Lookahead
Active Vision
Vision-Language Models
3D Gaussian Splatting
Knowledge Distillation
🔎 Similar Papers
No similar papers found.