Fisheye-VLA: Decoupling Coverage and Acuity for Manipulation with a Single Fisheye Camera

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究提出Fisheye-VLA,利用单个鱼眼相机解决操作中场景感知和局部反馈的问题,通过校准末端执行器投影和运动引导来追踪局部裁剪。
📝 Abstract
Manipulation requires both broad scene awareness and detailed local feedback, yet conventional camera rigs provide them through separate front and wrist cameras. We present Fisheye-VLA, a visual interface that brings these capabilities together using a single passive fisheye. A global view preserves the workspace, while local perspective crops direct detail toward the interaction. The key design question is where this local visual budget should go. We answer it through a controlled re-rendering study, comparing alternative crop directions on the same recorded observations. The study finds that end-effector-centered views capture most of the estimated benefit of a much larger candidate pool, motivating a compact allocation around both hands. Our interface uses calibrated end-effector projection and motion lead to track the crops, while a shared ray encoding preserves their spatial meaning as they move. Integrated with a pretrained VLA, it achieves 84% and 82% success in the two expanded tabletop regions, where some target placements extend beyond the front-camera coverage, and supports shelf and conveyor manipulation. Ablations show that local crops and their viewing directions become more important in the larger workspace regions. The results demonstrate that a single fisheye can support these manipulation tasks without physical wrist cameras.
Problem

Research questions and friction points this paper is trying to address.

manipulation
fisheye camera
scene awareness
local feedback
workspace
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fisheye Camera
End-Effector Projection
Local Crops
Shared Ray Encoding
Manipulation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ziang Ren
Ziang Ren
Columbia University
Computer VisionRobotics
Zike Yan
Zike Yan
PostDoc, Tsinghua University; PhD, Peking University
3D VisionRoboticsContinual Learning
R
Raymond Zhang
DeepCybo, University of Washington, Seattle, WA, USA
X
Xuguo He
DeepCybo, University of Washington, Seattle, WA, USA
Z
Zhongyu Li
Hong Kong Embodied AI Lab, The Chinese University of Hong Kong, Hong Kong SAR, China