CoRe-VLA: Preserving Cross-View Coordination in VLAs under Camera Shifts

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Vision-Language-Action (VLA) models suffer severe performance degradation under camera viewpoint shifts due to cross-view coordination failures. This work proposes CoRe-VLA, a plug-and-play framework that introduces the first fine-tuning-free and data-free point cloud rerendering adaptation scheme. By reconstructing scene point clouds and rerendering them from training viewpoints, CoRe-VLA restores observational consistency. It further incorporates R2C de-degradation techniques and an ETCA asynchronous execution alignment algorithm to eliminate adaptation discrepancies. Evaluated across five real-world robotic tasks and the LIBERO benchmark, the proposed method substantially enhances the deployment robustness of mainstream VLA models. Notably, it elevates the success rate of PI0.5 from 13.3% to 83.3% under a 1.6-meter viewpoint shift, demonstrating its effectiveness in mitigating spatial misalignment during robot deployment.
📝 Abstract
VLAs combine pretrained vision-language representations with action generation to enable language-guided control across diverse tasks, becoming a mainstream paradigm in embodied intelligence. However, multiple studies have reported VLA's substantial declines in task success under camera shifts, revealing a key vulnerability that limits reliable deployment. To address this vulnerability, existing methods collect paired observations of the same scene from different viewpoints to fine-tune the VLA or train visual adaptation modules. Unfortunately, they require additional data collection and VLA training costs. In this paper, we first identify \emph{cross-view coordination breakdown} under external camera shifts: the robot may rely too heavily on wrist-view cues and consequently execute subtasks in the wrong order when losing global view. Motivated by this, we propose CoRe-VLA, a plug-and-play framework requiring neither additional multi-view data collection nor VLA fine-tuning, which can incorporate with exsiting VLAs. It reconstructs a scene point cloud and renders the observation from the VLA's training viewpoint to restore cross-view coordination. In CoRe-VLA, Render-to-Camera (R2C) Restoration reduces rendering-induced visual degradation, while Execution-Trajectory-Conditioned Alignment (ETCA) reduces robot idle time and mitigates motion conflicts during asynchronous execution. Experiments on 5 real-robot tasks, LIBERO-100 and LIBERO-Plus demonstrate CoRe-VLA substantially improves task success across mainstream VLAs under camera shifts. For example, CoRe-VLA raises PI0.5's success rate from 13.3% to 83.3% at a 1.6m camera shift in real-robot environment.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
camera shifts
cross-view coordination breakdown
embodied intelligence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-View Coordination
Plug-and-Play Framework
Point Cloud Rendering
Render-to-Camera Restoration
Execution-Trajectory-Conditioned Alignment
🔎 Similar Papers
No similar papers found.
T
Tianhang Pan
Southeast University, Nanjing, China
X
Xuanhao Wang
Southeast University, Nanjing, China
Y
Yiwen Pang
Southeast University, Nanjing, China
B
Bo Zhou
Southeast University, Nanjing, China
J
Jun Yang
Southeast University, Nanjing, China; National Center of Technology Innovation for EDA, Nanjing, China
Min-Ling Zhang
Min-Ling Zhang
Professor, School of Computer Science and Engineering, Southeast University, China
Artificial IntelligenceMachine LearningData Mining
S
Shimin Di
Southeast University, Nanjing, China