Phone2Act: A Low-Cost, Hardware-Agnostic Teleoperation System for Scalable VLA Data Collection

📅 2026-05-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high cost and hardware dependency of existing vision-language-action (VLA) models, which rely on expensive, platform-specific demonstration data that lacks cross-platform reusability. The authors propose a low-cost, hardware-agnostic teleoperation system leveraging a smartphone and Google ARCore to achieve six-degree-of-freedom control. Built on a modular ROS 2 architecture, the system decouples control logic from hardware through pluggable bridge nodes, enabling seamless integration across diverse robotic platforms. It synchronizes multi-camera RGB video with robot state data to directly generate datasets in LeRobot format. Without requiring code modifications, the framework supports end-to-end workflows—from data collection to VLA model fine-tuning—spanning industrial collaborative robots to low-cost dual-arm setups. Fine-tuning the GR00T-N1.5 model on 130 task demonstrations achieves a 90% success rate on multi-stage pick-and-place tasks executed on a real Dobot CR5 robot.
📝 Abstract
Collecting diverse, high-quality manipulation data for Vision-Language-Action (VLA) model training remains prohibitively expensive for many research groups, as existing teleoperation frameworks rely on specialized hardware or are tightly coupled to specific robot platforms. We present Phone2Act, a low-cost, hardware-agnostic teleoperation framework that transforms a commodity smartphone into a 6-DoF robot controller via Google ARCore. Built on a modular ROS 2 architecture, Phone2Act decouples control logic from hardware specifics through interchangeable bridge nodes, supporting platforms from industrial cobots to low-cost bimanual arms without code modification. A Universal Recorder synchronizes multi-camera RGB streams with robot state feedback and exports demonstrations natively in the LeRobot dataset format, eliminating post-processing and enabling immediate VLA fine-tuning. We validate the framework by fine-tuning GR00T-N1.5 on 130 collected episodes, achieving a 90% success rate on a real-world multi-stage pick-and-place task deployed on a physical Dobot CR5.
Problem

Research questions and friction points this paper is trying to address.

teleoperation
Vision-Language-Action
data collection
robot manipulation
hardware-agnostic
Innovation

Methods, ideas, or system contributions that make the work stand out.

teleoperation
hardware-agnostic
Vision-Language-Action (VLA)
ROS 2
data collection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
O
Om Mandhane
Dept. of Automation & Robotics Engineering, Vivekanand Education Society’s Institute of Technology (VESIT), Mumbai, India
B
Bipin Yadav
Dept. of Automation & Robotics Engineering, Vivekanand Education Society’s Institute of Technology (VESIT), Mumbai, India
S
Sangeetha Prasanna Ram
Dept. of Automation & Robotics Engineering, Vivekanand Education Society’s Institute of Technology (VESIT), Mumbai, India
G
Gopalakrishnan Narayanan
Dept. of Automation & Robotics Engineering, Vivekanand Education Society’s Institute of Technology (VESIT), Mumbai, India