🤖 AI Summary
This work addresses the high barrier to entry in traditional robot programming by proposing a no-code framework that integrates mixed reality (MR) and vision-language models (VLMs) to enable intuitive, safe, and cross-platform robot task specification. Users define tasks through natural language commands, kinesthetic teaching trajectories, or digital twins within an MR environment. The system leverages a VLM to generate structured action plans from these multimodal inputs and employs MR as a safety verification layer, allowing virtual preview and embodied perceptual validation before execution. The approach unifies programming across diverse robotic morphologies—including fixed manipulators, mobile bases, and humanoid platforms—thereby achieving secure, intuitive, and platform-agnostic translation from natural language instructions to executable robot actions.
📝 Abstract
ORCESTRA is a mixed-reality system for programming robot digital twins through no-code waypoint teaching and language-guided control. In a passthrough mixed-reality workspace, users place robot twins on real surfaces, teach trajectories, save robot-relative episodes, or issue spoken/typed commands that a vision-language model converts into structured digital-twin plans. Both interaction modes share a backend for metric grounding, embodiment-aware validation, preview, confirmation, and digital-twin execution. The system supports heterogeneous robot embodiments, including fixed-base manipulators, a mobile base, and a humanoid robot, demonstrating MR validation as a safety layer for language-guided robot programming before physical deployment.