🤖 AI Summary
This study addresses the challenges of missing scene semantics, low observation quality, and unnatural interaction in VR-based teleoperation by proposing a robot-agnostic immersive collaborative teleoperation framework. The method constructs a unified semantic Gaussian map (Gaussian-TSDF) adaptable to multiple platforms and integrates a Qwen3-VL large vision-language model agent to enable open-vocabulary instance management and natural interaction. Additionally, a shadow tracking mechanism is introduced to bridge localization interruptions. Experimental results demonstrate that the system achieves an 81.24% tool matching rate, improves rendering PSNR by 2–8 dB, and attains an 86.7% success rate for navigation requests.
📝 Abstract
A photorealistic 3D view tells a teleoperator where a robot is, but not what the scene contains, how well each object has been observed, or how to turn pointing and speech into robot action. CognitiveReality turns a robot's RGB-D stream into a live, semantically indexed Gaussian-TSDF map shared by an operator in virtual reality and a tool-using language agent. One mapper binary serves any platform through configuration alone: it ingests poses from robot SLAM, joint kinematics, motion capture or an inline visual tracker, bridges localization outages with a shadow tracker and keyframe-anchored PnP, and maintains open-vocabulary instance identities with per-object quality at 2 Hz. Speech and controller rays are grounded against persistent scene objects through validated typed tools and operator-confirmed robot actions. In the controlled agent evaluation, the deployed local Qwen3-VL-8B router reaches 81.24\% tool exact match, while merge-aware replay correctly redirects 101 absorbed object identifiers. On robot data CognitiveReality exceeds a Gaussian-plus-SDF baseline by 2-8 dB; pose error through 5-40 s SLAM outages stays within 1-8 cm. Deployed live on two quadrupeds, the agent executed 26 of 30 navigation requests and 20 of 20 re-observation requests, raising object quality by 2-5 dB.