🤖 AI Summary
This work addresses the challenge of inflexible visual systems in high-mix, low-volume manufacturing, where introducing new objects typically demands extensive labeled data and model retraining. The paper proposes FreeZe, a zero-shot 6D object pose estimation method that achieves plug-and-play localization without requiring object-specific fine-tuning or annotations. By integrating semantic features from foundation models such as DINOv2 and GeDi with geometric information from CAD models, FreeZe attains millimeter-level accuracy even under severe occlusion. The approach substantially enhances robotic generalization in unstructured environments, supports deployment on industrial edge platforms like NVIDIA Jetson Thor, reaches Technology Readiness Level 6, and secured first place in the BOP Challenge 2024.
📝 Abstract
The transition toward high-mix low-volume manufacturing demands flexibility in robotic manipulation. However, conventional vision systems remain a bottleneck, requiring extensive data collection and model retraining whenever a new object is introduced to the production line. To overcome this rigidity, we present xperception, a zero-shot 6D pose estimation technology that eliminates the need for object-specific fine-tuning and laborious data annotation. By directly utilizing typical CAD models and integrating the rich semantic features of foundation models (e.g. DINOv2, GeDi), xperception achieves millimeter-accurate 6D pose estimation. xperception showed robustness against severe occlusions in industrial tasks like bin picking and is engineered for deployment on industrial edge hardware, such as NVIDIA Jetson Thor. Validated at a TRL of 6, the core methodology behind xperception is based on the FreeZe algorithm, which won the international BOP Challenge 2024, paving the way for scalable, plug-and-play robotic automation in unstructured high-mix low-volume manufacturing industries.