🤖 AI Summary
This study addresses the challenge of precisely executing geometric tasks with microrobots under varying linguistic instructions, component configurations, and focal shifts. We propose a semantics-to-physics mapping framework that parses natural language commands into constrained geometric operators while leveraging frozen perception models for open-vocabulary understanding. By employing confidence weighting and dynamic programming to optimize focal plane trajectories, and adopting a zero-new-label configuration that replaces image-level fusion with trajectory-space integration, the method substantially reduces annotation dependence. Experimental results demonstrate that our approach decreases the RMSE to 6.28 pixels (a 56.4% error reduction), shortens training time to 15 minutes, and achieves a 92.9% target region coverage rate. These findings confirm that the proposed framework effectively enhances multi-view geometric calibration accuracy and operational robustness for microrobot manipulation.
📝 Abstract
Microscopic robots require accurate task geometry despite changes in language, parts, and focus. We present a semantic-to-physical framework that maps instructions to constrained geometric operators, reuses frozen open-vocabulary perception, and integrates locally reliable focal-plane trajectories by confidence weighting and dynamic programming. Calibrated multi-view geometry connects 2-D paths to physical execution. Prompt, unseen-part, and geometry reconfiguration tests yield 6.30-6.59-pixel RMSE. Relative to part-specific U-Net training with 20-100 labels, the proposed zero-new-label configuration takes 15 rather than 72-165 min. Across nine part-illumination conditions, trajectory-space integration reduces RMSE from 14.41 to 6.28 pixels (56.4%) and P95 error from 20.07 to 8.13 pixels (59.5%) compared with image-first multi-focus fusion. An ablation isolates the roles of confidence and path-wise selection. In representative robot experiments, target-region coverage improves from 83.5% to 92.9%. Dispensing provides a measurable physical trace, not a task-specific limitation of the method.