From Language to Motion: Task-Conditioned Focal-Stack Trajectory Integration for Microscopic Robots

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of precisely executing geometric tasks with microrobots under varying linguistic instructions, component configurations, and focal shifts. We propose a semantics-to-physics mapping framework that parses natural language commands into constrained geometric operators while leveraging frozen perception models for open-vocabulary understanding. By employing confidence weighting and dynamic programming to optimize focal plane trajectories, and adopting a zero-new-label configuration that replaces image-level fusion with trajectory-space integration, the method substantially reduces annotation dependence. Experimental results demonstrate that our approach decreases the RMSE to 6.28 pixels (a 56.4% error reduction), shortens training time to 15 minutes, and achieves a 92.9% target region coverage rate. These findings confirm that the proposed framework effectively enhances multi-view geometric calibration accuracy and operational robustness for microrobot manipulation.
📝 Abstract
Microscopic robots require accurate task geometry despite changes in language, parts, and focus. We present a semantic-to-physical framework that maps instructions to constrained geometric operators, reuses frozen open-vocabulary perception, and integrates locally reliable focal-plane trajectories by confidence weighting and dynamic programming. Calibrated multi-view geometry connects 2-D paths to physical execution. Prompt, unseen-part, and geometry reconfiguration tests yield 6.30-6.59-pixel RMSE. Relative to part-specific U-Net training with 20-100 labels, the proposed zero-new-label configuration takes 15 rather than 72-165 min. Across nine part-illumination conditions, trajectory-space integration reduces RMSE from 14.41 to 6.28 pixels (56.4%) and P95 error from 20.07 to 8.13 pixels (59.5%) compared with image-first multi-focus fusion. An ablation isolates the roles of confidence and path-wise selection. In representative robot experiments, target-region coverage improves from 83.5% to 92.9%. Dispensing provides a measurable physical trace, not a task-specific limitation of the method.
Problem

Research questions and friction points this paper is trying to address.

microscopic robots
task geometry
focal-stack trajectory
language-to-motion
open-vocabulary perception
Innovation

Methods, ideas, or system contributions that make the work stand out.

Focal-Stack Trajectory Integration
Semantic-to-Physical Framework
Open-Vocabulary Perception
Zero-Shot Configuration
Multi-View Geometry
💼 Related Jobs
No related jobs found.
Junjie Xie
Junjie Xie
National university of defense technology
Computer Science
C
Chuxuan He
Zhejiang Gongshang University
Junkai Huang
Junkai Huang
Cornell University
H
Heng Zhang
Safe AI Lab, Carnegie Mellon University
A
Angen Ye
Institute of Automation, Chinese Academy of Sciences
Y
Yujia Song
Institute of Automation, Chinese Academy of Sciences
Yuqing Li
Yuqing Li
East China Normal University
Deep Learning Theory
Pengsong Zhang
Pengsong Zhang
Ph.D. Candidate, University of Toronto
RoboticsAgentComputer visionReinforcement learningAI4S
D
Dapeng Zhang
Institute of Automation, Chinese Academy of Sciences