Institution profile

Association Mobsya

Research institutioneurope · ch
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

RoboPace: Contact-Aware Time-Optimal Retiming for Action-Chunk Policies

Oct 07, 2026

This study addresses the challenge that directly inheriting human demonstration temporal profiles in robot learning often leads to an imbalance between contact safety and execution efficiency. To this end, it proposes a training-free online retiming layer that eliminates the need for policy retraining. By preserving geometric paths while integrating contact prediction with action chunking, the method dynamically adjusts execution speed based on predicted contact states, thereby unifying contact-dependent velocity limits with kinodynamic constraints. Experimental evaluations on a dual-arm robotic platform demonstrate that, compared to a conservative slow-execution baseline, the proposed approach reduces task completion time by half while significantly improving success rates and maintaining high reliability.

0 citationsRead paper

YUBI-STAG: Contact and Semantic-Rich Alignment for VLAs via Automated Video-Language Grounding

Oct 07, 2026

This study addresses the challenge of aligning instructions with physical manipulation in Vision-Language-Action (VLA) models, which arises from the lack of fine-grained interaction semantics in demonstration data. To this end, we propose YUBI-STAG, a spatiotemporal annotation framework that integrates contact-object segmentation, vision-language model (VLM) reasoning, and knowledge distillation to automatically annotate robot manipulation videos with detailed contact, object, and action labels. Furthermore, we distill an efficient model, YUBI-VLM, capable of recovering action structures solely from wrist-view observations. Experimental results demonstrate that our approach maintains high annotation accuracy while significantly reducing inference costs, effectively enhancing the instruction-following capabilities of bimanual robots and their fine-grained control over object identity and spatial relationships.

0 citationsRead paper

Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations

Aug 07, 2026

This work addresses the implicit misalignment between language instructions and behavioral trajectories in multimodal robot demonstration data by proposing MMPF, a training-free auditing framework. Treating each modality as an expert, MMPF estimates task label distributions through a combination of local neighborhood consistency and global prototype similarity, then fuses multimodal information via prediction entropy weighting to achieve high-accuracy mismatch detection and label correction. As the first method specifically designed to handle data errors where actions are correct but instructions are erroneous, MMPF achieves state-of-the-art performance in instruction-trajectory matching (ITM) detection and correction on both the LIBERO benchmark and real-world robot datasets. Experimental results demonstrate that MMPF significantly enhances downstream policy learning, with physical robot trials validating an effective trade-off between filtering and relabeling strategies.

0 citationsRead paper
Recent publications

Latest Papers

RoboPace: Contact-Aware Time-Optimal Retiming for Action-Chunk Policies

Oct 07, 2026

This study addresses the challenge that directly inheriting human demonstration temporal profiles in robot learning often leads to an imbalance between contact safety and execution efficiency. To this end, it proposes a training-free online retiming layer that eliminates the need for policy retraining. By preserving geometric paths while integrating contact prediction with action chunking, the method dynamically adjusts execution speed based on predicted contact states, thereby unifying contact-dependent velocity limits with kinodynamic constraints. Experimental evaluations on a dual-arm robotic platform demonstrate that, compared to a conservative slow-execution baseline, the proposed approach reduces task completion time by half while significantly improving success rates and maintaining high reliability.

0 citationsRead paper

YUBI-STAG: Contact and Semantic-Rich Alignment for VLAs via Automated Video-Language Grounding

Oct 07, 2026

This study addresses the challenge of aligning instructions with physical manipulation in Vision-Language-Action (VLA) models, which arises from the lack of fine-grained interaction semantics in demonstration data. To this end, we propose YUBI-STAG, a spatiotemporal annotation framework that integrates contact-object segmentation, vision-language model (VLM) reasoning, and knowledge distillation to automatically annotate robot manipulation videos with detailed contact, object, and action labels. Furthermore, we distill an efficient model, YUBI-VLM, capable of recovering action structures solely from wrist-view observations. Experimental results demonstrate that our approach maintains high annotation accuracy while significantly reducing inference costs, effectively enhancing the instruction-following capabilities of bimanual robots and their fine-grained control over object identity and spatial relationships.

0 citationsRead paper

Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations

Aug 07, 2026

This work addresses the implicit misalignment between language instructions and behavioral trajectories in multimodal robot demonstration data by proposing MMPF, a training-free auditing framework. Treating each modality as an expert, MMPF estimates task label distributions through a combination of local neighborhood consistency and global prototype similarity, then fuses multimodal information via prediction entropy weighting to achieve high-accuracy mismatch detection and label correction. As the first method specifically designed to handle data errors where actions are correct but instructions are erroneous, MMPF achieves state-of-the-art performance in instruction-trajectory matching (ITM) detection and correction on both the LIBERO benchmark and real-world robot datasets. Experimental results demonstrate that MMPF significantly enhances downstream policy learning, with physical robot trials validating an effective trade-off between filtering and relabeling strategies.

0 citationsRead paper