Institution profile

Tommoro Robotics

Industry research
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent Lookahead

Sep 29, 2026

This study addresses the vulnerability of Vision-Language-Action (VLA) models to visual distribution shifts caused by their reliance on spurious correlations. To mitigate this issue, we propose the DILL framework, which employs a task-domain dual encoder and domain-invariant latent lookahead prediction. By integrating contrastive learning with Gaussian decoupling regularization, DILL effectively disentangles task-relevant structures from domain-specific visual variations, thereby facilitating robust policy learning. Experimental results demonstrate that DILL achieves a success rate of 69.1% on the LIBERO-Plus benchmark, outperforming baseline methods by 11.4 percentage points. Furthermore, its generalization capability is validated through real-world robotic manipulation tasks.

0 citationsRead paper

Habilis-$β$: A Fast-Motion and Long-Lasting On-Device Vision-Language-Action Model

Feb 21, 2026

This work addresses the limitations of existing vision-language-action models, which are typically evaluated solely on single-task success rates and thus fail to capture the throughput and long-term reliability required for real-world deployment. To bridge this gap, the authors propose a vision-language-action model tailored for edge-based real-world scenarios, introducing the Productivity-Reliability Plane (PRP) evaluation framework grounded in continuous-operation protocols. Key innovations include language-agnostic pretraining on large-scale play data, cyclic task fine-tuning, phase-adaptive motion planning (ESPADA), rectified flow distillation, and classifier-free guidance. The model achieves 572.6 tasks per hour (TPH) with a mean time between interventions (MTBI) of 39.2 seconds in simulation, and 124 TPH with 137.4 seconds MTBI on real-world logistics tasks—significantly outperforming baselines and establishing state-of-the-art performance on the RoboTwin 2.0 benchmark.

0 citationsRead paper
Recent publications

Latest Papers

Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent Lookahead

Sep 29, 2026

This study addresses the vulnerability of Vision-Language-Action (VLA) models to visual distribution shifts caused by their reliance on spurious correlations. To mitigate this issue, we propose the DILL framework, which employs a task-domain dual encoder and domain-invariant latent lookahead prediction. By integrating contrastive learning with Gaussian decoupling regularization, DILL effectively disentangles task-relevant structures from domain-specific visual variations, thereby facilitating robust policy learning. Experimental results demonstrate that DILL achieves a success rate of 69.1% on the LIBERO-Plus benchmark, outperforming baseline methods by 11.4 percentage points. Furthermore, its generalization capability is validated through real-world robotic manipulation tasks.

0 citationsRead paper

Habilis-$β$: A Fast-Motion and Long-Lasting On-Device Vision-Language-Action Model

Feb 21, 2026

This work addresses the limitations of existing vision-language-action models, which are typically evaluated solely on single-task success rates and thus fail to capture the throughput and long-term reliability required for real-world deployment. To bridge this gap, the authors propose a vision-language-action model tailored for edge-based real-world scenarios, introducing the Productivity-Reliability Plane (PRP) evaluation framework grounded in continuous-operation protocols. Key innovations include language-agnostic pretraining on large-scale play data, cyclic task fine-tuning, phase-adaptive motion planning (ESPADA), rectified flow distillation, and classifier-free guidance. The model achieves 572.6 tasks per hour (TPH) with a mean time between interventions (MTBI) of 39.2 seconds in simulation, and 124 TPH with 137.4 seconds MTBI on real-world logistics tasks—significantly outperforming baselines and establishing state-of-the-art performance on the RoboTwin 2.0 benchmark.

0 citationsRead paper