Compact Task-Aligned Imitation Learning for Laboratory Automation

📅 2026-03-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes TVF-DiT, a lightweight imitation learning framework that overcomes the limitations of traditional lab automation—reliant on costly custom hardware—and existing imitation approaches burdened by large model sizes unsuitable for resource-constrained settings. TVF-DiT uniquely integrates a compact self-supervised vision foundation model with a diffusion-based policy by employing a small adapter to align visual and language representations and incorporating task prompts to enhance cross-modal alignment. The resulting diffusion Transformer-based action policy operates with fewer than 500 million parameters and achieves an average success rate of 86.6% across three real-world tasks: test tube cleaning, arrangement, and powder transfer. This performance significantly surpasses lightweight baselines while enabling real-time inference on low-memory GPUs.

Technology Category

Machine Learning: Imitation Learning & Inverse Reinforcement LearningComputer Vision: Diffusion Models for VisionIntelligent Robots: Manipulation

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Robotic laboratory automation has traditionally relied on carefully engineered motion pipelines and task-specific hardware interfaces, resulting in high design cost and limited flexibility. While recent imitation learning techniques can generate general robot behaviors, their large model sizes often require high-performance computational resources, limiting applicability in practical laboratory environments. In this study, we propose a compact imitation learning framework for laboratory automation using small foundation models. The proposed method, TVF-DiT, aligns a self-supervised vision foundation model with a vision-language model through a compact adapter, and integrates them with a Diffusion Transformer-based action expert. The entire model consists of fewer than 500M parameters, enabling inference on low-VRAM GPUs. Experiments on three real-world laboratory tasks - test tube cleaning, test tube arrangement, and powder transfer - demonstrate an average success rate of 86.6%, significantly outperforming alternative lightweight baselines. Furthermore, detailed task prompts improve vision-language alignment and task performance. These results indicate that small foundation models, when properly aligned and integrated with diffusion-based policy learning, can effectively support practical laboratory automation with limited computational resources.
Problem

Research questions and friction points this paper is trying to address.

laboratory automation
imitation learning
compact models
computational efficiency
robotic flexibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

compact imitation learning
foundation model alignment
Diffusion Transformer
laboratory automation
low-resource robotics
🔎 Similar Papers
No similar papers found.