Residual Modeling Closes the Regression and Generative Policy Gap in Robot Learning

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance gap between regression-based and generative policies in robot learning, which arises from heavy-tailed action residual distributions. We reveal that this discrepancy stems from state-dependent variations in residual scale and propose Heteroscedastic Student-t Regression (HT-Policies). By modeling input-dependent residual scales, our method suppresses interference from large-error gradients. It further repurposes pretrained flow matching networks as backbones to enable single-forward-pass action chunk prediction while supporting transfer learning for vision-language-action models. Across four simulation benchmarks and real-world evaluations, HT-Policies achieve success rates comparable to generative baselines while substantially improving both training and inference efficiency. This work establishes heteroscedastic regression as an effective and efficient alternative to generative approaches for robot policy learning.
📝 Abstract
Learning from demonstration has enabled impressive robot behaviors. A common choice for policy learning is to use diffusion or flow matching (Flow-Policies), which often outperforms direct action regression trained with mean squared error (MSE-Policies). This gap is commonly attributed to multimodal demonstrations. We revisit this gap from the perspective of statistical modeling: how action-prediction residuals shape policy optimization. Our analysis of real-world robot demonstration data reveals substantial state-dependent variation in residual scales and heavier-than-Gaussian tails. While both MSE-Policies and Flow-Policies exhibit heavy-tailed action residuals, their training gradients behave differently: MSE allocates more gradient magnitude to observations with large action residuals, which hurts optimization. Motivated by these findings, we introduce heteroscedastic Student-t action regression (HT-Policies), which learns input-dependent residual scales and reduces the influence of heavy tails. HT-Policies predict action chunks with a single feed-forward pass and can reuse pretrained flow-matching-based policy networks as the backbone. Across four simulation benchmarks and real-robot evaluations, HT-Policies achieves success rates competitive with generative policy baselines, both when trained from scratch and from pretrained vision-language-action and world-action models, despite being faster in training and inference. Together, these findings shed light on the practical advantages of generative objectives in robot learning from demonstrations and offer an efficient direct-regression alternative for a range of architectures and tasks. Project page: https://the-labone.github.io/regression-policy-project/
Problem

Research questions and friction points this paper is trying to address.

Robot Learning
Learning from Demonstration
Action Regression
Generative Policy
Residual Modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Heteroscedastic Regression
Student-t Distribution
Action Residuals
Flow Matching
Robot Learning