Score
Designs and implements fine-tuning procedures for inverse-dynamics models that map observed state transitions to actions, adapting pretrained inverse-dynamics models to target dynamics using supervised paired observation–action data. Works with small rollout datasets (minutes of interaction) to align model-predicted actions to the target dynamics and to evaluate finetuning performance and stability.
This work addresses the performance degradation of behavioral cloning under limited expert demonstrations and investigates the superior sample efficiency of Predictive Inverse Dynamics Models (PIDMs), whose underlying mechanism has remained unclear. We theoretically analyze PIDMs through the bias–variance tradeoff, revealing that while incorporating future state prediction introduces bias, it substantially reduces the variance in inverse dynamics estimation, thereby enhancing sample efficiency. By integrating a state predictor with an inverse dynamics model and leveraging additional data for both theoretical analysis and empirical validation, we demonstrate that PIDMs achieve comparable performance to behavioral cloning using only one-third of the demonstration data in 2D navigation tasks. In high-dimensional 3D environments with visual inputs, PIDMs reduce the required demonstration data by over 66%.
In robot policy fine-tuning, conventional pipelines—imitation learning (IL) pretraining followed by reinforcement learning (RL) fine-tuning—suffer from unstable exploration, low sample efficiency, and performance collapse. This paper proposes an alternating IL/RL fine-tuning paradigm. A gradient orthogonal separation mechanism decouples imitation and reinforcement objectives, with theoretical analysis revealing its intrinsic benefits for stability and sample efficiency. We further design a gradient subspace decomposition module that enables plug-and-play integration of behavior cloning with standard RL algorithms (e.g., PPO, SAC). Evaluated on 14 robotic manipulation and locomotion tasks, our method achieves an average 3.8× improvement in sample efficiency (up to 6.3×) and boosts task success rate from 12% to 88% (+76 percentage points), while robustly suppressing performance collapse under both sparse/dense and short-/long-horizon reward settings.
This work proposes a semi-supervised imitation learning approach based on an inverse dynamics model (IDM) in settings with limited action-labeled trajectories and abundant unlabeled trajectories. The IDM predicts actions from state transitions and can either serve directly as a policy (VM-IDM) or generate pseudo-labels for unlabeled data (IDM labeling). Theoretical analysis reveals that IDM outperforms behavioral cloning primarily due to its lower hypothesis class complexity and reduced stochasticity, leading to higher sample efficiency. Building on this insight, the authors enhance the LAPO algorithm and validate the proposed method within a unified video-action prediction (UVA) framework. Both theoretical analysis and empirical results consistently demonstrate the superior sample efficiency of the IDM-based approach.
Open-loop imitation learning has advanced modern autonomous driving policy architectures, but closed-loop deployment remains vulnerable to policy-induced distribution shift. Existing post-training paradigms exhibit fundamental trade-offs: closed-loop RL fine-tuning provides grounded feedback from executed actions but is constrained by the sparsity of informative events, whereas counterfactual fine-tuning provides dense supervision over candidate futures but inherits bias from imperfect future estimates. We introduce Counterfactual-to-Interactive Reinforcement Fine-Tuning (CRAFT), an on-policy framework that formulates closed-loop post-training as proxy-residual optimization. CRAFT uses group-normalized counterfactual advantages as a dense proxy for real closed-loop advantages and aligns this proxy with the closed-loop world through grounded residual correction from interaction-critical events. To stabilize adaptation, CRAFT regularizes the online policy toward an EMA teacher via asymmetric KL self-distillation. Theoretically, CRAFT decomposes the real closed-loop policy gradient into proxy and residual terms under the same visited-state distribution, reducing residual variance with an aligned proxy while mitigating proxy bias through grounded residual approximation. Empirically, CRAFT achieves the strongest closed-loop gains on Bench2Drive across hierarchical planning, vision-language-action, and vocabulary-scoring architectures. Ablations, scaling behavior, stability analyses, and transfer results further validate the complementary roles of dense counterfactual proxy and grounded residual correction. Project page: https://currychen77.github.io/CRAFT.
This study investigates learning dynamics in large language model (LLM) fine-tuning, focusing on three critical phenomena: (1) cross-task factual transfer and response repetition hallucinations in instruction tuning; (2) the “squeezing effect”—an anomalous decline in target response probability—under prolonged preference optimization; and (3) the superiority of on-policy over off-policy direct preference optimization (DPO). We propose an influence-function-based stepwise attribution method and construct a response-level influence-path model, providing the first unified explanation for these phenomena. We formally identify and characterize the squeezing effect as arising from excessive contraction of the policy distribution, thereby revealing the intrinsic advantage of on-policy training. Leveraging this insight, we design a novel alignment-optimization strategy. Experiments demonstrate that our framework significantly improves fine-tuning stability and alignment performance, offering both theoretical foundations and practical solutions for efficient, controllable LLM fine-tuning.
Behavioral cloning in robotic manipulation is prone to error accumulation and distribution shift, hindering its robust deployment in industrial settings. This work proposes a plug-and-play residual reinforcement learning fine-tuning framework that learns a residual policy on top of vision-language-action (VLA) model outputs, augmented with a human-in-the-loop mechanism to enable online correction of suboptimal actions and safe, efficient exploration. The approach introduces the first model-agnostic residual fine-tuning architecture, compatible with diverse VLA models, and achieves an average task success rate exceeding 95% after only 1.5 hours of real-robot online training. This significantly enhances the real-world deployability of behavioral cloning policies while maintaining compatibility with existing foundation models.
This study addresses the difficulty of pretrained robot policies in adapting to unknown physical environments due to their reliance on sparse rewards and limited use of geometric dynamic feedback. To overcome this, we propose SCOUT, a framework that couples action prediction with forward dynamics models to construct a shared belief latent space. By leveraging dynamics-aware meta-learning, it establishes a bilevel architecture comprising an inner loop for belief updating and an outer loop for action optimization. The core innovation lies in utilizing action-outcome feedback to update internal dynamics beliefs online, enabling rapid and robust adaptation without sparse rewards while avoiding catastrophic forgetting. Experiments demonstrate that SCOUT significantly accelerates online adaptation in simulated manipulation benchmarks and successfully validates robust sim-to-real transfer capabilities.
研究探讨了在微调大型行为模型时,非高斯先验是否优于标准高斯先验。结果表明,除了极少量数据外,非高斯先验并未带来性能提升。
本文提出ActObs方法,通过同时监督动作和观察令牌来改进强化学习初始化,从而提高代理在不同任务上的性能和探索能力。
This study addresses the problem of neural trajectory predictors violating physical constraints when control inputs are unobserved, and proposes MaDE, a post-processing operator that projects state transitions onto a feasible manifold. By inferring unknown controls and rectifying inequality constraints, MaDE ensures dynamical consistency. Its key innovations include training without ground-truth control signals, guaranteeing iteratively optimal feasibility via gradient-based correction, and serving as a plug-and-play module for arbitrary prediction models. Experiments demonstrate that MaDE drives dynamical residuals to near zero in simulation and reduces them to 0.0071 on real-world vehicle data, outperforming baselines. This strict physical compliance is achieved at the cost of only an approximately 1.8-fold increase in displacement error.