inverse dynamics finetuning

Designs and implements fine-tuning procedures for inverse-dynamics models that map observed state transitions to actions, adapting pretrained inverse-dynamics models to target dynamics using supervised paired observation–action data. Works with small rollout datasets (minutes of interaction) to align model-predicted actions to the target dynamics and to evaluate finetuning performance and stability.

inversedynamicsfinetuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the performance degradation of behavioral cloning under limited expert demonstrations and investigates the superior sample efficiency of Predictive Inverse Dynamics Models (PIDMs), whose underlying mechanism has remained unclear. We theoretically analyze PIDMs through the bias–variance tradeoff, revealing that while incorporating future state prediction introduces bias, it substantially reduces the variance in inverse dynamics estimation, thereby enhancing sample efficiency. By integrating a state predictor with an inverse dynamics model and leveraging additional data for both theoretical analysis and empirical validation, we demonstrate that PIDMs achieve comparable performance to behavioral cloning using only one-third of the demonstration data in 2D navigation tasks. In high-dimensional 3D environments with visual inputs, PIDMs reduce the required demonstration data by over 66%.

behavior cloningbias-variance tradeoffimitation learning

IN-RIL: Interleaved Reinforcement and Imitation Learning for Policy Fine-Tuning

May 15, 2025
DG
Dechen Gao
🏛️ University of California, Davis | Toyota InfoTech Labs

In robot policy fine-tuning, conventional pipelines—imitation learning (IL) pretraining followed by reinforcement learning (RL) fine-tuning—suffer from unstable exploration, low sample efficiency, and performance collapse. This paper proposes an alternating IL/RL fine-tuning paradigm. A gradient orthogonal separation mechanism decouples imitation and reinforcement objectives, with theoretical analysis revealing its intrinsic benefits for stability and sample efficiency. We further design a gradient subspace decomposition module that enables plug-and-play integration of behavior cloning with standard RL algorithms (e.g., PPO, SAC). Evaluated on 14 robotic manipulation and locomotion tasks, our method achieves an average 3.8× improvement in sample efficiency (up to 6.3×) and boosts task success rate from 12% to 88% (+76 percentage points), while robustly suppressing performance collapse under both sparse/dense and short-/long-horizon reward settings.

Combines IL and RL for stable policy fine-tuningImproves sample efficiency in robot learning tasksPrevents gradient interference in IL-RL optimization

This work proposes a semi-supervised imitation learning approach based on an inverse dynamics model (IDM) in settings with limited action-labeled trajectories and abundant unlabeled trajectories. The IDM predicts actions from state transitions and can either serve directly as a policy (VM-IDM) or generate pseudo-labels for unlabeled data (IDM labeling). Theoretical analysis reveals that IDM outperforms behavioral cloning primarily due to its lower hypothesis class complexity and reduced stochasticity, leading to higher sample efficiency. Building on this insight, the authors enhance the LAPO algorithm and validate the proposed method within a unified video-action prediction (UVA) framework. Both theoretical analysis and empirical results consistently demonstrate the superior sample efficiency of the IDM-based approach.

behavior cloninginverse dynamics modelpolicy learning

Open-loop imitation learning has advanced modern autonomous driving policy architectures, but closed-loop deployment remains vulnerable to policy-induced distribution shift. Existing post-training paradigms exhibit fundamental trade-offs: closed-loop RL fine-tuning provides grounded feedback from executed actions but is constrained by the sparsity of informative events, whereas counterfactual fine-tuning provides dense supervision over candidate futures but inherits bias from imperfect future estimates. We introduce Counterfactual-to-Interactive Reinforcement Fine-Tuning (CRAFT), an on-policy framework that formulates closed-loop post-training as proxy-residual optimization. CRAFT uses group-normalized counterfactual advantages as a dense proxy for real closed-loop advantages and aligns this proxy with the closed-loop world through grounded residual correction from interaction-critical events. To stabilize adaptation, CRAFT regularizes the online policy toward an EMA teacher via asymmetric KL self-distillation. Theoretically, CRAFT decomposes the real closed-loop policy gradient into proxy and residual terms under the same visited-state distribution, reducing residual variance with an aligned proxy while mitigating proxy bias through grounded residual approximation. Empirically, CRAFT achieves the strongest closed-loop gains on Bench2Drive across hierarchical planning, vision-language-action, and vocabulary-scoring architectures. Ablations, scaling behavior, stability analyses, and transfer results further validate the complementary roles of dense counterfactual proxy and grounded residual correction. Project page: https://currychen77.github.io/CRAFT.

autonomous drivingclosed-loop reinforcement learningcounterfactual reasoning

Learning Dynamics of LLM Finetuning

Jul 15, 2024
YR
Yi Ren
🏛️ University of British Columbia | Amii

This study investigates learning dynamics in large language model (LLM) fine-tuning, focusing on three critical phenomena: (1) cross-task factual transfer and response repetition hallucinations in instruction tuning; (2) the “squeezing effect”—an anomalous decline in target response probability—under prolonged preference optimization; and (3) the superiority of on-policy over off-policy direct preference optimization (DPO). We propose an influence-function-based stepwise attribution method and construct a response-level influence-path model, providing the first unified explanation for these phenomena. We formally identify and characterize the squeezing effect as arising from excessive contraction of the policy distribution, thereby revealing the intrinsic advantage of on-policy training. Leveraging this insight, we design a novel alignment-optimization strategy. Experiments demonstrate that our framework significantly improves fine-tuning stability and alignment performance, offering both theoretical foundations and practical solutions for efficient, controllable LLM fine-tuning.

Analyzes LLM learning dynamicsDescribes DPO's squeezing effectExplains finetuning-induced hallucinations

Latest Papers

What's happening recently
View more

Behavioral cloning in robotic manipulation is prone to error accumulation and distribution shift, hindering its robust deployment in industrial settings. This work proposes a plug-and-play residual reinforcement learning fine-tuning framework that learns a residual policy on top of vision-language-action (VLA) model outputs, augmented with a human-in-the-loop mechanism to enable online correction of suboptimal actions and safe, efficient exploration. The approach introduces the first model-agnostic residual fine-tuning architecture, compatible with diverse VLA models, and achieves an average task success rate exceeding 95% after only 1.5 hours of real-robot online training. This significantly enhances the real-world deployability of behavioral cloning policies while maintaining compatibility with existing foundation models.

behavior cloningcompounding errorsdistributional shift

This study addresses the difficulty of pretrained robot policies in adapting to unknown physical environments due to their reliance on sparse rewards and limited use of geometric dynamic feedback. To overcome this, we propose SCOUT, a framework that couples action prediction with forward dynamics models to construct a shared belief latent space. By leveraging dynamics-aware meta-learning, it establishes a bilevel architecture comprising an inner loop for belief updating and an outer loop for action optimization. The core innovation lies in utilizing action-outcome feedback to update internal dynamics beliefs online, enabling rapid and robust adaptation without sparse rewards while avoiding catastrophic forgetting. Experiments demonstrate that SCOUT significantly accelerates online adaptation in simulated manipulation benchmarks and successfully validates robust sim-to-real transfer capabilities.

action-outcome feedbackdynamics gaprobotic manipulation

This study addresses the problem of neural trajectory predictors violating physical constraints when control inputs are unobserved, and proposes MaDE, a post-processing operator that projects state transitions onto a feasible manifold. By inferring unknown controls and rectifying inequality constraints, MaDE ensures dynamical consistency. Its key innovations include training without ground-truth control signals, guaranteeing iteratively optimal feasibility via gradient-based correction, and serving as a plug-and-play module for arbitrary prediction models. Experiments demonstrate that MaDE drives dynamical residuals to near zero in simulation and reduces them to 0.0071 on real-world vehicle data, outperforming baselines. This strict physical compliance is achieved at the cost of only an approximately 1.8-fold increase in displacement error.

dynamics constraintsfeasibilityneural predictors

Hot Scholars

DZ

Danping Zou

Professor, Shanghai Jiao Tong University
Visual SLAMRobotic VisionVision-based navigation
SO

Shayegan Omidshafiei

Chief Scientist at FieldAI (previously: Google DeepMind, Google Research, MIT)
Artificial IntelligenceRoboticsMachine LearningReinforcement Learning
DF

David Fridovich-Keil

Assistant Professor, The University of Texas at Austin
optimal controldynamic gamesmotion planningrobotic safety
TW

Tyler Westenbroek

EECS PhD Student, UC Berkeley
ControlOptimizationGame TheoryHybrid Systems
FY

Feng Yu

University of Exeter
Efficient AIContinual LearningFederated LearningFoundation Model