Score
Designs and implements methods to adapt or fine‑tune existing control algorithms or pretrained controllers to new tasks or environments, including tuning low‑level parameters (e.g., PD gains) and updating codebooks or model weights to improve task‑specific control performance quickly. Builds training and adaptation procedures and evaluation analyses that measure transfer effectiveness, stability, speed of adaptation, and robustness when reusing pretrained behaviors.
This work systematically investigates the impact of position controller gains—traditionally selected based on task stiffness or compliance—on three learning paradigms: behavioral cloning, reinforcement learning from scratch, and sim-to-real transfer. Challenging conventional gain-tuning principles, the study advocates for a learnability-oriented gain selection strategy. Through extensive experiments across multiple tasks and robotic platforms, it reveals that behavioral cloning performs best with compliant, overdamped gains; reinforcement learning exhibits robustness to gain variations when hyperparameters are properly tuned; and stiff, overdamped gains significantly degrade sim-to-real transfer performance. These findings demonstrate that the optimal controller gain is dictated primarily by the learning paradigm rather than the intrinsic characteristics of the task itself, thereby overturning established practices in robotic control design.
Conventional low-level controllers for quadcopters require precise dynamical modeling and extensive parameter tuning, limiting generalization across platforms with substantial differences in mass, size, and actuator capabilities. Method: We propose a model-free, parameter-free learning-based low-level controller that jointly leverages imitation learning and deep reinforcement learning. It implicitly identifies system parameters online from sensor-action histories, enabling real-time adaptive control without explicit system identification. Contribution/Results: We introduce the first end-to-end latent-state system identification framework, achieving unprecedented dynamical generalization: in simulation, it adapts to unseen parameter combinations spanning 16× the training range; on physical hardware, it robustly handles 3.7× mass variation and >100× differences in propeller constants, while tolerating severe disturbances including payload asymmetry and single-motor failure. The controller has been successfully deployed on real drones, significantly enhancing the universality and engineering practicality of low-level flight control.
Traditional controllers struggle to generalize across systems with varying orders and dynamic characteristics. This work proposes a universal learning-based controller that constructs a dynamic state-space representation using a masked attention mechanism, integrating system label encoding, multi-scale temporal processing, and a mixture-of-experts architecture to enable a single neural network to uniformly control diverse linear and nonlinear systems. Notably, the approach is the first to adapt—without architectural modifications—to challenging dynamics such as unstable and non-minimum-phase systems, while supporting zero-shot generalization to unseen operating conditions. Trained on 25 system classes and 314,630 trajectories, the controller matches the performance of specialized LQI controllers and maintains robustness under previously unobserved conditions, including actuator saturation, noise, and disturbances.
Traditional robot design and control are typically decoupled, leading to morphologies poorly aligned with task requirements. This paper proposes a simulation-driven co-optimization framework for morphology and control, breaking the conventional “design-then-control” paradigm to enable task-oriented, end-to-end joint search. Our method employs gradient-free optimization to simultaneously evolve structural parameters and controller policies within a URDF-based multi-task reinforcement learning simulation environment. Key contributions include: (1) demonstrating that controller retraining significantly improves performance, yielding an average gain of 37%; and (2) revealing an inverse correlation between morphological complexity and controller training budget—providing theoretical justification for structural simplification under resource constraints. We validate the framework across four public simulation benchmarks, showing that co-optimization consistently yields more compact, robust, and task-adapted robot morphologies compared to sequential approaches.
Commercial recommendation systems must dynamically adapt to shifting task objectives (e.g., accuracy–diversity trade-offs), yet retraining models incurs prohibitive computational overhead, hindering online real-time adaptation. To address this, we propose a parameter-level, training-free, real-time controllable adaptation method. Our approach is the first to construct a controllable diffusion model directly in the parameter space and leverage classifier-free guidance to generate task-specific model parameters zero-shot and on-demand. It enables test-time adaptive inference and model-agnostic deployment. Evaluated on multiple public and real-world industrial datasets, our method significantly enhances multi-objective controllability while maintaining state-of-the-art recommendation performance. Parameter generation takes under one second—over 300× faster than full retraining—demonstrating unprecedented efficiency for dynamic recommendation adaptation.
This work addresses the long-standing isolation among research domains such as alignment training, model organisms, and toy models, which has hindered empirical cross-pollination and led to redundant exploration and inefficiency. For the first time, it systematically transfers supervised fine-tuning (SFT) practices across these domains by integrating cross-model output training, mixed-strategy data, and benign fine-tuning to rigorously evaluate the portability of key findings. The study demonstrates three successful transfer effects: enhanced behavioral generalization, mitigation of capability degradation, and the critical insight that preserving capabilities alone is insufficient to ensure robustness in subsequent training phases. These results underscore both the efficacy and limitations of reusing methodologies across domains, thereby fostering more synergistic development across disparate research areas.
This work addresses the fragmented landscape of post-training adaptation techniques, which suffer from inconsistent terminology and a lack of unified comparative or governance frameworks. To resolve this, the paper introduces the first six-dimensional taxonomy—spanning mechanism, objective, data requirements, persistence, structural scope, and model type—that systematically integrates mainstream approaches such as fine-tuning, retrieval augmentation, prompt engineering, model editing, and machine unlearning. This framework clarifies conceptual boundaries and reveals evolutionary and compositional relationships among methods. Beyond standardizing terminology, it enables standardized technical documentation, model change tracking, and AI governance analysis. The study further identifies critical challenges, including evaluation rigor, reproducibility, continual adaptation, multimodal alignment, and governance-aware workflows.
This work addresses the challenge of reference trajectory tracking for uncertain nonlinear systems with limited data by proposing a meta-learning-based control framework. The approach learns a shared dynamic representation from structurally similar source systems during an offline phase and enables rapid adaptation of the controller to a new target system using only a few online data samples. Innovatively adapting implicit Model-Agnostic Meta-Learning (iMAML) to the control domain, the method establishes a general bilevel optimization framework compatible with diverse learning algorithms while significantly reducing memory overhead and approximation error. Two implementation pathways—neural state-space models and deep Q-networks, corresponding respectively to explicit and implicit system identification—are evaluated through simulations and hardware experiments, consistently demonstrating superior control performance over baseline methods and confirming the framework’s effectiveness and practicality.
This work addresses the critical challenge in continual learning of mitigating catastrophic forgetting during downstream fine-tuning while preserving capabilities acquired during upstream training. The authors propose treating “robustness to subsequent fine-tuning” as a first-class objective in upstream training and systematically investigate data scheduling strategies across a three-stage pipeline—pretraining, post-training, and downstream fine-tuning. Their key finding is that early exposure to post-training data during pretraining—termed “early data exposure”—consistently outperforms pure post-training or conventional mixing strategies, yielding superior trade-offs between upstream knowledge retention and downstream task performance across model scales from 135M to 1B parameters. This approach complements regularization techniques such as replay and Dropout and, under fixed compute budgets, reveals an optimal data allocation scheme.