Score
Designs and implements online adaptation modules that use latent representations and action priors to update models or policies in real time; this includes building lightweight actor‑critic or Bayesian adapters, selective or sparse latent update mechanisms, and 'wa' adapter tuning to quickly correct calibration, perception, or contact errors during deployment.
This work investigates how agents can achieve controllable, continual evolution and adaptation with minimal human intervention. We model self-improving agents as operational scaffolds comprising a foundation model integrated with prompts, memory, tools, and control logic, and formalize self-improvement as self-triggered updates to either model parameters or scaffold components. Based on the update objectives and driving signals, we propose the first systematic taxonomy that unifies existing approaches, clarifies application scenarios and evaluation metrics, identifies key open challenges, and establishes a dynamically maintained repository of technical advances in the field.
To address the limited knowledge transfer of pretrained robotic policies during continual adaptation to novel tasks in dynamic home environments, this paper proposes the Online Meta-Learning Adapter (OMLA). OMLA is the first approach to embed meta-learning objectives directly into the online gradient updates of a lightweight, parameter-efficient fine-tuning (PEFT) adapter, enabling implicit cross-task knowledge reuse. Its plug-and-play architecture requires neither task identifiers nor historical data replay, supporting single-pass online adaptation in real-world settings. Experiments on both simulated and physical robot platforms demonstrate that OMLA achieves an average 23.6% improvement in task adaptation success rate over state-of-the-art baselines, while significantly accelerating convergence and enhancing final performance. This work establishes an efficient, scalable paradigm for continual autonomous learning in domestic service robotics.
To address low online adaptation efficiency and catastrophic forgetting in reinforcement learning agents facing abrupt environmental changes during dynamic deployment, this paper proposes an online meta-reinforcement learning framework integrating change-awareness and knowledge preservation. The method introduces: (1) a priority-based exploration sampling mechanism for rapid detection and response to environmental shifts; (2) a decomposable policy representation with parameter-isolated updates to selectively retain knowledge from prior tasks; and (3) a structured knowledge-preserving representation that disentangles shared and task-specific policy modules. Evaluated across diverse simulated environments featuring sudden dynamics shifts and goal drift, the approach achieves a 3.1× improvement in sample efficiency and reduces forgetting by 62% compared to state-of-the-art online adaptation baselines.
To address degraded model adaptability in online multistep time-series forecasting—caused by data distribution drift and delayed ground-truth feedback—this paper proposes ADAPT-Z. Methodologically, ADAPT-Z abandons conventional parameter fine-tuning and instead models the dynamics of latent factors. It introduces an adapter module that fuses current features with historical gradient information within a learned Z-space, enabling persistent tracking and incremental self-adaptation of feature representations. This design mitigates gradient mismatch induced by label delay and enhances robustness to non-stationary data. Empirical evaluation across multiple benchmark datasets demonstrates that ADAPT-Z significantly outperforms static baselines and state-of-the-art online learning methods, achieving superior generalization and sustained adaptive capability under streaming conditions.
Real-world robots must adapt in real time to gradual drift, transient disturbances, and abrupt structural changes in dynamically evolving environments. To address this, we propose an online Bayesian adaptive control framework tailored for nonlinear dynamics. Our method introduces a novel implicit changepoint detection mechanism grounded in data likelihood, decoupling offline representation learning from online closed-form Bayesian updating—enabling millisecond-scale relearning upon abrupt changes and continuous refinement under gradual drift, while preserving uncertainty calibration. By integrating latent-variable inference, adaptive regret analysis, and online probabilistic inference, the framework significantly enhances model robustness and responsiveness. Evaluated on inverted-pendulum simulations and real-world quadrotor experiments—including scenarios with swinging payloads and mid-air payload release—the approach achieves a 32% improvement in prediction accuracy, reduces disturbance recovery time by 47%, and lowers closed-loop trajectory tracking error by 58% relative to baseline methods.
Existing world models rely heavily on large-scale labeled action data and computationally expensive training, hindering rapid adaptation to novel environments with heterogeneous action spaces and scarce annotations. To address this, we propose a self-supervised framework that eliminates the need for explicit action labels: first, video representation learning implicitly extracts action representations from inter-frame dynamics; second, an autoregressive world model is constructed conditioned on these latent actions. This constitutes the first approach to integrate action modeling directly into the world model pretraining stage, enabling action-agnostic universal representation learning. Our method achieves cross-action-space transfer with only minimal environment interaction. Extensive experiments across multiple environments demonstrate substantial improvements in video prediction fidelity and visual planning performance, reduce fine-tuning costs by over 40%, and exhibit strong generalization across diverse action spaces.
This work addresses the limited cost-effectiveness of existing routing methods that merely assign simple tasks to small models without enhancing their capabilities. To overcome this, the authors propose a multi-cycle adaptation mechanism operating at the granularity of single inference calls. The approach leverages a teacher model to generate verification demonstrations from the small model’s failures, integrating skill distillation and LoRA fine-tuning to continuously improve its competence. Joint optimization is performed over a dynamic skill library, task-specific adapters, and a cost-calibrated routing policy, complemented by a verifier-supported fallback mechanism. Experiments show that Qwen2.5-Coder-1.5B achieves a pass rate increase from 28.7% to 49.7% on HumanEval+MBPP; the deployment strategy attains 88.3% of peak performance at only 60.8% of the cost; and Qwen3.5-2B matches the performance of an unadapted 4B model on TAU-2.
This work addresses the challenge of efficiently adapting and aligning large language models to multiple tasks without modifying their pretrained weights. The authors propose LARA, a lightweight adaptation method that freezes the backbone model and injects low-rank correction signals into the residual stream. LARA introduces token-level dynamic routing within the residual stream for the first time, enabling concurrent hosting and on-demand composition of multiple behavioral modules. It further incorporates a tunable interpolation coefficient γ to enable smooth control over behavior blending. With only 33 MB of additional overhead, LARA supports the simultaneous deployment of seven distinct behaviors on a 1.5B-parameter model, achieving performance comparable to LoRA on code fine-tuning and DPO tasks while significantly enhancing deployment flexibility and resource efficiency.