DaCe-DT: Data-Centric Offline Multi-Task Reinforcement Learning via Adaptive Prompts and Trajectory Correction for Heterogeneous Tasks

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance bottlenecks in offline multi-task reinforcement learning caused by heterogeneous data quality, verbose prompts, and fragmented trajectories. To overcome these challenges, we propose DaCe-DT, a framework that reconstructs the Decision Transformer from a data-centric perspective. Specifically, this method introduces three key innovations: a length-gated prompt masking mechanism, retrieval-augmented prompt construction, and value-adaptive return calibration, collectively enabling robust policy learning across heterogeneous tasks. Extensive evaluations on the Meta-World benchmark demonstrate that DaCe-DT significantly outperforms existing state-of-the-art methods, achieving performance improvements of 11.73% and 13.34% on optimal and suboptimal datasets, respectively.
📝 Abstract
Offline multi-task reinforcement learning (Offline MTRL) heavily depends on the quality and distribution of pre-collected data. However, existing methods mainly focus on algorithmic optimization, with less emphasis on data-level improvements to enhance learning ability and generalization performance. This paper, from a data perspective, reveals three key bottlenecks that limit Offline MTRL performance:(i) ineffective utilization of prompts length under diverse task complexities, and (ii) semantic irrelevance of randomly sampled prompt segments, (iii) misleading supervision induced by fragmented and discontinuous trajectories. To address these challenges, we propose DaCe-DT, a robust offline MTRL framework designed to be insensitive to heterogeneous task complexities and data quality, featuring length-gated prompt masking (LGPM), retrieval-augmented prompt construction (RAPC), and value-adaptive return calibration (VARC). Together, these mechanisms enable DaCe-DT to deliver data-centric prompt adaptation and trajectory refinement, resulting in robust multi-task generalization and stable policy learning amid heterogeneous offline data and tasks. Experimental results on Meta-World show that DaCe-DT consistently outperforms state-of-the-art methods, achieving an average improvement of 11.73% on optimal datasets and an improvement of 13.34% on suboptimal datasets, demonstrating its effectiveness in learning stably from imperfect data and improving overall multi-task performance.
Problem

Research questions and friction points this paper is trying to address.

Offline Multi-Task Reinforcement Learning
Data-Centric
Heterogeneous Tasks
Prompt Utilization
Trajectory Fragmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Offline Multi-Task Reinforcement Learning
Data-Centric
Prompt Adaptation
Trajectory Correction
Retrieval-Augmented