rl-based data generation

Design and build systems that use reinforcement learning to produce synthetic datasets of sequential behaviors, trajectories, and action–experience episodes by training policies or goal-conditioned agents to explore, diversify, and generate labeled episodes. These systems output executable action sequences, demonstrations, or replay buffers intended for downstream training, evaluation, or data augmentation.

rl-baseddatageneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of interpretability in AI strategies for complex games by proposing a novel method for generating executable programmatic policies. Inspired by human cognitive mechanisms, the approach integrates reinforcement learning with program synthesis to transform implicitly learned agent policies into executable programs based on action sequences. Experimental evaluations conducted in chess and grid-world environments demonstrate that the proposed method effectively extracts and generates efficient action sequences directly from data. Consequently, this approach significantly enhances the transparency and interpretability of AI decision-making processes. By bridging the gap between opaque policy representations and human-readable programmatic logic, this work establishes a new paradigm for game data analysis and explainable artificial intelligence.

Explainable RepresentationsGame-playing StrategiesReinforcement Learning

Performance Comparisons of Reinforcement Learning Algorithms for Sequential Experimental Design

Mar 07, 2025
YZ
Yasir Zubayr Barlas
🏛️ The University of Manchester | City St George's, University of London

This work systematically investigates the generalization capabilities of reinforcement learning (RL) in sequential experimental design. Addressing robustness bottlenecks under model misspecification and distributional shift, we propose a unified RL framework integrating dropout regularization and ensemble strategies across prominent algorithms—including PPO, SAC, and DQN. To our knowledge, this is the first systematic cross-task and parameter-perturbation generalization evaluation of diverse RL methods in this domain. Empirical results demonstrate that ensemble SAC and dropout-regularized PPO substantially enhance robustness: information gain improves by 18–32% over standard baselines under model mismatch and prior shift. Our core contribution lies in rigorously establishing the critical role of regularization and ensembling in improving generalization for RL-driven experimental design, while providing a reproducible, high-performance paradigm for constructing robust policy agents.

Assess generalization of policies across varying statistical properties.Compare effectiveness of dropout and ensemble methods in training.Evaluate reinforcement learning algorithms for sequential experimental design.

This work addresses the challenges in post-training multi-turn interactive tool-using agents, which are hindered by the difficulty of scaling high-quality synthetic data and the inefficiency caused by noisy user simulations in reinforcement learning. The authors propose a unified framework that integrates self-evolving synthetic data generation with validator-based reinforcement learning. Specifically, a multi-agent system produces tool-augmented dialogues equipped with executable verifiers, and a closed-loop self-evolution mechanism enhances data reliability. Training proceeds through staged trajectory-level group relative policy optimization (GRPO-style) with a novel verifiable reward mechanism and dynamic filtering strategy, enabling efficient, annotation-free learning. Evaluated on the tau²-bench, the model achieves 73.0% and 98.3% pass¹ rates on the Airline and Telecom tasks, respectively, matching or surpassing current state-of-the-art methods.

multi-turn interactionpost-trainingreinforcement learning

Data Augmentation for Instruction Following Policies via Trajectory Segmentation

Feb 25, 2025
NH
Niklas Hopner
🏛️ University of Amsterdam | Vrije Universiteit Amsterdam

To address the limited generalization capability of instruction-following agents caused by scarce annotated trajectory data, this paper proposes Play Segmentation (PS), a probabilistic model that automatically discovers high-quality, instruction-aligned trajectory segments from large-scale unlabeled gameplay or simulation traces. PS performs fine-grained semantic segmentation of long trajectories without assuming fixed-length segments and requires only a small number of short instruction examples. It integrates probabilistic graphical modeling, trajectory–instruction alignment learning, and a semi-supervised training framework, seamlessly embedding into imitation learning pipelines. Experiments on both game-playing and robotic manipulation tasks demonstrate that policies enhanced with PS achieve performance comparable to baselines trained on twice the volume of human-annotated data; in contrast, random sampling degrades performance significantly. These results validate PS’s effectiveness and practicality for scalable, low-supervision instruction grounding.

Extract labelled segments from unannotated play trajectories.Improve instruction-following policy via augmented dataset.Limited data pairs instructions with agent trajectories.

This work addresses the limitations of traditional reinforcement learning, which relies on sparse and opaque reward signals from game engines and requires extensive trial-and-error interaction with the environment. The authors propose a novel approach that leverages vision-language models (VLMs) to automatically annotate human-interpretable reward signals from gameplay video datasets. Using these semantically meaningful rewards, a conditional agent is trained via offline reinforcement learning to execute behaviors aligned with high-level instructions. This method represents the first application of VLMs to generate interpretable rewards, eliminating dependence on environment-provided rewards and dense online interaction. The approach enhances policy interpretability while simplifying the training pipeline. Empirical results demonstrate the feasibility of this paradigm and highlight current challenges and limitations.

offline RLReinforcement Learningreward engineering

Latest Papers

What's happening recently
View more

This work proposes an unsupervised learning approach grounded in bounded rationality, wherein agents develop internally consistent behavioral policies in the absence of explicit reward signals through a “cross-objective” mechanism that couples action prediction with outcome prediction. By integrating random walks and graph structural information, the method encourages agents to autonomously congregate within homogeneous regions during node classification tasks. Experimental results demonstrate that agents exhibit unexpected yet logically coherent behavioral patterns, successfully accomplishing classification while simultaneously revealing the critical role of observer interpretation in the emergence of such behaviors. These findings offer a novel perspective on policy self-organization in unsupervised reinforcement learning.

behavior strategiesbounded rationalityreinforcement learning

This work addresses the challenge of explainable reinforcement learning in resource-constrained environments by proposing an experience-based learning model grounded in state-transition graphs. The model explicitly constructs a graph encoding both utility values and evidence counts, integrating a global feedback mechanism to enable transparent policy modeling and interpretable decision-making. By embedding explainability directly into the learning process, the approach maintains low computational overhead while achieving performance on the OpenAI Gym Atari Breakout benchmark comparable to that of certain neural network–based methods, thereby demonstrating its effectiveness and practicality in settings with limited computational resources.

global feedbackinterpretable learningreinforcement learning

To address the challenge in continual reinforcement learning (CRL) of balancing knowledge stability and policy plasticity in dynamic environments, this paper proposes a novel CRL framework leveraging an external self-evolving demonstration library. Methodologically, it explicitly encodes prior knowledge as executable demonstrations and organizes them into a retrievable, autonomously updated external memory bank to directly guide agent exploration. Additionally, it introduces a curriculum-based progressive strategy that enables smooth transition from demonstration-guided behavior to autonomous policy optimization. The framework is seamlessly integrated with mainstream algorithms such as PPO and SAC. Empirical evaluation on 2D navigation and MuJoCo locomotion benchmarks demonstrates substantial improvements: +21.4% average performance gain, enhanced cross-task transferability, superior anti-catastrophic forgetting capability, and 37% higher training efficiency.

Addresses continual RL in dynamic environmentsBalances stability and plasticity via demonstrationsEnhances knowledge transfer and reduces forgetting

This work systematically investigates the underutilized potential of experience replay in reinforcement learning-based post-training of large language models, challenging the prevailing assumption that fresh online data generation is indispensable. By carefully balancing data staleness, sample diversity, and computational cost, the authors design an efficient replay buffer mechanism that effectively substitutes strict online sampling. Their approach demonstrates, for the first time, that experience replay can significantly reduce inference-time computational overhead while maintaining or even improving model performance and effectively preserving policy entropy. This finding offers a compelling alternative to costly online data collection, suggesting that strategic reuse of historical interactions can sustain training efficacy without compromising behavioral diversity or learning stability.

computational costExperience ReplayLLM post-training

This work addresses the challenge of learning robotic control policies from suboptimal, noisy, or imperfect demonstrations by proposing a trajectory refinement method based on Temporal Behavior Trees (TBTs). It introduces, for the first time, the use of TBTs to automatically correct demonstration trajectories that violate task specifications, thereby generating a logically consistent and interpretable dataset. Leveraging this refined data, the approach constructs a potential function to shape reward signals for reinforcement learning, enabling task-consistent policy learning without requiring an explicit model of system dynamics. Experimental results in grid-based navigation and continuous single- and multi-agent obstacle avoidance tasks demonstrate that the proposed method significantly improves both data efficiency and policy performance.

imitation learningimperfect demonstrationsreinforcement learning

Hot Scholars

ST

Sebastian Trimpe

Professor, RWTH Aachen University
ControlMachine LearningNetworked SystemsRobotics
MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing
HZ

Hao Zhou

Bytedance
Computer VisionMultimodal AIVideo UnderstandingSign Language Processing