Score
Design, build, and analyze decision-making policies from logged interaction data by estimating counterfactual policy value, constructing reliable offline targets, and performing off-policy evaluation, optimization, and reinforcement learning, including distillation of learned policies into deployable agents. Also create and manage on-policy rollouts and training workflows, perform online fine-tuning and policy iteration, and design policy architectures, integration, and deployment pipelines to transition learned policies from offline training into operational use.
Offline reinforcement learning (RL) suffers from a gap between theory and practice. This paper systematically characterizes its fundamental solvability boundary, revealing how function approximation capacity and data coverage assumptions fundamentally govern learning performance. Methodologically, we integrate tools from approximation theory and distribution matching analysis, complemented by explicit counterexample construction and derivation of sufficient conditions, to rigorously establish necessary and sufficient theoretical conditions for successful offline RL. These abstract conditions are then concretized into actionable constraints on algorithmic generalization capability and empirical data quality. Our results not only explain the failure mechanisms of existing methods under low-coverage regimes but also yield design principles that balance theoretical rigor with practical feasibility. The work provides an operational theoretical framework to guide algorithmic innovation for realistic challenges—including high generalization demands and severely limited data coverage—thereby bridging foundational analysis and scalable offline RL deployment.
In offline-to-online reinforcement learning (O2O-RL), policies are first safely trained offline using previously collected datasets and then further fine-tuned for tasks via limited online interactions. In a typical O2O-RL pipeline, candidate policies trained with offline RL are evaluated via either off-policy evaluation (OPE) or online evaluation (OE). The policy with the highest estimated value is then deployed and continually fine-tuned. However, this setup has two main issues. First, OPE can be unreliable, making it risky to deploy a policy based solely on those estimates, whereas OE may identify a viable policy with substantial online interaction, which could have been used for fine-tuning. Second--and more importantly--it is also often not possible to determine a priori whether a pretrained policy will improve with post-deployment fine-tuning, especially in non-stationary environments. As a result, procedures committing to a single deployed policy are impractical in many real-world settings. Moreover, a naive remedy that exhaustively fine-tunes all candidates would violate interaction budget constraints and is likewise infeasible. In this paper, we propose a novel adaptive approach for policy selection and fine-tuning under online interaction budgets in O2O-RL. Following the standard pipeline, we first train a set of candidate policies with different offline RL algorithms and hyperparameters; we then perform OPE to obtain initial performance estimates. We next adaptively select and fine-tune the policies based on their predicted performance via an upper-confidence-bound approach thereby making efficient use of online interactions. We demonstrate that our approach improves upon O2O-RL baselines with various benchmarks.
This work addresses a key challenge in offline-to-online reinforcement learning: how to efficiently select and fine-tune policies under a limited online interaction budget while avoiding performance degradation due to algorithmic or hyperparameter sensitivity. The paper introduces the first active policy selection framework tailored for such constrained settings. By constructing an upper confidence bound based on a local linear performance prediction model, the method dynamically balances resource allocation between online evaluation and fine-tuning, adaptively identifying the most promising policies for optimization. This approach overcomes the limitations of conventional strategies that either deploy a single policy or uniformly distribute the budget across candidates. Empirical results across multiple environments demonstrate that the proposed method significantly outperforms existing baselines, achieving more efficient utilization of scarce online interaction resources.
Offline pretraining often suffers from rapid degradation and poor exploration during early online reinforcement learning. To address this, we propose a policy expansion mechanism that treats the frozen offline policy as a fixed behavioral prior, dynamically coordinating it with a learnable online policy. Our key contribution is the first adaptive dual-policy architecture, integrating a policy-ensemble-based gating mechanism with behavioral distribution matching constraints. This ensures the offline policy remains unupdated while continuously guiding exploration, while enabling the online policy to incrementally acquire novel behaviors. Evaluated on multiple continuous-control benchmarks, our method significantly improves sample efficiency and final performance, avoids initial performance collapse, and achieves more stable convergence—outperforming standard fine-tuning and policy distillation baselines across all metrics.
Offline reinforcement learning (RL) faces two critical challenges in safety-critical medical settings: (1) policy uninterpretability undermines clinical trust, and (2) offline policy evaluation—particularly importance sampling—is highly sensitive to behavior policy mismatch. To address these, we propose an interpretable tree-structured behavioral cloning framework that automatically clusters patient states to generate intuitive, traceable decision rules. We incorporate a high-frequency action prior to constrain the policy space, enhancing overlap with the behavior policy and improving offline evaluation stability. Furthermore, we jointly estimate the behavior policy via the tree model, integrating behavioral cloning with importance sampling for both robust policy extraction and reliable evaluation. Evaluated on real-world clinical datasets for rheumatoid arthritis and sepsis, our method achieves significant improvements over current clinical practice: +12.3% in cumulative reward, full decision-path interpretability via visualization, and 47% reduction in evaluation variance—demonstrating superior performance, transparency, and assessment robustness.
Offline reinforcement learning (Offline RL) faces two key challenges: poor policy generalization and distributional shift—existing methods tend to overestimate out-of-distribution actions and require hyperparameter re-tuning for cross-task or cross-dataset transfer. To address this, we propose a dynamic policy switching mechanism at evaluation time, which adaptively fuses an offline RL policy with a behavior cloning policy based on dual uncertainty estimates: epistemic uncertainty (quantified via Monte Carlo Dropout) and aleatoric uncertainty (estimated via data density modeling). Our approach enables zero-shot cross-task transfer without parameter tuning and supports safe, zero-sample fine-tuning from offline to online settings. Evaluated on standard Offline RL benchmarks, it consistently outperforms state-of-the-art methods, achieving faster fine-tuning convergence and superior final performance.
通过分析160,000次训练运行,研究了离线策略学习中的算法性能、超参数敏感性及环境适应性问题,提出了基于数据集的推荐系统以提高研究可靠性。
LLM-based agents are increasingly deployed in real-world applications through tool-use APIs, yet training them for specific environments remains fundamentally difficult: real-world applications provide no pre-defined tasks or verifiers, no faithful simulators, and limited budget for large-scale environment interaction. In this paper, we propose \textbf{AgentBrew}, an offline training framework that learns effective tool-use policies from a single batch of raw interaction trajectories, without task verifiers or iterative on-policy rollouts. The agent first explores the target environment to collect a raw trajectory corpus without quality filtering. To extract training signal from this noisy corpus, \emph{retrospective task inference} reconstructs an aligned instruction for each trajectory based on its actual outcome, and \emph{PMI-Based credit assignment} decomposes the trajectory's total information about the inferred instruction into additive per-action credits via pointwise mutual information (PMI). These credits weight the policy training objective, amplifying informative actions while suppressing ineffective ones. On three real-world MCP applications (GitHub, Notion, PostgreSQL), AgentBrew improves Qwen3-32B by +8.7 Acc / +9.7 Score on average, surpassing Qwen3-235B (+2.3 / +4.4) and outperforming rejection sampling (+5.9 / +10.3). These results demonstrate that fine-grained offline learning can recover useful supervision from raw trajectories that filtering-based approaches would discard. The code is available at https://github.com/alphatogo/AgentBrew
This work addresses the challenges of policy learning in contextual bandits with extremely large action spaces, where inefficient exploration, high variance of importance weights, and optimization difficulties commonly arise. To improve exploration efficiency in online settings, the authors propose two approaches—mixed-effects Thompson Sampling (meTS) and diffusion Thompson Sampling (dTS)—that explicitly model dependencies among actions. For offline settings, they introduce a latent-variable-based method, sDM, which integrates a differentiable pessimism mechanism with a concave policy-weighted log-likelihood objective to mitigate extrapolation bias and variance issues. Theoretical analysis yields regret bounds that scale with the effective number of actions, and empirical results demonstrate that the proposed methods significantly enhance both stability and performance of policy learning in large action spaces.
This study addresses the challenges faced by reinforcement learning (RL) agents in early-stage autonomous cyber defense, including high exploration costs, weak decision-making capabilities, and behavioral instability. To overcome these limitations, the work proposes an online policy distillation framework that leverages a prompt-engineered large language model (LLM) specialized in cybersecurity as a teacher policy. The framework efficiently transfers knowledge from the LLM to a lightweight RL agent containing only 64,910 parameters. Evaluated in multi-scale CybORG network environments with 4 to 12 hosts, the distilled agent closely replicates the teacher’s performance and significantly outperforms baseline RL methods. Despite a five-order-of-magnitude reduction in parameter count, the agent maintains robust defensive capabilities, demonstrating the feasibility of deploying state-of-the-art security models in resource-constrained settings.
本文提出Brain API,一种意图感知控制平面,通过决策工件解决现有系统中缺乏意图级决策治理的问题。