Score
Designs, trains, and evaluates model-free reinforcement-learning agents that learn scheduling policies from interaction rather than from a system model. These policies select which tasks to run and when to run them to balance task execution against resource-preservation (e.g., survival) under unknown or changing resource profiles, providing tunable execution–survival trade-offs without requiring prior energy or system models.
This study addresses the problem of dynamically allocating prediction tasks among capacity-constrained agents—whether human or artificial—to maximize collective performance. It introduces, for the first time, a theoretical formulation of task assignment under explicit capacity constraints and proposes a context-aware sequential exploration–exploitation learning framework. This framework integrates multi-agent capability modeling with optimized task–agent matching strategies. Empirical evaluations demonstrate that the proposed approach significantly outperforms non-contextual baselines across tabular, image, and text prediction tasks, and is effective in collaborative settings involving both large language models and human agents.
To address the low sample efficiency, poor safety guarantees, and limited interpretability of model-free reinforcement learning (RL), this paper proposes a hybrid framework that synergistically integrates model-free and model-based approaches. Specifically, it embeds model predictive control (MPC) into the policy optimization pipeline, leveraging differentiable and interpretable dynamic and constraint models to encode safety priors. It further combines Bayesian optimization–driven policy search with offline RL to mitigate model mismatch. An end-to-end trainable adaptive model jointly captures system dynamics, cost functions, and constraints. Experimental results demonstrate that the method improves sample efficiency by up to 2.3×, significantly enhances decision safety and interpretability, and validates the effectiveness and generalization capability of the “model-guided + data-driven” paradigm in complex, uncertain environments.
Current large language models are constrained by fixed context windows, limiting their ability to handle highly complex tasks. This work proposes a reinforcement learning–based recursive agent training framework that enables agents, during inference, to autonomously decide whether and how to recursively invoke themselves, dynamically decomposing tasks and delegating subtasks. The approach achieves, for the first time, adaptive recursion and coordination among agents at inference time, effectively circumventing context length limitations. It significantly enhances generalization and reasoning efficiency on tasks far exceeding the complexity encountered during training, while maintaining higher training efficiency and achieving lower overall inference latency compared to single-agent systems.
This paper addresses the NP-hard problem of offline, non-preemptive mixed-criticality (MC) real-time scheduling on heterogeneous multicore platforms. We propose the first systematic reinforcement learning (RL)-based solution, modeling scheduling as a Markov decision process and employing PPO and DQN algorithms to jointly maximize overall task completion rate while guaranteeing schedulability of high-criticality tasks. Our approach innovatively overcomes the performance limitations of conventional static analysis and heuristic methods, enabling adaptive handling of dynamic workloads and processor-speed heterogeneity. Evaluated on 100,000 synthetic instances and real-world datasets, our method achieves an average task completion rate of 80% (85% for high-criticality tasks), rising to 94% (93% for high-criticality tasks) under stable scenarios—significantly outperforming state-of-the-art offline MC schedulers.
This study investigates whether purely model-free reinforcement learning (RL) agents can spontaneously develop planning capabilities and elucidates the underlying mechanisms. Method: Focusing on the Sokoban task, we propose a concept-driven interpretability framework integrating Test-Input Concept Activation Vectors (TCAV), representation intervention, causal attribution, and reverse engineering of planning algorithms. Contribution/Results: We provide the first mechanistic evidence for implicit planning in model-free RL. Specifically: (1) The DRC agent autonomously constructs an implicit planning logic resembling parallel bidirectional search—without any explicit planning module; (2) its decisions rely on learned conceptual representations to generate internal plans and perform causal reasoning; (3) planning-relevant representations exhibit a causal relationship with test-time computational resources, and increased computation yields significant gains in planning performance. This work challenges the conventional assumption that “model-free implies no planning,” establishing a verifiable and intervenable planning mechanism within model-free RL.
This study investigates whether the substantial training cost of deep reinforcement learning (DRL) in carbon-aware flow shop scheduling can be justified by long-term performance gains through policy generalization. To this end, the authors propose a framework that integrates DRL with dynamic algorithm configuration, training policies on small, simple instances and transferring them to unseen, complex instances for online parameter adaptation. Experimental results demonstrate that the proposed approach significantly outperforms baseline methods—such as static parameter tuning—on out-of-distribution, complex scenarios. These findings validate the strong generalization capability of DRL policies and confirm that the initial investment in training yields sustained performance benefits in practical applications.
This work proposes a scheduling approach for dynamic job shop scheduling problems under uncertainty caused by stochastic job arrivals and unexpected machine failures. The randomness of job arrivals and machine breakdowns is modeled using Gamma and Weibull distributions, respectively. To ensure policy optimization remains within the feasible action space, two action-masking mechanisms—non-gradient and gradient-based—are integrated with a Maskable Proximal Policy Optimization algorithm. Evaluated on standard dynamic JSSP benchmarks, the proposed method significantly outperforms conventional heuristics and dispatching rules, demonstrating superior performance in minimizing makespan, along with strong robustness and good scalability.
This work proposes a policy-based deep reinforcement learning hyper-heuristic framework to address the challenge of dynamically selecting effective dispatching rules for the job shop scheduling problem (JSSP). The framework enables an agent to adaptively switch among low-level dispatching rules based on the current system state. It introduces two key innovations: an action pre-filtering mechanism to ensure unbiased evaluation of heuristics and a commitment mechanism to regulate the frequency of rule switching. By integrating both deterministic and stochastic policy selection strategies, the approach significantly outperforms conventional heuristics, metaheuristics, and existing neural network–based scheduling methods on standard JSSP benchmarks, demonstrating superior scheduling performance and training stability.
Traditional reinforcement learning reward mechanisms struggle to precisely encode temporal constraints, limiting their applicability in time-sensitive tasks. To address this, we propose the Temporal Reward Machine (TRM), the first framework to embed timed automata semantics directly into reward modeling—enabling programmable reward logic such as delay penalties and prompt incentives under both discrete- and continuous-time semantics. Methodologically, TRM integrates temporal abstraction Q-learning, timed automaton specification, and a counterfactual imagination heuristic to achieve model-agnostic, efficient policy optimization. Experiments demonstrate that TRM significantly improves policy performance under strict temporal constraints on standard RL benchmarks. Ablation studies confirm that the counterfactual heuristic accelerates convergence by 42% and increases constraint satisfaction rate by 31%. Our core contribution is the establishment of the first semantically rigorous, interpretable, and scalable time-aware reward modeling paradigm.
This work addresses the challenge of dynamically coordinating context, prompts, and tool invocations for large language model (LLM) agents under resource constraints in multi-turn interactions. It introduces a Stackelberg game formulation into LLM resource scheduling and proposes a context-aware, repairable conditional policy framework: a controller sets quality–cost objectives, and an executor dynamically allocates resources accordingly. Strategy repair is achieved through learning conditional response models, optimizing the leader’s policy, and integrating real-API calibration with projection onto safe action sets. Theoretical analysis establishes guarantees on equilibrium existence, response stability, and environmental transferability. In 300-round real-API experiments, the repaired policy reduces token cost by 17.4% on average compared to a conservative baseline (p = 0.022) without significant degradation in output quality (p = 0.44).