model-free rl scheduling

Designs, trains, and evaluates model-free reinforcement-learning agents that learn scheduling policies from interaction rather than from a system model. These policies select which tasks to run and when to run them to balance task execution against resource-preservation (e.g., survival) under unknown or changing resource profiles, providing tunable execution–survival trade-offs without requiring prior energy or system models.

model-freerlscheduling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.06
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the problem of dynamically allocating prediction tasks among capacity-constrained agents—whether human or artificial—to maximize collective performance. It introduces, for the first time, a theoretical formulation of task assignment under explicit capacity constraints and proposes a context-aware sequential exploration–exploitation learning framework. This framework integrates multi-agent capability modeling with optimized task–agent matching strategies. Empirical evaluations demonstrate that the proposed approach significantly outperforms non-contextual baselines across tabular, image, and text prediction tasks, and is effective in collaborative settings involving both large language models and human agents.

agent expertisecapacity constraintsprediction tasks

To address the low sample efficiency, poor safety guarantees, and limited interpretability of model-free reinforcement learning (RL), this paper proposes a hybrid framework that synergistically integrates model-free and model-based approaches. Specifically, it embeds model predictive control (MPC) into the policy optimization pipeline, leveraging differentiable and interpretable dynamic and constraint models to encode safety priors. It further combines Bayesian optimization–driven policy search with offline RL to mitigate model mismatch. An end-to-end trainable adaptive model jointly captures system dynamics, cost functions, and constraints. Experimental results demonstrate that the method improves sample efficiency by up to 2.3×, significantly enhances decision safety and interpretability, and validates the effectiveness and generalization capability of the “model-guided + data-driven” paradigm in complex, uncertain environments.

Enhancing interpretability of RL-based control policiesEnsuring safe learning in autonomous decision-making systemsImproving sample efficiency in model-free RL agents

Current large language models are constrained by fixed context windows, limiting their ability to handle highly complex tasks. This work proposes a reinforcement learning–based recursive agent training framework that enables agents, during inference, to autonomously decide whether and how to recursively invoke themselves, dynamically decomposing tasks and delegating subtasks. The approach achieves, for the first time, adaptive recursion and coordination among agents at inference time, effectively circumventing context length limitations. It significantly enhances generalization and reasoning efficiency on tasks far exceeding the complexity encountered during training, while maintaining higher training efficiency and achieving lower overall inference latency compared to single-agent systems.

Context Window ExtensionInference-time ScalingRecursive Agents

Improving Offline Mixed-Criticality Scheduling with Reinforcement Learning

Apr 04, 2025
ME
Muhammad El-Mahdy
🏛️ The American University in Cairo | Pontificia Universidad Católica de Chile

This paper addresses the NP-hard problem of offline, non-preemptive mixed-criticality (MC) real-time scheduling on heterogeneous multicore platforms. We propose the first systematic reinforcement learning (RL)-based solution, modeling scheduling as a Markov decision process and employing PPO and DQN algorithms to jointly maximize overall task completion rate while guaranteeing schedulability of high-criticality tasks. Our approach innovatively overcomes the performance limitations of conventional static analysis and heuristic methods, enabling adaptive handling of dynamic workloads and processor-speed heterogeneity. Evaluated on 100,000 synthetic instances and real-world datasets, our method achieves an average task completion rate of 80% (85% for high-criticality tasks), rising to 94% (93% for high-criticality tasks) under stable scenarios—significantly outperforming state-of-the-art offline MC schedulers.

Improves task completion rates in dynamic real-time systemsSolves non-preemptive NP-hard scheduling in mixed-criticality systemsUses RL to prioritize high-critical tasks efficiently

Interpreting Emergent Planning in Model-Free Reinforcement Learning

Apr 02, 2025
TB
Thomas Bush
🏛️ University of Cambridge | FAR AI | Mila | University of Montreal

This study investigates whether purely model-free reinforcement learning (RL) agents can spontaneously develop planning capabilities and elucidates the underlying mechanisms. Method: Focusing on the Sokoban task, we propose a concept-driven interpretability framework integrating Test-Input Concept Activation Vectors (TCAV), representation intervention, causal attribution, and reverse engineering of planning algorithms. Contribution/Results: We provide the first mechanistic evidence for implicit planning in model-free RL. Specifically: (1) The DRC agent autonomously constructs an implicit planning logic resembling parallel bidirectional search—without any explicit planning module; (2) its decisions rely on learned conceptual representations to generate internal plans and perform causal reasoning; (3) planning-relevant representations exhibit a causal relationship with test-time computational resources, and increased computation yields significant gains in planning performance. This work challenges the conventional assumption that “model-free implies no planning,” establishing a verifiable and intervenable planning mechanism within model-free RL.

Causal effect of learned plans on agent behaviorInterpretability methodology for plan formation in SokobanMechanistic evidence of planning in model-free RL agents

Latest Papers

What's happening recently
View more

This study investigates whether the substantial training cost of deep reinforcement learning (DRL) in carbon-aware flow shop scheduling can be justified by long-term performance gains through policy generalization. To this end, the authors propose a framework that integrates DRL with dynamic algorithm configuration, training policies on small, simple instances and transferring them to unseen, complex instances for online parameter adaptation. Experimental results demonstrate that the proposed approach significantly outperforms baseline methods—such as static parameter tuning—on out-of-distribution, complex scenarios. These findings validate the strong generalization capability of DRL policies and confirm that the initial investment in training yields sustained performance benefits in practical applications.

carbon-aware schedulingcomputational costdeep reinforcement learning

This work proposes a scheduling approach for dynamic job shop scheduling problems under uncertainty caused by stochastic job arrivals and unexpected machine failures. The randomness of job arrivals and machine breakdowns is modeled using Gamma and Weibull distributions, respectively. To ensure policy optimization remains within the feasible action space, two action-masking mechanisms—non-gradient and gradient-based—are integrated with a Maskable Proximal Policy Optimization algorithm. Evaluated on standard dynamic JSSP benchmarks, the proposed method significantly outperforms conventional heuristics and dispatching rules, demonstrating superior performance in minimizing makespan, along with strong robustness and good scalability.

Dynamic Job Shop SchedulingMachine FailuresRandom Arrivals

This work proposes a policy-based deep reinforcement learning hyper-heuristic framework to address the challenge of dynamically selecting effective dispatching rules for the job shop scheduling problem (JSSP). The framework enables an agent to adaptively switch among low-level dispatching rules based on the current system state. It introduces two key innovations: an action pre-filtering mechanism to ensure unbiased evaluation of heuristics and a commitment mechanism to regulate the frequency of rule switching. By integrating both deterministic and stochastic policy selection strategies, the approach significantly outperforms conventional heuristics, metaheuristics, and existing neural network–based scheduling methods on standard JSSP benchmarks, demonstrating superior scheduling performance and training stability.

Dynamic SchedulingHyper-heuristicsJob Shop Scheduling Problem

About Time: Model-free Reinforcement Learning with Timed Reward Machines

Dec 19, 2025
AM
Anirban Majumdar
🏛️ Tata Institute of Fundamental Research | Max Planck Institute for Software Systems | University of Oxford

Traditional reinforcement learning reward mechanisms struggle to precisely encode temporal constraints, limiting their applicability in time-sensitive tasks. To address this, we propose the Temporal Reward Machine (TRM), the first framework to embed timed automata semantics directly into reward modeling—enabling programmable reward logic such as delay penalties and prompt incentives under both discrete- and continuous-time semantics. Methodologically, TRM integrates temporal abstraction Q-learning, timed automaton specification, and a counterfactual imagination heuristic to achieve model-agnostic, efficient policy optimization. Experiments demonstrate that TRM significantly improves policy performance under strict temporal constraints on standard RL benchmarks. Ablation studies confirm that the counterfactual heuristic accelerates convergence by 42% and increases constraint satisfaction rate by 31%. Our core contribution is the establishment of the first semantically rigorous, interpretable, and scalable time-aware reward modeling paradigm.

Enables expressive specifications with tunable reward logic for time-sensitive applications.Extends reward machines to incorporate precise timing constraints.Studies model-free RL frameworks to learn optimal policies under timing constraints.

This work addresses the challenge of dynamically coordinating context, prompts, and tool invocations for large language model (LLM) agents under resource constraints in multi-turn interactions. It introduces a Stackelberg game formulation into LLM resource scheduling and proposes a context-aware, repairable conditional policy framework: a controller sets quality–cost objectives, and an executor dynamically allocates resources accordingly. Strategy repair is achieved through learning conditional response models, optimizing the leader’s policy, and integrating real-API calibration with projection onto safe action sets. Theoretical analysis establishes guarantees on equilibrium existence, response stability, and environmental transferability. In 300-round real-API experiments, the repaired policy reduces token cost by 17.4% on average compared to a conservative baseline (p = 0.022) without significant degradation in output quality (p = 0.44).

computational budgetcontext managementLLM agents

Hot Scholars

GP

Giuseppe Prencipe

Dipartimento di Informatica, Universita' di Pisa
Digital HealthDistributed ComputingAutonomous Mobile Robots
PF

Paola Flocchini

University of Ottawa
Distributed ComputingDistributed AlgorithmsDynamic NetworksCellular Automata
NS

Nicola Santoro

Distingushed Research Professor, Carleton University
Distributed ComputingAlgorithmsMobile Computational EntitiesDynamic Systems
AG

Animesh Garg

Georgia Institute of Technology, University of Toronto
Robotic ManipulationRobot LearningReinforcement LearningMachine Learning
XX

Xun Xiao

Munich Research Center, Huawei Technologies
BlockchainDistributed NetworkingIn-network ComputingGame Theory