PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the sim-to-real gap in online reinforcement learning for industrial scheduling and the reliance of offline methods on high-quality data by proposing a hybrid paradigm that integrates online exploration with offline adaptation. The approach first pre-trains a generalizable policy through online interactions, followed by fine-tuning on offline data using a Kullback-Leibler divergence regularization constraint to align with the target distribution. This mechanism effectively restricts policy deviation and reduces sensitivity to data quality. Experimental results demonstrate that the proposed framework consistently achieves lower optimality gaps across diverse behavioral policy datasets, significantly outperforming standalone offline reinforcement learning approaches and existing baseline methods.
📝 Abstract
The Job Shop Scheduling Problem (JSSP) is a fundamental combinatorial optimization problem in industrial optimization. This work introduces Pretrained Offline Reinforcement Learning (PORL), a hybrid approach that combines simulation-based online pretraining with offline fine-tuning on production-specific data. Reinforcement learning through online interaction enables exploration of general scheduling strategies, but typically relies on simulation environments and may suffer from a simulation-to-reality gap. In contrast, offline RL avoids direct interaction with the environment by learning from historical data, but its performance is strongly influenced by dataset quality and coverage. PORL combines the strengths of both paradigms by first learning a general scheduling policy through online interaction and subsequently adapting it offline to a target distribution. A KL-divergence-based policy constraint is introduced to limit deviations from the pretrained policy during fine-tuning. The approach is evaluated on JSSP instances with distribution shift and datasets generated from heuristic, noisy-expert, and random behavioral policies. The results show that PORL consistently achieves lower optimality gaps than standalone offline RL and the considered general scheduling baselines. Furthermore, its advantage over standalone offline RL increases as dataset quality decreases, indicating reduced sensitivity to the quality and coverage of the available offline data. The results suggest that offline adaptation of pretrained policies is a promising approach for industrial scheduling environments where direct online exploration is impractical.
Problem

Research questions and friction points this paper is trying to address.

Job Shop Scheduling Problem
Offline Reinforcement Learning
Simulation-to-Reality Gap
Distribution Shift
Combinatorial Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pretrained Offline Reinforcement Learning
Job Shop Scheduling Problem
KL-divergence policy constraint
Simulation-to-reality gap
Offline fine-tuning
M
Mateo Toro Diz
Department of Industrial Engineering, Rosenheim University of Applied Sciences, Rosenheim, Germany
J
Jonathan Hoss
Department of Industrial Engineering, Rosenheim University of Applied Sciences, Rosenheim, Germany
Noah Klarmann
Noah Klarmann
Full Professor, Rosenheim Technical University of Applied Sciences
Artificial IntelligenceMachine LearningReinforcement Learning