When Every Simulation Counts: Value-Based Reinforcement Learning for Accelerated Photonics Inverse Design

📅 2026-07-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of high-dimensional parameter optimization in the inverse design of photonic crystal surface-emitting lasers (PCSELs), which traditionally relies on computationally expensive full-wave simulations. Under a strict simulation budget of only 83 evaluations, this work presents the first systematic assessment of Deep Q-Networks (DQN) and six of its value-based variants. A reproducible evaluation framework is introduced to disentangle algorithmic improvements from stochastic effects, alongside a reinforcement learning strategy that leverages reuse of simulation data. Experimental results demonstrate that Dueling DQN consistently outperforms all baselines and variants across random seeds, achieving a higher average quality factor, a 64% reduction in wavelength error, and a 47% increase in upward output power, thereby exhibiting superior robustness and sample efficiency.
📝 Abstract
Photonic-crystal surface-emitting lasers (PCSELs) can combine high-power operation with narrow-divergence surface emission, but optimizing coupled parameters requires costly full-wave simulations. Deep Q-network (DQN) optimization can reuse simulated transitions to guide edits, yet which value-learning mechanisms remain reliable under tight simulation budgets is unknown. We address this gap by comparing baseline DQN and six value-based variants for a seven-variable PCSEL design under a shared objective, simulator, 83-call budget, and four matched initializations. Beyond endpoints, we analyze sample efficiency, policy behavior, and physical response to separate learning gains from favorable starts or exploratory jumps. Dueling DQN is the only variant to improve all four seeds. Relative to the first evaluated designs, its selected structures increase the mean quality factor () from to (), reduce wavelength error by 64%, and increase upward power by 47%; compared with baseline DQN, they achieve a higher mean under the same budget. Other variants yield no consistent improvement; Double DQN reproduces baseline trajectories, while Rainbow-lite shows high upside but strong seed dependence. These results identify Dueling DQN as the most reliable configuration tested for simulation-budget-limited PCSEL inverse design and provide a reproducible framework for attributing algorithmic gains in scientific optimization. The source code is publicly available at https://github.com/Longying-Wen/PCSEL-RL.
Problem

Research questions and friction points this paper is trying to address.

photonic-crystal surface-emitting lasers
inverse design
simulation budget
value-based reinforcement learning
sample efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dueling DQN
inverse design
simulation-efficient optimization
photonic-crystal lasers
value-based reinforcement learning
🔎 Similar Papers
No similar papers found.