🤖 AI Summary
This study addresses the challenge of comparing machine learning-based optimal power flow (ML-OPF) methods caused by inconsistent evaluation criteria by proposing a unified benchmarking framework. Through standardized pipelines and multi-objective stress testing, representative algorithms are systematically evaluated under scalability, distribution shift, and resource constraints, with the framework released as an extensible open-source Python toolkit. The findings reveal that prediction accuracy is not a reliable indicator of operational feasibility, confirming that post-processing constraint repair outperforms pure predictive models. Furthermore, during data scaling, the trajectories of accuracy and violation rates decouple, while computational benefits exhibit diminishing returns. These insights provide critical practical guidance for ML-OPF system design.
📝 Abstract
Machine Learning (ML) methods promise a fast solution process for Optimal Power Flow (OPF). While inconsistent test cases, implementations, and evaluation metrics across existing studies make it challenging to determine which algorithmic advances are most critical for real-world deployment. To this end, we propose ML-OPF-Bench, a unified benchmark for AC- and DC-OPF that evaluates representative ML algorithms under a consistent pipeline, stress-tests them across system sizes, distribution shifts, and resource budgets, and ranks them with a multi-objective framework. We find that prediction accuracy alone is not a reliable indicator of operational feasibility. Under heavily loaded, congested conditions, even the strongest in-distribution performers lose their advantage, while feasibility is maintained largely by post-processing that enforces the target constraints rather than by the underlying pure ML predictor. Data scaling shows that prediction accuracy and constraint violations follow different trajectories, whereas compute scaling shows that returns diminish and that larger models do not consistently perform better. These results expose critical trade-offs among ML methods'speed, accuracy, and feasibility, and offer practical guidance for future ML-OPF design. We open-source the benchmark as an extensible Python package for integrating new learning-based OPF algorithms and evaluating them under the same standard as the existing baselines.