π€ AI Summary
This study addresses the challenge of rigorously evaluating large language model agents in power systems, where data confidentiality and the absence of chain-dependent evidence hinder assessment. To overcome this, we propose an interconnected heterogeneous data generation framework based on common dependency chains. This framework constructs a benchmark comprising a large-scale synthetic dataset and 300 evaluation questions to systematically test agentsβ autonomous retrieval and multi-step reasoning capabilities under constrained tool invocation, thereby bridging the evaluation gap for power-domain agents. Experimental results demonstrate that the best-performing model achieves a joint accuracy of only 74.2%, revealing performance bottlenecks across individual stages and providing critical guidance for reliable industrial-level deployment.
π Abstract
Large language model (LLM) agents offer new opportunities for automated analysis in industry. However, rigorous evaluation of such agents-for example, within power system scenarios-remains hindered: real operational data are confidential, and existing public resources fail to fully capture the chained dependencies and heterogeneous evidence. To address this gap, we propose PowerBench, comprising (1) a generation framework that derives interconnected heterogeneous operational data through a common dependency chain, and (2) a synthetic dataset generated by this framework. The dataset covers 761 devices across 100 device types, with 13.35 million hourly telemetry records spanning two years and 24,939 operational documents. Building on this dataset, we construct 300 questions across three task families that evaluate frontier LLMs' ability to complete analysis tasks that require autonomous evidence retrieval and reasoning across interconnected and heterogeneous data under restricted tool calls and time budgets. Results demonstrate that the evaluated frontier LLMs remain challenged on these tasks: the best model reaches only 74.2% joint accuracy. Our trace analysis further reveals that model performance varies across evidence discovery, content retrieval, tool use, reasoning over evidence, and answer submission. These findings provide detailed insights for evaluating LLM agents and guiding their reliable deployment in industry. The framework, dataset, and benchmark tasks are available at https://github.com/open-compass/PowerBench.