🤖 AI Summary
This study addresses the limitations of existing evaluations that obscure capability variations and fluctuations in large language models (LLMs) on quantum computing tasks. To this end, we construct a benchmark encompassing eleven task categories, including circuit construction, and introduce Item Response Theory alongside a multi-difficulty-level design to establish an automated code-execution verification paradigm requiring no human intervention. Our analysis reveals significant discrepancies between overall rankings and single-task performance, with rankings across four core tasks exhibiting no correlation. All datasets and source code are publicly released. This work effectively enhances the robustness and reliability of evaluating LLM capabilities in quantum computing.
📝 Abstract
We introduce QC-Stark, a benchmark for evaluating large language models (LLMs) on 11 quantum computing (QC) tasks, spanning circuit construction, debugging, compilation, error correction, and simulation. Across 2,750 evaluations (10 models $\times$ 11 tasks x 5 difficulty levels x 5 seeds), we find that overall rankings mask substantial per-task variation. The Spearman correlation between overall and per-task rankings is statistically insignificant for 4 out of the 11 tasks included in this benchmark. A 2-parameter Item Response Theory (IRT) model validates measurement quality, and prompt sensitivity analysis confirms ranking robustness across prompt conditions. All tasks are auto-verifiable via execution, thus not requiring any manual evaluation. We make the code and data publicly available on Huggingface.