QC-Stark: A Multi-Task Benchmark Revealing Capability Dissociations in LLMs Evaluated on Quantum Computing Tasks

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing evaluations that obscure capability variations and fluctuations in large language models (LLMs) on quantum computing tasks. To this end, we construct a benchmark encompassing eleven task categories, including circuit construction, and introduce Item Response Theory alongside a multi-difficulty-level design to establish an automated code-execution verification paradigm requiring no human intervention. Our analysis reveals significant discrepancies between overall rankings and single-task performance, with rankings across four core tasks exhibiting no correlation. All datasets and source code are publicly released. This work effectively enhances the robustness and reliability of evaluating LLM capabilities in quantum computing.
📝 Abstract
We introduce QC-Stark, a benchmark for evaluating large language models (LLMs) on 11 quantum computing (QC) tasks, spanning circuit construction, debugging, compilation, error correction, and simulation. Across 2,750 evaluations (10 models $\times$ 11 tasks x 5 difficulty levels x 5 seeds), we find that overall rankings mask substantial per-task variation. The Spearman correlation between overall and per-task rankings is statistically insignificant for 4 out of the 11 tasks included in this benchmark. A 2-parameter Item Response Theory (IRT) model validates measurement quality, and prompt sensitivity analysis confirms ranking robustness across prompt conditions. All tasks are auto-verifiable via execution, thus not requiring any manual evaluation. We make the code and data publicly available on Huggingface.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Quantum Computing
Benchmark
Capability Dissociation
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantum Computing Benchmark
Large Language Models
Multi-Task Evaluation
Item Response Theory
Capability Dissociation
🔎 Similar Papers
No similar papers found.