Agentic AI for Scientific Reasoning in Autonomous Quantum Sensing Experiments

πŸ“… 2026-07-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work presents the first integration of large language model (LLM) agents into nitrogen-vacancy (NV) center–based quantum sensing, establishing an end-to-end autonomous experimental workflow that fully automates the sequence from individual NV center selection and frequency calibration to Tβ‚‚* measurement and CPMG sequence validation. The approach combines persistent experimental logging, deterministic hardware control, photoluminescence-detected magnetic resonance (pODMR) analysis, signal modeling, and quantitative evaluation tools, and introduces two offline benchmarks to independently assess scientific reasoning capabilities. Experimental results demonstrate that the agent efficiently executes complex quantum sensing tasks, and incorporating expected signal computation substantially reduces false-positive rates under high reasoning demands, thereby validating the effectiveness of synergistically combining LLMs with deterministic code for autonomous scientific discovery.
πŸ“ Abstract
We implement an agentic AI workflow built around a large language model (LLM) agent for autonomous experiments with nitrogen-vacancy (NV) centers in diamond. NV centers are a widely used platform for quantum sensing, and the ability to control many measurements from a computer makes NV experiments a natural setting for autonomous workflows. We make two main contributions. First, we demonstrate an autonomous NV experiment workflow that combines persistent project records, quantitative calculation and data analysis tools, and deterministic experiment control. In one autonomous experiment, the agent selected a single NV center, calibrated its resonant frequency, measured \(T_2^\ast\) with Ramsey measurements, and added a Carr--Purcell--Meiboom--Gill (CPMG) measurement to check a weak feature that could be related to nearby \(^{13}\mathrm{C}\). Second, we introduce two offline benchmarks that evaluate the agent's reasoning separately from laboratory execution. We evaluated both benchmarks with GPT-5.4, GPT-5.5, and GPT-5.6 Sol. In the Ramsey checkpoint benchmark, greater reasoning effort generally improved recognition of a residual resonance calibration offset. By contrast, in the pulsed optically detected magnetic resonance (pODMR) data evaluation benchmark, pulse sequence information alone produced more false positive resonance judgments at higher reasoning effort. Requiring an expected signal calculation kept false positive rates low across all three models and reasoning settings. The results suggest a clear division of labor for autonomous experiments. The agent forms scientific hypotheses and uses quantitative tools to evaluate data, while deterministic code controls the hardware and enforces safety constraints.
Problem

Research questions and friction points this paper is trying to address.

Agentic AI
Quantum Sensing
Autonomous Experiment
Scientific Reasoning
NV Centers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic AI
Autonomous Quantum Sensing
Nitrogen-Vacancy Centers
Scientific Reasoning
Offline Benchmarking