Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System

๐Ÿ“… 2026-06-23
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge of distinguishing genuine errors from scoring artifacts in the evaluation of multimodal intelligent data analyst systems. To this end, it proposes a three-tier human-in-the-loop cascaded evaluation framework that integrates strict regex matching, large modelโ€“driven parser-agnostic lenient scoring, keyword-anchored answer extraction, and manual snippet verification. The framework is systematically applied to evaluate LAMBDA across 153 numerical tasks on DSGym. The study introduces two key innovations: an iterative prompting refinement (โ€œnudgeโ€) mechanism and variable-type metadata awareness, both of which substantially enhance scoring robustness. Experimental results demonstrate that the automated scorer achieves 100% precision, while lenient scoring attains a 97% recall; moreover, the nudge mechanism elevates scoring success rates from 36% to 97% and increases overall task pass rates from 16% to 46%.
๐Ÿ“ Abstract
Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This makes them more challenging to evaluate than single-turn LLM responses. It is therefore necessary to distinguish genuine disagreement between an agent's output and a ground-truth answer from grading artifacts. We investigate how reliably automated graders assess such a system and what strategies improve grading quality by applying LAMBDA, a multi-agent data-analysis system, on 153 numerical QRData tasks from DSGym. We develop and evaluate a three-layer human-AI grading cascade: strict regex matching, LLM-based lenient grading, and snippet-based human inspection, which combines non-GenAI and GenAI strategies with different failure profiles. Both automated graders achieve 100% observed precision (0/70 false positives). The lenient grader's recall is 97% against human labels. A keyword-anchored extraction pipeline raises the strict grader's recall by 60 percentage points over a last-number heuristic; the lenient grader is architecturally parser-independent. An iterative nudge mechanism raises grading run success from 36% to 97% and lenient-pass rates from 16% to 46%; comparing nudging with and without original-question re-injection shows that re-injection offers no benefit, confirming the nudge as an answer template cue. We further observe in this case study that variable type is the task metadata field most consistently associated with grading pipeline dynamics and observed outcome grades.
Problem

Research questions and friction points this paper is trying to address.

agentic data analysis
automated grading
evaluation reliability
grading artifacts
LLM-based assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

agentic evaluation
grading cascade
keyword-anchored extraction
iterative nudge mechanism
parser-independent grading