ChartBmkAgent: Harness-Governed Multi-Agent Construction of Chart QA Benchmarks from Sparse Error-Taxonomy Specifications

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lag in benchmark development for multimodal large language models and the misalignment between diagnostic objectives and generated content in automated evaluation. To this end, we propose a multi-agent collaborative framework governed by a central Harness mechanism. This approach introduces a novel Harness governance paradigm that integrates sparse error taxonomy specification mapping to construct chart question-answering diagnostic samples on demand, ensuring strict alignment with externally specified diagnostic objectives during data generation. Experimental results demonstrate that 86.4% of the generated samples satisfy the target requirements, effectively revealing capability disparities across different models. These findings validate both the feasibility and efficacy of targeted data generation for model evaluation.
📝 Abstract
Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observed capability gaps. Such investigation requires an expressive task format and an on-demand construction process: information-rich charts make chart question answering (Chart QA) suitable for probing coupled perception and reasoning. Automated Chart QA construction is intended to shorten the benchmark-development cycle by turning identified gaps into targeted samples on demand. Current methods, however, commonly separate target guidance from scratch generation: target-guided systems often require prepared data, charts, or templates, while scratch-generation systems primarily ensure artifact validity, without explicitly controlling whether newly synthesized requirements and content remain aligned with an externally specified diagnostic target. We introduce ChartBmkAgent, which turns an identified capability gap into targeted diagnostic evidence by constructing complete Chart QA samples from sparse error-taxonomy specifications. Throughout construction, a central harness governs specialized agents, requires stage-specific evidence of alignment with the original error category, and records the basis for each acceptance decision. On 300 taxonomy-wide samples, MLLM accuracies ranged from 32.7% to 84.3% with distinct category profiles, showing that generated samples reveal capability differences. Across three source-model comparisons, targeted follow-ups scored 50.0% versus 82.2% on matched controls ($p=8.96\times10^{-6}$); all six cross-model comparisons had the same direction, demonstrating targeted validation and diagnostic-data generation. Multiple evaluator models assessed whether each sample tested its specified error category; 86.4% met this criterion, providing empirical evidence of target preservation.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Chart Question Answering
Benchmark Construction
Capability Gap Diagnosis
Error Taxonomy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent System
Chart Question Answering
Benchmark Construction
Error Taxonomy
Harness-Governed Generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.