VisAudit: Evaluating Multimodal Agents for Visual Diagnosis and Repair

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of autonomous auditing capabilities in multimodal agents for visualization tasks, which hinders their ability to effectively identify, repair, and verify chart defects. To this end, it introduces VisAudit, a benchmark that generates defective instances with source code and data table evidence through controlled perturbations. The proposed approach designs an iterative mechanism that executes code and analyzes visual feedback to achieve a closed-loop "diagnose-repair-verify" autonomous auditing process. This work establishes the first three-track evaluation framework encompassing diagnosis, autonomous repair, and open-world verification, thereby filling the gap in end-to-end autonomous auditing assessment. Experiments on 2,200 samples reveal that even the strongest model attains only a 47.4% success rate in autonomous repair, exposing significant limitations in current techniques.
📝 Abstract
Multimodal agents are increasingly used for data visualization tasks but remain limited in autonomous review. Unlike humans, they may fail to recognize when a visualization is incorrect, determine what to change, repair it without disrupting correct content, and verify whether the intervention succeeded. Existing benchmarks largely evaluate predefined individual capabilities such as chart generation, instruction-guided editing, or defect detection, and therefore do not capture this gap in autonomous review. We introduce VisAudit, a benchmark for evaluating visualization diagnosis, repair, and verification. Given a rendered chart and configurable auxiliary evidence, including its source data table, intended text summary, and visualization code, an agent iteratively diagnoses potential defects, modifies and executes visualization code, inspects execution and visual feedback, and determines when no further intervention is needed. VisAudit defines three tracks spanning diagnosed repair, autonomous repair, and open-world verification, and contains 1,900 flawed instances across 21 chart types and 10 flaw categories, together with 300 initially correct charts. We construct the benchmark through controlled perturbations of validated source visualizations, with systematic verification and human-aligned quality control to ensure that injected defects are well-defined and recoverable from the available evidence. Experiments with leading multimodal models reveal a substantial gap from reliable autonomous review: the strongest evaluated model fully recovers only $47.4\%$ of flawed charts in the autonomous-repair setting.
Problem

Research questions and friction points this paper is trying to address.

multimodal agents
data visualization
autonomous review
benchmark evaluation
visual diagnosis and repair
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Agents
Visual Diagnosis and Repair
Autonomous Review Benchmark
Data Visualization
Iterative Code Execution
🔎 Similar Papers
No similar papers found.