IChart2Code: Benchmarking Multimodal Large Language Models for Interactive Chart Code Generation

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of interactive chart code generation tasks and browser-based execution verification in existing benchmarks by constructing an interactive chart benchmark comprising 377 tasks, alongside a pioneering end-to-end evaluation framework that supports interactive state transitions. Methodologically, it proposes TRAIL, a dual-agent framework incorporating a trajectory-guided self-correction mechanism. By integrating multimodal large language models with sandboxed browser automation testing, TRAIL establishes a closed-loop process for diagnosis and repair. Experimental results demonstrate that the proposed evaluator achieves an F1 score of 0.8844, while the TRAIL framework yields an average improvement of over three percentage points across four key metrics.
📝 Abstract
Interactive chart code generation requires models to reproduce a reference chart's appearance and underlying data and correctly implement the state changes triggered by specified user interactions. Existing chart-to-code benchmarks focus on static outputs and lack task representations or evaluation protocols for interaction specification, browser execution, and post-interaction verification. We introduce IChart2Code, a benchmark comprising 377 tasks across 20 chart forms and 13 data families, with 1209 interaction requirements in six families. Each task provides a reference screenshot, task-local data, and natural-language interaction requirements, with executable HTML/JavaScript code as the target output. We further develop a browser-based evaluation protocol with an Executability gate and three rubric-guided dimensions: Data Fidelity, Static Visual Correctness, and Interaction Correctness. The protocol tests runtime viability, consistency with task-local data, fidelity of the initial rendering to the reference screenshot, and interaction-induced state changes in a sandboxed browser. A rubric-guided MLLM judge evaluates task-specific items for the three scored dimensions using the collected browser observations and achieves an overall item-level F1 score of 0.8844 against adjudicated human labels. We also propose TRAIL, a trajectory-guided dual-agent framework for interactive chart code generation. An Inspector derives task-specific inspection checks, executes them in the browser, and uses the resulting trajectories to diagnose failures and produce structured repair feedback. Averaged across four MLLMs, TRAIL improves the four evaluation dimensions over direct prompting by 7.89, 4.98, 3.45, and 4.47 percentage points, respectively.
Problem

Research questions and friction points this paper is trying to address.

interactive chart code generation
multimodal large language models
benchmark
evaluation protocol
chart-to-code
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interactive Chart Code Generation
Multimodal Benchmark
Browser-based Evaluation Protocol
Trajectory-guided Dual-agent Framework
MLLM Judge
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xu Zhang
School of Big Data and Software Engineering, Chongqing University, Chongqing, China
H
Hongzhang Zheng
School of Big Data and Software Engineering, Chongqing University, Chongqing, China
Z
Zhili Huang
School of Big Data and Software Engineering, Chongqing University, Chongqing, China
Y
Yaoyi Wang
School of Big Data and Software Engineering, Chongqing University, Chongqing, China
L
Ling Xu
School of Big Data and Software Engineering, Chongqing University, Chongqing, China
Sheng Huang
Sheng Huang
School of Big Data and Software Engineering, Chongqing University
Pattern RecognitionMachine learningComputer Vision