🤖 AI Summary
This study addresses the absence of interactive chart code generation tasks and browser-based execution verification in existing benchmarks by constructing an interactive chart benchmark comprising 377 tasks, alongside a pioneering end-to-end evaluation framework that supports interactive state transitions. Methodologically, it proposes TRAIL, a dual-agent framework incorporating a trajectory-guided self-correction mechanism. By integrating multimodal large language models with sandboxed browser automation testing, TRAIL establishes a closed-loop process for diagnosis and repair. Experimental results demonstrate that the proposed evaluator achieves an F1 score of 0.8844, while the TRAIL framework yields an average improvement of over three percentage points across four key metrics.
📝 Abstract
Interactive chart code generation requires models to reproduce a reference chart's appearance and underlying data and correctly implement the state changes triggered by specified user interactions. Existing chart-to-code benchmarks focus on static outputs and lack task representations or evaluation protocols for interaction specification, browser execution, and post-interaction verification. We introduce IChart2Code, a benchmark comprising 377 tasks across 20 chart forms and 13 data families, with 1209 interaction requirements in six families. Each task provides a reference screenshot, task-local data, and natural-language interaction requirements, with executable HTML/JavaScript code as the target output. We further develop a browser-based evaluation protocol with an Executability gate and three rubric-guided dimensions: Data Fidelity, Static Visual Correctness, and Interaction Correctness. The protocol tests runtime viability, consistency with task-local data, fidelity of the initial rendering to the reference screenshot, and interaction-induced state changes in a sandboxed browser. A rubric-guided MLLM judge evaluates task-specific items for the three scored dimensions using the collected browser observations and achieves an overall item-level F1 score of 0.8844 against adjudicated human labels. We also propose TRAIL, a trajectory-guided dual-agent framework for interactive chart code generation. An Inspector derives task-specific inspection checks, executes them in the browser, and uses the resulting trajectories to diagnose failures and produce structured repair feedback. Averaged across four MLLMs, TRAIL improves the four evaluation dimensions over direct prompting by 7.89, 4.98, 3.45, and 4.47 percentage points, respectively.