🤖 AI Summary
This study addresses the disconnect between visual understanding and code generation evaluation in multimodal large language models by proposing FigCodeBench, a unified benchmark. The framework comprises 6,194 instances organized within a three-tier complexity hierarchy and introduces a multidimensional evaluation protocol encompassing visual fidelity and syntactic isomorphism, which demonstrates strong alignment with human preferences. Furthermore, it presents an integrated reasoning mechanism combining structured visual modeling with cross-lingual generation. Evaluations across 24 mainstream models reveal a nonlinear performance cliff phenomenon and significant deficiencies in declarative languages, offering critical insights for optimizing multimodal code generation.
📝 Abstract
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedicated assessment of their unification, i.e., how a model can perceive complex visual structures and synthesize them into precise, executable code. Moreover, current visual code generation benchmarks often rely on simplified layouts within single programming environments, falling short of evaluating true unified multimodal reasoning. To bridge this gap, we propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation. We first design a systematic dataset construction pipeline, resulting in a total of 6,194 instances that cover 7 functional categories and 4 types of programming languages. We further categorize figure reproduction into three tiers with visual and code complexity modeling, specifically targeting complex structural reasoning, varying aspect ratios, and dense geometric constraints. We introduce a multi-dimensional evaluation protocol, encompassing visual fidelity and syntactic isomorphism, that aligns highly with the Mean Machine Opinion Score (MMOS) and human preferences. Based on our framework, we conducted extensive experiments on 24 widely used proprietary and open-source MLLMs (e.g., Gemini 3.1 Pro, GPT-5.4, and Kimi-K2.5), where we observed a universal, non-linear performance cliff across different programming languages and difficulty scenarios for all models, and gained several insights, such as the significant metric decline in rigid declarative languages.