From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the disconnect between visual understanding and code generation evaluation in multimodal large language models by proposing FigCodeBench, a unified benchmark. The framework comprises 6,194 instances organized within a three-tier complexity hierarchy and introduces a multidimensional evaluation protocol encompassing visual fidelity and syntactic isomorphism, which demonstrates strong alignment with human preferences. Furthermore, it presents an integrated reasoning mechanism combining structured visual modeling with cross-lingual generation. Evaluations across 24 mainstream models reveal a nonlinear performance cliff phenomenon and significant deficiencies in declarative languages, offering critical insights for optimizing multimodal code generation.
📝 Abstract
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedicated assessment of their unification, i.e., how a model can perceive complex visual structures and synthesize them into precise, executable code. Moreover, current visual code generation benchmarks often rely on simplified layouts within single programming environments, falling short of evaluating true unified multimodal reasoning. To bridge this gap, we propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation. We first design a systematic dataset construction pipeline, resulting in a total of 6,194 instances that cover 7 functional categories and 4 types of programming languages. We further categorize figure reproduction into three tiers with visual and code complexity modeling, specifically targeting complex structural reasoning, varying aspect ratios, and dense geometric constraints. We introduce a multi-dimensional evaluation protocol, encompassing visual fidelity and syntactic isomorphism, that aligns highly with the Mean Machine Opinion Score (MMOS) and human preferences. Based on our framework, we conducted extensive experiments on 24 widely used proprietary and open-source MLLMs (e.g., Gemini 3.1 Pro, GPT-5.4, and Kimi-K2.5), where we observed a universal, non-linear performance cliff across different programming languages and difficulty scenarios for all models, and gained several insights, such as the significant metric decline in rigid declarative languages.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Figure Reproduction
Visual Code Generation
Benchmark Evaluation
Multimodal Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Large Language Models
Figure Reproduction
Benchmark
Visual Code Generation
Evaluation Protocol
🔎 Similar Papers
Zijian Chen
Zijian Chen
Shanghai Jiao Tong University | Shanghai AI Laboratory
Image/Video Quality AssessmentLarge Multi-modal Models
Z
Zhengyu Chen
Institute of Image Communication and Information Processing, Shanghai Jiao Tong University, Shanghai, 200240, China
B
Bohan Liang
Shanghai Artificial Intelligence Laboratory, Shanghai, 200030, China
L
Lirong Deng
Macao Polytechnic University, Macao, 999078, China
Y
Yushuo Zheng
Institute of Image Communication and Information Processing, Shanghai Jiao Tong University, Shanghai, 200240, China; Shanghai Artificial Intelligence Laboratory, Shanghai, 200030, China
Y
Yanwei Jiang
Institute of Image Communication and Information Processing, Shanghai Jiao Tong University, Shanghai, 200240, China; Shanghai Artificial Intelligence Laboratory, Shanghai, 200030, China
Q
Qi Jia
Shanghai Artificial Intelligence Laboratory, Shanghai, 200030, China
K
Kaiwei Zhang
Shanghai Artificial Intelligence Laboratory, Shanghai, 200030, China
Wenjun Zhang
Wenjun Zhang
City University of Hong Kong
Thin film technologynanomaterials and nanodevices
Guangtao Zhai
Guangtao Zhai
Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI EvaluationDisplays