Faithful Chart Generation for Multimodal Deep Research: Frame-Evidence Co-Adaptation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the mismatch between chart values and evidence in multimodal research, as well as the inflexibility of predefined visualization plans, by proposing the FECA framework. Inspired by data frame theory, this method shifts chart generation from fixed-plan verification to evidence-based adaptive visual reasoning. Through a bidirectional perception mechanism and an iterative interaction algorithm, the framework dynamically couples visual frames with retrieved evidence, enabling real-time adjustment or discarding of unreliable visual representations. Experiments conducted on 100 real-world research topics demonstrate that FECA significantly improves the numerical fidelity of generated charts while effectively preserving overall report quality and chart utility.
📝 Abstract
Analytical charts in multimodal deep research encode quantitative claims, requiring every visualized value to be faithfully grounded in supporting evidence. Unlike retrieved images that mainly provide contextual information, charts require numerical fidelity: visualized values should not only match retrieved evidence quantitatively but also preserve its original meaning and scope. However, achieving such fidelity remains challenging because current systems usually construct visualization plans before knowing what quantitative evidence can actually be retrieved from the web. As a result, predefined plans may require entities, temporal ranges, or comparison dimensions that the retrieved evidence only partially supports. Existing approaches mainly address this issue through post-hoc verification after chart plans are fixed, enabling unsupported values to be identified but leaving the underlying visual frames unchanged. To address this challenge, we propose Frame-Evidence Co-Adaptation (FECA), an evidence-adaptive visual planning framework for multimodal deep research. Inspired by the bidirectional sensemaking process in Data-Frame Theory, FECA models chart generation as an iterative interaction between visual frames and retrieved evidence. Each visual frame is adaptive: the frame guides evidence acquisition, while retrieved evidence determines whether the frame should be accepted, revised, or dropped before rendering. By coupling visualization planning with evidence availability, FECA shifts chart generation from fixed-plan verification to adaptive evidence-grounded visual reasoning. Experiments on 100 real-world research topics show that FECA substantially improves numerical fidelity while preserving report quality and chart utility.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Deep Research
Chart Generation
Numerical Fidelity
Evidence Grounding
Visual Planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Frame-Evidence Co-Adaptation
Multimodal Deep Research
Numerical Fidelity
Visual Reasoning
Data-Frame Theory
🔎 Similar Papers
No similar papers found.
Y
Yuxin Yue
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences, Beijing, China
Y
Yingchen Zhang
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences, Beijing, China
Ruqing Zhang
Ruqing Zhang
Institute of Computing Technology, Chinese Academy of Sciences
Information RetrievalNatural Language ProcessingLarge Language Models
Jiafeng Guo
Jiafeng Guo
Professor, Institute of Computing Techonology, CAS
Information RetrievalMachine LearningText AnalysisNeuIR
Maarten de Rijke
Maarten de Rijke
University of Amsterdam & ICAI
Information retrieval
Xueqi Cheng
Xueqi Cheng
Ph.D. student, Florida State University
Data miningLLMGNNComputational social science