🤖 AI Summary
This study addresses the mismatch between chart values and evidence in multimodal research, as well as the inflexibility of predefined visualization plans, by proposing the FECA framework. Inspired by data frame theory, this method shifts chart generation from fixed-plan verification to evidence-based adaptive visual reasoning. Through a bidirectional perception mechanism and an iterative interaction algorithm, the framework dynamically couples visual frames with retrieved evidence, enabling real-time adjustment or discarding of unreliable visual representations. Experiments conducted on 100 real-world research topics demonstrate that FECA significantly improves the numerical fidelity of generated charts while effectively preserving overall report quality and chart utility.
📝 Abstract
Analytical charts in multimodal deep research encode quantitative claims, requiring every visualized value to be faithfully grounded in supporting evidence. Unlike retrieved images that mainly provide contextual information, charts require numerical fidelity: visualized values should not only match retrieved evidence quantitatively but also preserve its original meaning and scope. However, achieving such fidelity remains challenging because current systems usually construct visualization plans before knowing what quantitative evidence can actually be retrieved from the web. As a result, predefined plans may require entities, temporal ranges, or comparison dimensions that the retrieved evidence only partially supports. Existing approaches mainly address this issue through post-hoc verification after chart plans are fixed, enabling unsupported values to be identified but leaving the underlying visual frames unchanged. To address this challenge, we propose Frame-Evidence Co-Adaptation (FECA), an evidence-adaptive visual planning framework for multimodal deep research. Inspired by the bidirectional sensemaking process in Data-Frame Theory, FECA models chart generation as an iterative interaction between visual frames and retrieved evidence. Each visual frame is adaptive: the frame guides evidence acquisition, while retrieved evidence determines whether the frame should be accepted, revised, or dropped before rendering. By coupling visualization planning with evidence availability, FECA shifts chart generation from fixed-plan verification to adaptive evidence-grounded visual reasoning. Experiments on 100 real-world research topics show that FECA substantially improves numerical fidelity while preserving report quality and chart utility.