🤖 AI Summary
This study investigates whether structured text can substitute for vision in chart reasoning, explicitly disentangling errors arising from missing representational information versus insufficient solver capability. We propose a diagnostic protocol that contrasts original images, model-extracted text, and source structures, integrating edge interventions with token efficiency analysis to precisely localize structural bottlenecks during the textualization process. Experimental results demonstrate that answer-relevant topology predicts accuracy more effectively than global topology. Furthermore, source structure achieves 87% accuracy, whereas both direct visual processing and learned textual representations fall below 30%. Notably, editing a single critical edge suffices to reduce accuracy to zero. This work establishes a fine-grained evaluation framework for multimodal chart reasoning.
📝 Abstract
Can structured text replace vision for diagram reasoning? A wrong answer after textualization can arise because the representation omits information the question needs, or because the solver fails to use information that is present. We introduce a diagnostic protocol to distinguish these explanations. Using the same solver model and generation settings, we compare three input conditions: the original image, question-blind structure extracted by a vision-language model, or gold structure derived from the diagram source. Validity-triggered recovery tests truncation and schema failure, question-relevant fidelity measures preservation of answer-critical structure, and matched edge interventions test the effect of error location. On a reserved holdout of 240 public FlowGen diagrams, evaluated under a frozen protocol, gold structure reaches 87% accuracy while direct vision and learned text both remain below 30%. The aggregate comparison includes source-derived relation labels that may not be printed in the image and uses different learned and gold graph encodings, so it does not isolate extraction error alone. Retrying only invalid extractions makes nearly every public representation schema-valid yet leaves accuracy essentially unchanged. The public learned-text deficit relative to gold more than doubles with structural difficulty. Question-relevant topology predicts correctness better than whole-graph topology. In an exposed intervention study, a single answer-relevant edge edit reduces the primary solver's original-answer accuracy to near zero, while matched irrelevant edits largely preserve it. Supplied structure requires fewer solving tokens than vision, but learned acquisition removes this advantage at single use. These comparisons motivate evaluating acquired text by the answer-relevant evidence it preserves and by the solver's ability to use that representation.