🤖 AI Summary
This study addresses the inefficiency of multimodal large language models in interpreting scientific figures, which typically rely on manual text-image alignment by users. To overcome this limitation, we propose FigAct, a framework that introduces the novel concept of an "active canvas" to simulate human explanatory behavior by transforming static figures into question-guided dynamic visual demonstrations. Specifically, FigAct employs a hierarchical search strategy to optimize element localization and trains the FigAct-8B model using task-specific rewards. Experimental results demonstrate that the proposed framework reduces token consumption by approximately 40-fold while generating clearer and more intuitive visual explanations. Consequently, FigAct significantly enhances both comprehension efficiency and interactive experience for scientific figure understanding.
📝 Abstract
Scientific figures are designed to communicate information visually, yet MLLMs typically explain them by translating their visual content back into text. This requires readers to manually map the resulting explanations back to the figure. Inspired by how people present visual information, we introduce FigAct, a framework that transforms static scientific figures into question-conditioned visual presentations by acting directly on their existing graphical elements. Like a human presenter, FigAct generates a sequence of short narrations, grounds each narration in the corresponding visual evidence, and applies visual actions to guide the viewer's attention. We develop a hierarchical search strategy for efficient element localization, reducing token usage by approximately 40$\times$. We further train FigAct-8B using three task-specific rewards for grounding accuracy, search efficiency, and rendering quality. We further build a human-verified benchmark from figures in real-world scientific papers to evaluate the ability of MLLMs to generate grounded visual explanations. Our results demonstrate the effectiveness of FigAct and show that treating scientific figures as presentation canvases makes explanations clearer and easier to follow.