🤖 AI Summary
This study addresses the challenges of disconnected links, content hallucination, and layout distortion when adapting machine learning paper flowcharts across different canvas dimensions. To overcome these issues, this work proposes a three-stage multi-agent pipeline encompassing parsing, styling, and layout. The method introduces a novel explicit connection verification mechanism to prevent silent link disconnections, integrating vision-language model (VLM) visual feedback with deterministic constraint checking to output editable draw.io XML for high-fidelity cross-scale reflow. Benchmark evaluations demonstrate that the proposed approach achieves a content fidelity of 68.6%, substantially outperforming existing methods, which range from 11.2% to 41.4%. These performance gains are further corroborated by human evaluation, confirming the effectiveness of the pipeline in preserving both structural integrity and semantic accuracy during flowchart adaptation.
📝 Abstract
Pipeline figures in ML papers must be repurposed across many canvases, including paper columns, 16:9 slides, portrait posters, 1:1 social teasers, 9:16 phone previews. Each format imposes a different aspect ratio on the same computational graph, where any silently broken connection misrepresents the method. We formulate aspect-ratio-adaptive flowchart relayout as a distinct task: given a raster flowchart and a target ratio, produce a structurally faithful, hallucination-free, editable layout. Existing methods fail characteristically: image-to-image models stretch blocks and reject extreme ratios, text-to-image agentic systems hallucinate content, and parse-then-render systems mis-route edges. We propose an agentic pipeline factored into Parse, Style, and Layout stages, each pairing a main agent with a critic that combines deterministic constraint checks with VLM visual feedback so connectivity is explicitly checked and prevented from being silently broken. Outputs are draw.io-editable mxGraph XML. On a curated benchmark of 100 flowcharts at five aspect ratios, evaluated by Gemini 3.1 Pro and validated against human judgments, our method reaches 68.6% Content Fidelity versus 11.2-41.4% for prior work. Project page: https://onefigureeverycanvas.vercel.app/