CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs

πŸ“… 2026-10-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the tendency of multimodal large language models to over-rely on textual priors due to sparse visual cues in geometric diagram understanding, resulting in misinterpretations that contradict visual evidence. To overcome this, we introduce a pioneering training-free visual intervention paradigm at inference time, proposing a criticality-driven visual intervention framework. Specifically, we develop Geometric Constraint Local Relation Reconstruction (GCLR) to reconstruct spatial relationships and an Adaptive Visual Steering Operator (AVSO) to identify critical layers for correcting attention distributions. Together, these components effectively suppress interference from textual priors while enhancing the weighting of visual evidence. Experimental results demonstrate that our method achieves F1 scores of 85.58 and 82.84 on the PGPS9K and PGDP5K datasets, respectively, significantly advancing geometric understanding capabilities.
πŸ“ Abstract
Despite significant progress in visual tasks by Multimodal Large Language Models (MLLMs), geometric diagram understanding remains challenging due to the presence of sparse visual cues and ambiguous symbol-primitive associations. MLLMs may therefore rely on textual priors, producing interpretations that conflict with visual evidence. We introduce the training-free Criticality-Driven Visual Intervention Framework (CVIF), an inference-time method that localizes critical layers and executes visual interventions during the transition from evidence aggregation to semantic decoding. At these layers, a Geometry-Constrained Local Relation Reconstruction (GCLR) module selects and weights vertex-centered visual evidence, while an Adaptive Visual Steering Operator (AVSO) redistributes attention mass toward the selected tokens. Experiments on PGPS9K and PGDP5K show that CVIF raises Overall F1 from 77.85 to 85.58 and from 75.23 to 82.84, respectively, establishing a novel inference-time visual intervention paradigm.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Geometric Diagram Understanding
Sparse Visual Cues
Symbol-Primitive Associations
Textual Priors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Criticality-Driven Visual Intervention
Geometry-Constrained Local Relation Reconstruction (GCLR)
Adaptive Visual Steering Operator (AVSO)
Inference-time Paradigm
Geometric Diagram Understanding
πŸ”Ž Similar Papers
J
Jiahui Kang
School of Computer Science and Technology, Xi’an Jiaotong University
B
Bifan Wei
School of Computer Science and Technology, Xi’an Jiaotong University
Lingling Zhang
Lingling Zhang
Assistant Professor, Xi'an Jiaotong University
Computer visionFew-shot learningZero-shot learning
Tianwen Jiang
Tianwen Jiang
Harbin Institute of Technology
Knowledge GraphInformation ExtractionNatural Language Processing
Q
Qiuyong Xiao
Tencent Hy AI Data
J
Jihong Zhang
Tencent Hy AI Data
J
Jun Liu
School of Computer Science and Technology, Xi’an Jiaotong University