GeoThought: A Dataset for Enhancing Mathematical Geometry Reasoning in Vision-Language Models

๐Ÿ“… 2025-10-23
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Large language models (LLMs) excel at textual mathematical reasoning but exhibit substantial performance degradation on visual geometric reasoning tasks, primarily due to the challenges of image-based geometric understanding, multi-step spatial reasoning, and the scarcity of large-scale, diverse, and reasoning-annotated geometric datasets. Method: We introduce GeoThoughtโ€”the first large-scale, diverse geometric reasoning dataset featuring explicit chain-of-thought and reflective reasoning steps, systematically covering hierarchical geometric reasoning processes. Our approach integrates vision-language description generation, multimodal large language model (MLLM) architecture, and error-correcting chain-of-thought training. Contribution/Results: The resulting GeoThought-MLLM achieves state-of-the-art performance on both in-domain and cross-domain geometric reasoning benchmarks. Error analysis demonstrates that explicit reflection mechanisms effectively mitigate conceptual misclassifications and spatial relationship misunderstandings, significantly enhancing geometric semantic comprehension.

Technology Category

Knowledge Representation and Reasoning: Geometric, Spatial, and Temporal ReasoningMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Large Vision Models

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
๐Ÿ“ Abstract
Large language models (LLMs) have demonstrated strong reasoning capabilities in text-based mathematical problem solving; however, when adapted to visual reasoning tasks, particularly geometric problem solving, their performance substantially declines because geometric problems present unique challenges. Specifically, these challenges stem from two key factors: first, the intrinsic complexity of geometry requiring detailed image comprehension and multi-step reasoning, and second, the limitations of existing datasets which lack sufficient scale, diversity, and explicit reasoning traces, consequently hindering effective model training. To address these challenges, we developed the GeoThoughts dataset, a comprehensive geometric reasoning corpus with two subsets: Geo-Thought-6K with 6,243 samples and its augmented version Geo-Thought-Augmented-10K containing 10,834 samples. Each entry includes visual descriptions, step-by-step solutions, explicit reasoning chains, reflection steps, and final answers. Using this dataset, we developed GeoThought-MLLM, a mathematical reasoning multimodal model that generates detailed thinking processes during problem-solving. Our model outperforms existing benchmarks in geometric tasks, demonstrating that training with our Chain-of-Thought dataset improves geometric reasoning capabilities across both in-domain and out-of-domain settings. Finally, we analyze failure cases and observe that errors primarily arise from incorrect interpretation of mathematical concepts or spatial misjudgment. By invoking CoT to correct these mistakes, the model produces correct answers.
Problem

Research questions and friction points this paper is trying to address.

Enhancing geometric reasoning in vision-language models
Addressing limitations in existing geometry datasets
Improving multi-step visual reasoning with explicit chains
Innovation

Methods, ideas, or system contributions that make the work stand out.

Developed GeoThought dataset for geometry reasoning
Created multimodal model generating step-by-step solutions
Used Chain-of-Thought training to improve performance
๐Ÿ”Ž Similar Papers
No similar papers found.
N
Nannan Shi
Baidu Inc., Beijing, China
C
Chuanyu Qin
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
S
Shipeng Song
Baidu Inc., Beijing, China
M
Man Luo
Intel Lab, Intel, Santa Clara, USA