spatial-geometric reasoning

Design, build, or analyze models, representations, and algorithms that represent, infer, and manipulate spatial and geometric relations—including 2D/3D geometry, epipolar constraints, and multi‑hop relational chains—to answer queries, predict geometry‑dependent outcomes, or verify spatial constraints. This work covers structured and symbolic formulations as well as neuro‑symbolic and stepwise/latent‑step inference methods, and emphasizes geometry‑aware representations and transparent, explainable scene reasoning.

spatial-geometricreasoning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.21
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$216K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks

Oct 29, 2025
XZ
Xu Zheng
🏛️ HKUST(GZ) | South China University of Technology | INSAIT | Sofia University "St. Kliment Ohridski" | Shanghai Jiao Tong University | University of Pisa | University of Trento | HKUST

A systematic survey and open-source evaluation benchmark for spatial reasoning in large multimodal models (MLLMs) remain absent. Method: This work introduces the first comprehensive taxonomy of multimodal spatial reasoning tasks, covering 2D/3D scene understanding, spatial relation modeling, and embodied intelligence applications. We propose MM-SpatialBench—a scalable, modular, open-source evaluation platform—unifying diverse tasks including visual question answering, 3D localization, and navigation. The framework integrates MLLMs with post-training optimization, cross-modal interpretability analysis, and joint reasoning over multimodal sensor data (e.g., vision, audio, egocentric video). Contribution/Results: Experiments demonstrate significant improvements in model generalization and structured spatial reasoning capabilities on complex tasks. MM-SpatialBench establishes a standardized infrastructure for advancing multimodal spatial cognition research.

Examining spatial understanding through emerging modalitiesIntroducing open benchmarks for model evaluationReviewing multimodal spatial reasoning tasks with large models

Must-Read Papers

Most classic and influential ideas
View more

Towards Geometry Problem Solving in the Large Model Era: A Survey

Jun 03, 2025
YZ
Yurui Zhao
🏛️ National University of Defense Technology | Chinese University of Hong Kong

Geometric Problem Solving (GPS) has long suffered from automation bottlenecks due to its dual requirements of spatial understanding and rigorous logical reasoning, compounded by fragmented benchmarks, inconsistent evaluation protocols, and disjointed methodological approaches. To address these challenges, this work introduces the first unified 3D analytical framework for GPS tailored to the large-model era, spanning benchmark construction, multimodal (text-and-diagram) parsing, and reasoning paradigms. We propose an automated benchmark generation methodology and a novel interpretable neuro-symbolic reasoning approach that tightly integrates large language models, multimodal perception, symbolic reasoning, graph-structured modeling, and principled evaluation design. Our analysis systematically clarifies the field’s fragmentation, identifies fundamental technical bottlenecks, and delivers the first comprehensive roadmap for advancing geometric intelligence—enabling applications in education, computer-aided design (CAD), and computational geometry.

Automating geometry problem solving with AIIntegrating spatial and logical reasoning challengesUnifying fragmented methodologies and benchmarks

GeoThought: A Dataset for Enhancing Mathematical Geometry Reasoning in Vision-Language Models

Oct 23, 2025
NS
Nannan Shi
🏛️ Baidu Inc. | Institute of Information Engineering, Chinese Academy of Sciences | Intel Lab, Intel

Large language models (LLMs) excel at textual mathematical reasoning but exhibit substantial performance degradation on visual geometric reasoning tasks, primarily due to the challenges of image-based geometric understanding, multi-step spatial reasoning, and the scarcity of large-scale, diverse, and reasoning-annotated geometric datasets. Method: We introduce GeoThought—the first large-scale, diverse geometric reasoning dataset featuring explicit chain-of-thought and reflective reasoning steps, systematically covering hierarchical geometric reasoning processes. Our approach integrates vision-language description generation, multimodal large language model (MLLM) architecture, and error-correcting chain-of-thought training. Contribution/Results: The resulting GeoThought-MLLM achieves state-of-the-art performance on both in-domain and cross-domain geometric reasoning benchmarks. Error analysis demonstrates that explicit reflection mechanisms effectively mitigate conceptual misclassifications and spatial relationship misunderstandings, significantly enhancing geometric semantic comprehension.

Addressing limitations in existing geometry datasetsEnhancing geometric reasoning in vision-language modelsImproving multi-step visual reasoning with explicit chains

This work addresses the limitations of current large language models in geometric problem solving, which often rely on a single chain-of-thought and struggle to effectively integrate diagram understanding, symbolic manipulation, and multi-step logical reasoning. The authors propose MARS-GPS, a novel framework that introduces multi-chain-of-thought (Multi-CoT) voting with self-verification for geometric reasoning. It generates multiple parallel reasoning paths, validates numerical results via Python code execution, ranks paths using token-level entropy as a confidence measure, and aggregates answers through multi-stage voting. Evaluated on Geometry3K, the method achieves 88.8% accuracy—surpassing the previous state of the art by 11%—and demonstrates consistent performance gains of up to 6.0% as the number of parallel paths increases from 1 to 16, substantially overcoming the constraints of conventional neural or symbolic approaches.

Chain-of-ThoughtGeometric Problem SolvingLarge Language Models

Bridging Formal Language with Chain-of-Thought Reasoning to Geometry Problem Solving

Aug 12, 2025
TY
Tianyun Yang
🏛️ Shenzhen Research Institute of Big Data | The Chinese University of Hong Kong | Shenzhen International Center for Industrial and Applied Mathematics

Large Vision-Language Models (LVLMs) face critical limitations in Geometry Problem Solving (GPS): unreliable diagram understanding, opaque reasoning processes, and insufficient intermediate steps in formal program generation. To address these, we propose an interpretable reasoning framework that *interleaves* natural-language chain-of-thought reasoning with executable formal code generation, yielding progressive, verifiable reasoning paths. We further introduce a symbolic computation solver–guided reinforcement learning paradigm, integrated with supervised fine-tuning, to train a Qwen2.5-VL-7B model on a novel 11K-sample synthetic GPS dataset. Experiments demonstrate state-of-the-art performance on standard GPS benchmarks—achieving up to a 15% absolute accuracy gain—outperforming both same-scale and significantly larger models (e.g., Qwen2.5-VL-72B). Moreover, our generated reasoning traces are more concise and formally verifiable, enhancing transparency and trustworthiness.

Addressing unreliable diagram interpretation in vision language modelsEnhancing transparency and accuracy in solver-executable code generationImproving geometry problem solving via formal language and reasoning

Geometric Reasoning in the Embedding Space

Apr 02, 2025
JH
Jan Hrula
🏛️ Czech Technical University in Prague | University of Ostrava

This work investigates the geometric reasoning capabilities of Graph Neural Networks (GNNs) and Transformers in embedding space, focusing on reconstructing implicit 2D geometric structures—specifically, predicting spatial coordinates and recovering underlying shapes from point sets defined by discrete geometric constraints on a 2D grid. Method: We propose a geometry-aware GNN architecture explicitly designed for geometric reasoning. Crucially, it operates without explicit coordinate supervision, relying solely on relational graph structure. Contribution/Results: We demonstrate, for the first time, that the learned node embeddings spontaneously organize into a low-dimensional subspace preserving neighborhood relationships—effectively recovering the latent 2D grid topology. Quantitatively, our GNN significantly outperforms Transformer baselines in both prediction accuracy and scalability. Qualitative analysis confirms that the embedding space faithfully encodes geometric structure, providing strong evidence of implicit geometric modeling capacity. These findings establish a novel, interpretable paradigm for spatial reasoning grounded in learned embeddings.

Compare GNN and Transformer performanceLearn hidden figures in embedding spacePredict spatial positions from geometric constraints

Latest Papers

What's happening recently
View more

This work addresses the challenge of simulating human-like multi-step logical reasoning with auxiliary constructions in geometric problem solving by proposing a novel framework that integrates mathematical reasoning with procedural representations. The approach employs program code as an intermediate visual representation, decoupling discovery reasoning from code generation in a latent space and structuring the reasoning manifold through supervised fine-tuning. The study demonstrates that hierarchical syntactic code structures effectively encode rich mathematical semantics, offering greater expressiveness than purely visual representations. Experimental results show that the proposed method significantly enhances geometric reasoning performance while yielding clearer and more interpretable multi-step derivations.

geometry problem solvinginterleaved reasoningmathematical reasoning

Large language models lack native support for continuous spatial representations, hindering their capacity for genuine geometric reasoning. This work proposes the Spatial Language Model (SLM), which for the first time integrates learnable spatial representations as a first-class modality directly into the model’s reasoning process. By leveraging a multimodal architecture, atomic geometric operations, and training with spatial instruction alignment, SLM achieves a paradigm shift from symbolic matching to true geometric reasoning. Accompanied by a newly curated spatial instruction dataset and the SpatialEval benchmark, empirical evaluations demonstrate that SLM substantially outperforms existing approaches—whether based on prompt engineering or textual abstraction—across tasks involving spatial attributes, distance, topology, and relative positioning.

continuous spatial representationsgeometric reasoninglarge language models

Existing vision-language models lack verifiable intermediate states in geometric reasoning, making it difficult to ensure precise spatial relationships. This work proposes a "propose–draw–verify" iterative framework that externalizes geometric reasoning through agent-based interaction with the GeoGebra constraint engine. By explicitly expressing hypotheses on an executable canvas and obtaining structured feedback, the approach grounds reasoning in a shared state validated by algebraic constraints. The method enables independent auditing of construction fidelity and measurement faithfulness, achieving 95.9% predicate-level and 84.0% strict problem-level accuracy on the GeoGoal benchmark. It yields performance gains of up to 4.1% and 16.4% in planar and solid geometry tasks, respectively, and attains GenExam-math rendering scores of 68.2% (strict) and 90.5% (lenient).

constraint satisfactionexternalizationgeometry reasoning

Hot Scholars

MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
HT

Hao Tang

Peking University
computer vision
XL

Xiaodan Liang

Professor of Computer Science, Sun Yat-sen University, MBZUAI, CMU, NUS
Computer visionEmbodied AIMachine learning
FT

Federico Tombari

Google, TU Munich
Computer VisionMachine Learning3D Perception
VB

Valts Blukis

NVIDIA
RoboticsArtificial IntelligenceMachine LearningNatural Language Processing