Fine-Tuning VLM for Enhancing AI's Spatial Intelligence: Understanding 3D and 2D Rotations

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of vision-language models in spatial reasoning tasks, particularly 2D and 3D rotation recognition. By fine-tuning a Gemma-4 Mixture-of-Experts (MoE) architecture on a multi-object rotation dataset, this work enhances the model's spatial intelligence. The research reveals that salient linear features, rather than recognizable objects, are more effective in improving angle detection accuracy, and proposes a novel paradigm for optimizing 2D estimation without requiring explicit coordinate systems. Experimental results demonstrate that the fine-tuned MoE model significantly outperforms general-purpose baselines in predicting both rotation axes and angles, substantially advancing overall performance on spatial reasoning tasks.
📝 Abstract
Spatial intelligence is a fundamental skill in multiple domains, such as Science, Technology, Engineering, and Mathematics (STEM), Medicine, Architecture, and Construction. Recent studies indicate that Vision-Language Models (VLMs) still face limitations in spatial reasoning, which inhibits artificial intelligence (AI) from performing practical spatial tasks. Using multiple object-rotation datasets developed for training and evaluation, our experiments demonstrated promising improvements in both 2D and 3D rotation detection. Fine-tuned Google DeepMind-built Gemma-4 mixture-of-experts (MoE) models significantly outperformed fine-tuned Gemma-4 generalist models in predicting rotations defined by both their axes and angles. Fine-tuning also substantially improved angle estimation for 2D representation without requiring an explicit coordinate system. Furthermore, identifiable objects did not improve angle-detection accuracy; instead, objects with prominent linear features showed improved performance.
Problem

Research questions and friction points this paper is trying to address.

Spatial Intelligence
Vision-Language Models
Spatial Reasoning
3D Rotation
2D Rotation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Intelligence
Vision-Language Models
Mixture-of-Experts
Fine-Tuning
Rotation Detection
🔎 Similar Papers
No similar papers found.