projective geometry

Using geometric representations and algebraic relations of cameras and scenes (rays, homographies, projective transforms) to express and manipulate viewpoints, disentangle pose parameters, and synthesize training supervision for planar alignment and related tasks.

projectivegeometry

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey

May 20, 2025
SC
Seunghyuk Cho
🏛️ POSTECH | Australian National University

Planar Geometric Problem Solving (PGPS) serves as a critical benchmark for evaluating geometric reasoning in multimodal large language models, yet no systematic survey exists. This paper introduces the first unified classification framework for PGPS methods—grounded in the encoder-decoder paradigm—and systematically analyzes state-of-the-art approaches across three dimensions: model architecture, output format, and benchmark design. Our analysis identifies two fundamental challenges: (1) visual-symbol mapping hallucination during encoding, and (2) data leakage risks inherent in current benchmarks. Through multimodal reasoning diagnostics, architectural abstraction, and benchmark vulnerability assessment, we clarify key technical bottlenecks and propose concrete future directions: scalable encoding schemes, leakage-resistant benchmark construction, and formal verification protocols. This survey provides both theoretical foundations and practical guidance for advancing geometric reasoning in vision-language models.

Addressing hallucination and data leakage issues in PGPS benchmarksAssessing multi-modal reasoning in vision-language models via PGPSLack of comprehensive survey on recent PGPS research advancements

Must-Read Papers

Most classic and influential ideas
View more

Decoupled Geometric Parameterization and its Application in Deep Homography Estimation

May 22, 2025
YH
Yao Huang
🏛️ Donghua University | Zhejiang University | Czech Technical Univeresity in Prague | Huawei Technologies | Shanghai Jiao Tong University | Chinese Academy of Sciences

Planar homographies possess eight degrees of freedom, yet conventional four-corner offset parameterizations lack geometric interpretability and require solving an 8×9 linear system to recover the homography matrix. To address this, we propose a decoupled geometric parameterization based on the Similarity–Kernel–Similarity (SKS) decomposition, explicitly factoring the homography into two orthogonal four-dimensional parameter groups: similarity transformations and kernel transformations. Crucially, we establish, for the first time, an analytical linear mapping between kernel parameters and angular offsets, enabling direct, closed-form generation of the homography matrix without linear system solving. Evaluated on deep homography estimation tasks, our method achieves accuracy comparable to four-corner regression while significantly enhancing parameter interpretability and inference efficiency. This work introduces a novel paradigm for homography modeling that unifies geometric meaning with computational advantages.

Improving direct homography estimation via decoupled geometric parametersLack of geometric interpretability in homography parameterizationNeed for solving linear systems to compute homography matrix

This work addresses the limitation of existing vision-language-action (VLA) models that disregard known camera geometry in multi-camera setups, leading to visual representations misaligned with the true 3D space. To resolve this, the authors propose a camera-aware geometric module that injects calibrated geometric information into the visual token stream without altering the pre-trained VLA action space. The approach leverages intrinsic-conditioned ray embeddings, Projected Positional Encoding (PRoPE), and a bidirectional cross-view fusion mechanism. Notably, it requires neither depth sensors nor manual annotations, instead utilizing confidence-gated geometric supervision derived from a π³X teacher model. The method consistently improves performance across LIBERO, RoboCasa24, RoboTwin2.0, and real-world robotic platforms, with particularly pronounced gains on tasks sensitive to spatial reasoning and object relationships.

Geometric Inductive BiasMulti-camera GeometryRobot Manipulation

This work addresses the limitations of existing feedforward view synthesis methods, which rely on Plücker ray representations that are highly sensitive to camera coordinate systems, resulting in poor cross-view geometric consistency. To overcome this, the authors propose a projection-conditioning strategy that replaces raw ray inputs with 2D projection cues from the target view, effectively reformulating the task as a stable image-to-image translation problem. A tailored masked autoencoder pretraining mechanism is introduced to leverage large-scale uncalibrated data under this new conditioning paradigm. The proposed approach significantly enhances model robustness and view consistency, achieving state-of-the-art performance across multiple novel view synthesis benchmarks. Notably, it outperforms ray-based baselines by a clear margin on geometric consistency metrics, demonstrating the effectiveness of decoupling geometry representation from explicit ray parameterization.

camera encodingcoordinate sensitivitygeometric consistency

MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision

Oct 24, 2024
RW
Ruicheng Wang
🏛️ USTC | Microsoft Research | Harvard | Tsinghua University

This work addresses the problem of open-domain single-image 3D geometric reconstruction. To resolve global scale and translation ambiguities, we propose an affine-invariant 3D point cloud representation. Methodologically, we design an optimal point cloud alignment solver and a multi-scale local geometric consistency loss to mitigate the inherent ambiguity of monocular geometric supervision. Our approach integrates affine-invariant representation learning, robust point cloud registration, and end-to-end training on a hybrid large-scale dataset. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple unseen benchmarks. It significantly improves accuracy and generalization in monocular 3D point cloud reconstruction, depth estimation, and field-of-view prediction. By eliminating the need for camera calibration or explicit metric priors, our framework establishes a new paradigm for uncalibrated single-image geometric understanding.

Enhancing geometry learning with novel global and local supervisionsPredicting affine-invariant 3D point maps without global scale ambiguityRecovering 3D geometry from monocular open-domain images

Traditional 3D scene editing suffers from heavy reliance on manual object repositioning, expert modeling, and extensive annotated data. To address this, we propose a natural language–driven zero-shot 3D editing framework. Our method formalizes spatial semantics using conformal geometric algebra (CGA), which is integrated as an interpretable, verifiable semantic mapping language within the reasoning chain of a large language model (LLM). The framework jointly leverages CGA-based geometric representation, zero-shot LLM instruction parsing, real-time 3D simulation, and standard graphics pipeline interfaces—requiring no domain-specific fine-tuning or human modeling intervention. Experiments demonstrate that our approach reduces system response latency by 16% and improves task success rate by 9.6% over Euclidean-space baselines; notably, it achieves a 100% perfect execution rate on typical practical queries.

Integrates LLMs with CGA for 3D scene editingReduces manual effort in object repositioning tasksTranslates natural language to precise spatial transformations

Latest Papers

What's happening recently
View more

The field of 3D vision suffers from fragmented data representations, learning paradigms, and benchmarking protocols, leading to a lack of unified understanding regarding efficiency, fidelity, and scalability. This work proposes the first cohesive conceptual framework that integrates geometric representations—such as point clouds, meshes, voxels, and 3D Gaussians—with diverse learning paradigms—including 2D-supervised learning, implicit neural representations, and 4D modeling—and connects them to real-world application scenarios. By constructing a structured knowledge graph of 3D vision, the study systematically relates dataset design, supervision mechanisms, and task requirements, clarifying the trade-offs between efficiency and fidelity and charting pathways for multimodal geometric grounding. This framework offers systematic guidance for reconstruction, generation, and dynamic scene modeling, advancing the field toward a unified and efficient paradigm.

3D visionbenchmark fragmentationdata representation

This study addresses the challenge of achieving high-precision camera guidance and alignment for multiple rectangular planar regions under extremely limited annotation—requiring only a single labeled image. To this end, the authors propose a geometry-centric intra-image navigation framework that leverages homography as the central organizing variable to unify modeling, alignment, and evaluation. The method integrates intra-image augmentation to generate synthetic training data and employs a two-stage inference mechanism—comprising global detection followed by local refinement—alongside a Stable Warp training strategy. This approach substantially improves alignment accuracy even with low-resolution inputs and enables sparse keypoint localization together with sample-level confidence estimation. The work establishes a robust foundation for geometry-driven camera guidance and self-supervised learning in unconstrained video settings.

camera guidancegeometric alignmenthomographic navigation

Existing robotic manipulation approaches rely on vision-language or video models whose representations are constrained by semantic priors or 2D assumptions, limiting their ability to accurately capture the 3D geometric structures essential for physical interaction. This work proposes the Vision-Geometry-Action (VGA) model, which, for the first time, leverages a pretrained 3D world model as its backbone to directly map visual inputs to geometric actions, bypassing indirect modeling pathways based on language or 2D videos. By incorporating a progressive voxel modulation module and a joint training strategy, VGA outperforms state-of-the-art baselines such as π₀.₅ and GeoVLA in simulation and demonstrates exceptional zero-shot viewpoint generalization in real-world settings, significantly surpassing current methods.

3D RepresentationGeneralizable ControlGeometric Consistency

Existing text-to-image generation models struggle to precisely control camera viewpoints through natural language. This work proposes a parameterized viewpoint token that enables viewpoint-conditioned image synthesis by jointly fine-tuning diffusion models with geometric supervision, 3D rendering, and photorealism-enhanced data. The proposed viewpoint representation disentangles geometry from appearance, generalizes to unseen object categories, and explicitly models 3D camera structure within the text-to-visual latent space. Experimental results demonstrate that the method significantly improves the accuracy of camera viewpoint control while preserving high image quality and prompt fidelity, achieving state-of-the-art performance.

3D camera structurecamera controlgeometric representation

Existing methods for 3D scene graph generation struggle to distinguish between directional and viewpoint-invariant relationships, leading to inaccurate relation predictions under viewpoint variations. This work proposes the Transformation-Aware Disentanglement (TAD) framework, which explicitly decomposes relation reasoning into a viewpoint-stable branch and a direction-sensitive branch based on the transformation properties of predicates, and fuses both for multi-label predicate prediction. TAD introduces viewpoint-invariant object representations, transformation-aware relational descriptors, and group-aware auxiliary supervision, enabling robust 3D scene graph generation without relying on rotation-based data augmentation. Evaluated on the 3DSSG dataset, TAD significantly outperforms existing approaches under viewpoint perturbations while maintaining state-of-the-art performance on standard benchmarks.

3D Scene Graph Generationpredicate heterogeneityrelation transformation

Hot Scholars

LG

Luca Giuzzi

Associate Professor, Università di Brescia
Incidence GeometryPolar SpacesGrassmann spacesCoding theory
PS

Paolo Santonastaso

PhD student, Università degli studi della Campania "Luigi Vanvitelli"
geometria finita
CK

Chaya Keller

School of Computer Science, Ariel University
CombinatoricsDiscrete and Computational Geometry
BA

Boris Aronov

Professor of Computer Science, Tandon School of Engineering, New York University
Computational GeometryCombinatorial GeometryDiscrete GeometryAlgorithms