pixel-to-world decomposition

Designs and evaluates multi-stage models or pipelines that decompose the mapping between image pixels and world-coordinate representations into explicit substeps, including both forward and reverse (bidirectional) mappings. Builds per-step supervised mappings and consistency constraints that enforce joint agreement across coordinate frames and substeps to improve geometric and multi-task prediction quality.

pixel-to-worlddecomposition

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Geometric Constraints in Deep Learning Frameworks: A Survey

Mar 19, 2024
VK
Vibhas K Vats
🏛️ Indiana University Bloomington

This paper addresses the fragmented and unsystematic modeling of geometric constraints in deep learning by proposing the first unified taxonomy of geometric constraints tailored for modern deep learning frameworks. Methodologically, it systematically integrates multi-view geometry, epipolar constraints, camera calibration models, self-supervised geometric consistency losses, and differentiable rendering to establish a three-dimensional classification framework spanning modeling principles, integration strategies, and optimization objectives. The contributions are threefold: (1) clarifying the applicability boundaries and failure mechanisms of over one hundred geometric constraints across vision tasks such as depth estimation; (2) uncovering key design paradigms for synergistic co-design of geometric priors and neural architectures; and (3) identifying principled pathways to overcome three core challenges—dynamic scenes, textureless regions, and cross-domain generalization.

Compare geometry-enforcing constraints in deep learning for depth estimationPropose taxonomy for geometry constraints in modern deep learning frameworksSurvey geometry-inspired deep learning frameworks for vision tasks

Must-Read Papers

Most classic and influential ideas
View more

Existing 3D reconstruction methods struggle to balance generalization and practicality due to either inefficient per-scene optimization or reliance on category-specific training. This work proposes a feed-forward, output-representation-agnostic framework for 3D reconstruction, systematically addressing five core challenges: feature enhancement, geometry awareness, model efficiency, data augmentation, and temporal modeling. By unifying the analysis of image backbones, multi-view fusion mechanisms, and geometric priors—and integrating major datasets and evaluation benchmarks—it establishes a standardized benchmarking protocol. The study transcends differences in geometric representations, formulates a problem-driven, generalizable modeling paradigm, and outlines promising future directions in scalability, evaluation metrics, and world modeling.

3D scene representationcross-scene generalizationfeed-forward 3D reconstruction

Existing CAD learning approaches discretize B-Rep models into triangle meshes, thereby discarding the analytical surface representations and topological information essential for consistent instance-level analysis. This work proposes STEP-Parts, a deterministic pipeline that directly extracts geometric instance partitions from native STEP B-Rep data. The method defines partitions based on intrinsic B-Rep topology, merges faces using analytical surface types and near-tangent plane continuity criteria, and transfers labels to triangulated meshes via face-to-mesh correspondence mapping. STEP-Parts ensures boundary consistency across varying triangulations and processes the DeepCAD subset of the ABC dataset—comprising approximately 180,000 models—in under six hours. The resulting labels significantly enhance performance in implicit reconstruction-segmentation tasks and point cloud networks. Code and precomputed labels are publicly released.

Boundary RepresentationsCAD processinggeometric partitioning

The field of 3D vision suffers from fragmented data representations, learning paradigms, and benchmarking protocols, leading to a lack of unified understanding regarding efficiency, fidelity, and scalability. This work proposes the first cohesive conceptual framework that integrates geometric representations—such as point clouds, meshes, voxels, and 3D Gaussians—with diverse learning paradigms—including 2D-supervised learning, implicit neural representations, and 4D modeling—and connects them to real-world application scenarios. By constructing a structured knowledge graph of 3D vision, the study systematically relates dataset design, supervision mechanisms, and task requirements, clarifying the trade-offs between efficiency and fidelity and charting pathways for multimodal geometric grounding. This framework offers systematic guidance for reconstruction, generation, and dynamic scene modeling, advancing the field toward a unified and efficient paradigm.

3D visionbenchmark fragmentationdata representation

This work addresses the problem of efficiently merging multiple fine-tuned models into a unified multitask model without retraining. The authors formalize model merging as a convex quadratic program over residual updates, achieving theoretically optimal fusion by calibrating inputs and outputs to minimize calibration error in the output space. This study provides the first formal optimality guarantees for model merging, introduces an interpretable diagnostic metric based on residual energy, and unifies existing heuristic approaches within a single theoretical framework as special cases. Experimental results demonstrate that the proposed method matches or surpasses current techniques in single-layer settings and consistently improves performance across language and vision benchmarks in multilayer merging scenarios. Furthermore, the quality of merged models can be accurately predicted using a small calibration set.

fine-tuned modelsmodel mergingmulti-task learning

MOWA: Multiple-in-One Image Warping Model

Apr 16, 2024
KL
Kang Liao
🏛️ Nanyang Technological University

Existing image warping methods require task-specific model training, exhibiting poor generalization and limited adaptability to diverse camera models or custom distortions. To address this, we propose MOWA—a unified, multi-task image warping model capable of handling six distinct warping tasks within a single architecture. Our approach introduces three key innovations: (1) a region-pixel two-level motion disentanglement mechanism for enhanced geometric modeling accuracy; (2) a lightweight point-based classifier that dynamically generates task-aware prompts for conditional feature modulation; and (3) end-to-end differentiable warping networks jointly optimized with multi-scale motion estimation. Extensive experiments demonstrate that MOWA consistently outperforms dedicated state-of-the-art models across all six tasks. Moreover, it exhibits strong cross-camera generalization and zero-shot transfer capability, enabling robust adaptation without task-specific retraining.

Generalizes across different camera models and manipulationsSolves multiple image warping tasks with one modelUses multi-level motion estimation for accurate warping

Latest Papers

What's happening recently
View more

While existing multi-frame models achieve cross-frame consistency, their single-frame accuracy often lags behind that of single-frame methods. Through systematic ablation studies, this work demonstrates that data diversity and quality are critical for 3D geometry estimation and reveals that commonly used loss functions may inadvertently suppress performance. To address these issues, the authors propose CARVE, a novel approach integrating a high-resolution network architecture, joint sequence- and frame-level supervision, a consistency loss, and alignment between depth maps and camera parameters. CARVE achieves state-of-the-art and robust performance across multiple benchmarks in tasks including point cloud reconstruction, video depth estimation, and estimation of camera pose and intrinsics.

3D reconstructiondepth estimationmulti-frame consistency

This work addresses the high computational cost of high-fidelity 3D generation, which typically relies on large-scale data and models while underutilizing the rich semantic and structural priors embedded in discriminative 3D foundation models. To bridge this gap, the authors propose ROAD, a novel framework that, for the first time, transfers priors from discriminative 3D foundation models into a diffusion Transformer. ROAD introduces a reciprocal objective alignment mechanism that effectively reconciles the heterogeneity between generative and discriminative latent spaces through global semantic compression and optimal micro-structure matching—formulated as bipartite graph matching. Remarkably, without increasing inference overhead, ROAD achieves generation quality on par with the industrial baseline Step1X-3D using only 1.5% of the training data, substantially reducing both training cost and computational requirements.

3D shape generationcomputational costdiscriminative priors

This work addresses the challenges of missing correspondences and incomplete geometric information in point cloud reconstruction from partially observed multi-view inputs. The authors propose a training-free optimization method that jointly recovers the 3D point cloud and its cross-view projection mappings. Built upon an extended multi-view synchronized embedding framework, the approach integrates variable projection, geometric constraints, and visibility modeling, making it applicable to both fixed and variable projection settings without requiring category-specific priors. Experiments on ShapeNet and Pix3D demonstrate that the method robustly reconstructs partial multi-view point clouds, consistently outperforming existing non-learning baselines across Chamfer distance, Earth Mover’s Distance (EMD), and Reconstruction Overlap Accuracy (ROA) metrics.

3D point cloud reconstructioncross-view correspondencemissing points

Traditional visual navigation struggles to balance global geometric consistency with topological generalization, limiting its performance in complex environments. This work proposes a novel map representation based on pixel-level relative 3D connectivity, which constructs a pixel correspondence graph in a relative coordinate frame from image sequences and generates a “WayPixel Costmap” for planning and control. By preserving high-fidelity geometric information without requiring global geometric consistency, the approach overcomes the limitations of conventional topological graphs and dense reconstructions. Experimental results demonstrate that the method significantly outperforms image-level and object-level representations across four simulated tasks and real-world scenarios, validating its accuracy and practicality for visual navigation.

3D map representationgeometric consistencypixel-relative connectivity

This work addresses the limitations of existing panoramic stitching methods, which rely on pairwise feature matching and often fail to maintain multi-view geometric consistency in complex scenes characterized by weak textures, large disparities, or repetitive patterns, leading to misalignments and distortions. To overcome these challenges, the authors propose a photogrammetry-driven global alignment framework that leverages estimated camera poses to align images in 3D space. They introduce a novel 3D-aware Transformer architecture that explicitly models multi-view geometric consistency through joint feature optimization and cross-view information aggregation. Key contributions include the first formulation of multi-view consistency in 3D space, a Transformer-based 3D-aware stitching network, and the creation of the first large-scale real-world panoramic stitching dataset. Experiments demonstrate that the proposed method significantly outperforms state-of-the-art approaches in both alignment accuracy and visual quality, particularly exhibiting superior robustness and consistency in challenging scenarios.

3D photogrammetrygeometric consistencyimage distortion

Hot Scholars

ST

Siyu Tang

ETH Zürich
computer visionmachine learning
ZL

Zhengqin Li

Meta
Computer visioncomputer graphicsmachine learning
BZ

Bohan Zeng

PhD student, Peking University
Data-Centric AIComputer VisionDiffusion Model3D
WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
KJ

Kui Jiang

Harbin Institute of Technology
computer visionimage processingdeep learning