Score
Designs and implements models and training pipelines that regress 3D pose values or offset vectors for individual nodes or keypoints in a structured representation, producing localized pose estimates or corrections. Also involves developing loss functions, architectures, and evaluation procedures to enable fine-grained pose adjustments and robustness to large pose offsets.
Existing 3D human pose estimation methods suffer from severe overfitting to single datasets, exhibiting poor generalization across viewpoints, environments, and camera configurations. Method: We introduce the first standardized cross-dataset evaluation framework—covering four major benchmarks with dynamic extensibility—and establish a fair, unified-interface evaluation paradigm. The framework features standardized data loading and preprocessing pipelines, dual-metric assessment (MPJPE and PA-MPJPE), and modular design for seamless integration of diverse models. Contribution/Results: Through systematic re-evaluation of 18 state-of-the-art methods (>100 new results), we quantitatively reveal that model performance degrades by 30–60% under cross-domain settings; moreover, preprocessing choices and data configuration critically govern generalization capability. This work shifts the field’s focus from single-dataset optimization toward robust modeling for real-world deployment.
This work addresses the geometric inaccuracies in sparse-view filament-based wireframe 3D printing caused by deformation. We propose a Gaussian alignment framework anchored on parametric curves, which constrains Gaussian kernels to these curves to yield a compact, geometry-aware wireframe representation that substantially reduces reconstruction ambiguity under sparse observations. By integrating differentiable rendering with neural deformation field estimation, our method achieves globally consistent deformation alignment and drives the co-evolution of a digital twin model. This enables dynamic updating of printing paths within a closed-loop adaptive control system, robustly compensating for deformations during robotic wireframe fabrication and significantly enhancing both manufacturing accuracy and adaptability for complex structures.
This work addresses the bottleneck in general object pose estimation—its reliance on hard-to-acquire CAD models—by investigating the feasibility of substituting them with image-reconstructed 3D models. To this end, we introduce the first pose-estimation-oriented 3D reconstruction quality benchmark, built upon the YCB-V dataset and featuring calibrated reconstructions aligned with ground-truth poses. The benchmark integrates geometric reconstruction pipelines (COLMAP for SfM/MVS), learning-based methods (PixelNeRF, iMAP), and the BOP evaluation framework. Key contributions are: (1) the first reconstruction quality benchmark explicitly designed for pose estimation; (2) empirical evidence that traditional geometric methods outperform learning-based approaches in the accuracy–speed trade-off; (3) discovery that standard reconstruction metrics (e.g., Chamfer distance) exhibit weak correlation with pose estimation accuracy; and (4) demonstration that most image-reconstructed models support high-accuracy pose estimation, albeit systematically underperforming CAD models. Code and benchmark are publicly released.
To address the demand for high-fidelity, diverse 3D asset generation and flexible editing, this paper introduces SLAT—a structured 3D implicit representation that jointly encodes sparse 3D mesh topology and multi-view visual foundation model features, enabling unified decoding into multiple 3D formats (e.g., radiance fields, 3D Gaussians, explicit meshes). Methodologically, SLAT pioneers a 2B-parameter Transformer architecture based on Rectified Flow for large-scale 3D latent-space modeling—the first of its kind. We curate a high-quality dataset of 500K 3D assets and perform end-to-end training. SLAT supports text- and image-conditioned generation, achieving state-of-the-art performance in fidelity, diversity, and editability. It enables real-time local 3D editing and dynamic output format switching. All code, models, and data are publicly released.
Existing pose transfer methods rely on target pose graph inputs and cannot directly generate human images aligned with semantic textual descriptions. This paper introduces the first end-to-end text-driven pose synthesis framework, enabling direct generation of high-fidelity pose-conditioned images from natural language. Our approach employs a three-stage decoupled modeling pipeline: (1) text encoding, (2) differentiable pose representation learning and optimization, and (3) neural rendering. We construct DF-PASS—the first fine-grained text-pose alignment dataset—built upon DeepFashion. The framework integrates a CLIP-based text encoder, implicit pose representation learning, and a differentiable pose optimization module. Extensive qualitative and quantitative evaluations demonstrate significant improvements over state-of-the-art baselines, validating both the effectiveness and practicality of text-to-pose image synthesis.
This work addresses the common trade-off in efficient camera pose estimation, where speed is often achieved at the expense of accuracy, and proposes a hybrid paradigm that integrates neural network-based initializations with classical Structure-from-Motion (SfM) optimization. The approach maintains high reconstruction accuracy while significantly reducing the number of required feature points. To systematically evaluate the efficacy of sparse matching and neural initialization in guiding bundle adjustment, the authors construct a novel SfM benchmark tailored for novel view synthesis. Experiments demonstrate that merely lowering feature density can accelerate conventional SfM pipelines, yet the combination of neural priors with traditional optimization achieves the best balance between efficiency and accuracy. The publicly released benchmark aims to advance research in high-precision, efficient SfM methods.
This work addresses the poor performance of conventional post-training quantization (PTQ) on 3D geometric models, which stems from complex feature distributions and high calibration costs. To overcome these limitations, the authors propose TAPTQ, a tail-aware PTQ framework that integrates progressive coarse-to-fine calibration, an efficient quantization range optimization based on ternary search, and a module-level compensation mechanism guided by Tail Relative Error (TRE). This approach simultaneously enhances quantization accuracy and substantially reduces calibration overhead. Experimental results demonstrate that TAPTQ significantly outperforms existing PTQ methods on both VGGT and Pi3 benchmarks, achieving superior accuracy with notably lower computational cost.
This study addresses the significant performance degradation of existing pose estimation algorithms in challenging scenarios such as trampoline gymnastics, which involve extreme body postures and unconventional camera viewpoints. To tackle this issue, the authors introduce STP, the first synthetic pose dataset specifically designed for trampoline gymnastics, created by fitting a parametric human body model to noisy motion capture data and rendering photorealistic multi-view images. Building upon this dataset, they fine-tune the ViTPose model and integrate multi-view 3D triangulation to achieve high-accuracy 2D and 3D pose estimation. Experimental results demonstrate that the proposed approach achieves state-of-the-art 2D pose accuracy on real-world trampoline sequences and reduces the 3D mean per-joint position error (MPJPE) by 12.5 mm, representing a 19.6% improvement over the original ViTPose model.