π€ AI Summary
This study addresses the instability of high-frequency details and cross-view appearance inconsistencies in head avatar reconstruction by proposing a unified optimization framework. The core innovations include a novel frequency-aware progressive optimization strategy that leverages curriculum learning to achieve stable coarse-to-fine reconstruction, and a canonical space consensus mechanism that eliminates the need for additional rendering. By integrating geometric visibility filtering with feature aggregation, this mechanism establishes efficient cross-view consistency constraints. Experimental results on the NeRSemble dataset demonstrate that the proposed method significantly improves both reconstruction quality and multi-view consistency, validating its generalizability and transferability.
π Abstract
We propose Fresco++, a unified optimization framework for fine-grained and view-consistent head avatar reconstruction. Head avatar optimization is typically driven by per-view image supervision, which can lead to premature fitting of unstable high-frequency details and inconsistent local appearance across viewpoints. Fresco++ addresses these challenges by regulating both the progression of visual detail and the formation of cross-view supervision during optimization. For frequency-aware optimization, a progressive curriculum first stabilizes low-frequency structures and then introduces high-frequency constraints to recover fine facial and hair details without amplifying spurious responses at early stages. For cross-view optimization, we introduce Canonical Group Consensus, which associates local observations through shared canonical surface regions and establishes correspondence across different viewpoints. Geometric and visibility-aware screening removes unreliable observations, while the remaining multi-view evidence is aggregated in feature space to form a consensus target for supervising the current rendering. This design enforces local consistency without relying on a specific image-space parameterization and avoids additional rendering of the auxiliary view. Together, the frequency curriculum and canonical consensus provide stable optimization from coarse structures to fine details while maintaining coherent appearance across viewpoints. Extensive experiments on NeRSemble demonstrate improved reconstruction quality and cross-view consistency, while evaluations across diverse avatar representations further confirm the generality and transferability of Fresco++.