π€ AI Summary
This work addresses the challenges of geometric discontinuities, semantic misalignment, and identity degradation commonly encountered in thermal-to-visible face translation. To overcome these issues, the authors propose a multimodal conditional latent diffusion framework that effectively integrates depth maps and text prompts as complementary priors. The approach employs a dual-branch cross-attention fusion module, a gated textβvisual feature alignment mechanism, and spatial feature transformation to preserve identity information while ensuring coherent and semantically consistent synthesis. Evaluated on the MCXFace and SpeakingFaces datasets, the method substantially outperforms existing state-of-the-art techniques, achieving up to a 48.3% reduction in FID and an 8.9% improvement in Rank-1 identification accuracy.
π Abstract
Thermal-to-visible face translation presents fundamental challenges including geometric discontinuities, semantic attribute mismatches, and identity degradation. We propose MTVDiff, a novel multimodal latent diffusion framework that synergistically integrates depth and textual information to address these limitations while preserving identity characteristics. The MTVDiff framework presents three core technical contributions: (1) a Dual-Branch Cross-Attention Fusion (DBCAF) module for multi-scale thermal-depth feature extraction and fusion; (2) a Gated Text-to-Visual Feature Alignment mechanism for semantically-guided generation; and (3) Spatial Feature Transformations (SFT) for adaptive multimodal prior integration. Extensive experiments on the MCXFace and SpeakingFaces datasets demonstrate that our multimodal approach significantly outperforms existing GAN-based and diffusion-based approaches across multiple metrics, achieving substantial improvements in both image quality and face verification performance, with FID reductions of up to 48.3% and Rank-1 accuracy improvements of up to 8.9\%. Our work provides a robust solution for face recognition systems operating under varying illumination conditions and advances the state-of-the-art in cross-spectral facial image translation through effective multimodal integration.