MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation

πŸ“… 2026-07-22
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenges of geometric discontinuities, semantic misalignment, and identity degradation commonly encountered in thermal-to-visible face translation. To overcome these issues, the authors propose a multimodal conditional latent diffusion framework that effectively integrates depth maps and text prompts as complementary priors. The approach employs a dual-branch cross-attention fusion module, a gated text–visual feature alignment mechanism, and spatial feature transformation to preserve identity information while ensuring coherent and semantically consistent synthesis. Evaluated on the MCXFace and SpeakingFaces datasets, the method substantially outperforms existing state-of-the-art techniques, achieving up to a 48.3% reduction in FID and an 8.9% improvement in Rank-1 identification accuracy.
πŸ“ Abstract
Thermal-to-visible face translation presents fundamental challenges including geometric discontinuities, semantic attribute mismatches, and identity degradation. We propose MTVDiff, a novel multimodal latent diffusion framework that synergistically integrates depth and textual information to address these limitations while preserving identity characteristics. The MTVDiff framework presents three core technical contributions: (1) a Dual-Branch Cross-Attention Fusion (DBCAF) module for multi-scale thermal-depth feature extraction and fusion; (2) a Gated Text-to-Visual Feature Alignment mechanism for semantically-guided generation; and (3) Spatial Feature Transformations (SFT) for adaptive multimodal prior integration. Extensive experiments on the MCXFace and SpeakingFaces datasets demonstrate that our multimodal approach significantly outperforms existing GAN-based and diffusion-based approaches across multiple metrics, achieving substantial improvements in both image quality and face verification performance, with FID reductions of up to 48.3% and Rank-1 accuracy improvements of up to 8.9\%. Our work provides a robust solution for face recognition systems operating under varying illumination conditions and advances the state-of-the-art in cross-spectral facial image translation through effective multimodal integration.
Problem

Research questions and friction points this paper is trying to address.

thermal-to-visible face translation
geometric discontinuities
semantic attribute mismatches
identity degradation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Diffusion
Thermal-to-Visible Translation
Cross-Attention Fusion
Text-to-Visual Alignment
Spatial Feature Transformation
πŸ”Ž Similar Papers
Z
Zhiyuan Xia
School of Computer Science and Engineering, Southeast University, Nanjing, China; Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China
H
Haojie Li
Monash University, Australia
J
Jingyu Lin
Monash University, Australia
Y
Yiguo Qiao
School of Computer Science and Engineering, Southeast University, Nanjing, China; Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China
Cunjian Chen
Cunjian Chen
Monash University
Generative AIComputer VisionDeep Learning