Face and Voice Cross-modal Association with Learning Convex Feature Embedding

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high false positive and false negative rates in cross-modal association between faces and voices, which stem from modality heterogeneity. To mitigate this issue, the authors propose a joint learning framework that integrates convex hull feature embedding with a cross-modal attention mechanism. By compactly aggregating cross-modal features of the same identity within a unified embedding space and incorporating deep metric learning, the method effectively narrows the semantic gap across modalities. Experimental results on the VoxCeleb dataset demonstrate that the proposed approach significantly outperforms state-of-the-art methods in cross-modal verification, matching, and retrieval tasks, achieving substantial reductions in error rates.
📝 Abstract
Face-and-voice association learning is one of the most challenging tasks in deep learning. In this paper, we propose a simple but powerful cross-modal feature embedding method for the association of faces and voices. Previous work has studied cross-modal association tasks to establish the correlation between voice clips and facial images. These works have addressed cross-modal discrimination but underestimate the importance of handling heterogeneity in inter-modal features between audio and video, resulting in a lot of false positives and false negatives. To tackle the problem, the proposed method learns the embeddings of cross-modal features by making another feature exist between cross-modal features, facilitating the voice and face features of the same person to be embedded in a convex hull. Moreover, the incorporation of cross-modal attention mechanisms with convex embedding techniques represents a highly effective strategy for the attenuation of false positives and false negatives, accomplished via the minimization of inter-class discrepancies. We exhaustively evaluated our method for cross-modal verification, matching, and retrieval tasks on the large-scale VoxCeleb dataset. Extensive experimental results demonstrate that the proposed method achieves notable improvements over existing state-of-the-art methods.
Problem

Research questions and friction points this paper is trying to address.

cross-modal association
face-voice matching
feature heterogeneity
false positives
false negatives
Innovation

Methods, ideas, or system contributions that make the work stand out.

convex feature embedding
cross-modal association
face-voice matching
cross-modal attention
heterogeneous modality
🔎 Similar Papers
No similar papers found.