🤖 AI Summary
This study addresses the insufficient spatial consistency in multimodal binaural audio generation by extending the AudioX model, proposing a unified text/video-to-binaural audio generation framework based on conditional embeddings. Methodologically, it introduces a perceptual encoder for feature alignment and designs a condition-dependent low-rank transformation to optimize latent representations. The framework further incorporates LoRA fine-tuning, an ILTD supervision objective, and multimodal alignment strategies to enhance generation quality. Experimental results demonstrate that the proposed approach outperforms the ViSAGe baseline across multiple metrics, significantly improving SpatialCLAP scores and subjective listening quality. This work establishes an efficient new paradigm for cross-modal spatial audio synthesis.
📝 Abstract
We present PLACE, a method that extends the pretrained any-to-audio model AudioX for binaural generation from arbitrary combinations of text, video, and optional audio prompts. PLACE augments video conditioning with Perception Encoder Core features, aligns text and video representations to derive spatial cues, and applies a conditioning-dependent low-rank transformation to the generated latent. The adapter is supervised via decoded-audio interaural level and time difference objectives. Trained on MRSAudio, PLACE improves most metrics over ViSAGe on FAIR-Play and achieves an improved SpatialCLAP score over SpatialSonic on the BEWO-1M Single Static test split. Listener evaluations favor PLACE for video-to-audio and out-of-distribution text-to-audio generation, demonstrating flexible multimodal control and improved spatial consistency.