TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of systematic evaluation of self-consistency between generation and understanding capabilities in existing unified audio models. We propose TORUS, the first benchmark framework specifically designed for evaluating self-consistency in native unified audio models, encompassing five task families across speech, sound, and music. TORUS introduces 48 three-stage tests comprising 432 six-option questions to establish a cross-modal consistency benchmark spanning audio generation, editing, and comprehension. The framework integrates multi-stage question answering and cross-task combinations, alongside a cascaded baseline that fuses state-of-the-art specialized models for comparison. Experimental results reveal that the best unified model achieves only 50.5% self-consistency accuracy—significantly below the cascaded baseline’s 63.2% (against a random chance level of 16.7%)—highlighting substantial deficiencies in tasks such as audio editing.
📝 Abstract
Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered: do the two heads of a unified model agree about the same audio? Current practice evaluates each capability in isolation on specialized benchmarks, and never asks whether a model can make sense of its own generations. We present TORUS, the first self-coherence test for audio-native unified models. TORUS comprises 48 three-stage self-coherence tests carrying 432 six-option questions spanning speech, sound and music across five task families. We holistically evaluate five open unified models alongside a Cascaded Baseline that combines state-of-the-art specialized generation, editing and understanding models. The best unified model answers 50.5% of questions against the Cascaded Baseline's 63.2% and a 16.7% chance floor. Models struggle on audio editing. Among the evaluated audio models (specialized and unified), we observe limited self-coherence, and thus position self-coherence as an essential test for future audio systems.
Problem

Research questions and friction points this paper is trying to address.

self-coherence
unified audio models
audio understanding
audio generation
audio editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-coherence
unified audio models
audio generation
audio understanding
TORUS benchmark
🔎 Similar Papers