C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the tendency of multimodal large models to over-rely on dominant modalities during cross-modal reasoning, which undermines their compositional and counterfactual reasoning capabilities. To systematically diagnose the root causes of such failures, the authors introduce the C³PO benchmark, comprising 3,404 multimodal samples spanning video, audio, image, and text, automatically generated via a pipeline grounded in 25 logical templates. The benchmark incorporates paired information-consistent and information-conflicting (IC/CC) structures and a four-tier evaluation framework. Experiments reveal a significant performance gap: human accuracy reaches 88.64%, while the best model achieves only 73.17%. Analysis shows that 86–95% of model failures stem from textual modality dominance and excessive attention concentration. Notably, entropy in mid-level cross-modal attention effectively predicts reasoning success, indicating that modality structural roles—not mere compositional strategies—primarily determine performance.
📝 Abstract
Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a dominant modality, resulting in brittle cross-modal reasoning. We introduce C$^3$PO, a benchmark of 3,404 samples spanning video, audio, image, and text, evaluating two abilities: information composition (fusing dispersed evidence) and counterfactual conflict (resolving deliberate contradictions). C$^3$PO's paired IC/CC structure and four-tier design enable targeted diagnosis of when and why cross-modal reasoning fails. Built through a fully automatic pipeline using 25 logically grounded templates, C$^3$PO reveals that while humans achieve 88.64% accuracy, the best model (Gemini-3.1-Pro) reaches only 73.17%, with open-source models collapsing under conflict. Through attention probes, we find 86-95% of failures stem from modality dominance: models commit to one modality while ignoring contradictory evidence, concentrating 87-95% of attention on text. Mid-layer attention entropy predicts correctness-sustained exploration succeeds, premature collapse fails. The 56-point accuracy gap between equally complex templates reveals that performance depends on modalities' structural roles in conflict resolution, not combinations. These findings show multimodal perception does not guarantee robust reasoning; architectures must enable sustained cross-modal attention to avoid premature
Problem

Research questions and friction points this paper is trying to address.

cross-modal reasoning
modality dominance
multimodal conflict
information composition
counterfactual reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-modal reasoning
counterfactual conflict
modality dominance
attention entropy
multimodal benchmark
🔎 Similar Papers