Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of end-to-end accuracy in compositional tasks for multimodal large language models, which fails to disentangle intrinsic capability deficits from upstream cascading errors. To this end, we propose a probabilistic causal decomposition framework that isolates these two failure modes through controlled interventions. We introduce the N-Score and S-Score metrics to quantify the necessity and sufficiency of prerequisite dependencies, respectively, and construct the CADET benchmark—spanning perceptual, spatial, and other categories—for systematic evaluation. Experimental results demonstrate that providing correct prerequisites eliminates 54% of errors in cognitive tasks, while a single critical prerequisite captures 84% of the overall performance gain. These findings effectively reveal latent patterns of systemic failures obscured by conventional end-to-end evaluations.
📝 Abstract
End-to-end accuracy on compositional tasks records how often MLLMs fail, but cannot distinguish whether a failure reflects an intrinsic deficit in the targeted capability or a cascading error from an upstream prerequisite. We propose a causal decomposition framework that isolates these two failure modes through controlled interventions on the prerequisite dependencies of each task. Our capability metrics (NC, IC, RC) score each task under unassisted, correct, or incorrect prerequisites to diagnose where failures arise; contribution metrics (N-Score, S-Score), adapted from probabilities of causation, quantify each prerequisite's necessity and sufficiency to determine why. We instantiate the framework in CADET, a diagnostic benchmark of 10 composite tasks decomposed into 46 unit tasks with over 33,000 human-annotated questions spanning perception, spatial, temporal, and cognitive categories. Diagnosing frontier MLLMs with our framework uncovers systematic patterns that end-to-end accuracy obscures. Capability-wise, supplying correct prerequisites eliminates 54\% of errors on cognitive tasks, lifting them from weakest to above spatial and temporal. Prerequisite-wise, causal contributions are concentrated in a few critical prerequisites, and supplying the single most important one alone captures 84\% of the gain from supplying all prerequisites.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Compositional Tasks
Failure Diagnosis
Cascading Errors
Capability Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Task Decomposition
Failure Diagnosis
Multimodal Large Language Models
Probabilities of Causation
CADET Benchmark