🤖 AI Summary
This study addresses the challenge that responses of multimodal geometric alignment scores to modality degradation cannot be adequately explained by perturbation magnitude alone, which accounts for only a small fraction of variance. To overcome this limitation, we propose Directional Geometric Response (DGR) theory, departing from conventional scalar perspectives. By leveraging Gramian volume gradient projections and first-order Taylor expansions, DGR integrates operating points, magnitudes, and directions to precisely model geometric volume variations. The framework is validated through experiments employing frozen embeddings with controlled audio-visual noise injection. Our findings demonstrate that directional dependence constitutes the primary driver of multimodal geometric responses. DGR achieves out-of-sample R² values ranging from 0.838 to 0.969 and ranking accuracy exceeding 0.864, significantly outperforming direction-agnostic baselines.
📝 Abstract
Geometric alignment scores based on Gram determinants provide a compact way to model higher-order consistency among modalities, yet how such scores respond to modality degradation is poorly understood. This paper asks whether the response of a multimodal geometric score is determined primarily by the magnitude of the perturbation-induced displacement. Using frozen cohorts from MSR-VTT (N=878) and DiDeMo (N=980), we apply controlled video blur and audio noise and analyze the response in the relational geometry on which the score is defined. Displacement magnitude explains at most 15% of the out-of-sample variance in the absolute response, and magnitude-matched pairs respond systematically differently, so scalar magnitude does not organize the response. The closed-form first-order expansion of the Gramian volume yields the Directional Geometric Response (DGR): the projection of the displacement onto the local volume gradient, which jointly captures the clean operating point, displacement magnitude, and displacement direction. The absolute first-order DGR term explains the observed response with out-of-sample R^2 of 0.838-0.969, matched-magnitude ranking accuracies of 0.864-0.963, and response-sign accuracies of 0.909-0.989, whereas the tested direction-free alternatives remain weak or unstable under the corresponding evaluation protocols. A pre-specified gain-normalization candidate, V/(g_V+eps), fails its predictability and clean-order gates. DGR uses the observed degraded-state displacement and is therefore an explanatory quantity, not a deployment-time predictor: geometric response depends on where the representation operates, how far degradation moves the relational geometry, and in which direction it moves.