Refinement Symmetry in Multimodal Transformers

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the non-conservation of content contributions during representation splitting in multimodal Transformers, caused by attention weights varying with token count. Drawing on refinement symmetry theory, this work proves that partition invariance necessitates a linear local mass factor, and accordingly proposes a metric-weighted mechanism that bounds attention error via physical coupling to suppress distribution drift during visual token duplication and merging while preserving contextual consistency. Experiments on Qwen2.5-Omni-7B and the MVBench benchmark validate the approach: three-fold token replication yields zero prediction change, and two-fold video token merging improves accuracy by 0.93%, significantly enhancing cross-partition robustness.
📝 Abstract
Attention weights depend on token counts, which change with the representation of a signal. We study refinement symmetry: splitting a representation while preserving content, position, visible context, and total mass should preserve its contribution. Building on proportional and quadrature attention, we show that split invariance forces the local mass factor to be linear for any fixed positive attention kernel, provided that factor is nondecreasing. For changed representations, a physical coupling bounds attention error by separating feature change from weight reallocation. In Qwen2.5-Omni-7B, duplicating half the visual tokens threefold changes 255 of 3,586 MVBench answers under standard attention; measure weighting preserves every answer under matched visibility. Under natural frame resampling, it reduces distributional drift. At twofold merging of a frozen video encoding, a five-seed evaluation shows an all-partition-correct accuracy gain of 1.04 percentage points over global count weighting (average group mass) and 0.93 points over standard attention. The advantage over global count also holds on WorldSense but depends on the compression budget. The result is a representation principle with a measured benefit in robustness across partitions.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Transformers
Attention weights
Refinement symmetry
Token representation
Robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Refinement Symmetry
Multimodal Transformers
Measure Weighting
Split Invariance
Attention Robustness
🔎 Similar Papers
No similar papers found.