🤖 AI Summary
This study addresses the inherent non-injectivity of neural networks, which irreversibly discards input distinctions and creates a mismatch between information bottlenecks and downstream task requirements. To this end, we propose the Null Space Residual (NSR) framework, which performs null space analysis on linear operators to extract task-relevant information eliminated during mapping. By integrating member-level encoded gated attention mechanisms with residual connections, NSR enables adaptive feature fusion, thereby systematically compensating for information loss induced by non-injective mappings. Experimental results demonstrate that the proposed framework achieves up to a 31.51 mIoU improvement in token merging tasks and attains 100% training accuracy alongside significant performance gains in graph aggregation tasks.
📝 Abstract
Non-injective mappings in neural networks map distinct inputs to the same representation, thereby implicitly inducing equivalence relations in the input space. However, the input differences eliminated by these mappings may still be required by downstream tasks, creating a mismatch between operator-induced indistinguishability and task-required distinctions. For non-injective linear operators realized in the current forward pass, their null spaces exactly characterize these invisible input variations. We propose Task-Relevant Null-Space Residuals (NSR), a general residual framework for non-injective linear mappings. NSR combines null-space component extraction from pre-mapping representations, member-level encoding and gating, and application-specific integration to exploit potentially task-relevant information under downstream supervision while preserving the original aggregation or merging rules. We evaluate NSR in two structurally different settings: token merging and graph aggregation. In token merging, NSR achieves higher semantic segmentation performance than the corresponding compressed baselines in 34 out of 36 evaluated configurations, with a maximum observed gain of 31.51 mIoU points under strong compression. In graph aggregation, NSR achieves 100% training accuracy on Tree-NeighborsMatch at depths d=2--6 across three backbones, alongside gains on heterophilic node classification and molecular graph regression. Together, these results support null-space residuals as a practical complement to non-injective linear mappings, enabling downstream models to learn from input distinctions invisible in the original operator's output.