PAIQ: Patch-Aligned Semantic Injection via Residual Rotation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent difficulty of existing visual encoders in simultaneously preserving semantic abstraction and spatial detail. To this end, it proposes the PAIQ framework, which introduces a novel rotation-based orthogonal residual injection mechanism that orthogonally fuses SigLIP semantic features into DINOv3 spatial representations. This mechanism is theoretically proven to enhance patch separability under similar semantics. By incorporating cross-encoder matching and joint source assignment techniques, the approach keeps the pretrained models frozen while exclusively fine-tuning the projection and fusion parameters. Extensive evaluations demonstrate that PAIQ significantly mitigates hallucinations in image captioning and visual question answering tasks, achieving an average improvement of approximately 2.9 points in correctness score over state-of-the-art baselines.
📝 Abstract
Language-aligned and self-supervised visual encoders offer complementary strengths in semantic abstraction and spatial detail. Harnessing this complementarity requires enriching local features while retaining distinctions between semantically related patches. We introduce PAIQ, a patch-aligned semantic injection framework that combines content-based cross-encoder matching with orthogonally constrained residual updates. Using DINOv3 patch features as the spatial base, PAIQ aggregates complementary SigLIP features through joint source allocation and injects the aggregate--base differences through a shared orthogonal transformation Q. This rotation adapts update directions while preserving residual norms and pairwise angles. For fixed projected features, we derive conditions for patch separability under similar semantic aggregates and show that rotation adds a nonnegative separation term over direct interpolation when the aggregate is shared. Only the projection and fusion parameters are trained; both visual encoders and the language model remain frozen, and fusion retains 196 visual tokens. Across diverse language backbones, PAIQ yields broad gains in judge-assessed correctness and reductions in hallucination severity over single-encoder interfaces on image description and visual question answering. On the 2B and 9B Qwen backbones, this compact interface outperforms the strongest evaluated fusion or token-compression baselines by about 2.9 correctness points on average.
Problem

Research questions and friction points this paper is trying to address.

visual encoder fusion
patch-level semantics
multimodal hallucination
visual question answering
complementary feature integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Patch-Aligned Semantic Injection
Residual Rotation
Orthogonal Transformation
Cross-Encoder Matching
Visual-Language Fusion
💼 Related Jobs
No related jobs found.
P
Pinze Ren
Tsinghua University
Y
Yuwei Zhang
Beijing University of Posts and Telecommunications
H
Hao Chen
Tsinghua University
L
Linghao Meng
National University of Singapore
Chang Li
Chang Li
Tsinghua University
AudioGenerative ModelsRepresentation Learning
Qiankun Li
Qiankun Li
Research Fellow@NTU, Ph.D.@USTC
MLLMAI4HealthComputer VisionPattern RecognitionTrustworthy AI