Spatio-Semantic Expert Routing Architecture with Mixture-of-Experts for Referring Image Segmentation

📅 2026-03-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing referring expression segmentation methods employ a uniform refinement strategy that struggles to balance diverse semantic and spatial reasoning demands, often resulting in fragmented outputs, ambiguous boundaries, or misidentified targets—particularly when the pretrained backbone is frozen. To address this, this work proposes SERA, a novel architecture featuring a two-stage, lightweight, expression-aware expert refinement mechanism. SERA integrates an expression-conditioned adapter (SERA-Adapter) into the backbone to enhance spatial consistency and boundary precision, and introduces a geometry-preserving expert transformation (SERA-Fusion) applied to visual features prior to multimodal fusion. By combining mixture-of-experts with a lightweight routing scheme, SERA achieves substantial improvements in spatial localization and boundary delineation while fine-tuning fewer than 1% of parameters—limited to normalization layers and bias terms—consistently outperforming strong baselines on standard benchmarks, especially in high-precision spatial reference tasks.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Mixture of Experts (MoE)Knowledge Representation and Reasoning: Geometric, Spatial, and Temporal Reasoning

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Referring image segmentation aims to produce a pixel-level mask for the image region described by a natural-language expression. Although pretrained vision-language models have improved semantic grounding, many existing methods still rely on uniform refinement strategies that do not fully match the diverse reasoning requirements of referring expressions. Because of this mismatch, predictions often contain fragmented regions, inaccurate boundaries, or even the wrong object, especially when pretrained backbones are frozen for computational efficiency. To address these limitations, we propose SERA, a Spatio-Semantic Expert Routing Architecture for referring image segmentation. SERA introduces lightweight, expression-aware expert refinement at two complementary stages within a vision-language framework. First, we design SERA-Adapter, which inserts an expression-conditioned adapter into selected backbone blocks to improve spatial coherence and boundary precision through expert-guided refinement and cross-modal attention. We then introduce SERA-Fusion, which strengthens intermediate visual representations by reshaping token features into spatial grids and applying geometry-preserving expert transformations before multimodal interaction. In addition, a lightweight routing mechanism adaptively weights expert contributions while remaining compatible with pretrained representations. To make this routing stable under frozen encoders, SERA uses a parameter-efficient tuning strategy that updates only normalization and bias terms, affecting less than 1% of the backbone parameters. Experiments on standard referring image segmentation benchmarks show that SERA consistently outperforms strong baselines, with especially clear gains on expressions that require accurate spatial localization and precise boundary delineation.
Problem

Research questions and friction points this paper is trying to address.

referring image segmentation
spatio-semantic reasoning
frozen pretrained models
boundary precision
expression grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Referring Image Segmentation
Expert Routing
Parameter-Efficient Tuning
Vision-Language Alignment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Alaa Dalaq
King Fahd University of Petroleum and Mineral, Dhahran, 31261, Saudi Arabia
M
Muzammil Behzad
King Fahd University of Petroleum and Mineral, Dhahran, 31261, Saudi Arabia; SDAIA-KFUPM Joint Research Center for Artificial Intelligence, Dhahran, 31261, Saudi Arabia