query-based segmentation

Design and implement transformer-based segmentation architectures that use learnable query tokens to predict per-query instance- or class-aligned masks, including decoder heads and mask embeddings as in mask transformers and Mask2Former. Build training and evaluation pipelines to align queries to targets, incorporate global context for disambiguation, and analyze mask prediction quality and metrics (e.g., mIoU) relative to transformer baselines.

query-basedsegmentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.61
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work proposes TokenMask, a novel segmentation framework that departs from conventional query-based Vision Transformer approaches which rely on explicit reconstruction of image-space feature maps—a process that incurs substantial computational redundancy and hinders deployment. Instead, TokenMask operates entirely in the query token space, generating mask logits directly through token affinity and performing interpolation in logit space. By integrating a ViT backbone, a token-space mask head, and TensorRT FP16 inference, the method significantly reduces both computational and memory overhead across multiple datasets and segmentation tasks while preserving accuracy. Notably, it achieves substantial acceleration on the Jetson AGX Orin platform, offering an efficient and streamlined architecture well-suited for embedded vision applications.

efficient segmentationembedded visionmask prediction

Rethinking Decoders for Transformer-based Semantic Segmentation: A Compression Perspective

Nov 05, 2024
QW
Qishuai Wen
🏛️ Beijing University of Posts and Telecommunications

To address the limited interpretability, low computational efficiency, and insufficient theoretical foundations of Transformers in semantic segmentation, this paper proposes DEPICT—a novel decoder grounded in principal component analysis (PCA). DEPICT is the first to formulate semantic segmentation as a low-rank reconstruction problem, enabling a fully attention-based architecture with theoretically justified design principles. It comprises three key components: embedding-refinement self-attention, class-aware cross-attention, and dot-product mask generation—collectively endowing both self- and cross-attention mechanisms with explicit, semantically meaningful principal component interpretations. This yields a transparent, interpretable “white-box” decoder. On ADE20K, DEPICT surpasses the black-box baseline Segmenter despite using fewer parameters, while demonstrating superior robustness. These results validate the effectiveness and generalizability of low-rank compression–inspired decoder design for semantic segmentation.

Image ClassificationRegion SegmentationTransformer Efficiency

This work addresses a critical misalignment in existing query-based masked Transformers, where training objectives do not reflect inference goals, leading to high-confidence queries that may correspond to low-quality masks and the discarding of superior intermediate predictions. To resolve this, the authors propose iFAN, an inference-aware training framework that aligns query ranking with mask quality via Adjusted Probability-Mask Ranking (APMR) and transfers high-quality predictions from intermediate layers to the final layer through Cross-Layer Self-Distillation (CLSD). Notably, iFAN requires no modification to the inference pipeline and consistently improves performance across diverse benchmarks—yielding average gains of 1.20 PQ, 1.30 AP, and 0.63 mIoU on COCO, ADE20K, and Cityscapes—while being compatible with various architectures, backbone scales, and input resolutions, all with negligible increases in parameters, computational cost, or latency.

inference-aware learningmask transformersquery competition

TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters

Oct 30, 2024
HW
Haiyang Wang
🏛️ Max Planck Institute for Informatics | Peking University | Google

Scaling Transformer models is prohibitively expensive due to fixed-parameter linear projection layers; architectural modifications necessitate full retraining. Method: We propose TokenFormer, the first architecture introducing *parameter tokenization*, which models model parameters as learnable tokens and replaces all linear layers with token-parameter self-attention—unifying parameter and input token representations in a shared latent space. Contribution/Results: Our method enables zero-shot, progressive parameter expansion without retraining, overcoming classical scaling bottlenecks. Without altering network topology, we scale model parameters from 124M to 1.4B while matching the performance of fully trained baselines, achieving substantial training cost reduction. The code and models are publicly released.

Dependence on fixed parameters requiring full retrainingHigh computational cost of scaling Transformer modelsLack of efficient progressive scaling for large models

MSDNet: Multi-Scale Decoder for Few-Shot Semantic Segmentation via Transformer-Guided Prototyping

Sep 17, 2024
AF
Amirreza Fateh
🏛️ Iran University of Science and Technology | IUST

To address the challenges of detail loss and high computational overhead in few-shot semantic segmentation, this paper proposes an efficient and accurate Transformer-based approach. Methodologically, it introduces (1) a novel spatial Transformer decoder coupled with a context-aware mask generation module; (2) a multi-scale hierarchical decoding mechanism that fuses intermediate-layer global features to enhance fine-grained localization; and (3) prototype-guided multi-scale feature pyramid decoding, integrated with spatial-attention-driven support-query relational modeling. With only 1.5M parameters, the model achieves state-of-the-art performance on both PASCAL-5ⁱ and COCO-20ⁱ benchmarks under 1-shot and 5-shot settings. It simultaneously delivers superior accuracy, strong generalization across diverse base/novel classes, and high inference efficiency—outperforming all existing methods while maintaining architectural compactness and computational tractability.

Computational EfficiencyDetail PreservationFew-shot Semantic Segmentation

Latest Papers

What's happening recently
View more

This work addresses the challenge of fairly evaluating the true performance of diverse backbone architectures in image segmentation, which has been hindered by inconsistencies in decoders, training strategies, and pretraining protocols. To this end, the authors propose LUMA (Lightweight Universal Mask Adapter), a lightweight and architecture-agnostic decoder based on cross-attention, enabling unified training recipes across backbones. Using LUMA, they conduct the first systematic benchmark under identical conditions, evaluating 20 backbones—including ViTs, CNNs, and MoEs—across 11 pretraining strategies and multiple resolutions. Their experiments reveal that the choice of pretraining objective exerts a far greater influence on segmentation performance than architectural design. Moreover, LUMA matches the accuracy of the state-of-the-art efficient ViT-based segmenter EoMT at lower computational cost and demonstrates that so-called “efficient” token mixers offer no advantage at high resolutions, with standard ViTs consistently occupying the Pareto frontier in throughput–accuracy trade-offs.

benchmarkingimage segmentationmodel comparison

This study addresses the inherent complexity of Transformer internal mechanisms, which impedes intuitive understanding of how geometric representations translate into algorithmic logic. We construct a minimal Transformer with an embedding dimension of two and propose a "geometry as algorithm" framework. Through comprehensive two-dimensional visualization of attention patterns, residual streams, and decision boundaries, this approach directly interprets learned geometric structures as stepwise algorithmic logic. The proposed method enables explicit visualization and interpretability analysis of the model's entire internal computation pipeline. Furthermore, we successfully reproduce and dissect the predictive mechanisms on numerical sequence tasks. Ultimately, this work provides an innovative pedagogical and experimental platform for advancing mechanistic interpretability research in deep learning.

informational geometryminimal transformersnext-token prediction

研究使用SEWN模型测试了Transformer中令牌稀疏路由的有效性,通过学习门控机制将令牌分配到轻量级或全容量处理,验证了不同令牌需要不同计算资源的假设。

Adaptive ComputationEfficient TransformersSparse Token Routing

Transformer models in federated learning are vulnerable to gradient-based attacks, as gradients from positional encodings can be exploited to reconstruct original inputs and leak sensitive information. To address this, this work proposes a unified Masked Jigsaw Puzzle (MJP) framework that enhances privacy and robustness by randomly shuffling input tokens and replacing original positional encodings with learnable embeddings at unknown positions. This disrupts local spatial structure and encourages the model to learn more robust feature representations. MJP is the first approach to unify the jigsaw puzzle mechanism across both vision and language Transformers. Experiments demonstrate consistent improvements in downstream performance on ImageNet-1K image classification and sentiment analysis tasks on Yelp and Amazon text datasets, while simultaneously providing strong defense against gradient inversion attacks.

federated learninggradient attackposition embedding

Standard supervised training often struggles to learn effective query-key attention patterns in Transformer-based sequence classification tasks, particularly failing to induce a preference for neighboring positions. This work demonstrates through systematic ablation studies and simplified theoretical analysis that self-pretraining (SPT), driven by a masked reconstruction objective, enables the model to acquire such localized attention structures from random initialization, substantially improving optimization dynamics. Without relying on external data, SPT significantly outperforms purely supervised training on benchmarks such as the Long-Range Arena. The performance gains are primarily attributed to the model’s enhanced ability to learn interactions among nearby tokens, highlighting proximity-aware attention as a key mechanism underlying SPT’s effectiveness.

attention patternsmasked token predictionself-pretraining

Hot Scholars

HX

Hongli Xu

University of Science and Technology of China
Software Defined NetworkCooperative CommunicationSensor Networks
YF

Yuqian Fu

Research Scientist, INSAIT, Sofia University | ETH Zurich | Fudan University
Transfer LearningDomain AdaptationMultimodal LearningEgoCentric Vision
JC

Jialei Chen

Nagoya University
Semantic SegmentationComputer Vision
BB

Benjamin Busam

Technical University of Munich
PhotogrammetryComputer VisionMachine LearningSensor Fusion