mask-based tokenization

Design, build, and evaluate methods that convert between segmentation masks and discrete token representations—tokenizing masks into spatially grounded token sequences, predicting masked tokens (including autoregressive masked-token prediction), and aligning tokenized patches with positional embeddings. Implement token-conditioned decoders and projections (token-to-mask decoding, token splatting decoders, SAM-like decoders) that reconstruct pixel-level masks from token embeddings, fuse multi-scale visual features, and support token-conditioned sequence generation, mask-token alignment, and token-based reasoning for segmentation.

mask-basedtokenization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes SemTok, a semantic-driven one-dimensional tokenizer that addresses the limitations of existing vision tokenizers, which often rely on fixed 2D grids and prioritize pixel-level reconstruction at the expense of compact global semantics. SemTok compresses images into high-level discrete semantic tokens and introduces a masked autoregressive generative framework. Its key innovations include a 2D-to-1D semantic tokenization strategy, a semantic alignment constraint mechanism, and a two-stage generative training paradigm. Experimental results demonstrate that SemTok achieves state-of-the-art performance in image reconstruction, delivering higher fidelity under extremely compact token representations and significantly enhancing downstream generative capabilities.

2D-to-1D mappingimage reconstructionlatent space

Fast SAM2 with Text-Driven Token Pruning

Dec 24, 2025
AM
Avilasha Mandal
🏛️ University of Electronic Science and Technology of China | Indian Institute of Technology | Harbin Institute of Technology

To address the high computational and memory overhead of SAM2 in video object segmentation caused by dense visual tokens, this paper proposes a text-guided post-encoding visual token pruning framework. Without modifying the original model architecture, the method dynamically evaluates and retains critical tokens by jointly leveraging local visual context, text–visual semantic alignment, and uncertainty modeling. It introduces, for the first time, user- or auto-generated textual prompts into the post-encoding importance scoring process, enabling semantic-aware early sparsification prior to temporal propagation. Key design elements include a lightweight routing mechanism, multi-source importance fusion, and decoupling of the image encoder from the memory module. Experiments across multiple benchmarks demonstrate a 42.50% speedup in inference latency and a 37.41% reduction in GPU memory consumption, while maintaining near-equivalent J&F scores—significantly enhancing the deployability of Transformer-based video segmentation models in real-time and edge-computing scenarios.

Maintains segmentation accuracy while improving inference speed and efficiencyPrunes irrelevant visual tokens using text guidance before temporal propagationReduces computational and memory costs in SAM2 video segmentation

This work proposes a unified autoregressive segmentation framework based on run-length encoding (RLE) to address the lack of a common architecture for image and video semantic segmentation and the inefficiency of long-sequence generation. By discretizing segmentation masks into RLE token sequences and leveraging an enhanced Pix2Seq architecture for autoregressive generation, the method introduces a novel token compression strategy that substantially reduces sequence length. The approach naturally supports temporal modeling in video segmentation and can be extended to panoptic segmentation by incorporating instance-level information. Experimental results on two benchmark datasets demonstrate performance comparable to state-of-the-art methods, confirming the effectiveness and versatility of the proposed framework.

panoptic segmentationrun length encodingsemantic segmentation

Morphing Tokens Draw Strong Masked Image Models

Dec 30, 2023
TK
Taekyung Kim
🏛️ NAVER AI Lab

In masked image modeling (MIM) pretraining of Vision Transformers (ViTs), tokenization and local masking induce spatially inconsistent reconstruction supervision, degrading representation discriminability. This work is the first to systematically identify and address this spatial inconsistency issue, proposing Dynamic Token Morphing (DTM): a context-aware, dynamic token aggregation mechanism that generates spatially coherent reconstruction targets. DTM introduces no additional parameters or computational overhead and is plug-and-play across diverse MIM frameworks. On ImageNet-1K and ADE20K, DTM achieves significant gains over state-of-the-art MIM methods—yielding lower training loss and more stable convergence. When transferred to downstream tasks such as iNaturalist, it delivers consistent performance improvements. The core contribution is the first lightweight, parameter-free, and framework-agnostic solution specifically designed to resolve the spatial inconsistency problem in MIM.

Addresses spatial inconsistency in masked image modeling supervisionImproves discriminative representation learning in Vision TransformersReduces training costs while enhancing MIM performance

To address the performance trade-off between multimodal understanding and generation caused by mismatched visual granularity, this paper proposes TokenFlow—the first unified image tokenizer. Its core innovation is a dual-codebook vector quantization architecture: a semantic codebook captures high-level semantics, while a pixel codebook preserves fine-grained texture; both are decoupled during learning yet jointly aligned via a shared index mechanism. This design enables discrete token-driven, integrated modeling of multimodal understanding and autoregressive image generation. Experiments demonstrate that TokenFlow achieves an average 7.2% improvement over LLaVA-1.5 (13B) on multimodal understanding benchmarks; attains an FID of 0.63 for 384×384 image reconstruction; and scores 0.55 on GenEval for 256×256 autoregressive generation—comparable to SDXL. TokenFlow effectively mitigates granularity conflict and advances unified visual representation learning.

Bridging gap between multimodal understanding and generationDecoupling semantic and pixel-level feature learningImproving performance in visual understanding and generation tasks

Latest Papers

What's happening recently
View more

This work addresses the instability in representation alignment during diffusion model training, which arises from the mismatch between noisy inputs and clean image features, leading models to over-rely on complete token sets. To mitigate this alignment discrepancy, the authors propose MaskAlign, the first approach that operates from the perspective of token subsets. MaskAlign dynamically aligns representations using randomly masked token subsets and introduces a lightweight pre-mask token mixing module to encourage cross-token information sharing prior to masking. Integrated with a self-supervised visual encoder and diffusion Transformer training, MaskAlign significantly enhances generation quality and alignment robustness while maintaining computational efficiency, thereby improving the model’s generalization under perturbations of varying token subsets.

clean-image featuresdiffusion modelsnoisy inputs

This work proposes TokenMask, a novel segmentation framework that departs from conventional query-based Vision Transformer approaches which rely on explicit reconstruction of image-space feature maps—a process that incurs substantial computational redundancy and hinders deployment. Instead, TokenMask operates entirely in the query token space, generating mask logits directly through token affinity and performing interpolation in logit space. By integrating a ViT backbone, a token-space mask head, and TensorRT FP16 inference, the method significantly reduces both computational and memory overhead across multiple datasets and segmentation tasks while preserving accuracy. Notably, it achieves substantial acceleration on the Jetson AGX Orin platform, offering an efficient and streamlined architecture well-suited for embedded vision applications.

efficient segmentationembedded visionmask prediction

This work addresses the challenge of balancing accuracy, instruction-following capability, and inference efficiency in multimodal large language models for segmentation tasks. The authors propose STAMPlus, a structured fully masked prediction framework that decouples autoregressive dialogue from non-autoregressive mask prediction, enabling parallel generation of multiple explicitly ID-tagged target masks within a single forward pass. Key innovations include a <SEG> trigger mechanism, image-aligned mask token fusion, hybrid attention-based classification, and shared multi-class mask space binding, collectively supporting diverse segmentation scenarios—such as referring expression, open-vocabulary semantics, instance-aware parsing, and small-object remote sensing—while overcoming single-target limitations. Experiments demonstrate state-of-the-art performance across multiple benchmarks, preserved general-purpose multimodal instruction-following ability, and a significant reduction in inference latency for 12-class segmentation from 13.50 seconds to 5.16 seconds.

MLLM-based segmentationmulti-target segmentationmultimodal instruction following

This work addresses the computational bottleneck in vision-language models during inference caused by processing a large number of visual tokens, noting that existing pruning methods fail to account for the varying functional roles of redundant tokens. Building on EmbedLens-based token role identification, the study reveals that prevailing pruning strategies implicitly favor certain token roles, yet this preference shows no direct correlation with downstream performance. The authors propose a role-aware pruning strategy that deliberately preserves specific non-“alive” tokens—such as those with weaker semantic alignment—and demonstrate that doing so can maintain or even enhance model performance. These findings underscore the critical influence of functional token roles on importance assessment and offer a novel perspective for designing efficient vision-language models.

computational bottleneckredundant tokenstoken roles

Hot Scholars

LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
JB

Jinbin Bai

National University of Singapore
Machine LearningContent CreationGenerative Modeling
MM

Ming-Ming Cheng

Professor of Computer Science, Nankai University
Computer VisionComputer GraphicsVisual AttentionSaliency
XW

Xinyu Wang

PhD student, McGill University
Large Language ModelRetrieval Augmented GenerationQuantization
HS

Hyunjung Shim

Associate Professor, KAIST
Computer visionmachine learning