Score
Design, build, and evaluate methods that convert between segmentation masks and discrete token representations—tokenizing masks into spatially grounded token sequences, predicting masked tokens (including autoregressive masked-token prediction), and aligning tokenized patches with positional embeddings. Implement token-conditioned decoders and projections (token-to-mask decoding, token splatting decoders, SAM-like decoders) that reconstruct pixel-level masks from token embeddings, fuse multi-scale visual features, and support token-conditioned sequence generation, mask-token alignment, and token-based reasoning for segmentation.
This work proposes SemTok, a semantic-driven one-dimensional tokenizer that addresses the limitations of existing vision tokenizers, which often rely on fixed 2D grids and prioritize pixel-level reconstruction at the expense of compact global semantics. SemTok compresses images into high-level discrete semantic tokens and introduces a masked autoregressive generative framework. Its key innovations include a 2D-to-1D semantic tokenization strategy, a semantic alignment constraint mechanism, and a two-stage generative training paradigm. Experimental results demonstrate that SemTok achieves state-of-the-art performance in image reconstruction, delivering higher fidelity under extremely compact token representations and significantly enhancing downstream generative capabilities.
To address the high computational and memory overhead of SAM2 in video object segmentation caused by dense visual tokens, this paper proposes a text-guided post-encoding visual token pruning framework. Without modifying the original model architecture, the method dynamically evaluates and retains critical tokens by jointly leveraging local visual context, text–visual semantic alignment, and uncertainty modeling. It introduces, for the first time, user- or auto-generated textual prompts into the post-encoding importance scoring process, enabling semantic-aware early sparsification prior to temporal propagation. Key design elements include a lightweight routing mechanism, multi-source importance fusion, and decoupling of the image encoder from the memory module. Experiments across multiple benchmarks demonstrate a 42.50% speedup in inference latency and a 37.41% reduction in GPU memory consumption, while maintaining near-equivalent J&F scores—significantly enhancing the deployability of Transformer-based video segmentation models in real-time and edge-computing scenarios.
This work proposes a unified autoregressive segmentation framework based on run-length encoding (RLE) to address the lack of a common architecture for image and video semantic segmentation and the inefficiency of long-sequence generation. By discretizing segmentation masks into RLE token sequences and leveraging an enhanced Pix2Seq architecture for autoregressive generation, the method introduces a novel token compression strategy that substantially reduces sequence length. The approach naturally supports temporal modeling in video segmentation and can be extended to panoptic segmentation by incorporating instance-level information. Experimental results on two benchmark datasets demonstrate performance comparable to state-of-the-art methods, confirming the effectiveness and versatility of the proposed framework.
In masked image modeling (MIM) pretraining of Vision Transformers (ViTs), tokenization and local masking induce spatially inconsistent reconstruction supervision, degrading representation discriminability. This work is the first to systematically identify and address this spatial inconsistency issue, proposing Dynamic Token Morphing (DTM): a context-aware, dynamic token aggregation mechanism that generates spatially coherent reconstruction targets. DTM introduces no additional parameters or computational overhead and is plug-and-play across diverse MIM frameworks. On ImageNet-1K and ADE20K, DTM achieves significant gains over state-of-the-art MIM methods—yielding lower training loss and more stable convergence. When transferred to downstream tasks such as iNaturalist, it delivers consistent performance improvements. The core contribution is the first lightweight, parameter-free, and framework-agnostic solution specifically designed to resolve the spatial inconsistency problem in MIM.
To address the performance trade-off between multimodal understanding and generation caused by mismatched visual granularity, this paper proposes TokenFlow—the first unified image tokenizer. Its core innovation is a dual-codebook vector quantization architecture: a semantic codebook captures high-level semantics, while a pixel codebook preserves fine-grained texture; both are decoupled during learning yet jointly aligned via a shared index mechanism. This design enables discrete token-driven, integrated modeling of multimodal understanding and autoregressive image generation. Experiments demonstrate that TokenFlow achieves an average 7.2% improvement over LLaVA-1.5 (13B) on multimodal understanding benchmarks; attains an FID of 0.63 for 384×384 image reconstruction; and scores 0.55 on GenEval for 256×256 autoregressive generation—comparable to SDXL. TokenFlow effectively mitigates granularity conflict and advances unified visual representation learning.
This work addresses the instability in representation alignment during diffusion model training, which arises from the mismatch between noisy inputs and clean image features, leading models to over-rely on complete token sets. To mitigate this alignment discrepancy, the authors propose MaskAlign, the first approach that operates from the perspective of token subsets. MaskAlign dynamically aligns representations using randomly masked token subsets and introduces a lightweight pre-mask token mixing module to encourage cross-token information sharing prior to masking. Integrated with a self-supervised visual encoder and diffusion Transformer training, MaskAlign significantly enhances generation quality and alignment robustness while maintaining computational efficiency, thereby improving the model’s generalization under perturbations of varying token subsets.
This work proposes TokenMask, a novel segmentation framework that departs from conventional query-based Vision Transformer approaches which rely on explicit reconstruction of image-space feature maps—a process that incurs substantial computational redundancy and hinders deployment. Instead, TokenMask operates entirely in the query token space, generating mask logits directly through token affinity and performing interpolation in logit space. By integrating a ViT backbone, a token-space mask head, and TensorRT FP16 inference, the method significantly reduces both computational and memory overhead across multiple datasets and segmentation tasks while preserving accuracy. Notably, it achieves substantial acceleration on the Jetson AGX Orin platform, offering an efficient and streamlined architecture well-suited for embedded vision applications.
This work addresses the challenge of balancing accuracy, instruction-following capability, and inference efficiency in multimodal large language models for segmentation tasks. The authors propose STAMPlus, a structured fully masked prediction framework that decouples autoregressive dialogue from non-autoregressive mask prediction, enabling parallel generation of multiple explicitly ID-tagged target masks within a single forward pass. Key innovations include a <SEG> trigger mechanism, image-aligned mask token fusion, hybrid attention-based classification, and shared multi-class mask space binding, collectively supporting diverse segmentation scenarios—such as referring expression, open-vocabulary semantics, instance-aware parsing, and small-object remote sensing—while overcoming single-target limitations. Experiments demonstrate state-of-the-art performance across multiple benchmarks, preserved general-purpose multimodal instruction-following ability, and a significant reduction in inference latency for 12-class segmentation from 13.50 seconds to 5.16 seconds.
This work addresses the computational bottleneck in vision-language models during inference caused by processing a large number of visual tokens, noting that existing pruning methods fail to account for the varying functional roles of redundant tokens. Building on EmbedLens-based token role identification, the study reveals that prevailing pruning strategies implicitly favor certain token roles, yet this preference shows no direct correlation with downstream performance. The authors propose a role-aware pruning strategy that deliberately preserves specific non-“alive” tokens—such as those with weaker semantic alignment—and demonstrate that doing so can maintain or even enhance model performance. These findings underscore the critical influence of functional token roles on importance assessment and offer a novel perspective for designing efficient vision-language models.