structured attention masking

Design, build, and evaluate attention-masking schemes that constrain token interactions according to explicit structure—spatial regions, temporal segments, or semantic boundaries—so as to prevent cross-region interference, preserve fine-grained alignment cues, and enable token-level parallel generation. This includes implementing hard and soft masks, progressive soft-masked cross-attention, region- and boundary-aware token routing, and attention-masking strategies that integrate with prompting or cross-attention while targeting low or zero extra compute overhead.

structuredattentionmasking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing multimodal large language models (MLLMs) employ static cross-modal tokenization, limiting their ability to emulate human-like, context-sensitive integration of multimodal information. To address this, we introduce— for the first time in MLLMs—the cognitive science principle of *chunking* into tokenizer design, proposing an adaptive cross-modal tokenization framework. Our method comprises three core components: differentiable dynamic boundary learning, hierarchical multi-granularity representation, and vision-language alignment-guided attention. This framework departs from conventional fixed-tokenization paradigms by enabling semantic-driven, context-aware token segmentation. Evaluated on visual question answering (VQA) and complex scene description tasks, our approach achieves absolute improvements of 7.8% and 5.3%, respectively. Moreover, error patterns and attention distributions align significantly more closely with human cognitive behavior. Our work establishes a novel paradigm for developing human-inspired multimodal understanding models.

Bridging human cognitive chunking and MLLM token representation gapsEnhancing MLLMs with dynamic cross-modal tokenization for cognitive alignmentOvercoming static tokenization limits in multimodal human-like processing

This work addresses the instability in representation alignment during diffusion model training, which arises from the mismatch between noisy inputs and clean image features, leading models to over-rely on complete token sets. To mitigate this alignment discrepancy, the authors propose MaskAlign, the first approach that operates from the perspective of token subsets. MaskAlign dynamically aligns representations using randomly masked token subsets and introduces a lightweight pre-mask token mixing module to encourage cross-token information sharing prior to masking. Integrated with a self-supervised visual encoder and diffusion Transformer training, MaskAlign significantly enhances generation quality and alignment robustness while maintaining computational efficiency, thereby improving the model’s generalization under perturbations of varying token subsets.

clean-image featuresdiffusion modelsnoisy inputs

Local Representative Token Guided Merging for Text-to-Image Generation

Jul 17, 2025
ML
Min-Jeong Lee
🏛️ Korea University

Stable diffusion models suffer from low inference efficiency due to the quadratic computational complexity of self-attention. Existing token merging methods fail to adequately model the locality and semantic importance of cross-modal attention in text-to-image generation, thus struggling to balance efficiency and generation quality. To address this, we propose a local representative token-guided token merging method: we introduce the novel concept of *local representative tokens*, integrating dynamic window partitioning with similarity-based adaptive token selection to identify the most representative tokens within context-aware local regions. This strategy is model-agnostic and requires no architectural modifications. Experiments demonstrate that our method achieves a 6.2% reduction in FID while significantly improving CLIP Score, all without compromising inference speed—effectively reconciling high-fidelity generation with computational efficiency.

Balances visual quality and computational efficiency in text-to-imageImproves token merging for attention-based image generation modelsReduces quadratic complexity in stable diffusion attention operations

This work addresses the insufficient characterization of the multiscale dynamic properties of cross-attention in existing diffusion models, which limits training-free controllable generation. Treating cross-attention during diffusion as a spatiotemporal signal in latent space, the study reveals—for the first time—a stable time–frequency evolution pattern throughout the denoising process. Building on this insight, the authors propose a plug-and-play inference-time intervention method that enables continuous scale control without modifying prompts or model parameters. The approach combines Fourier-domain attention log-modulation, radial frequency band reweighting, timestep-aligned scheduling, and an adaptive gating mechanism based on token assignment entropy. Evaluated on Stable Diffusion, the method effectively redistributes the attention spectrum, significantly enhancing visual editing quality while preserving semantic consistency, and demonstrates that entropy primarily serves as an adaptive gain rather than an independent control dimension.

attention dynamicscross-attentiondiffusion models

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

Dec 11, 2025
YF
Yuchen Feng
🏛️ Institute of Information Engineering, Chinese Academy of Sciences | Baidu Inc.

To address the limited visual perception capability of multimodal large language models (MLLMs), this paper proposes a novel framework that dynamically adjusts visual token resolution within a single forward pass—inspired by human “saccadic” visual scanning. The method comprises two key components: (1) a layer-wise, attention-guided saliency scanning strategy that adaptively focuses computation on semantically critical regions; and (2) a plug-and-play Token Super-Resolution (TokenSR) module enabling dynamic expansion or pruning of token-level computational resources. Crucially, the approach requires no additional training or architectural modification, enhancing visual representation quality efficiently during inference. Evaluated on multiple vision-language understanding benchmarks—including MMBench, OCRBench, and TextVQA—the method consistently outperforms strong baselines, demonstrating that dynamic resolution control meaningfully improves both visual perception fidelity and downstream multimodal reasoning performance.

Dynamically allocates computation to salient visual tokensEnhances visual perception in multimodal language modelsImproves efficiency and adaptability in multimodal understanding

Latest Papers

What's happening recently
View more

This work addresses the limitation of existing linear attention methods, which rely on predefined spatial layouts to compress image tokens, thereby constraining information aggregation to coordinate positions rather than semantic content. To overcome this, the paper introduces Representative Attention (RPAttention), a novel representation-driven token compression mechanism that dynamically generates semantic representative tokens to enable spatially agnostic global interactions. RPAttention adopts a lightweight Gather-Interact-Distribute paradigm, integrating competitive similarity-based routing, interaction among representative tokens in a compact latent space, and query-driven cross-attention. This design maintains linear computational complexity while substantially enhancing semantic alignment. Experimental results demonstrate consistent and significant performance gains across image classification, object detection, and semantic segmentation tasks.

global attentionlinear attentionsemantic organization

This work addresses the lack of a unified theoretical foundation in existing attention mask designs. It establishes, for the first time, a formal connection between attention masks and partially ordered structures, proving that information flow in sufficiently deep multi-layer Transformers converges to a Hasse diagram. The mask design problem is thereby reformulated as finding the minimal common supergraph of such Hasse diagrams, yielding a general framework that derives attention masks directly from task families. Leveraging this framework, the authors propose two novel mechanisms—Block Two-Stream Attention and Butterfly Attention—and derive block-wise causal masks and fully supervised bidirectional masks that guarantee consistency between training and inference. Empirical results validate both the effectiveness and generality of the proposed approach.

attention masksHasse diagramsinformation flow

This work proposes TokenMask, a novel segmentation framework that departs from conventional query-based Vision Transformer approaches which rely on explicit reconstruction of image-space feature maps—a process that incurs substantial computational redundancy and hinders deployment. Instead, TokenMask operates entirely in the query token space, generating mask logits directly through token affinity and performing interpolation in logit space. By integrating a ViT backbone, a token-space mask head, and TensorRT FP16 inference, the method significantly reduces both computational and memory overhead across multiple datasets and segmentation tasks while preserving accuracy. Notably, it achieves substantial acceleration on the Jetson AGX Orin platform, offering an efficient and streamlined architecture well-suited for embedded vision applications.

efficient segmentationembedded visionmask prediction

Existing training-free samplers for masked diffusion language models decide token submission positions sequentially, overlooking the tendency of high-confidence predictions to emerge as contiguous segments, thereby limiting parallelization efficiency. This work proposes CLAD (Confidence-guided Localized Aggregation Decoding), a novel decoding strategy that first identifies continuous high-confidence segments—termed Confidence-Induced Clusters (CICs)—via confidence-guided clustering and then leverages self-attention maps to assess inter-cluster dependencies, enabling conflict-aware, cluster-level parallel submission. CLAD introduces, for the first time, a cluster-wise parallelism mechanism that requires no modification to model training. Evaluated on LLaDA and Dream models, it achieves speedups of 1.77× to 8.47× while largely preserving generation quality comparable to original sequential decoding across most tasks.

commitment unitsconfidence spansmasked diffusion language models

Hot Scholars

XS

Xiaoyu Shen

Eastern Institute of Technology, Ningbo
language modelmulti-modal learningreasoning
PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
YW

Yunnan Wang

Department of Computer Science and Engineering, Shanghai Jiao Tong University
Computer VisionMultimodal Representation Learning
YM

Yunpu Ma

Ludwig Maximilian University of Munich
Foundation ModelsAgentic AITemporal Knowledge GraphQuantum AI