Score
Design, build, and evaluate attention-masking schemes that constrain token interactions according to explicit structure—spatial regions, temporal segments, or semantic boundaries—so as to prevent cross-region interference, preserve fine-grained alignment cues, and enable token-level parallel generation. This includes implementing hard and soft masks, progressive soft-masked cross-attention, region- and boundary-aware token routing, and attention-masking strategies that integrate with prompting or cross-attention while targeting low or zero extra compute overhead.
The quadratic time and memory complexity of Transformer self-attention severely hinders efficient long-context modeling. This paper presents a systematic survey and reconstruction of efficient attention mechanisms for large language models, proposing the first unified taxonomy encompassing both linearization paradigms (e.g., kernel-based approximations and fast weight dynamics) and sparsification paradigms (e.g., fixed patterns, block-wise routing, and clustering-driven selection). It innovatively integrates algorithmic design with hardware-aware optimization, clarifying integration pathways for purely efficient attention and hybrid architectures in large-scale pretraining. Furthermore, it establishes a comprehensive reference framework spanning theoretical analysis, algorithmic implementation, and engineering deployment. The work delivers a systematic design paradigm and practical guidelines for scalable long-context language models.
Existing multimodal large language models (MLLMs) employ static cross-modal tokenization, limiting their ability to emulate human-like, context-sensitive integration of multimodal information. To address this, we introduce— for the first time in MLLMs—the cognitive science principle of *chunking* into tokenizer design, proposing an adaptive cross-modal tokenization framework. Our method comprises three core components: differentiable dynamic boundary learning, hierarchical multi-granularity representation, and vision-language alignment-guided attention. This framework departs from conventional fixed-tokenization paradigms by enabling semantic-driven, context-aware token segmentation. Evaluated on visual question answering (VQA) and complex scene description tasks, our approach achieves absolute improvements of 7.8% and 5.3%, respectively. Moreover, error patterns and attention distributions align significantly more closely with human cognitive behavior. Our work establishes a novel paradigm for developing human-inspired multimodal understanding models.
This work addresses the instability in representation alignment during diffusion model training, which arises from the mismatch between noisy inputs and clean image features, leading models to over-rely on complete token sets. To mitigate this alignment discrepancy, the authors propose MaskAlign, the first approach that operates from the perspective of token subsets. MaskAlign dynamically aligns representations using randomly masked token subsets and introduces a lightweight pre-mask token mixing module to encourage cross-token information sharing prior to masking. Integrated with a self-supervised visual encoder and diffusion Transformer training, MaskAlign significantly enhances generation quality and alignment robustness while maintaining computational efficiency, thereby improving the model’s generalization under perturbations of varying token subsets.
Stable diffusion models suffer from low inference efficiency due to the quadratic computational complexity of self-attention. Existing token merging methods fail to adequately model the locality and semantic importance of cross-modal attention in text-to-image generation, thus struggling to balance efficiency and generation quality. To address this, we propose a local representative token-guided token merging method: we introduce the novel concept of *local representative tokens*, integrating dynamic window partitioning with similarity-based adaptive token selection to identify the most representative tokens within context-aware local regions. This strategy is model-agnostic and requires no architectural modifications. Experiments demonstrate that our method achieves a 6.2% reduction in FID while significantly improving CLIP Score, all without compromising inference speed—effectively reconciling high-fidelity generation with computational efficiency.
This work addresses the insufficient characterization of the multiscale dynamic properties of cross-attention in existing diffusion models, which limits training-free controllable generation. Treating cross-attention during diffusion as a spatiotemporal signal in latent space, the study reveals—for the first time—a stable time–frequency evolution pattern throughout the denoising process. Building on this insight, the authors propose a plug-and-play inference-time intervention method that enables continuous scale control without modifying prompts or model parameters. The approach combines Fourier-domain attention log-modulation, radial frequency band reweighting, timestep-aligned scheduling, and an adaptive gating mechanism based on token assignment entropy. Evaluated on Stable Diffusion, the method effectively redistributes the attention spectrum, significantly enhancing visual editing quality while preserving semantic consistency, and demonstrates that entropy primarily serves as an adaptive gain rather than an independent control dimension.
To address the limited visual perception capability of multimodal large language models (MLLMs), this paper proposes a novel framework that dynamically adjusts visual token resolution within a single forward pass—inspired by human “saccadic” visual scanning. The method comprises two key components: (1) a layer-wise, attention-guided saliency scanning strategy that adaptively focuses computation on semantically critical regions; and (2) a plug-and-play Token Super-Resolution (TokenSR) module enabling dynamic expansion or pruning of token-level computational resources. Crucially, the approach requires no additional training or architectural modification, enhancing visual representation quality efficiently during inference. Evaluated on multiple vision-language understanding benchmarks—including MMBench, OCRBench, and TextVQA—the method consistently outperforms strong baselines, demonstrating that dynamic resolution control meaningfully improves both visual perception fidelity and downstream multimodal reasoning performance.
This work addresses the limitation of existing linear attention methods, which rely on predefined spatial layouts to compress image tokens, thereby constraining information aggregation to coordinate positions rather than semantic content. To overcome this, the paper introduces Representative Attention (RPAttention), a novel representation-driven token compression mechanism that dynamically generates semantic representative tokens to enable spatially agnostic global interactions. RPAttention adopts a lightweight Gather-Interact-Distribute paradigm, integrating competitive similarity-based routing, interaction among representative tokens in a compact latent space, and query-driven cross-attention. This design maintains linear computational complexity while substantially enhancing semantic alignment. Experimental results demonstrate consistent and significant performance gains across image classification, object detection, and semantic segmentation tasks.
This work addresses the lack of a unified theoretical foundation in existing attention mask designs. It establishes, for the first time, a formal connection between attention masks and partially ordered structures, proving that information flow in sufficiently deep multi-layer Transformers converges to a Hasse diagram. The mask design problem is thereby reformulated as finding the minimal common supergraph of such Hasse diagrams, yielding a general framework that derives attention masks directly from task families. Leveraging this framework, the authors propose two novel mechanisms—Block Two-Stream Attention and Butterfly Attention—and derive block-wise causal masks and fully supervised bidirectional masks that guarantee consistency between training and inference. Empirical results validate both the effectiveness and generality of the proposed approach.
This work proposes TokenMask, a novel segmentation framework that departs from conventional query-based Vision Transformer approaches which rely on explicit reconstruction of image-space feature maps—a process that incurs substantial computational redundancy and hinders deployment. Instead, TokenMask operates entirely in the query token space, generating mask logits directly through token affinity and performing interpolation in logit space. By integrating a ViT backbone, a token-space mask head, and TensorRT FP16 inference, the method significantly reduces both computational and memory overhead across multiple datasets and segmentation tasks while preserving accuracy. Notably, it achieves substantial acceleration on the Jetson AGX Orin platform, offering an efficient and streamlined architecture well-suited for embedded vision applications.
Existing training-free samplers for masked diffusion language models decide token submission positions sequentially, overlooking the tendency of high-confidence predictions to emerge as contiguous segments, thereby limiting parallelization efficiency. This work proposes CLAD (Confidence-guided Localized Aggregation Decoding), a novel decoding strategy that first identifies continuous high-confidence segments—termed Confidence-Induced Clusters (CICs)—via confidence-guided clustering and then leverages self-attention maps to assess inter-cluster dependencies, enabling conflict-aware, cluster-level parallel submission. CLAD introduces, for the first time, a cluster-wise parallelism mechanism that requires no modification to model training. Evaluated on LLaDA and Dream models, it achieves speedups of 1.77× to 8.47× while largely preserving generation quality comparable to original sequential decoding across most tasks.