attention-guided span selection

Design and implement components that score, localize, and select contiguous spans in input sequences using attention weights and instruction-following attention scoring to identify high-attention hotspots. Build attention-based locators and span selection modules that retain top-k high-attention spans and reduce the candidate set for downstream inspection or analysis.

attention-guidedspanselection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.06
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms

Feb 21, 2025
FC
Feiyang Chen
🏛️ Shanghai Jiao Tong University | Microsoft Research | Peking University

Existing attention optimization frameworks require extensive manual tuning for heterogeneous hardware and struggle to accommodate novel attention variants or hardware configurations. To address this, we propose a modular, programmable attention computation framework. Our approach introduces: (1) a decomposable attention operator design, enabling flexible composition of arbitrary attention variants; and (2) a unified intermediate representation (IR) coupled with a multi-backend auto-scheduler based on programmable kernel templates, facilitating algorithm-hardware co-optimization. The framework automatically adapts attention implementations across diverse model architectures and hardware platforms without manual re-tuning. Experimental evaluation demonstrates up to 10× speedup over state-of-the-art solutions on non-mainstream hardware configurations, including specialized accelerators and emerging processor architectures. The framework achieves both generality and efficiency while significantly reducing deployment overhead. Open-source implementation is publicly available.

Automates kernel optimization with programmable templatesEnables scalable deployment with minimal manual tuningOptimizes attention mechanisms across diverse hardware platforms

This work addresses a fundamental paradox in hybrid recurrent-attention architectures: content routing seeks to bypass the computational expense of attention, yet relies on attention-derived representations to function effectively. Through over twenty controlled experiments, the study systematically investigates the representational conditions necessary for efficient routing and reveals, for the first time, that attention not only serves as a computational mechanism but also constructs a low-dimensional, routable subspace by encoding pairwise matching outcomes into representations. Empirical results demonstrate that a single-layer softmax attention module can induce a subspace of approximately 34 dimensions, achieving 98.4% routing accuracy—dramatically higher than the 1.2% obtained without attention. This subspace cannot be replicated by random projections or contrastive learning, while classical non-learning methods such as Bloom filters attain only 90.9% accuracy, underscoring the irreplaceable role of attention in shaping effective routing representations.

attention mechanismcontent-based routinghybrid sequence models

This work systematically investigates the interplay between efficient attention mechanisms—such as sliding window attention—and full attention in modern hybrid architectures, where the functional roles of these components remain poorly understood. Through extensive experiments, mechanistic analysis, and architectural ablation studies, the authors demonstrate that long-range information retrieval is predominantly handled by full attention layers, while efficient attention modules significantly shape the optimization trajectory. Building on these insights, they propose applying NoPE (No Positional Encoding) exclusively to full attention layers, which substantially enhances performance on long-context tasks with minimal degradation on short-context benchmarks. The study further uncovers a “large-window inertia” phenomenon, empirically validating that small-window sliding attention paired with NoPE-equipped full attention achieves superior efficiency and effectiveness.

efficient attentionfull attentionhybrid architectures

This work addresses the high computational complexity—typically O(L²)—of standard causal self-attention over long sequences and its inability to support random access to arbitrary context positions. The authors propose a multi-step attention architecture that reformulates causal self-attention as a learnable, stepwise search process. By integrating span selection and local attention mechanisms inspired by skip-list structures, the method achieves sub-quadratic complexity of O(L^{1+1/N}) for the first time while preserving non-exclusive attention patterns and enabling random access. Implemented as a two-step attention pipeline within a 30B sparse Mixture-of-Experts (MoE) model, the approach demonstrates strong empirical performance: on a single B200 GPU, it achieves decoding speeds of 114 tokens/s at 1M context length and 80 tokens/s at 10M. Furthermore, it successfully handles 256K-context tasks in Needle-in-a-Haystack (NIAH) evaluations, confirming the learnability of its routing mechanism and effectiveness in ultra-long-context modeling.

long-context attentionmulti-step attentionrandom context access

What are you sinking? A geometric approach on attention sink

Aug 04, 2025
VR
Valeria Ruscio
🏛️ Sapienza University of Rome

This work investigates the fundamental nature of Attention Sinks (AS): whether they are mere architectural byproducts or reflect geometric principles underlying the construction of stable coordinate systems in high-dimensional representation spaces. Method: We systematically analyze attention maps across diverse Transformer architectures and conduct ablation studies—particularly on positional encoding schemes—to characterize AS from a geometric reference frame perspective. Contribution/Results: We demonstrate that AS spontaneously emerge early in training and constitute an optimal solution for establishing a stable geometric reference frame in high-dimensional space. We identify and categorize three canonical AS structures: centralized, distributed, and bidirectional. Our findings establish the universality and functional necessity of AS, revealing their role in underpinning the intrinsic stability of attention mechanisms. Moreover, this geometric interpretation provides principled, geometry-driven guidance for model design—including positional encoding strategies and special token placement—thereby bridging representational geometry with architectural engineering.

Analyzes attention sink patterns in transformer attention mapsExplores impact of architecture components on reference frame typesIdentifies geometric principles behind reference frame establishment

Latest Papers

What's happening recently
View more

This study addresses the significant performance degradation in existing visual token pruning methods caused by early text guidance, which restricts the identification of answer-relevant regions. To overcome this limitation, we propose a training-free two-stage pruning strategy that innovatively decouples vision-guided pruning from delayed text-guided re-selection. Specifically, the method first performs preliminary pruning using visual encoder attention, and subsequently refines the final token set via text-to-vision attention at intermediate decoder layers. Extensive experiments across three models and eight benchmarks demonstrate that our approach recovers an average of 11.10 and 16.84 percentage points in performance at 80% and 90% pruning ratios, respectively, while effectively reducing inference latency.

Computational CostText-guided SelectionVision-Language Model

This work addresses the entanglement of routing and filtering functions in conventional attention mechanisms, which leads to structural opacity and optimization challenges. The authors propose S-D Attention, which explicitly decouples the interaction matrix into a low-rank routing component and a symmetric filtering component, thereby clearly distinguishing these two mechanisms for the first time. They further uncover that routing self-organizes into a spectral cascade phenomenon in deep networks. Leveraging this insight, they achieve stable training without layer normalization and validate their approach using linear attention variants (e.g., ELU+1) and effective rank analysis. Experiments show that linearizing the first seven layers of a 125M-parameter model incurs less than 5% perplexity degradation, while cascade-informed architectures reduce attention parameters by 47%–65% with only a 3.9%–8.4% increase in perplexity.

attentionattention decompositionfiltering

This work addresses the memory bottleneck of Compressed Sparse Attention (CSA), which is constrained by the excessive size of intermediate score tensors and thus limited to single-GPU memory, hindering its application to long sequences. The paper introduces the first CSA indexer that avoids fully materializing intermediate tensors, leveraging a streaming partition-merge top-k algorithm to efficiently generate attention indices without instantiating the full score matrix. Integrated with a Triton-based chunked top-k dispatcher, a TileLang-pipelined attention kernel, and key sharding, the proposed method achieves remarkable scalability: on an NVIDIA H200 GPU, it processes sequences up to 1,048,576 tokens with a peak memory footprint of only 6.21 GB—32× longer than prior approaches—while maintaining high fidelity, with an average recall of 1.0000 (minimum 0.9980).

Compressed Sparse AttentionGPU memorylong-context

This study addresses the limited instruction-awareness of text embedding models by proposing Attention Relay, a training-free cross-model attention transfer mechanism. This method directly transfers attention weights from large language models (LLMs), such as Qwen3 and Llama 3.1, to general-purpose Transformer-based embedding models, endowing them with instruction-following capabilities without requiring additional training. Experimental results demonstrate that this mechanism is consistently effective across diverse combinations of LLMs and embedding architectures. It significantly enhances the ability of embedding representations to focus on instruction-relevant content. Consequently, this work establishes an efficient, training-free paradigm for constructing instruction-aware embedding models, offering a practical solution to bridge the functional gap between generative LLMs and representation learning systems.

attention weightsinstruction followinglarge language models

Hot Scholars

OS

Ozgur Sinanoglu

Professor of Electrical and Computer Engineering, New York University Abu Dhabi
Hardware Security
JK

Johann Knechtel

New York University Abu Dhabi
Electronic Design Automation3D IntegrationHardware Security
PM

Parvin Mousavi

School of Computing, Queen's University
medical imagingimage guided interventionssystems biologybioinformatics
CH

Cong Hu

Jiangnan University
deep learning、machine learning、 computer vision、pattern recognition