task-conditioned prefix tuning

Design and train small, task-conditioned prefix modules made of learned token embeddings or quantized feature tokens that are prepended to model inputs or injected into attention key/value streams; these prefixes are conditioned on task identity or external lexical signals and used to modulate encoder and attention representations so the base pretrained model weights remain fixed.

task-conditionedprefixtuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Enhancing Transformers Through Conditioned Embedded Tokens

May 19, 2025
HS
Hemanth Saratchandran
🏛️ Australian Institute for Machine Learning | University of Adelaide

Transformer attention mechanisms suffer from inherent ill-conditioning, leading to suboptimal gradient optimization and slow training convergence. To address this, we establish, for the first time, a theoretical connection between the condition number of the embedding matrix and that of the attention block. Based on this insight, we propose *Conditioned Embedding Tokens* (CET): an adaptive reparameterization of the embedding space that mitigates ill-conditioning at its source—without altering the model architecture or attention computation. CET introduces only lightweight, learnable enhancements to the token embeddings, substantially improving numerical stability and gradient flow quality. Extensive experiments across four major domains—image classification, object detection, instance segmentation, and natural language processing—demonstrate consistent performance gains on diverse state-of-the-art architectures, including ViT, DETR, Mask R-CNN, and BERT. These results validate CET’s strong generalizability and practical utility.

Addresses ill-conditioning in transformer attention blocksEnhances transformer performance across multiple applicationsImproves gradient-based optimization for efficient training

Weak long-range dependency modeling and hierarchical structural degradation in large language models (LLMs) impair generation coherence and prediction stability. To address this, we propose a context-aware structured dependency encoding mechanism. Our method explicitly incorporates syntactic and semantic dependency relations into token embedding initialization—rather than relying on attention mechanisms to infer them dynamically—and requires no external syntactic annotations or auxiliary training objectives. It comprises four key components: dependency-weighted attention, structured embedding initialization, multi-layer dependency preservation, and a lightweight Transformer encoding module. Evaluated across multiple language benchmarks, our approach significantly reduces perplexity, improves long-sequence dependency alignment accuracy, mitigates abrupt phrase switching, and maintains training feasibility and architectural compatibility with standard Transformer-based LLMs.

Coherence in Text GenerationHierarchical Structure PreservationStability in Prediction

This work addresses the issue of ill-conditioned attention matrices in standard Transformers, which hinder optimization and reduce training efficiency. The authors introduce, for the first time, a learnable preconditioned attention module that integrates a conditioning matrix within each attention head to substantially lower the condition number of the attention matrix, thereby improving its optimization landscape. Designed as a generic plug-and-play component, this approach is compatible with various Transformer variants and consistently enhances both training efficiency and model performance across diverse tasks—including image classification, object detection, instance segmentation, long-sequence modeling, and language modeling—demonstrating its broad applicability and effectiveness.

attention mechanismcondition numberill-conditioned matrices

Vector Arithmetic in Concept and Token Subspaces

Nov 22, 2025
SF
Sheridan Feucht
🏛️ Northeastern University

This work investigates the structured separation of semantic and surface-level information within the hidden states of large language models (LLMs). Focusing on Llama-2-7b, the authors identify concept-inducing and token-inducing attention heads that respectively encode high-level semantics and low-level morphological features. Leveraging this observation, they construct orthogonal semantic and token subspaces from the model’s internal representations. By applying attention-weight-based transformations, they extract subspace projections and, for the first time, enable disentangled vector arithmetic directly on hidden states: semantic analogies (e.g., “Athens − Greece + China = Beijing”) achieve 80% nearest-neighbor accuracy in the concept subspace—substantially surpassing the 47% attained on raw hidden states—while the token subspace precisely recovers lexical and morphological properties. This reveals an intrinsic geometric separability in LLM representations, establishing a novel paradigm for interpretable modeling and controllable reasoning.

Enhancing word analogy accuracy through transformed hidden statesIdentifying semantic subspaces in LLMs using concept induction headsRevealing surface-level token relationships via token induction heads

Causal attention in decoder-only LLMs introduces sentence embedding encoding bias—early tokens cannot attend to subsequent content, leading to information loss and error propagation. To address this, we propose Hierarchical Token Prepending (HTP), a training-free, plug-and-play method that dynamically prepends the sentence embedding generated by the previous Transformer layer to the input sequence of each subsequent layer. This enables early tokens to indirectly capture full-sentence semantics via cross-layer feature re-injection. HTP is the first approach to achieve sentence-level feature re-injection under causal masking without modifying model architecture or parameters, and it is fully compatible with any prompt-based embedding method and autoregressive LLMs (e.g., Llama, Qwen). Extensive evaluation on multiple STS semantic similarity benchmarks and downstream classification tasks demonstrates consistent and significant improvements over strong baselines, with negligible inference overhead.

Biased sentence encoding in decoder-only LLMs due to causal attentionEnhancing semantic understanding without additional inference costNeed for training-free method to improve sentence embeddings

Latest Papers

What's happening recently
View more

本文提出了一种条件初始化方法,通过优化注意力层的谱特性来改进Transformer模型的训练稳定性和收敛速度。

attention layeroptimization biastraining dynamics

This work addresses the limitation of conventional Transformers, which employ uniform-dimensional key/value (KV) representations for all historical tokens, thereby neglecting the differing contributions of local and distant context to prediction. The authors propose Distance-Adaptive Representation (DAR), a method that retains full-dimensional KV representations within a local window while using reduced-dimensional representations for tokens beyond this window. This approach provides the first systematic validation that local and global tokens can be effectively represented with different dimensions, challenging the prevailing assumption that KV dimensions must remain consistent across all tokens. Experiments on decoder-only architectures demonstrate that DAR matches the performance of full-dimensional baselines across pretrained models ranging from 70M to 410M parameters and a 1B fine-tuned model, significantly outperforms global dimensionality reduction schemes, and offers potential reductions in KV cache memory during inference.

attention mechanismdimensionality allocationKV cache

This work investigates how attention mechanisms enhance the ability of sequence models to recover latent signals from noisy observations in the high-dimensional limit. Leveraging random matrix theory, the authors analyze the spectral properties of the sample covariance matrix after attention-weighted pooling, characterizing its eigenvalue distribution, outlier eigenvalues, and the alignment between eigenvectors and the true signal. The analysis reveals two BBP-type phase transitions governing signal recovery and demonstrates that optimal attention weights correspond to the leading eigenvector of a location-dependent kernel matrix. Moreover, a specific causal self-attention mechanism is shown to yield deterministic harmonic weights that outperform simple averaging. Combining free multiplicative convolution, high-dimensional spectral analysis, and Gaussian mixture embedding models, the proposed framework exhibits excellent agreement with finite-dimensional experiments, quantitatively elucidating the advantage of attention in boosting signal-to-noise ratio and recovery performance.

attention mechanismhigh-dimensional statisticsrandom matrices

This work addresses the challenges of influence decay in long-context prefixes during generation and the linearly growing computational cost of attention with sequence length. The authors propose a training-free attention state memory mechanism that externalizes the prefix into a lightweight, dynamically updatable store of precomputed attention states. By leveraging query-prefix attention precomputation and memory-efficient lookup, the method enables efficient reuse of prefix information during inference. This approach is the first to compress long-prefix information into a flexible external memory without requiring any training. On the ManyICLBench benchmark, it surpasses in-context learning accuracy under memory budgets of 1K–8K tokens—achieving a 1.36× reduction in inference latency at 8K—and outperforms full-attention RAG on the NBA benchmark using only 20% of the memory.

attention efficiencycontext memorizationinference latency

Hot Scholars

ZS

Zhan Su

University of Montreal;MILA
PEFT approachesLLMsInformation retrieval
JY

Jian-Yun Nie

university of montreal
information retrievalnatural language processing
YH

Yulan He

Professor, King's College London; Turing AI Fellow
Natural Language ProcessingLarge Language ModelsAI for education and health
PY

Peng Yan

Research Assistant of ZHAW, PhD student of UZH
Deep LearningTransfer LearningIntelligent Algorithm