soft-token injection

Designs and implements learned continuous (soft) tokens that are inserted into sequence-model inputs to integrate new information or act as query handles, covering token-based integration and token-based querying within transformer-style architectures. These methods specify how such tokens are parameterized and trained end-to-end to fit a fixed context length while providing a compact, low-memory/compute mechanism for adding or retrieving information from the model.

soft-tokeninjection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Law of Next-Token Prediction in Large Language Models

Aug 24, 2024
HH
Hangfeng He
🏛️ University of Rochester | University of Pennsylvania

The black-box nature of large language models (LLMs) severely limits their interpretability. Method: We conduct the first systematic empirical and formal analysis of per-layer contribution to next-token prediction across dominant architectures—including Transformer, RWKV, and Mamba—quantifying gradient flows, hidden-state evolution, and cross-architectural consistency. Contribution/Results: We discover and formally verify a universal linear scaling law: contextualized token embeddings’ predictive capacity increases approximately linearly with layer depth, implying near-uniform per-layer contribution to next-token prediction. This pattern holds consistently across models, training scales, and tasks. Our findings reveal a fundamental, quantifiable structure in LLM internal information flow—establishing a novel interpretability paradigm grounded in empirical regularity. The result provides theoretical foundations for principled model scaling laws, improved pretraining objectives, and controllable information-flow engineering in LLMs.

Establishing universal law across diverse model architecturesQuantifying layer contributions to prediction accuracyUnderstanding internal token embedding learning in LLMs

TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters

Oct 30, 2024
HW
Haiyang Wang
🏛️ Max Planck Institute for Informatics | Peking University | Google

Scaling Transformer models is prohibitively expensive due to fixed-parameter linear projection layers; architectural modifications necessitate full retraining. Method: We propose TokenFormer, the first architecture introducing *parameter tokenization*, which models model parameters as learnable tokens and replaces all linear layers with token-parameter self-attention—unifying parameter and input token representations in a shared latent space. Contribution/Results: Our method enables zero-shot, progressive parameter expansion without retraining, overcoming classical scaling bottlenecks. Without altering network topology, we scale model parameters from 124M to 1.4B while matching the performance of fully trained baselines, achieving substantial training cost reduction. The code and models are publicly released.

Dependence on fixed parameters requiring full retrainingHigh computational cost of scaling Transformer modelsLack of efficient progressive scaling for large models

Retrofitting (Large) Language Models with Dynamic Tokenization

Nov 27, 2024
DF
Darius Feher
🏛️ University of Cambridge

Existing pre-trained language models rely on static subword tokenizers, leading to suboptimal multilingual efficiency and imbalanced cross-lingual performance. To address this, we propose the first dynamic tokenization framework tailored for pre-trained language models: it dynamically identifies high-frequency subword sequences in each input and merges them on-the-fly; a lightweight hypernetwork then instantaneously generates token embeddings, enabling input-adaptive subword boundary decisions. Our method integrates a BPE-inspired intra-batch merging algorithm and is compatible with both encoder (e.g., XLM-R) and decoder (e.g., Mistral-7B) architectures. Evaluated across 14 languages, it achieves over 20% average sequence length reduction for XLM-R with less than 2% performance degradation; for English decoding, it shortens sequences by 6%, accelerates inference, and significantly improves multilingual fairness.

Dynamic tokenization improves inference speed and fairnessMethod reduces token sequence lengths without major performance lossStatic tokenizers reduce efficiency and language capabilities

This work investigates why Transformers exhibit both strong and weak performance on symbolic reasoning tasks. Method: We introduce Production System Language (PSL)—a fully mechanistically interpretable, Turing-complete symbolic language—and design an exact compiler that maps PSL programs to Transformer weights. Our approach integrates production-system modeling, the semantics-agnostic Templatic Generation (TGT) benchmark, and an intrinsically interpretable architecture to enable end-to-end, traceable compilation of symbolic programs into Transformer parameters. Contributions/Results: (1) The first 100% mechanistically interpretable Transformer implementation for symbolic processing; (2) zero-shot abstract reasoning on TGT, without task-specific training; (3) mechanistic insights into in-context learning (ICL), revealing both its inherent symbolic operations and fundamental limitations; and (4) a verifiable, architecture-level roadmap for enhancing large language models’ symbolic capabilities.

Developing a symbolic programming language for transformersEnhancing transformer capabilities in abstract symbol manipulationUnderstanding symbol processing mechanisms in transformers

TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection

Nov 05, 2024
WW
Wei Wu
🏛️ University of Science and Technology of China | Tsinghua University | Alibaba Cloud Computing | The Hong Kong University of Science and Technology

Large language models (LLMs) suffer from both performance degradation and high latency—due to quadratic computational complexity—when processing ultra-long contexts. To address this, we propose a training-free, dynamic token-level KV cache selection mechanism. Our approach introduces the first token-level quantification of KV importance based on discontinuous attention sparsity, coupled with a per-head soft-voting strategy for fine-grained cache pruning. We further design a dedicated Selection Cache structure and a customized high-efficiency dot-product kernel to jointly optimize accuracy and inference speed. Experiments demonstrate that our method maintains state-of-the-art performance on long-context tasks while accelerating attention computation by up to 23.84× and reducing end-to-end latency by 2.28× compared to existing approaches.

Addresses performance degradation in LLMs with long sequences.Enables efficient long-context inference without sacrificing accuracy.Reduces excessive inference times due to quadratic attention complexity.

Latest Papers

What's happening recently
View more

This study addresses the computational bottlenecks of byte-level language models caused by excessively long sequences and the absence of explicit textual abstraction. To overcome these limitations, this work proposes a tokenizer-free architecture that leverages a Token Hyper-position training strategy alongside hash embeddings, enabling standard Transformers to model efficiently at the byte level. The findings demonstrate that additional computation can effectively substitute fixed tokenizers, allowing models to spontaneously construct local context representation mechanisms. Furthermore, the proposed approach surpasses subword-based models in performance at large scales. Notably, by exploiting the non-uniform uncertainty inherent in generated outputs, the method facilitates speculative decoding, achieving a 3.4-fold improvement in acceptance rates.

Byte Language ModelsEmergent AbstractionsTokenizer-free

Existing theories of in-context learning typically assume that demonstration examples in prompts are independent and identically distributed, thereby overlooking the pervasive temporal correlations present in real-world sequences. This work develops a theoretically tractable model based on linear attention, integrating a linear regression theory sandbox with an actual Transformer architecture to systematically investigate how temporal dependencies within prompts affect in-context learning. We reveal for the first time that such temporal correlations induce an “effective context length,” rendering correlated prompts equivalent to shorter i.i.d. ones. Moreover, we find that when queries are temporally aligned with the context, Softmax attention substantially outperforms linear attention, highlighting the critical importance of aligning attention mechanisms with task-specific structural properties.

architectural mismatchattention mechanismseffective context length

This work addresses the limitation of existing large language model (LLM)-based recommender systems in effectively leveraging fine-grained non-textual features—such as continuous numerical values and dense embeddings—due to their native architecture being designed exclusively for discrete textual tokens. To overcome this, the authors propose a soft token fusion framework that, for the first time, maps heterogeneous features into the LLM’s embedding space through standard token interfaces. An interactive fusion module is further introduced to enhance feature integration. Integrated within a parameter-shared dual-tower LLM retrieval architecture, the proposed method significantly outperforms current LLM-based baselines on three Amazon recommendation benchmarks, demonstrating that interactive fusion is more effective than simple concatenation for combining multimodal signals.

embedding featuresheterogeneous signalsLLM-based recommender systems

This study investigates whether language models inherently require trainable word embedding tables by systematically exploring whether fixed token encodings can sustain core modeling capabilities. Based on a decoder-only Transformer architecture, we compare three input interfaces—learned embeddings, canonical binary encoding, and GF(2) invertible recoding—employing a de-parameterized input projection alongside a fixed 16-bit token encoding scheme. This work is the first to disentangle architectural necessity from empirical utility in this context. Notably, at the 1.7B parameter scale, removing approximately 100 million input parameters still enables the fixed-encoding model to achieve a substantial performance of 52.4% on HellaSwag. These results demonstrate that independent, trainable word vectors are not strictly necessary for effective language modeling, thereby establishing a feasibility benchmark for parameter-free input representations.

fixed token codesinput embedding tablelanguage models

Hot Scholars

PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
SG

Shivank Garg

UG Student Artificial Intelligence and Data Science,IIT Roorkee
Deep LearningGenerative AIAI Security
BX

Bingxuan Xu

Beijing university of posts and telecommunications
semantic communicationdiffusion model
QL

Qingfeng Li

Institute of Automation, Chinese Academy of Science
GenAI
ZL

Zequn Liu

Microsoft Research AI4Science, Asia