identify extremal words

Designs and implements methods to score, rank, and select tokens or words by their extremality within a model’s representation or output space, identifying top and bottom tokens. Builds characterizations of the corresponding extremal codewords and derives minimal subsets of tokens that suffice to reconstruct or reproduce those extreme behaviors.

identifyextremalwords

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of providing language models with verifiable and unforgeable identity credentials without exposing their internal parameters. It introduces, for the first time, a model-signing mechanism based on the ranking of top-k output tokens, relying solely on ordinal information rather than exact probability values. Theoretical analysis demonstrates that forging such a signature is NP-hard, thereby ensuring computational security against polynomial-time adversaries. By integrating geometric constraints inherent in language model outputs, feasibility analysis of token rankings, and complexity-theoretic arguments—and further substantiated through model extraction experiments—the study confirms the method’s efficacy: while attackers can leverage ranking data to approximate final-layer parameters, they cannot successfully forge valid signatures. Moreover, selecting an appropriately small top-k value preserves both the uniqueness and robustness of the signature while safeguarding model confidentiality.

API securitylanguage model signaturesmodel stealing

Large language models (LLMs) excel at code generation, yet the impact of their compressed variants—such as those produced via quantization or knowledge distillation—on token representations of programming languages remains poorly understood, hindering deployment quality. This work systematically investigates how LLM tokenizers encode programming languages and introduces a novel “cold-start probability” analysis method that operates without explicit prompting. By integrating lexical distribution analysis, keyword coverage, and multidimensional evaluation metrics, the study provides the first comprehensive characterization of the subtle effects of compression strategies—including quantization, knowledge distillation, model scaling, and task-specific fine-tuning—on code token representations. The findings offer both theoretical grounding and empirical guidance for deploying high-quality, efficient code generation models.

code generationcompressed LLMsdistillation

Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.

Compression-optimized tokenization benefits low-resource languagesImproved performance in multilingual NLP tasksOptimal BPE configuration reduces token count

Retrofitting (Large) Language Models with Dynamic Tokenization

Nov 27, 2024
DF
Darius Feher
🏛️ University of Cambridge

Existing pre-trained language models rely on static subword tokenizers, leading to suboptimal multilingual efficiency and imbalanced cross-lingual performance. To address this, we propose the first dynamic tokenization framework tailored for pre-trained language models: it dynamically identifies high-frequency subword sequences in each input and merges them on-the-fly; a lightweight hypernetwork then instantaneously generates token embeddings, enabling input-adaptive subword boundary decisions. Our method integrates a BPE-inspired intra-batch merging algorithm and is compatible with both encoder (e.g., XLM-R) and decoder (e.g., Mistral-7B) architectures. Evaluated across 14 languages, it achieves over 20% average sequence length reduction for XLM-R with less than 2% performance degradation; for English decoding, it shortens sequences by 6%, accelerates inference, and significantly improves multilingual fairness.

Dynamic tokenization improves inference speed and fairnessMethod reduces token sequence lengths without major performance lossStatic tokenizers reduce efficiency and language capabilities

Latest Papers

What's happening recently
View more

This work investigates whether large language models must rely on trainable input embedding tables. To address this, the authors propose a novel approach that entirely eliminates trainable input embeddings by employing fixed 16-dimensional binary token codes, combined with a zero-parameter dimensional expansion and an invertible affine recoding mechanism over a finite field, while retaining the standard trainable output projection. Evaluated on a 32-layer decoder model, this method achieves validation perplexity comparable to the baseline (2.36 vs. 2.44), reducing input parameters by approximately 67.1 million. A vocabulary-independent variant of the approach also performs closely (2.39). This study presents the first demonstration that large-scale language models can completely dispense with trainable input embeddings without sacrificing performance.

binary token codesembedding tableinput embedding

This work investigates how different sequence representations—such as bytes, characters, and subwords—affect the information acquisition capacity of Transformer models under a fixed context window, a question that remains poorly understood. From an information-theoretic perspective, the paper introduces the notion of “fragmentation” and formally demonstrates that it inherently increases the log-loss of the optimal finite-context model. It establishes theoretical guarantees linking tokenization compression rates to the reliability of source context coverage, thereby constructing the first information-theoretic framework for representation selection in finite-context settings. Through Markov source modeling and comparative analysis of various tokenization strategies—including BPE, WordPiece, and byte-level methods—the study reveals the theoretical underpinnings of performance differences observed in models like ByT5 and CANINE, and proposes practical metrics to evaluate the effective context coverage of real-world tokenizers.

finite-context predictionfragmentationrepresentation

This study systematically investigates the impact of programming language choice on token consumption during code generation by reasoning agents. Through controlled experiments, it evaluates token efficiency across five state-of-the-art models solving problems of equivalent difficulty in Python, Java, Rust, and OCaml. The work introduces a multidimensional trajectory analysis framework integrating trace re-execution, test-result vector abstraction, intermediate solution annotation, and natural language analysis. It presents the first quantitative assessment of cross-language token usage disparities among agents, revealing that language familiarity significantly influences generation behavior: in less familiar languages, agents are more prone to producing uncompilable code, redundantly modifying already correct solutions, and resorting to Python prototyping to circumvent direct implementation in the target language.

coding agentsmultilingual agentsprogramming languages

Existing language model decoding methods lack a unified theoretical framework and often rely on heuristic hyperparameter tuning. This work formulates the decoding process as a regularized optimization problem over the probability simplex and provides a unified interpretation of multiple mainstream decoding strategies through the introduction of KL-divergence anchoring and analysis of optimality conditions. Building upon this framework, we propose a novel Best-of-K sampler that substantially improves generation quality under high-temperature settings. Experimental results demonstrate that our method achieves an 18.6% absolute accuracy gain on the MATH500 benchmark when applied to the Qwen2.5-Math-7B model, confirming its effectiveness and broad applicability.

decodinglanguage modelsoptimization

Hot Scholars

SI

Shunsuke Inenaga

Professor, Department of Informatics, Kyushu University
Algorithms and Data StructuresString AlgorithmsCompressionCombinatorics on Words
TM

Takuya Mieno

The University of Electro-Communications
Stringology