selective token supervision

Design and evaluate methods that select specific tokens within text sequences to receive supervision or ground-truth signals, including algorithms for semantic token selection and for filtering noisy tokens produced by external generators. Build training procedures and analyses that use those selected tokens to compute loss and propagate gradients, transfer learned signals to unsupervised tokens, and trade off supervision sparsity against representational compactness and task performance.

selectivetokensupervision

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Unsupervised text representation learning (TRL) suffers from limited representational quality due to the absence of explicit supervision. To address this, we propose Text2Token, the first unsupervised TRL framework based on token-level target prediction. Text2Token constructs high-quality synthetic token distributions via dual pathways—data-driven and model-derived—and employs a large language model backbone to predict token-level distributions, enabling fine-grained alignment between text embeddings and salient semantic units. Crucially, the representation space and vocabulary space are jointly optimized during training, leading to improved convergence toward superior solutions. Evaluated on the MTEB v2 benchmark, Text2Token achieves performance on par with the state-of-the-art contrastive method LLM2Vec, demonstrating that generative token prediction constitutes an effective and competitive paradigm for unsupervised text representation learning.

Aligns vocabulary and representation spaces toward optimal training solutionConstructs target token distribution using data-driven and model-derived methodsDevelops unsupervised generative framework for text representation learning

Discrete tokenizers lack a systematic, cross-task survey. Method: We propose the first unified analytical framework covering generation, understanding, recommendation, and information retrieval; introduce a hierarchical decomposition paradigm for tokenizer submodules; establish a cross-task taxonomy; and conduct a horizontal comparison of representative approaches—including VQ-VAE, SoundStream, K-means tokenization, semantic hashing, and cross-modal alignment—through the lenses of information theory, representation learning, and structured modeling. Contribution/Results: We identify three core challenges: semantic alignment, cross-modal generalization, and the efficiency–accuracy trade-off. Furthermore, we deliver a reusable evaluation dimension matrix and an open challenge map, providing both theoretical foundations and practical guidelines for designing next-generation tokenizers that are robust, interpretable, and cross-modal.

Guide future research on tokenizers for AI advancementReview design, applications, and challenges of tokenizersSurvey of discrete tokenizers in AI systems

This work addresses a key limitation of conventional large language model training, which relies on token-level next-token prediction and consequently struggles to distinguish between semantically equivalent expressions that differ in surface form, leading to a bias toward superficial patterns rather than deep semantic understanding. To overcome this, the authors propose elevating the training objective to the conceptual level by introducing a systematic concept-level supervision signal through a concept mapping framework—e.g., unifying surface variants like “mom” and “mother” under a shared concept such as MOTHER. Their approach integrates a concept alignment loss with a multi-surface aggregation strategy, encouraging the model to prioritize semantic correctness. Experiments demonstrate that the resulting concept-aware models achieve lower perplexity, superior performance across multiple NLP benchmarks, and enhanced robustness in domain transfer scenarios.

concept-level traininglarge language modelsnext-token prediction

Where is the signal in tokenization space?

Aug 16, 2024
RL
Renato Lui Geh
🏛️ University of California, Los Angeles

This work challenges the conventional assumption in large language models (LLMs) that the probability of a text string equals the probability of its canonical tokenization, revealing that non-canonical tokenizations of the same string encode underutilized semantic and structural signals. We first prove that, under autoregressive LLMs, both finding the most probable tokenization and computing marginal probabilities across all tokenizations are NP-hard. To address this, we propose an efficient approximation algorithm based on dynamic programming with aggressive pruning, compatible with diverse architectures including Transformers and State Space Models (SSMs). Empirically, aggregating marginal probabilities over non-canonical tokenizations—without modifying model parameters or training—yields consistent performance gains across multiple LLM evaluation benchmarks (e.g., LM Evaluation Harness and HELM subsets). These results demonstrate that the tokenization space harbors exploitable latent probabilistic structure, offering a novel, architecture-agnostic avenue for improving LLM inference.

Aggregating non-canonical tokenization probabilities improves LLM benchmark performanceFinding the most likely tokenization for autoregressive LLMs is computationally hardMarginal probability over all tokenizations often matches canonical probability

Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.

Compression-optimized tokenization benefits low-resource languagesImproved performance in multilingual NLP tasksOptimal BPE configuration reduces token count

Latest Papers

What's happening recently
View more

本文探讨了语言模型中分词策略对模型训练的影响,通过实验表明输出分词决定了模型的学习难度与内部表示,并指出当前研究对此重视不足。

Autoregressive ModelsModel PerformanceNumeric Reasoning

This work proposes three key techniques to enhance computational efficiency in large language model training while preserving or even improving performance. Selective Ground Truth training (SGT) achieves a 67% reduction in loss using only 15% of the original supervision signal. A depth-compression strategy combining layer averaging with unrolling attains the performance of larger models under a 2.5× parameter compression ratio. Additionally, a Mixture-of-Efficient-Experts (MoEE) architecture integrates dual experts to lower validation loss to 2.789. Together, these methods enable the construction of a highly efficient and high-performing Korean foundation model that significantly reduces computational overhead without compromising—indeed, sometimes enhancing—model efficacy.

compute-efficient language modelsdepth compressionMixture of Experts

This study investigates the fundamental limits of language models in learning the true data-generating process of natural language solely from observed text sequences, particularly when unobserved contextual factors—such as facts, intentions, or social settings—influence generation. By distinguishing between the full conditional generative process, the marginalized text-only process, and the distribution learned by the model, the work introduces criteria based on local sufficient statistics and conditional mutual information to delineate when next-token prediction remains valid. It demonstrates that standard training implicitly assumes stationarity and ergodicity, assumptions often violated in heterogeneous corpora, and proves that marginal models succeed only when observed prefixes are approximately sufficient for latent variables. The paper further interprets retrieval-augmented generation (RAG) and tool use as mechanisms that restore conditional sufficiency, thereby transcending conventional language modeling paradigms.

conditional sufficiencyergodicitylanguage models

This study addresses the challenge of aligning generated content with user intent in large language and vision-language models by presenting a systematic review of decoding methods during the inference phase. The research categorizes emerging approaches into three paradigms: token-level guidance, sequence-level generation, and parallel acceleration, while establishing a dedicated resource repository. As the first comprehensive survey of decoding strategy evolution, this work elucidates the critical role of these methods in enhancing both generation efficiency and alignment quality. Furthermore, it outlines future research directions, providing essential theoretical foundations and practical references for optimizing model inference performance.

Decoding MethodsInference-time ControlLarge Language Models

Hot Scholars

PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
ZL

Zuchao Li

Wuhan University
Natural Language ProcessingMachine Learning
XW

Xinchao Wang

National University of Singapore
Machine LearningAIComputer VisionImage Processing
SZ

Sihang Zhou

NUDT
Machine LearningMedical Image AnalysisInformation Fusion
XS

Xiaoyu Shen

Eastern Institute of Technology, Ningbo
language modelmulti-modal learningreasoning