language modeling

Designs and implements models that assign probabilities to sequences of natural-language tokens and generate text, including choices of architecture, tokenization, training objectives, and decoding algorithms. Builds evaluation and analysis pipelines for metrics (perplexity, likelihood, calibration), qualitative generation assessment, and investigation of learned representations and failure modes.

languagemodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$233K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Large language models face multifaceted uncertainties across prompt formulation, text generation, and downstream interpretation, yet lack a unified framework for modeling and quantifying these uncertainties. This work proposes the first formal, unified framework that conceptualizes these interrelated stages as coupled autoregressive processes, structured through a sampling tree. Within this framework, various sources of uncertainty are characterized via filtering mechanisms and objective functions. The approach not only reveals commonalities and intrinsic connections among existing uncertainty quantification methods but also subsumes diverse prior techniques under a coherent theoretical lens. Furthermore, it identifies previously unexplored dimensions of uncertainty, thereby opening new avenues for future research in reliable and interpretable language model deployment.

interpretationlarge language modelsprompting

Low-Perplexity LLM-Generated Sequences and Where To Find Them

Jul 02, 2025
AW
Arthur Wuhrmann
🏛️ École Polytechnique Fédérale de Lausanne | HES-SO Valais-Wallis

Understanding how large language models (LLMs) memorize training data is critical for ensuring transparency, accountability, privacy, and fairness. This paper introduces a systematic provenance tracing method that combines low-perplexity sequence detection, deduplication, alignment, and scalable, efficient text-matching algorithms to precisely attribute generated content to its original training corpus. Experiments reveal that a substantial fraction of high-probability generations cannot be localized in the source corpus—uncovering a widespread “unmapped memory” phenomenon wherein LLMs reproduce training data without detectable lexical or structural correspondence to existing provenance techniques. The study provides the first quantitative characterization of the distribution of traceable versus untraceable low-perplexity sequences, empirically confirming the coexistence of direct memorization and implicit reproduction. These findings establish a novel methodology and empirical foundation for data provenance analysis, copyright assessment, and model auditing.

Analyzing low-perplexity sequences in LLM outputsTracing model-generated text to training data sourcesUnderstanding verbatim recall distribution in LLM behavior

Locally Typical Sampling

Feb 01, 2022
CM
Clara Meister
🏛️ ETH Zürich | University of Cambridge

Existing probabilistic language generators achieve strong performance on metrics like perplexity but often produce text lacking coherence, fluency, and diversity. To address this, we propose Locally Typical Sampling—a novel decoding strategy that formalizes principles of efficiency and robustness from human linguistic communication as an information-theoretic criterion based on the expected conditional entropy. At each decoding step, the method retains only tokens whose log-probabilities lie within a dynamically computed neighborhood of the model’s current conditional entropy, enabling lightweight, parameter-free, and adaptive probability truncation. Unlike prior methods, it requires no additional training or hyperparameter tuning. Empirical evaluation on summarization and story generation tasks shows that Locally Typical Sampling significantly reduces repetition while maintaining fluency and coherence comparable to nucleus (top-p) and top-k sampling. Results are validated through both automated metrics and human evaluation.

Addressing dullness and repetition in high-probability text outputsEnhancing efficiency and error-minimization in language generationImproving coherence and fluency in probabilistic language generation

Looking beyond the next token

Apr 15, 2025
AT
Abitha Thankaraj
🏛️ Carnegie Mellon University

Conventional causal language models assume token-level autoregressive generation conditioned solely on preceding context, fundamentally misaligning with human writing and reasoning—where goal specification precedes content generation. Method: We propose Trelawney, a data reordering technique that implicitly injects long-horizon goal signals via sequence-level permutation, without altering model architecture or training procedure. This enables standard Transformer training to spontaneously acquire goal-directed generation capabilities. Contribution/Results: Trelawney yields interpretable, goal-conditioned generation and enables goal-guided reasoning algorithms. It achieves significant performance gains across planning, algorithmic reasoning, and story generation benchmarks—demonstrating, for the first time, zero-cost extension of language models’ goal modeling capacity while preserving standard training paradigms.

Addressing mismatch between human writing and token predictionEnabling long-term goal generation in language modelsImproving planning and reasoning without architectural changes

The Foundations of Tokenization: Statistical and Computational Concerns

Jul 16, 2024
JL
Juan Luis Gastaldi
🏛️ ETH Zürich | City University of New York

This paper addresses the lack of theoretical foundations for tokenization in natural language processing (NLP), systematically investigating its impact on the statistical estimation consistency of language models. While prior work relies predominantly on empirical analysis, we introduce the first unified formal framework grounded in the category of random mappings to rigorously characterize the modeling essence of tokenizers. Our key contributions are: (1) necessary and sufficient conditions for tokenizers to preserve statistical estimation consistency; (2) a four-dimensional theoretical analysis framework—covering inconsistency, ambiguity, finiteness, and sequentiality; and (3) principled, verifiable tokenizer design criteria derived from the integration of category theory, statistical learning theory, and formal language theory. This work establishes the first rigorous mathematical foundation for representation reliability in neural language modeling.

Addressing statistical and computational concerns in tokenizer designAnalyzing tokenization's impact on language model consistencyUnderstanding theoretical foundations of tokenization in NLP

Latest Papers

What's happening recently
View more

本文从概率角度探讨大型语言模型,通过自回归条件分布和最大似然估计训练模型,分析了KL散度在文本生成中的作用,并讨论了扩散模型。

Kullback--Leibler DivergenceLarge Language ModelsText Generation

This work addresses the challenge of interpreting the contribution of input tokens to outputs in large language model generation by proposing the first model-agnostic probabilistic attribution method. The approach models text generation as a stochastic process and leverages Bayes’ rule to infer the conditional probability of a response given a prompt. Attribution scores are defined via the logarithm of probability ratios, while conditional entropy is introduced to quantify context sensitivity and generation uncertainty. Experiments across eight mainstream models and seven prompt categories demonstrate that the method effectively identifies anomalous generations, token-sensitive regions, and unstable behaviors, substantially enhancing users’ awareness and understanding of generative uncertainty.

interpretabilitylarge language modelsprobabilistic attribution

This study addresses whether generative AI systems, by virtue of memorizing training data, produce outputs that constitute legally actionable “copies” of copyrighted works. Integrating insights from the memory mechanisms and probabilistic generation behaviors of large language models, the paper offers the first systematic interdisciplinary analysis arguing that copyright law should adopt a functional standard to determine whether an AI model contains a “copy.” The research demonstrates that current legal frameworks typically recognize copying only when specific protected works can be readily extracted from the model, thereby exposing significant limitations in the applicability of existing doctrines in the AI era. Building on this finding, the work proposes targeted legal reforms to better align copyright enforcement with the technical realities of modern generative systems.

copyrightgenerative AILLMs

This work addresses the lack of systematic and reproducible frameworks for natural language processing (NLP) research and deployment in low-resource languages by proposing an open-source, end-to-end practical pipeline that spans the full modern NLP workflow. Built around a unified corpus across twelve experimental stages, the framework integrates subword tokenization, vectorization, large model fine-tuning, retrieval-augmented generation, and reinforcement learning from human feedback, all implemented within the Hugging Face ecosystem using openly licensed models to avoid reliance on proprietary APIs. The project delivers the first publicly available tokenizers, embeddings, lexicons, and transliteration benchmarks for languages such as Tajik and Tatar, forming a comprehensive educational curriculum tailored for advanced undergraduates, graduate students, and practitioners, thereby significantly advancing reproducible research and capacity building in low-resource language NLP.

large language modelslow-resource languagesNatural Language Processing

This study addresses the lack of standardized evaluation protocols in machine-generated text detection, which hinders fair comparison of model performance. The authors systematically evaluate 15 detection methods across diverse datasets comprising both human-written and machine-generated English texts, covering six detector families and seven generative models. Employing a multi-dataset cross-validation framework and multiple evaluation metrics, they find that no single detector consistently outperforms others across all scenarios—most excel only in specific settings—and overall performance degrades significantly on novel, human-authored texts from high-stakes domains. The work highlights the strong dependence of detector efficacy on training and evaluation data as well as metric choice, exposing critical blind spots in current evaluation paradigms and underscoring the decisive role of methodological decisions in shaping empirical conclusions.

dataset biasevaluation metricsgenerative language models

Hot Scholars

NH

Nizar Habash

Professor of Computer Science, New York University Abu Dhabi
Natural Language ProcessingComputational LinguisticsArtificial Intelligence
TP

Tiago Pimentel

ETH Zurich
Natural Language ProcessingComputational LinguisticsInformation TheoryMachine Learning
OP

Oiwi Parker Jones

Applied Artificial Intelligence and Clinical Neurosciences, University of Oxford
AINeuroscienceDeep LearningSpeech Recognition
SW

Shinji Watanabe

Carnegie Mellon University
Speech recognitionSpeech processingSpeech enhancementSpeech translation