Score
Designs and implements models that assign probabilities to sequences of natural-language tokens and generate text, including choices of architecture, tokenization, training objectives, and decoding algorithms. Builds evaluation and analysis pipelines for metrics (perplexity, likelihood, calibration), qualitative generation assessment, and investigation of learned representations and failure modes.
Large language models face multifaceted uncertainties across prompt formulation, text generation, and downstream interpretation, yet lack a unified framework for modeling and quantifying these uncertainties. This work proposes the first formal, unified framework that conceptualizes these interrelated stages as coupled autoregressive processes, structured through a sampling tree. Within this framework, various sources of uncertainty are characterized via filtering mechanisms and objective functions. The approach not only reveals commonalities and intrinsic connections among existing uncertainty quantification methods but also subsumes diverse prior techniques under a coherent theoretical lens. Furthermore, it identifies previously unexplored dimensions of uncertainty, thereby opening new avenues for future research in reliable and interpretable language model deployment.
Understanding how large language models (LLMs) memorize training data is critical for ensuring transparency, accountability, privacy, and fairness. This paper introduces a systematic provenance tracing method that combines low-perplexity sequence detection, deduplication, alignment, and scalable, efficient text-matching algorithms to precisely attribute generated content to its original training corpus. Experiments reveal that a substantial fraction of high-probability generations cannot be localized in the source corpus—uncovering a widespread “unmapped memory” phenomenon wherein LLMs reproduce training data without detectable lexical or structural correspondence to existing provenance techniques. The study provides the first quantitative characterization of the distribution of traceable versus untraceable low-perplexity sequences, empirically confirming the coexistence of direct memorization and implicit reproduction. These findings establish a novel methodology and empirical foundation for data provenance analysis, copyright assessment, and model auditing.
Existing probabilistic language generators achieve strong performance on metrics like perplexity but often produce text lacking coherence, fluency, and diversity. To address this, we propose Locally Typical Sampling—a novel decoding strategy that formalizes principles of efficiency and robustness from human linguistic communication as an information-theoretic criterion based on the expected conditional entropy. At each decoding step, the method retains only tokens whose log-probabilities lie within a dynamically computed neighborhood of the model’s current conditional entropy, enabling lightweight, parameter-free, and adaptive probability truncation. Unlike prior methods, it requires no additional training or hyperparameter tuning. Empirical evaluation on summarization and story generation tasks shows that Locally Typical Sampling significantly reduces repetition while maintaining fluency and coherence comparable to nucleus (top-p) and top-k sampling. Results are validated through both automated metrics and human evaluation.
Conventional causal language models assume token-level autoregressive generation conditioned solely on preceding context, fundamentally misaligning with human writing and reasoning—where goal specification precedes content generation. Method: We propose Trelawney, a data reordering technique that implicitly injects long-horizon goal signals via sequence-level permutation, without altering model architecture or training procedure. This enables standard Transformer training to spontaneously acquire goal-directed generation capabilities. Contribution/Results: Trelawney yields interpretable, goal-conditioned generation and enables goal-guided reasoning algorithms. It achieves significant performance gains across planning, algorithmic reasoning, and story generation benchmarks—demonstrating, for the first time, zero-cost extension of language models’ goal modeling capacity while preserving standard training paradigms.
This paper addresses the lack of theoretical foundations for tokenization in natural language processing (NLP), systematically investigating its impact on the statistical estimation consistency of language models. While prior work relies predominantly on empirical analysis, we introduce the first unified formal framework grounded in the category of random mappings to rigorously characterize the modeling essence of tokenizers. Our key contributions are: (1) necessary and sufficient conditions for tokenizers to preserve statistical estimation consistency; (2) a four-dimensional theoretical analysis framework—covering inconsistency, ambiguity, finiteness, and sequentiality; and (3) principled, verifiable tokenizer design criteria derived from the integration of category theory, statistical learning theory, and formal language theory. This work establishes the first rigorous mathematical foundation for representation reliability in neural language modeling.
本文从概率角度探讨大型语言模型,通过自回归条件分布和最大似然估计训练模型,分析了KL散度在文本生成中的作用,并讨论了扩散模型。
This work addresses the challenge of interpreting the contribution of input tokens to outputs in large language model generation by proposing the first model-agnostic probabilistic attribution method. The approach models text generation as a stochastic process and leverages Bayes’ rule to infer the conditional probability of a response given a prompt. Attribution scores are defined via the logarithm of probability ratios, while conditional entropy is introduced to quantify context sensitivity and generation uncertainty. Experiments across eight mainstream models and seven prompt categories demonstrate that the method effectively identifies anomalous generations, token-sensitive regions, and unstable behaviors, substantially enhancing users’ awareness and understanding of generative uncertainty.
This study addresses whether generative AI systems, by virtue of memorizing training data, produce outputs that constitute legally actionable “copies” of copyrighted works. Integrating insights from the memory mechanisms and probabilistic generation behaviors of large language models, the paper offers the first systematic interdisciplinary analysis arguing that copyright law should adopt a functional standard to determine whether an AI model contains a “copy.” The research demonstrates that current legal frameworks typically recognize copying only when specific protected works can be readily extracted from the model, thereby exposing significant limitations in the applicability of existing doctrines in the AI era. Building on this finding, the work proposes targeted legal reforms to better align copyright enforcement with the technical realities of modern generative systems.
This work addresses the lack of systematic and reproducible frameworks for natural language processing (NLP) research and deployment in low-resource languages by proposing an open-source, end-to-end practical pipeline that spans the full modern NLP workflow. Built around a unified corpus across twelve experimental stages, the framework integrates subword tokenization, vectorization, large model fine-tuning, retrieval-augmented generation, and reinforcement learning from human feedback, all implemented within the Hugging Face ecosystem using openly licensed models to avoid reliance on proprietary APIs. The project delivers the first publicly available tokenizers, embeddings, lexicons, and transliteration benchmarks for languages such as Tajik and Tatar, forming a comprehensive educational curriculum tailored for advanced undergraduates, graduate students, and practitioners, thereby significantly advancing reproducible research and capacity building in low-resource language NLP.
This study addresses the lack of standardized evaluation protocols in machine-generated text detection, which hinders fair comparison of model performance. The authors systematically evaluate 15 detection methods across diverse datasets comprising both human-written and machine-generated English texts, covering six detector families and seven generative models. Employing a multi-dataset cross-validation framework and multiple evaluation metrics, they find that no single detector consistently outperforms others across all scenarios—most excel only in specific settings—and overall performance degrades significantly on novel, human-authored texts from high-stakes domains. The work highlights the strong dependence of detector efficacy on training and evaluation data as well as metric choice, exposing critical blind spots in current evaluation paradigms and underscoring the decisive role of methodological decisions in shaping empirical conclusions.