The Probabilistic Structure of Large Language Models

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文从概率角度探讨大型语言模型,通过自回归条件分布和最大似然估计训练模型,分析了KL散度在文本生成中的作用,并讨论了扩散模型。
📝 Abstract
This paper presents a probabilistic perspective on large language models (LLMs), developed with the aim of bringing together, in a single self-contained account, tools that are usually treated separately across the literature. LLMs are described through probability measures on the set of sequences of tokens, specified via their autoregressive conditional distributions. Training is formulated as a maximum-likelihood estimation problem, addressed by stochastic gradient methods, while text generation is viewed as the sequential simulation of the resulting stochastic process. The role of the asymmetry of the Kullback--Leibler divergence in text generation is examined in relation with characteristic phenomena such as hallucination and the distinction between statistical plausibility and truth. As a complementary illustration of the same viewpoint, we also discuss diffusion models, built around the score function, which cast generation not as sequential token prediction but as the simulation of a reverse-time stochastic process transforming noise into data both in discrete and continuous time.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Kullback--Leibler Divergence
Text Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

probabilistic perspective
autoregressive conditional distributions
maximum-likelihood estimation
stochastic gradient methods
diffusion models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Adnan Aboulalaâ