Next-token functional estimation

📅 2026-09-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种留窗估计方法来解决序列中下一个点与已观测点间函数估计问题,特别是在时间依赖性存在时,该方法比传统的留一法更有效。
📝 Abstract
Suppose we observe the first $n$ points of a sequence of random variables having length $n+1$, and wish to estimate a functional of the unobserved final point and the empirical measure of the $n$ observed training points. Such next-token functionals include the probability that the next token is novel (also known as the surprise probability), the tail probability of the minimum distance between the next token and training points, and the test error of a classifier trained on the observed points. All of these quantities are classically estimated by the leave-one-out method, which is inconsistent under temporal dependence. We propose a leave-a-window-out estimator, which deletes a window of length $τ$ after each index before forming the empirical measure and reduces to leave-one-out at $τ= 1$. Under natural assumptions, we show that the error of our estimator decays at a parametric rate for any stationary $β$-mixing process that also admits a Marton coupling. Our results thus cover several natural functionals on a large class of stochastic processes. We complement these upper bounds with a sharp minimax lower bound for estimating the surprise probability on mixing Markov chains. Simulations on Markov chains, moving-average processes, and autoregressive processes show that our estimator succeeds in many scenarios where leave-one-out and add-constant baselines fail.
Problem

Research questions and friction points this paper is trying to address.

next-token functional
surprise probability
tail probability
test error
leave-a-window-out estimator
Innovation

Methods, ideas, or system contributions that make the work stand out.

leave-a-window-out estimator
next-token functional estimation
parametric rate of convergence
temporal dependence
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Milind Nakul
Schools of Industrial and Systems Engineering and Electrical and Computer Engineering, Georgia Institute of Technology
Vidya Muthukumar
Vidya Muthukumar
Georgia Institute of Technology
machine learning theoryonline decision-makinggame theory
A
Ashwin Pananjady
Schools of Industrial and Systems Engineering and Electrical and Computer Engineering, Georgia Institute of Technology