Belief-Trajectory Energy: Measuring the Path to a Prediction

๐Ÿ“… 2026-10-04
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the limitation of existing large language models that output only final predictions while discarding intermediate layer-wise trajectory information. We propose Belief Trajectory Energy (BTE), a metric that quantifies internal prediction shifts across layers to characterize input data. Furthermore, we construct a theoretical framework grounded in Fisher-Rao geometry to map these internal states into interpretable difficulty signals and fingerprint features. By integrating Transformer layer-wise analysis, shared prediction space mapping, and deep learning representation techniques, our approach effectively leverages latent belief dynamics. Experimental evaluations demonstrate that the proposed method achieves a 99.8% macro AUROC in humanโ€“machine review detection and attains 95.6% accuracy in eight-way generator attribution, highlighting its efficacy in forensic text analysis.
๐Ÿ“ Abstract
Large language models (LLMs) progressively revise their predictions across Transformer layers, yet we typically observe only the final output, discarding the trajectory through which it is formed. We introduce Belief-Trajectory Energy(BTE), a model-grounded measure that characterizes an input through the layerwise predictive revisions it induces in a model. By mapping intermediate states into a shared predictive space, BTE provides a principled measure of belief change that can be summarized as either a scalar or a structured depth profile. Theoretically, we show that local BTE corresponds to predictive revision under the Fisher-Rao geometry, while the sequence of revisions captures information beyond the initial-to-final belief change. Empirically, scalar BTE provides a model-relative signal of difficulty across diverse reasoning tasks, while richer BTE representations support human-LLM review detection and fine-grained generator attribution, reaching up to $0.998$ macro-AUROC and $95.6\%$ eight-way attribution accuracy. Further analysis shows that BTE develops throughout pretraining and is selectively reshaped by targeted training, demonstrating that the resulting measurement reflects what the scoring model has learned. Together, our results establish belief trajectories as a principled model-grounded signal and suggest a broader perspective in which learned models can themselves serve as instruments for characterizing the data they process. More demonstrations can be found at https://yingjiahao14.github.io/BTE-web/.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Belief Trajectory
Intermediate Representations
Predictive Revision
Generator Attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Belief-Trajectory Energy
Large Language Models
Predictive Revision
Generator Attribution
Fisher-Rao Geometry
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
J
Jiahao Ying
Fudan University
W
Wei Tang
University of Science and Technology of China
B
Boxian Ai
Fudan University
Y
Yaoning Wang
Fudan University
Haotian Chen
Haotian Chen
University of California, Los Angeles
Political EconomyNon-market StrategyAmerican Politics
W
Wenhe Sun
Fudan University
C
Caijun Xu
Fudan University
H
Haozhan Cai
Fudan University
C
Changyi Xiao
Fudan University
Yixin Cao
Yixin Cao
Fudan University
Natural Language ProcessingKnowledge EngineeringMulti-modal data processing