Lingtai: What Concept Geometry Reveals--and Does Not Reveal--About LLM Inference

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of training-free, real-time monitoring of large language model reasoning by proposing a training-free conceptual telemetry layer. Methodologically, it generates structured signals via residual stream projection and introduces a novel label- and optimization-free approach to constructing concept anchors. By integrating domain-specific anchor banks with elastic alignment techniques, the method reveals the geometric structure of predictive uncertainty under task conditioning. Experimental results demonstrate that this approach incurs minimal decoding overhead (<1.6%) and establishes a robust correlation between telemetry signals and uncertainty, while simultaneously confirming that such signals do not serve as reliable indicators of correctness.
📝 Abstract
Observing what a large language model computes during autoregressive inference--online and without training probes--remains difficult. We introduce Lingtai, a training-free concept telemetry layer: at each generation step, residual states are projected onto a domain-specific bank of named concept anchors, constructed without labeled concept examples, outcome labels, gradient fitting, or activation-space optimization, producing a structured per-step concept-coordinate signal. Across code generation and grade-school mathematical reasoning, this signal exhibits a robust association with predictive uncertainty: the association survives problem-identity and token-position controls and is not attributable to a single token type, is not explained by a simple correct/incorrect mixture on GSM8K, and is not reproduced by matched random anchors; it is markedly weaker or direction-inconsistent in K-means and PCA projections. Two structures emerge: a recurring uncertainty-linked activity signal whose functional geometry is task-conditioned (distinct activity-entropy shapes on HumanEval, MBPP, and GSM8K), and an execution-specific trajectory identity with strong local inertia but weak re-instantiation invariance--under completion-only elastic alignment, corruption at k=32 (approximately a median quarter of the completion) on the matched re-execution subset still retrieves the archived episode at 62.0%, while a fresh execution retrieves it only 11.7-16.0% of the time. Finally, a matched audit finds no evidence that the scalar concept-activity signal used here supplies a stable correctness coordinate under the tested protocol; we therefore treat correctness as externally supplied. Telemetry adds 0.7-1.6% per-token decode overhead for the 161-anchor code implementation, with unchanged generated tokens.
Problem

Research questions and friction points this paper is trying to address.

LLM inference
concept geometry
autoregressive generation
predictive uncertainty
training-free telemetry
Innovation

Methods, ideas, or system contributions that make the work stand out.

training-free concept telemetry
residual state projection
predictive uncertainty
concept anchors
trajectory identity
🔎 Similar Papers