🤖 AI Summary
This study addresses how to determine, from observed data during in-context learning, whether the cost of expanding a hypothesis space is justified. It formulates model expansion as a costly sequential decision problem within a structure-revision environment and introduces a "value boundary" theory that supersedes conventional evidence thresholds, revealing that optimal policies are shaped by costs, horizons, and queries. Comparative experiments are conducted via Bayesian computation, exact solutions, and Transformer training. The results demonstrate that Transformers can replicate this value boundary. Furthermore, the analysis reveals that while existing large language models exhibit failure-sensitive signals, they lack cost-benefit trade-off capabilities, and post-training merely alters prior distributions without optimizing decision criteria.
📝 Abstract
Learning systems adapt quickly inside a familiar family of models. The harder step comes earlier: deciding, from observations that could be noise, an exception, a change within the family or structure outside it, whether opening a richer family is worth its cost. We treat this as a costly sequential decision: prediction failure must be turned into structural evidence, evidence into a value of expansion, and value into action. The Structural Revision Environment produces matched failures from each source, varies the price of expansion and the remaining horizon independently of the evidence, and admits exact Bayesian calculations and an exact normative solution of the one-shot decision. Its solution shows that revision is a value boundary and not an evidence threshold: one history has different optimal actions under different prices, horizons and announced queries, the boundary between local repair and expansion is set by the inputs a rule predicts and a repair cannot cover, and belief in the richer family crosses long before the decision does. Transformers trained in the environment reproduce this boundary from utility alone. Language models of three post-training lineages carry a failure-sensitive signal in their predictions that is not reflected in their revision decisions, and given the gain of expanding they read it without weighing it against price and horizon. Three models allowed to reason weigh the stated gain in the reference's proportions and still do not turn the history into an estimate of what expansion would buy. Controlled post-training of the meta-trained learners moves the prior and the sharpness of predictions, and neither moves the criterion.