When Should an In-Context Learner Expand Its Hypothesis Space?

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses how to determine, from observed data during in-context learning, whether the cost of expanding a hypothesis space is justified. It formulates model expansion as a costly sequential decision problem within a structure-revision environment and introduces a "value boundary" theory that supersedes conventional evidence thresholds, revealing that optimal policies are shaped by costs, horizons, and queries. Comparative experiments are conducted via Bayesian computation, exact solutions, and Transformer training. The results demonstrate that Transformers can replicate this value boundary. Furthermore, the analysis reveals that while existing large language models exhibit failure-sensitive signals, they lack cost-benefit trade-off capabilities, and post-training merely alters prior distributions without optimizing decision criteria.
📝 Abstract
Learning systems adapt quickly inside a familiar family of models. The harder step comes earlier: deciding, from observations that could be noise, an exception, a change within the family or structure outside it, whether opening a richer family is worth its cost. We treat this as a costly sequential decision: prediction failure must be turned into structural evidence, evidence into a value of expansion, and value into action. The Structural Revision Environment produces matched failures from each source, varies the price of expansion and the remaining horizon independently of the evidence, and admits exact Bayesian calculations and an exact normative solution of the one-shot decision. Its solution shows that revision is a value boundary and not an evidence threshold: one history has different optimal actions under different prices, horizons and announced queries, the boundary between local repair and expansion is set by the inputs a rule predicts and a repair cannot cover, and belief in the richer family crosses long before the decision does. Transformers trained in the environment reproduce this boundary from utility alone. Language models of three post-training lineages carry a failure-sensitive signal in their predictions that is not reflected in their revision decisions, and given the gain of expanding they read it without weighing it against price and horizon. Three models allowed to reason weigh the stated gain in the reference's proportions and still do not turn the history into an estimate of what expansion would buy. Controlled post-training of the meta-trained learners moves the prior and the sharpness of predictions, and neither moves the criterion.
Problem

Research questions and friction points this paper is trying to address.

in-context learning
hypothesis space expansion
structural revision
sequential decision-making
prediction failure
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Context Learning
Hypothesis Space Expansion
Sequential Decision Making
Structural Revision Environment
Value Boundary
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Weihan Li
Department of Mechanical Engineering, The University of Tokyo, Tokyo, Japan
X
Xinlei Chen
School of Mechanical Engineering and Automation, Harbin Institute of Technology, Shenzhen, China
Junhao Wu
Junhao Wu
Towson university
Computer VisionCryo emMedical image
Tianshi Zheng
Tianshi Zheng
HKUST
Natural Language ProcessingLogical InferenceScientific DiscoveryResearch Agent