🤖 AI Summary
This study addresses the challenge that reward signals alone cannot independently determine computational processes, investigating how pretraining and mid-training synergistically provide adaptive mechanisms. Methodologically, it constructs finite-sample Adam optimization trajectories and introduces task-agnostic source observations to eliminate training ambiguity, theoretically establishing a novel paradigm wherein prediction acquires execution capabilities while reward learns task-specific utilization. Technically, the approach integrates sequential state computation, in-context memory, and Qwen2.5 model fine-tuning. Experimental evaluations across multi-world benchmarks demonstrate that models incorporating correct source supervision achieve an 82.61% success rate, significantly outperforming control groups. These results effectively establish a theoretical bridge between information acquisition and reward-guided learning.
📝 Abstract
A reward can identify a correct answer while leaving the computation needed for new inputs undetermined. We study how pretraining and midtraining supply the information and computation that make reward adaptation effective. In sequential state computation and contextual memory, we characterize mechanisms that agree on every training reward yet demand different held-out answers. Task-independent source observations resolve this ambiguity. We construct finite sampled Adam paths from specified random initializations through source prediction and reward adaptation in the same parameters, proving how prediction acquires execution or retrieval and rewards learn their task-specific use. Experiments with pretrained Qwen2.5 checkpoints test this division of labor. Across eight worlds, Sequential models trained with correct source and first-operation supervision reach 82.61% success, versus 44.15% for a private-random source control. Memory replay preserves retrieval during reward adaptation, and an independent eight-world confirmation achieves 75.32% task success versus 49.86% after matched alternative-retrieval training. GSM8K and HotpotQA separate accuracy at reward entry, subsequent gain and final performance. Together, these results connect information acquisition, executable computation and reward-guided task learning.