FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between format-based feedback and downstream learning utility in trajectory replay by proposing an auditable full-trajectory replay framework. Methodologically, a learner-aware selector is introduced to optimize post-training effectiveness, while an optimizer-aware virtual update constructs a magnitude-aware utility surface, enabling training-independent baseline comparisons. Reliability is further ensured through normalized gradient alignment, metadata cross-fitting calibration, and frozen auditing contracts. Experimental results on the GSM8K dataset demonstrate that the proposed approach significantly improves generation quality to 0.6624, outperforming both uniform sampling and format-feedback baselines. Overall, this work establishes an efficient and auditable data selection paradigm for model post-training.
📝 Abstract
Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache-level correctness and downstream learner utility are distinct objectives. We formalize this selection-to-learning gap and introduce FAER as an auditable full-trajectory replay framework. Its training-free fixed selector is a protocol baseline; FAER-UTILITY is the learner-aware selector fitted on disjoint calibration blocks. The normalized gradient alignment is reported as a baseline, while a disposable optimizer-aware virtual update supplies a magnitude-aware utility surface. The audit contract freezes observed fields and replay traces before evaluation labels are joined. On GSM8K with Qwen2.5-1.5B-Instruct, the matched learner study reports quality 0.6329 for the fixed selector, compared with 0.5482 for uniform and 0.6037 for format-feedback under 128 updates. Metadata-only cross-fitted calibration reaches $0.6476\!\pm\!0.0139$ over eight seeds (median 0.6481; paired 95% interval $[+0.079,+0.122]$) at 63,276 target-run tokens; its recorded full cost is 189,642 tokens and 3.48 GPU-hours including calibration. The completed FAER-UTILITY row reaches 0.6624 at 62,844 target-run tokens and 4.26 GPU-hours. Format-feedback selects records with correctness 0.6953, compared with 0.3594 for the fixed selector, despite the different downstream ranking. The completed comparison surfaces report the learner-aware ablation, same-seed gap, policy-optimization rows, and strict zero-shot transfer.
Problem

Research questions and friction points this paper is trying to address.

trajectory replay
language model post-training
selection-to-learning gap
learner utility
auditable framework
Innovation

Methods, ideas, or system contributions that make the work stand out.

trajectory replay
utility-aligned selection
auditable framework
post-training
virtual update
🔎 Similar Papers
No similar papers found.
M
Miaobo Hu
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
S
Shuhao Hu
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Xiaobo Guo
Xiaobo Guo
Dartmouth College
machine learningdeep learningnatural language processingsocia mediapropagantion
Xin Wang
Xin Wang
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Biomedical Engineering
Bokun Wang
Bokun Wang
Texas A&M University
Machine LearningArtificial IntelligenceMultimodal Machine Learning
T
Tianshu Fu
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
D
Daren Zha
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
J
Jun Xiao
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China