Trending
Score
Score
+20 in 12 mo
96
12 mo agoNow
Training strategies that reduce mismatch from teacher forcing and chunk-wise autoregressive denoising by approximating inference-time autoregressive feedback (rollouts) without slow sequential sampling, enabling parallelized training while preserving effective sequence length.