Diffusion Transformers are Provably Optimal In-context Generators
This study addresses the issue that task uncertainty under few-shot demonstrations can readily induce distributional bias in generative modeling. To mitigate this, the authors employ a Diffusion Transformer (DiT) architecture that integrates score estimation with an attention-based aggregation mechanism for diffusion sampling. Theoretically, they demonstrate that DiT is capable of capturing residual task uncertainty and learning a predictive distribution that faithfully reflects such uncertainty, rather than merely estimating a single task output. As a key contribution, this work elucidates the intrinsic mechanisms by which DiT handles task uncertainty and establishes minimax optimal convergence rates on test tasks within Hölder classes. These findings provide a rigorous theoretical foundation for few-shot generative modeling.