๐ค AI Summary
This study investigates how generating multiple responses influences the sample complexity of learning from demonstrations under suboptimal demonstration settings. By adopting a finite reward class model, this work employs a greedy multiplicative weights learning algorithm combined with matching upper and lower bound techniques to systematically analyze the statistical benefits of multiple responses when demonstration quality is unknown. It reveals a qualitative improvement mechanism whereby multiple responses reduce the worst-case sample complexity dependence from 1/ฮตยฒ to 1/ฮต, while decoupling two independent benefit pathways. Through establishing tight upper and lower bounds, this work demonstrates that optimal sample efficiency can be achieved without prior assumptions regarding demonstration quality, thereby significantly reducing the data requirements for effective learning.
๐ Abstract
Many generative systems return multiple candidate responses and are evaluated according to the best one. Recent work shows that, when demonstrations are optimal, pass@$k$ can reduce the sample complexity of learning from demonstrations by a logarithmic factor in $k$. We ask what happens when the demonstrator is not assumed to be optimal. We find that multiple responses provide a qualitatively stronger benefit in this setting. In a finite reward-class model with no reward feedback, moving from pass@$1$ to any pass@$k$ with $k\ge2$ changes the worst-case dependence on target accuracy from $1/\varepsilon^2$ to $1/\varepsilon$, uniformly over demonstrator quality. Under standard evaluation, where an unknown reward is fixed before training, increasing $k$ provides an additional and distinct benefit: the optimal dependence on a reward class of size $N$ improves from $\log N$ to $\log N/\log k$. We further show that these two effects can be separated. Under robust evaluation, where one learned policy must compete with the demonstrator simultaneously for every reward in the class, the fast $1/\varepsilon$ dependence persists, while the $1/\log k$ improvement can disappear. We establish matching upper and lower bounds in the corresponding regimes and give a greedy multiplicative-weights learner achieving the upper bounds without any assumption on demonstrator quality.