🤖 AI Summary
This work addresses the longstanding challenge in optimization of simultaneously achieving stability and scalability in high-dimensional, long-horizon training settings: gradient-based methods often suffer from instability, while gradient-free approaches struggle to scale to large parameter spaces. To overcome this, the paper proposes a bilevel meta-learning algorithm wherein an inner loop performs adaptive updates in the high-dimensional parameter space, guided efficiently by an outer loop that optimizes a small set of meta-parameters via zeroth-order methods. This architecture effectively decouples the complexity of high-dimensional learning from the control of learning dynamics, thereby circumventing the dual limitations of temporal horizon and dimensionality inherent in existing approaches. As a result, the method substantially enhances the stability, efficiency, and robustness of large-scale models during prolonged training.
📝 Abstract
Gradient descent scales well to large models, but becomes unstable over long time horizons. Gradient-free optimizers can scale to arbitrary timespans, but are hobbled by high dimensions. Since learning occurs in large models over long timescales, neither of these approaches is likely to produce traits which can accelerate the learning process. Instead, we propose a meta-learning algorithm in which the agent learns to modify its own weights and biases. Our algorithm consists of an inner loop, wherein the agent performs some high-dimensional optimization upon itself, and an outer loop, wherein we perform some low-dimensional optimization upon the inner loop. Since the outer loop handles very few parameters, standard zeroth-order methods may be used.