🤖 AI Summary
This study addresses the inherent tension in automated kernel design between the limited expressiveness imposed by fixed grammars and the validity issues arising from unconstrained generation. To resolve this, the work reformulates kernel design as an open-ended model discovery task, leveraging large language model-based coding agents to synthesize programs while enforcing syntactic validity through construction contracts. Furthermore, a quality-diversity archive coupled with novelty screening is introduced to balance predictive performance against structural variation, enabling interpretable and general-purpose kernel search. The discovered kernels outperform deep kernel baselines in black-box optimization tasks. Following manual refinement, they achieve a 5.7% reduction in prediction error and a 7.8% decrease in optimization regret, demonstrating the efficacy of the proposed framework.
📝 Abstract
Kernels encode the inductive bias of a wide range of machine learning models, yet automated kernel design faces a fundamental dilemma. A fixed grammar of base kernels and operators guarantees validity but limits the search to structures expressible by those building blocks. Conversely, unrestricted programs remove this limitation but no longer guarantee validity. In our stress tests, 22-58% of LLM-generated kernels that pass numerical checks on random inputs fail when evaluated at different scales or dimensions. We propose Kernel Autoresearch (Kernaut), which treats kernel design as open-ended model discovery. Coding agents write kernels as programs, while construction contracts ensure that every accepted kernel is valid. A quality-diversity archive retains high-performing kernels with distinct behaviors, and novelty screening steers agents toward functionally new candidates. Our experiments demonstrate that the discovered kernels encode reusable inductive biases that generalize to unseen tasks. On held-out black-box optimization families, a discovered kernel outperforms a meta-learned deep kernel trained on the same episodes. Furthermore, kernels discovered from ten enzyme-kinetic rate laws achieve lower error than tuned ARD and deep kernel baselines on five unseen mechanisms. The discovered kernels are also interpretable programs that human researchers can refine: a human-refined version of one further reduces the held-out predictive error by 5.7% and optimization regret by 7.8%.