🤖 AI Summary
Existing intrinsic motivation objectives—such as novelty search and the free energy principle—struggle to distinguish between learnable and unlearnable surprising information, often failing in noisy or static environments. This work introduces the concept of “learnable novelty” and, for the first time, formalizes it as a unified measure of intelligence. The authors develop a closed-form estimator based on differentiable reservoir computing that autonomously discovers computational structures, organizes representations, and drives exploration—all without supervision. Empirically, the method successfully reproduces the complexity ordering of cellular automata in an unsupervised manner (identifying Rule 110 as optimal), guides neural automata to generate soliton-like structures, induces category-wise self-organization of MNIST representations, and outperforms baseline approaches in 9 out of 10 reinforcement learning environments without suffering catastrophic collapse.
📝 Abstract
Intelligence appears under different names in different fields: as data compression in statistics and machine learning, as universal computation in dynamical systems, and as adaptive behavior in agents. Each field carries its own objective, and the two most influential drives often fail in mirror image: novelty search, which seeks surprise, is transfixed by a noisy television screen, while the free-energy principle, which avoids surprise, is most content in a dark room. Both failures have a single cause: each objective treats as one quantity the surprise a learner can convert into knowledge and the surprise it never can. Here we show that the learnable part of that information, which we call learnable novelty, yields the seemingly disparate projections of intelligence, and we give a closed-form estimator of it built on a cheap and differentiable reservoir computer. Used as a measure, with no supervision of any kind, the estimator recovers decades of complexity classification, ranking the Turing-complete rule~110 highest among the elementary cellular automata. Used as an objective, its gradient carries a neural cellular automaton from simple dynamics into a regime of solitons, the traveling, colliding structures by which rule~110 computes, as well as organizes the representation of an image encoder around the ten digit classes of MNIST, fully unsupervised: no label ever enters training. Handed to a reinforcement-learning agent as an intrinsic reward, it supplies the exploration that task rewards lack, improving on the task baseline in nine of ten environments and collapsing in none. Complexity generation, abstraction, and exploration, ordinarily pursued with unrelated objectives in separate fields, thus emerge from ascent on one differentiable quantity, and the projections of intelligence gain a common quantitative footing.