Information Structure in Mappings: An Approach to Learning, Representation, and Generalisation

📅 2025-05-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The absence of a unified symbolic and quantifiable structural description of neural network representation spaces impedes principled understanding of their generalization mechanisms. Method: We introduce the first scalable information-theoretic framework, featuring structural primitives for mapping architecture and an efficient vector-space entropy estimation algorithm—scalable to models ranging from millions to 12B parameters. Contribution/Results: Our framework enables unified characterization of representational structure evolution, its correlation with generalization performance, and structural commonalities across paradigms—including multi-agent reinforcement learning, sequence models, and large language models (LLMs). Empirically, we uncover a deep structural parallelism between linguistic constraints and neural generalization structure, and establish an interpretable causal chain: “design choices → emergent structure → observed performance.” This work provides a new paradigm for quantitative analysis and controllable design of neural representations.

Technology Category

Machine Learning: Deep Neural Architectures and Foundation ModelsNatural Language Processing: (Large) Language ModelsSearch and Optimization: Learning to Search

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web dataUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Despite the remarkable success of large large-scale neural networks, we still lack unified notation for thinking about and describing their representational spaces. We lack methods to reliably describe how their representations are structured, how that structure emerges over training, and what kinds of structures are desirable. This thesis introduces quantitative methods for identifying systematic structure in a mapping between spaces, and leverages them to understand how deep-learning models learn to represent information, what representational structures drive generalisation, and how design decisions condition the structures that emerge. To do this I identify structural primitives present in a mapping, along with information theoretic quantifications of each. These allow us to analyse learning, structure, and generalisation across multi-agent reinforcement learning models, sequence-to-sequence models trained on a single task, and Large Language Models. I also introduce a novel, performant, approach to estimating the entropy of vector space, that allows this analysis to be applied to models ranging in size from 1 million to 12 billion parameters. The experiments here work to shed light on how large-scale distributed models of cognition learn, while allowing us to draw parallels between those systems and their human analogs. They show how the structures of language and the constraints that give rise to them in many ways parallel the kinds of structures that drive performance of contemporary neural networks.
Problem

Research questions and friction points this paper is trying to address.

Lack unified notation for neural network representational spaces
Need methods to describe representation structure and generalization
Develop quantitative methods to analyze mapping structures in models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantitative methods for mapping structural primitives
Information theoretic analysis of learning models
Novel entropy estimation for large-scale models