From Optimal Policies to Individual Differences: Rethinking Reinforcement Learning for Biology

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a key limitation in existing reinforcement learning approaches, which often overlook individual differences when modeling biological behavior, focusing instead on optimal policies or population averages. To overcome this constraint, the work introduces a biologically interpretable framework that integrates methods from multiple subfields of reinforcement learning to construct a computational model capable of generating diverse individual behaviors. By systematically synthesizing technical strategies that support behavioral diversity, the research establishes a novel paradigm for modeling individual variation in biological agents. This paradigm effectively narrows the gap between simulated and real-world biological behaviors, offering both a theoretical foundation and practical guidance for future research in biologically plausible behavior modeling.
📝 Abstract
Reinforcement learning (RL) is primarily known as a computational method for optimizing control tasks, but it is increasingly used to explain biological behavior. While RL successfully captures key aspects of biology, a major gap remains: between-agent behavioral variability. Consistent individual differences naturally permeate biological populations, yet RL models typically present only the single best individual or the population average. Addressing this gap requires moving beyond current practices to generate behavioral diversity using biologically plausible mechanisms. Here, we examine approaches from various subfields of RL and outline potential paths forward to close the gap between biology and simulation.
Problem

Research questions and friction points this paper is trying to address.

reinforcement learning
individual differences
behavioral variability
biological behavior
agent diversity
Innovation

Methods, ideas, or system contributions that make the work stand out.

reinforcement learning
individual differences
behavioral diversity
biological plausibility
computational modeling
P
Patrick Govoni
Institute for Theoretical Biology, Humboldt Universität zu Berlin, Berlin, Germany
P
Palina Bartashevich
Faculty of Life Sciences, Thaer-Institute for Agricultural and Horticultural Sciences, Humboldt Universität zu Berlin, Berlin, Germany; Science of Intelligence, Cluster of Excellence, Berlin, Germany
C
Clémence Bergerot
Institute for Theoretical Biology, Humboldt Universität zu Berlin, Berlin, Germany
V
Valerii Chirkov
Institute for Theoretical Biology, Humboldt Universität zu Berlin, Berlin, Germany; Science of Intelligence, Cluster of Excellence, Berlin, Germany
Valentin Lecheval
Valentin Lecheval
Humboldt Universität zu Berlin
collective behaviourcollective motion
Pawel Romanczuk
Pawel Romanczuk
Institute for Theoretical Biology, Humboldt Universität zu Berlin
Collective BehaviorActive MatterBiological PhysicsComplex SystemsStatistical Physics