🤖 AI Summary
This work addresses the challenge of constructing biologically plausible place-cell-like spatial representations from visual input alone—without motor cues or reward signals—thereby enabling cognitive map formation under minimal sensory constraints. We propose the Visual Place-Cell Encoding (VPCE) model: high-dimensional visual features are extracted via ResNet, then clustered using K-means; spatial selectivity emerges spontaneously through radial basis function–based similarity computation, yielding activation patterns exhibiting spatial proximity tuning, orientation alignment, and boundary sensitivity. Critically, we demonstrate for the first time that unsupervised clustering of raw visual appearance features alone yields spatial codes statistically indistinguishable from canonical hippocampal place cell properties (p < 0.001). The model maintains robust spatial discrimination even under dynamic environmental changes (e.g., wall addition/removal). By bypassing reliance on path integration or reinforcement learning, VPCE establishes a novel, vision-driven mechanism for cognitive mapping.
📝 Abstract
This paper presents the Visual Place Cell Encoding (VPCE) model, a biologically inspired computational framework for simulating place cell-like activation using visual input. Drawing on evidence that visual landmarks play a central role in spatial encoding, the proposed VPCE model activates visual place cells by clustering high-dimensional appearance features extracted from images captured by a robot-mounted camera. Each cluster center defines a receptive field, and activation is computed based on visual similarity using a radial basis function. We evaluate whether the resulting activation patterns correlate with key properties of biological place cells, including spatial proximity, orientation alignment, and boundary differentiation. Experiments demonstrate that the VPCE can distinguish between visually similar yet spatially distinct locations and adapt to environment changes such as the insertion or removal of walls. These results suggest that structured visual input, even in the absence of motion cues or reward-driven learning, is sufficient to generate place-cell-like spatial representations and support biologically inspired cognitive mapping.