🤖 AI Summary
This study addresses the underexplored integration of binary neural networks (BNNs) with event cameras by proposing the first fully binarized event-based vision system. Methodologically, we introduce PBEV, a novel binary representation that adapts modern deep BNN architectures to event data, alongside an RGB cross-modal pretraining strategy to enhance model accuracy. The primary contributions lie in enabling the direct application of BNNs to event streams and demonstrating the substantial benefits of cross-modal transfer learning for neuromorphic data. Experimental results show that the proposed system achieves 90.58% accuracy on the N-Caltech101 dataset while reducing computational cost by 7.5× compared to full-precision models, effectively balancing high efficiency with strong predictive performance.
📝 Abstract
Binary Neural Networks (BNNs) enable efficient deep learning deployment on resource constrained devices with weights and activations compressed to one bit, substantially reducing model size and inference cost. Event cameras offer complementary advantages, including low latency, high dynamic range, and low power consumption, by capturing asynchronous streams of events rather than dense image frames. Despite their shared emphasis on efficiency, the combination of these technologies remains largely unexplored. This work aims at adapting and evaluating modern deep BNN architectures on event data. We also show that cross-modal pretraining from RGB data can improve the classification accuracy of BNNs on neuromorphic datasets. We introduce the Polar-wise Binary Event Volume (PBEV), a binary representation that enables event-camera data to be processed directly by BNNs and represents a step toward fully binarized event-based vision systems. Best evaluated BNN on N-Caltech101 classification benchmarks shows 90.58% accuracy with 7.5x less operations than their full-precision counterparts.