🤖 AI Summary
This study addresses the limited robustness of RGB models due to their over-reliance on high-frequency textures and the challenge of evaluating inductive biases in event cameras. It presents the first systematic demonstration of a "shape-first" mechanism induced by event data, proposing a novel edge-based paradigm for model robustification. Methodologically, cross-domain knowledge distillation is employed to reshape RGB representations using event data, while spectral analysis and early-layer feature processing suppress texture dependence and enhance shape perception. Experiments show that this approach significantly improves robustness against high-frequency noise and color invariance, establishing shape bias as an effective prior for downstream tasks. The source code is publicly available.
📝 Abstract
Convolutional neural networks trained on ImageNet are known to exhibit a strong preference for local high-frequency texture, an inductive bias that translates into fragile robustness against distribution shifts in real-world environments. Event cameras, in contrast, record only changes in scene brightness and are therefore well suited to capturing contour information; however, due to the absence of diagnostic benchmarks in the event domain, the inductive bias that event-camera data instills in vision models has remained underexplored. In this work, we use knowledge distillation from the event domain to the RGB domain so as to exploit the rich evaluation toolkit available in the RGB domain and systematically dissect this inductive bias. Our experiments show that distillation from the event domain induces, in the RGB domain, color invariance, shape bias, and robustness to high-frequency noise. We identify the underlying mechanism as the model suppressing its dependence on high-frequency texture while acquiring a stronger dependence on edge-based object shape. This hypothesis is supported by changes in how color and spatial information are processed at the early layers, together with a spectral trade-off in which robustness to the absence of high-frequency components coexists with vulnerability to contamination of the relied-upon frequency bands and to disruption of geometric structure. We further show that this inductive bias differs from existing robustification methods and that it functions as a useful prior for diverse downstream tasks in which shape and contour information contribute alongside other cues. The code is available at https://github.com/snskysk/event2rgb-distillation .