🤖 AI Summary
This work addresses the limitations of existing approaches to pain recognition from heterogeneous 3D modalities—such as facial videos and functional near-infrared spectroscopy (fNIRS)—which rely on modality-specific architectures and handcrafted inductive biases. To overcome these constraints, we propose the first unified tokenization framework that maps raw signals and their time-frequency representations into a shared token space, enabling end-to-end learning while preserving both spatiotemporal and time-frequency structures. By eliminating the need for modality-customized models and avoiding manual inductive biases, our method significantly enhances generalization and deployment flexibility. Evaluated on the AI4Pain benchmark, the approach achieves state-of-the-art performance and supports efficient CPU/GPU inference, making it suitable for real-time pain assessment.
📝 Abstract
Pain is a complex and pervasive phenomenon affecting a large percentage of the population, and accurate assessment is essential for effective clinical management and intervention. Computational pain recognition systems enable continuous monitoring, support clinical decision-making, and help mitigate pain-related distress and functional decline. This study introduces a unified tokenization framework for heterogeneous 3D modalities in pain recognition that provides a single processing pipeline across behavioral and brain-activity 3D data, without requiring separate architectures for each modality or handcrafted inductive biases. The framework preserves spatial, temporal, and time--frequency structure while mapping diverse inputs into a shared token space. Extensive experiments show that the proposed approach effectively processes facial videos and fNIRS data in both raw-signal and spectrogram-based representations. On the AI4Pain benchmark dataset, the proposed framework achieves state-of-the-art performance while maintaining high computational efficiency and enabling real-time assessment on both GPU and CPU hardware.