Towards semantic reconstruction of individual words from fnirs using clip loss

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing fNIRS decoders that neglect the geometric structure of embedding spaces by proposing a contrastive learning-based neural semantic reconstruction method. Specifically, a Bi-LSTM network is employed to map fNIRS signals into GloVe or T5 word embedding spaces, and a CLIP-inspired contrastive loss is introduced to replace the conventional mean squared error (MSE) objective, thereby enhancing geometric consistency within the embedding space. Experimental results demonstrate that the Bi-LSTM decoder trained with the CLIP objective achieves the most stable performance, effectively reconstructing semantic information from neural activity elicited by perceived words. These findings validate the potential and advantages of contrastive learning objectives for fNIRS-based semantic decoding.
📝 Abstract
Semantic reconstruction maps neural activity to a word-embedding space, recovering the meaning of a perceived word instead of selecting it from a fixed vocabulary. Functional near-infrared spectroscopy (fNIRS) carries semantic information suitable for this mapping. However, most fNIRS decoders are trained with a squared-error objective that fits each word independently and ignores the geometry of the embedding space. To address this limitation, we evaluate a contrastive loss based on the contrastive-language-image-pretraining (CLIP) loss, as an alternative to mean-squared-error (MSE) for reconstructing perceived words from fNIRS. We compare the two objectives by training a bidirectional long short-term memory (Bi-LSTM) decoder to map fNIRS signals to word embeddings. We use GloVe-50 and T5 word embeddings as targets, across three fNIRS datasets recorded under a shared paradigm pairing each word image with its spoken name. Performance is measured with a pairwise matching score and open-vocabulary top-$k$ retrieval. The Bi-LSTM trained with CLIP is the most consistent decoder across experiments. T5 produces higher matching scores, whereas every significant retrieval result uses GloVe-50. These results support the use of contrastive objectives as a promising direction for fNIRS semantic decoding and motivate validation on larger datasets.
Problem

Research questions and friction points this paper is trying to address.

semantic reconstruction
fNIRS
contrastive loss
word embedding
neural decoding
Innovation

Methods, ideas, or system contributions that make the work stand out.

fNIRS
CLIP loss
semantic reconstruction
contrastive learning
Bi-LSTM
🔎 Similar Papers
No similar papers found.
S
Santiago Posso-Murillo
Department of Electrical and Computer Engineering, University of Kentucky
N
Nathan Palladino
Department of Psychology, Wheaton College
B
Ben Pyykkonen
Department of Psychology, Wheaton College
D
Dan Y. Han
Departments of Neurology, Neurosurgery, and Physical Medicine & Rehab., University of Kentucky
L
Luis G. Sanchez-Giraldo
Department of Electrical and Computer Engineering, University of Kentucky
Jihye Bae
Jihye Bae
Assistant Professor in ECE, University of Kentucky
signal processingmachine learningbrain machine interfacesEEG analysis and source imaging