Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of annotation scarcity, underutilization of non-swallowing data, and spatial bias in semi-supervised keypoint localization for Videofluoroscopic Swallowing Studies (VFSS). To this end, we construct the VFSSKep dataset and propose the S3KL framework. Methodologically, we enhance anatomical representation by extending annotations to critical regions such as the soft palate. A block-shuffling mechanism is designed to eliminate coordinate memorization bias, thereby improving the utilization of unlabeled data. Furthermore, high-resolution structural cue extraction is integrated with consistency learning to achieve precise localization. Experimental results demonstrate that our approach surpasses fully supervised baselines using only 25% of annotated data, establishing a new state-of-the-art in semi-supervised VFSS keypoint detection.
📝 Abstract
Videofluoroscopic Swallowing Study (VFSS) is one of the gold standard for diagnosing swallowing disorders, providing dynamic X-ray imaging of the swallowing process. Automated kinematic analysis in VFSS relies fundamentally on precise anatomical keypoint localization. However, existing studies focus on limited keypoints (e.g., cervical vertebrae or the hyoid) and overlook critical regions such as the soft palate, while annotating only active swallowing segments and ignoring abundant non-swallowing data, resulting in poor data efficiency. Moreover, leveraging this unlabeled data via standard semi-supervised learning is suboptimal, as generic methods are prone to spatial bias. In medical X-rays with fixed layouts, models tend to memorize absolute coordinates rather than understanding anatomical structures. To tackle these challenges, we introduce VFSSKep, a novel dataset that extends annotations to the soft palate and incorporates large-scale unlabeled data. We further propose S$^3$KL, a Structure-aware Semi-Supervised Keypoint Localization framework designed to overcome spatial bias. It integrates a Structure-Aware Learning strategy to extract high-resolution structural cues for structure-aware representation learning, and a Structural Representation Consistency Learning strategy with block shuffling to enforce invariant structural recognition. Experiments show our method achieves state-of-the-art semi-supervised performance, even with unlabeled and 25% labeled data surpassing fully supervised learning with 100% labeled data. Code and data will be made publicly available at: https://github.com/kaai520/S3KL.
Problem

Research questions and friction points this paper is trying to address.

Videofluoroscopic Swallowing Study
Keypoint Localization
Semi-supervised Learning
Spatial Bias
Anatomical Structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Keypoint Localization
Semi-Supervised Learning
Structure-Aware Representation
Videofluoroscopic Swallowing Study
Spatial Bias
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kai Zhou
South China University of Technology
C
Chuanshen Chen
South China University of Technology
R
Runhao Zeng
Shenzhen MSU-BIT University
M
Meng Dai
The Third Affiliated Hospital of Sun Yat-sen University
Yifan Yang
Yifan Yang
South China University of Technology
Neural renderingImage synthesisophthalmologic
Jinwu Hu
Jinwu Hu
South China University of Technology; Pazhou Lab
Large Language ModelsComputer VisionReinforcement Learning
D
Daiyuan Li
South China University of Technology
Mingkui Tan
Mingkui Tan
South China University of Technology
Machine LearningLarge-scale Optimization
F
Fei Liu
South China University of Technology