Enhanced Knowledge Distillation for Detection Transformer via Teacher Prediction Refinement

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决DETR在边缘设备部署时的高计算成本问题,提出了一种通过改进教师预测质量来增强知识蒸馏的方法(TPRD),从而提高学生模型的学习效果。
📝 Abstract
Detection Transformers (DETRs) achieve strong performance in object detection but remain challenging to deploy on edge devices due to their high computational cost. Existing DETR distillation methods mainly focus on aligning distillation points, while largely overlooking the quality of the teacher's supervision itself. We observe that due to stage-wise non-monotonic prediction behavior in DETRs, well-localized or correctly classified predictions from earlier stages may degrade in later ones, and some negative predictions become increasingly overconfident. As a result, relying solely on the current stage's predictions yields inaccurate and inconsistent supervision. To address this issue, we propose Teacher Prediction Refinement Distillation (TPRD), a plug-and-play module that refines teacher predictions before distillation by exploiting stage-wise prediction information. TPRD improves supervision quality through Positive Prediction Correction (PPC), which corrects degraded positive predictions by restoring more accurate ones from earlier stages, ensuring reliable localization and classification signals, and Negative Prediction Suppression (NPS) suppresses the influence of overconfident negatives, preventing them from providing misleading supervision to the student. To preserve informative dark knowledge, we further introduce Maximum Dark Knowledge Preservation (MDKP), which selectively refines target-class logits while retaining non-target relations. Extensive experiments on MS COCO and PASCAL VOC demonstrate the effectiveness and robustness of the proposed method. Our code is available at https://github.com/xingyitong1/TPRD.
Problem

Research questions and friction points this paper is trying to address.

Detection Transformers
high computational cost
teacher supervision
non-monotonic prediction
inconsistent supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Teacher Prediction Refinement Distillation
Positive Prediction Correction
Negative Prediction Suppression
Maximum Dark Knowledge Preservation
🔎 Similar Papers
No similar papers found.
Y
Yitong Xing
School of Computer Science, Shanghai Jiao Tong University, Shanghai 200240, China
Yuhao Cheng
Yuhao Cheng
Shanghai Jiao Tong University
Computer VisionFaceEmbodied AI
Yanping Li
Yanping Li
Shanghai Jiao Tong University
Y
Yichao Yan
School of Computer Science, Shanghai Jiao Tong University, Shanghai 200240, China