🤖 AI Summary
This study addresses the issue in video gaze prediction where unfocused prior signals, when blindly fused, degrade model performance. To overcome this, we propose FocusGate, a gated ensemble framework that leverages shape statistics from defocus maps to dynamically select reliable frames for on-demand prior fusion. By introducing an abstention mechanism and median rank normalization, zero-parameter priors can be automatically deactivated, effectively eliminating the adverse effects of unconditional fusion. Built upon the NTIRE 2026 champion baseline, our method incurs only a 1% latency overhead while significantly improving the AUC metric across four supervised models in three distinct video domains. Notably, even when deployed independently, FocusGate outperforms mainstream approaches.
📝 Abstract
Video gaze prediction is led by gaze-trained models, yet gaze-free priors carry signal those models have not absorbed, if one knows when to trust them. We propose FocusGate, a gated ensemble of gaze-free priors whose members may abstain. A per-frame gate reads three shape statistics of a defocus map and selects the frames on which the estimator is above chance on average, so rejected frames reduce to the base exactly, while midrank normalisation lets an all-zero prior abstain at zero parameters. Gated fusion is significantly positive on film, sports and web video, whereas unconditional fusion is harmful on sports and null on web. Added to four supervised predictors, the NTIRE 2026 champion among them, FocusGate improves all sixteen model-domain cells in shuffled AUC, fifteen significantly, one domain pre-registered and scored once, while adding only 1% to the champion's latency. Alone, it surpasses TASED-Net and UNISAL in shuffled AUC on film with a 16-frame causal mean.