Mouse-Guided Gaze: Semi-Supervised Learning of Intention-Aware Representations for Reading Detection

📅 2025-09-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
During magnified reading, frequent viewport shifts fragment gaze data, severely impeding reading intention recognition. To address this, we propose a semi-supervised learning framework that leverages mouse trajectories as weak supervision to guide gaze point representation learning; introduces a viewport compensation mechanism to restore the spatial continuity of text; and jointly models gaze velocity prediction, mouse trajectory regression, and multi-view spatiotemporal dynamics to enable fine-grained discrimination between reading and scanning behaviors. Adopting a pretraining–fine-tuning paradigm, our method achieves up to a 7.5% improvement in F1 score across multiple benchmark datasets. It significantly enhances the robustness and generalizability of eye-movement understanding—particularly under low-interaction-bandwidth conditions where viewport manipulation is prevalent.

Technology Category

Computer Vision: Motion & TrackingMachine Learning: Multi-instance/Multi-view LearningNatural Language Processing: Sentence-level Semantics, Textual Inference, etc.

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labeling
📝 Abstract
Understanding user intent during magnified reading is critical for accessible interface design. Yet magnification collapses visual context and forces continual viewport dragging, producing fragmented, noisy gaze and obscuring reading intent. We present a semi-supervised framework that learns intention-aware gaze representations by leveraging mouse trajectories as weak supervision. The model is first pretrained to predict mouse velocity from unlabeled gaze, then fine-tuned to classify reading versus scanning. To address magnification-induced distortions, we jointly model raw gaze within the magnified viewport and a compensated view remapped to the original screen, which restores spatial continuity across lines and paragraphs. Across text and webpage datasets, our approach consistently outperforms supervised baselines, with semi-supervised pretraining yielding up to 7.5% F1 improvement in challenging settings. These findings highlight the value of behavior-driven pretraining for robust, gaze-only interaction, paving the way for adaptive, hands-free accessibility tools.
Problem

Research questions and friction points this paper is trying to address.

Understanding reading intent during magnified viewing for accessibility
Addressing gaze fragmentation from magnification-induced visual distortions
Learning intention-aware gaze representations with weak mouse supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semi-supervised learning with mouse-guided weak supervision
Joint modeling of raw gaze and compensated view for distortion
Behavior-driven pretraining for robust gaze-only interaction
🔎 Similar Papers
No similar papers found.