Score
Designs and implements end-to-end eye‑tracking data pipelines, including protocols and tools for gaze data collection, preprocessing (saccade/fixation detection, pupilometry calibration), and normalization across stimuli and users. Builds analytic and feature‑engineering methods to extract fixation- and gaze-based descriptors—spatial, temporal and statistical features, per-user and per-image vectors—and applies those features to analyze attention allocation and gaze-informed behavior.
This study addresses the limitations of traditional eye movement event detection, which relies heavily on programming expertise and is sensitive to data preprocessing and parameter tuning, thereby hindering deployment in non-specialist settings. To overcome these barriers, this work proposes the first large language model (LLM)-based, zero-code framework for eye-tracking analysis. The system enables end-to-end analytical pipelines—from raw data parsing and cleaning to event labeling (e.g., fixations and saccades) and algorithm optimization—through natural language instructions. It automatically infers data structure and integrates established algorithms such as I-VT and I-DT. Evaluated on public benchmarks, the method achieves accuracy comparable to conventional approaches while substantially lowering technical barriers, thus enhancing the accessibility and interpretability of eye movement research.
To address the limitations of conventional RGB cameras—low temporal resolution, high computational overhead, and difficulty capturing rapid pupil dynamics—in eye-tracking applications, this paper proposes a lightweight pupil tracking and identity authentication framework leveraging neuromorphic event cameras. Methodologically: (1) an adaptive event slicing algorithm dynamically accumulates asynchronous event streams to generate monocular event frames; (2) an end-to-end lightweight segmentation and tracking pipeline is developed; (3) short-term pupil kinematics—namely position, velocity, and acceleration—are first modeled as individual-specific biometric traits, with identity authentication performed efficiently via random forest classification. Experiments demonstrate a pupil tracking IoU of 92% (a >6% improvement), over threefold reduction in end-to-end latency, an authentication accuracy of 82%, and inference time of only 12 ms per frame.
Existing eye movement and gaze estimation research is hindered by the scarcity of large-scale, multi-scenario, high-precision publicly available datasets—especially in real-world VR/AR environments. To address this, we introduce the largest head-mounted device-collected eye image dataset to date (>20 million images), spanning diverse daily activities and VR/AR scenarios. It features the first multi-device synchronized acquisition and unified annotation of comprehensive eye-related attributes: 2D/3D eye landmarks, pupil/iris/eyelid segmentation masks, parametric 3D eyeball models, gaze vectors, and fine-grained eye movement types. We propose a geometrically constrained eyeball fitting and gaze estimation method, integrated with a semi-automatic labeling pipeline validated by domain experts. This dataset establishes the first real-world benchmark for eye movement analysis, yielding consistent improvements of 12–28% in eye movement estimation and gaze prediction accuracy across multiple state-of-the-art models.
Existing eye movement modeling approaches primarily focus on low-frequency features, neglecting user-specific fine-grained motion patterns embedded in high-frequency components (>30 Hz), thereby failing to generate identity-discriminative gaze sequences. This work introduces the first user-specific synthesis framework for high-frequency eye movements: it models individual high-frequency gaze characteristics as injectable “identity noise” and proposes a user-identity-guided conditional diffusion model. To jointly optimize identity discriminability and motion naturalness, we incorporate pretrained authentication embeddings and a spatial-domain identity fidelity loss. Evaluated on two public high-frequency eye-tracking datasets, our synthesized sequences are perceptually indistinguishable from real recordings. The framework successfully supports diverse downstream applications—including gaze imputation, super-resolution, animation driving, biometric identification, and context-aware modeling—demonstrating both strong identity preservation and dynamic plausibility.
This paper addresses privacy leakage risks arising from the correlation between eye-tracking data and visual stimuli in VR environments. We systematically survey full-stack VR eye-tracking technologies—from pupil detection and gaze estimation to cognitive modeling—published between 2012 and 2022, alongside their associated privacy threats. First, we establish the first cross-disciplinary survey framework bridging VR eye-tracking and privacy protection, identifying three privacy-centric research directions. Second, we propose a novel co-design paradigm integrating eye movement authentication with data anonymization, synergizing computer vision, human-computer interaction modeling, differential privacy, adversarial generation, and biometric encryption. Third, we clarify the technological evolution trajectory and privacy threat landscape, and introduce quantifiable evaluation metrics and an implementable defense roadmap. Our work provides both theoretical foundations and practical guidelines for developing secure and trustworthy VR systems. (149 words)
This study addresses the high costs and limited applicability of traditional eye-tracking-based vision screening, which relies on screen calibration and precision equipment. We propose a low-cost, screen-calibration-free gaze acquisition framework using a GC308 near-infrared camera, integrating the Gaze Quest and Orlosky pipelines. Notably, this work introduces inter-frame gaze vector angular variation as a novel core screening metric, replacing conventional fixation accuracy. Experimental evaluations confirm the data acquisition stability of this configuration, yielding a mean inter-frame gaze angle variation of 28.2 degrees and an effective sampling rate of approximately 8 FPS. These findings demonstrate the feasibility and effectiveness of employing low-cost hardware for vision screening without screen calibration, offering a scalable alternative for resource-constrained settings.
研究解决了几何眼动追踪系统中的校准误差问题,通过创建标准化数据集并评估多种校准方法,包括引入轻量级神经优化器,有效降低了校准误差。
This work proposes a low-cost, scalable real-time gaze tracking method that operates within standard web browsers using only an ordinary webcam, enabling accessible cognitive and clinical research. The system leverages a lightweight, open-source deep learning pipeline that integrates MediaPipe for facial landmark detection with a convolutional neural network inspired by iTracker, augmented by a user-specific fine-tuning mechanism. We evaluate two training strategies—transferring a model pretrained on mobile data versus training from scratch on desktop-collected data—and find comparable performance on the MPIIFaceGaze benchmark, with fine-tuning substantially reducing prediction error. In a Dot-Probe task, the system’s left–right gaze allocation closely aligns with outputs from the commercial SeeSo SDK, demonstrating practical validity. The approach balances transparency, scalability, and ease of deployment.
This study addresses the scarcity of real-world eye-tracking data, which is hindered by high annotation costs and privacy concerns, thereby limiting large-scale behavioral modeling research. To overcome this bottleneck, the authors propose an end-to-end synthetic data generation framework that combines authentic iris trajectory extraction with replay in a 3D eye movement simulator, enabling the first large-scale, high-fidelity, and automatically annotated eye-tracking video synthesis. By integrating headless browser automation with trajectory replay techniques, the method effectively circumvents the constraints of real data collection. The released dataset comprises 144 sessions totaling 12 hours of 25-fps synthetic eye-tracking videos, exhibiting highly faithful temporal dynamics (KS D < 0.14), and demonstrates strong validity in script-reading detection tasks.
研究通过收集19名参与者在真实场景下的眼动数据,构建了GazeDepth数据集,用于基于注视行为估计观察距离,支持距离感知交互。