Gaze-Anchored Social Net: Decoding Implicit Relations via Joint Modeling

πŸ“… 2026-07-24
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of modeling gaze behavior driven by implicit social intent, which existing methods struggle to capture due to their tendency to process individuals in isolation or assign social relationships post hoc. To overcome this limitation, we propose ANCHOR, a novel framework that treats social intent as a latent structural driver of gaze behavior and jointly models the distributions of visual attention and implicit social relations in a target-centric manner. Our approach integrates a relational attention mechanism, feature-level modulation, and a single-backbone multi-task architecture, complemented by a flat minima–guided co-optimization strategy to mitigate gradient conflicts between gaze localization and social reasoning. Evaluated on an extended benchmark featuring dense multi-person annotations and a new social influence ranking metric, ANCHOR achieves state-of-the-art performance and provides the first quantitative evidence that static gaze patterns can robustly disentangle and learn implicit social hierarchies.
πŸ“ Abstract
Human gaze does more than point to visual targets; it serves as a subtle indicator of social intent within static images, whereas standard models typically process individuals independently, treating gaze as an i.i.d. quantity or predicting social semantics in isolation. Recent multi-person methods attempt to address this but often treat social relations as rigid, post-hoc classifications decoupled from the gaze estimation process. This oversimplification fails to capture the nuanced nature of social intent, which acts as an underlying driver of gaze behavior rather than a secondary categorical output. We address these limitations by proposing ANCHOR, a target-centric paradigm designed to decode gaze-anchored social intent by modeling the joint distribution of visual attention and latent implicit relations. Our approach surfaces these dependencies as the latent structural scaffolding of gaze behavior. The architecture utilizes a relational attention mechanism to capture fine-grained interpersonal links, leveraging feature-wise modulation for efficient multi-person parsing from a single vision backbone. To stabilize the training of this coupled formulation, we implement an optimization synergy to resolve the inherent conflicts between spatial gaze accuracy and latent social reasoning. This approach ensures robust generalization by seeking stable, flat minima while simultaneously harmonizing competing task gradients. We validate our framework on an extended benchmark featuring dense multi-person annotations and novel social influence rankings. Our results demonstrate state-of-the-art performance and provide the first quantitative evidence that implicit social hierarchies can be robustly disentangled and learned directly from static gaze patterns.
Problem

Research questions and friction points this paper is trying to address.

gaze estimation
social relations
implicit intent
multi-person interaction
visual attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

gaze-anchored social intent
joint modeling
relational attention
implicit social hierarchy
optimization synergy
πŸ”Ž Similar Papers
No similar papers found.