🤖 AI Summary
This study addresses the auditing distortion in personalized news recommendation caused by the conflation of preference, exposure, and click data. We propose PEC, a hierarchical auditing framework that explicitly delineates five data boundaries. Leveraging large-scale mobile user behavior logs, the framework integrates weighted profile modeling with statistical baseline comparisons to systematically decouple and quantify the independent effects of these three factors. Our analysis reveals inherent limitations in single-metric evaluation: exposure exhibits limited predictive power for clicks, and different measurement dimensions yield significantly divergent conclusions regarding diversity. By innovatively establishing a reusable hierarchical auditing paradigm that underscores the irreplaceability of each data layer, this work effectively corrects traditional evaluation biases and provides empirical foundations for recommendation algorithm governance.
📝 Abstract
Personalized news platforms are often evaluated as if stated preferences, logged recommendation exposure, and click consumption form a single coherent pipeline. Collapsing these layers can distort audit conclusions: a platform may appear more aligned or diverse than observed click consumption supports, which can misdirect diversity governance or algorithmic intervention. We introduce a reusable Preference-Exposure-Consumption (PEC) audit framework that separates stated preference, observed weighted profile state, logged recommendation exposure, app-surface pathways, and click consumption under explicit observability boundaries. Using six months of logs from a deployed mobile news application, we audit 1,583 user profiles, 95,143 logged recommendation items, and 17,512 click events. Each trace type contributes distinct information; none directly substitutes for another. Preference-consumption alignment exceeds chance-based null baselines but captures only part of users' top-consumed category set. Logged recommendation lists contain clicked articles more often than a date-matched candidate-pool baseline predicts (11.29% vs 9.13%; top-5 lift 1.47x), but many clicks arrive through other app surfaces. Raw exposure-consumption diversity gaps shrink under count matching, yet concentration mismatch persists in the audit-eligible cohort (HHI gap 0.082). Together, the results show that audit conclusions change depending on whether platforms measure stated preference, logged exposure, or click consumption.