Identifying Representational Biases in Datasets Using PCA: A Max-Disparity Partition Framework

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种基于PCA的最大差异分区框架,通过无标签数据矩阵识别出最大表征差异的二元分区,并分析导致差异的具体特征。
📝 Abstract
Principal Component Analysis (PCA) minimises aggregate reconstruction error, which can inadvertently represent majority subgroups with substantially higher fidelity than minority subgroups. Fairness-aware extensions of PCA correct this disparity but require group labels as input. We address the logically prior question: given only a data matrix, which binary partition of the data suffers the greatest representational disparity under a shared PCA projection? We formalise this as the max-disparity partition problem and propose a greedy local-search algorithm, grounded in the Fiduccia-Mattheyses bipartitioning framework, that discovers the disparity-maximising partition without any predefined group labels. Two benchmark algorithms, a fixed-projection sorting baseline and a simulated-annealing variant, confirm that the greedy solution is empirically near-optimal. Having identified the partition, we attribute the disparity to specific features via PCA loading scores and association rule mining, enabling a practitioner to assess whether the disadvantaged group corresponds to a human-meaningful minority. On the Predict Students' Dropout and Academic Success dataset, representational disparity is driven predominantly by institutional and programmatic proxies for socioeconomic disadvantage, with gender emerging as a secondary but consistent contributor within the disadvantaged group. The discovered partition is then passed directly to Fair PCA, completing a detect-explain-mitigate pipeline.
Problem

Research questions and friction points this paper is trying to address.

PCA
representational disparity
binary partition
group labels
Fiduccia-Mattheyses bipartitioning
Innovation

Methods, ideas, or system contributions that make the work stand out.

PCA
max-disparity partition
Fiduccia-Mattheyses bipartitioning
association rule mining
representational disparity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Arjun KM
Department of Management Studies, Indian Institute of Science, Bangalore
S
Shashi Jain
Department of Management Studies, Indian Institute of Science, Bangalore