demographic benchmarking

Designs and implements analyses and benchmarks that compare metrics across demographic cohorts (e.g., age, tenure, role, recruitment and employment groups), using appropriate statistical tests including non-parametric methods to detect significant group differences. Produces group-level summaries and prioritized lists of cohorts with relatively higher or lower scores to inform targeted improvement interventions.

demographicbenchmarking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the absence of standardized metrics for quantifying the distributional alignment between survey samples and target populations across multidimensional demographic characteristics. To this end, it introduces the Global Representativeness Index (GRI), which—by incorporating total variation distance into survey methodology for the first time—establishes a symmetric [0,1] scoring framework to assess the fidelity of samples with respect to complex demographic structures. The GRI leverages benchmark demographic data from the United Nations and Pew Research Center and complements design effect to form a novel paradigm for sample quality evaluation, implemented via an open-source Python library. Validation across multiple international survey datasets reveals that even large-scale probability samples typically achieve fine-grained GRI scores below 0.36, underscoring substantial deficiencies in current surveys’ demographic representativeness.

demographic fidelityrepresentativenesssample quality

Demographic Benchmarking: Bridging Socio-Technical Gaps in Bias Detection

Jan 27, 2025
GG
Gemma Galdon Clavell
🏛️ Eticas AI

This study addresses fairness risks in AI recommendation systems arising from demographic imbalances. Methodologically, it introduces a novel, customizable controlled-dataset construction paradigm; establishes an “affected population vs. overall population” comparative analytical model; designs quantitative bias metrics and a dynamic drift detection mechanism; and integrates these components into the AI auditing platform ITACA. The contributions include: (1) the first operational definition and end-to-end lifecycle monitoring of fairness thresholds across multiple application scenarios; (2) support for training-data calibration, fairness-aware objective formulation, and post-deployment continuous auditing; and (3) real-world validation through Eticas.ai’s auditing practice, delivering actionable fairness guidelines for developers and enabling regulators to formulate verifiable, implementable AI governance policies.

Bias ReductionFairnessResponsible AI

Medical Data Pecking: A Context-Aware Approach for Automated Quality Evaluation of Structured Medical Data

Jul 03, 2025
IG
Irena Girshovitz
🏛️ Tel Aviv University | AI and Data Science Center of Tel Aviv University | Safra Center for Bioinformatics

Electronic health records (EHRs) are widely used in epidemiology and AI research, yet their data quality suffers from subgroup bias, systematic errors, and insufficient applicability assessment. To address these challenges, we propose a context-aware, automated medical data quality evaluation framework—the first to adapt software engineering principles of unit testing and coverage analysis to EHR validation. Our method integrates large language models (LLMs) for test case generation, medical knowledge anchoring, and research-context-driven data fitness analysis. Based on this, we develop MDPT, a tool comprising a test generator and executor. Evaluated on All of Us, MIMIC-III, and SyntheticMass datasets, MDPT generates 55–73 tests per cohort and detects 20–43 instances of data inconsistency or anomaly. The approach significantly improves both accuracy and interpretability in assessing EHR suitability for downstream research.

Automated quality evaluation of structured medical dataIdentifying data quality issues in Electronic Health RecordsImproving reliability of EHR data for research and AI

Traditional cancer detection models suffer from evaluation bias due to class imbalance, inter-population performance heterogeneity, and inconsistent patient-level predictions. To address these limitations, we propose CAT (Cohort-Attention-based Evaluation Framework), the first framework introducing patient-level confidence aggregation and entropy-driven distribution weighting. CAT redefines cohort-weighted sensitivity (CATSen), specificity (CATSpe), and their harmonic mean (CATMean). Leveraging multi-center stratified evaluation and adaptive threshold recalibration, CAT enables fair, interpretable, and population-robust assessment of medical AI systems. In multi-center cancer screening tasks, CAT reduces false-negative misclassification rates by 12.7%, significantly enhancing cross-population performance comparability and clinical applicability.

Addresses biased evaluation in cancer detection modelsEnhances fairness and reliability in medical screeningProposes Cohort-Attention Evaluation Metrics (CAT) framework

Latest Papers

What's happening recently
View more

This study investigates the impact of using raw scores versus demographically adjusted scores on classification accuracy and decision fairness in cognitive screening. Through theoretical analysis and empirical validation on the OASIS-3 dataset, it rigorously derives, for the first time, sufficient conditions under which raw scores outperform adjusted scores. The findings demonstrate that common adjustment methods—such as z-score normalization—not only can degrade classification performance under certain conditions but also do not necessarily enhance fairness, thereby challenging the widely held assumption that statistical correction inherently promotes equitable outcomes. This work provides a principled theoretical foundation and practical guidance for score selection in cognitive assessment protocols.

classification accuracycognitive screeningdemographic correction

Existing benchmarks for knowledge work evaluation largely adhere to traditional NLP task paradigms, failing to capture systems’ capabilities in real-world knowledge-intensive settings. This work proposes a three-step framework—explicitly defining work activities, establishing realistic test environments, and focusing evaluation on deliverable outputs—and derives 18 core knowledge work activities from the O*NET database. Innovatively integrating role responsibilities, local tool usage, and downstream usability into benchmark design, the approach establishes a coherent “work activity–test setup–scoring artifact” alignment. Validation through three case studies (GDPval, OfficeQA Pro, and APEX-SWE) exposes critical misalignments in current benchmarks between tasks, environments, and actual work objectives, offering a new paradigm for evaluating knowledge work systems in practical, application-oriented contexts.

benchmark designevaluationknowledge work

Hot Scholars

LQ

Lucy Qin

Georgetown University
securityprivacyapplied cryptography
LL

Lan Luo

Assistant Professor of Biostatistics, Rutgers University
Streaming dataonline statistical inferencemobile healthmediation analysis
YJ

Yohan Jo

Seoul National University
Natural Language ProcessingAgentsComputational PsychologyReasoning
SB

Sebastian Baltes

University of Bayreuth
software engineeringempirical software engineering
YY

Yumeng Yang

ShanghaiTech University, National University of Singapore
SpintronicsSilicon-on-insulator