Score
Designs and implements analyses and benchmarks that compare metrics across demographic cohorts (e.g., age, tenure, role, recruitment and employment groups), using appropriate statistical tests including non-parametric methods to detect significant group differences. Produces group-level summaries and prioritized lists of cohorts with relatively higher or lower scores to inform targeted improvement interventions.
This study addresses the absence of standardized metrics for quantifying the distributional alignment between survey samples and target populations across multidimensional demographic characteristics. To this end, it introduces the Global Representativeness Index (GRI), which—by incorporating total variation distance into survey methodology for the first time—establishes a symmetric [0,1] scoring framework to assess the fidelity of samples with respect to complex demographic structures. The GRI leverages benchmark demographic data from the United Nations and Pew Research Center and complements design effect to form a novel paradigm for sample quality evaluation, implemented via an open-source Python library. Validation across multiple international survey datasets reveals that even large-scale probability samples typically achieve fine-grained GRI scores below 0.36, underscoring substantial deficiencies in current surveys’ demographic representativeness.
This study addresses fairness risks in AI recommendation systems arising from demographic imbalances. Methodologically, it introduces a novel, customizable controlled-dataset construction paradigm; establishes an “affected population vs. overall population” comparative analytical model; designs quantitative bias metrics and a dynamic drift detection mechanism; and integrates these components into the AI auditing platform ITACA. The contributions include: (1) the first operational definition and end-to-end lifecycle monitoring of fairness thresholds across multiple application scenarios; (2) support for training-data calibration, fairness-aware objective formulation, and post-deployment continuous auditing; and (3) real-world validation through Eticas.ai’s auditing practice, delivering actionable fairness guidelines for developers and enabling regulators to formulate verifiable, implementable AI governance policies.
Electronic health records (EHRs) are widely used in epidemiology and AI research, yet their data quality suffers from subgroup bias, systematic errors, and insufficient applicability assessment. To address these challenges, we propose a context-aware, automated medical data quality evaluation framework—the first to adapt software engineering principles of unit testing and coverage analysis to EHR validation. Our method integrates large language models (LLMs) for test case generation, medical knowledge anchoring, and research-context-driven data fitness analysis. Based on this, we develop MDPT, a tool comprising a test generator and executor. Evaluated on All of Us, MIMIC-III, and SyntheticMass datasets, MDPT generates 55–73 tests per cohort and detects 20–43 instances of data inconsistency or anomaly. The approach significantly improves both accuracy and interpretability in assessing EHR suitability for downstream research.
Traditional cancer detection models suffer from evaluation bias due to class imbalance, inter-population performance heterogeneity, and inconsistent patient-level predictions. To address these limitations, we propose CAT (Cohort-Attention-based Evaluation Framework), the first framework introducing patient-level confidence aggregation and entropy-driven distribution weighting. CAT redefines cohort-weighted sensitivity (CATSen), specificity (CATSpe), and their harmonic mean (CATMean). Leveraging multi-center stratified evaluation and adaptive threshold recalibration, CAT enables fair, interpretable, and population-robust assessment of medical AI systems. In multi-center cancer screening tasks, CAT reduces false-negative misclassification rates by 12.7%, significantly enhancing cross-population performance comparability and clinical applicability.
This study investigates the impact of using raw scores versus demographically adjusted scores on classification accuracy and decision fairness in cognitive screening. Through theoretical analysis and empirical validation on the OASIS-3 dataset, it rigorously derives, for the first time, sufficient conditions under which raw scores outperform adjusted scores. The findings demonstrate that common adjustment methods—such as z-score normalization—not only can degrade classification performance under certain conditions but also do not necessarily enhance fairness, thereby challenging the widely held assumption that statistical correction inherently promotes equitable outcomes. This work provides a principled theoretical foundation and practical guidance for score selection in cognitive assessment protocols.
Existing benchmarks for knowledge work evaluation largely adhere to traditional NLP task paradigms, failing to capture systems’ capabilities in real-world knowledge-intensive settings. This work proposes a three-step framework—explicitly defining work activities, establishing realistic test environments, and focusing evaluation on deliverable outputs—and derives 18 core knowledge work activities from the O*NET database. Innovatively integrating role responsibilities, local tool usage, and downstream usability into benchmark design, the approach establishes a coherent “work activity–test setup–scoring artifact” alignment. Validation through three case studies (GDPval, OfficeQA Pro, and APEX-SWE) exposes critical misalignments in current benchmarks between tasks, environments, and actual work objectives, offering a new paradigm for evaluating knowledge work systems in practical, application-oriented contexts.