association testing

Design and implement statistical hypothesis tests that assess associations between genetic variants (genotypes) and phenotypes, including tests conditioned on or adjusted for departures from Hardy–Weinberg equilibrium. This work includes deriving appropriate null distributions and asymptotics, computing association statistics and p-values that account for HWE, controlling for covariates and confounders, estimating effect sizes and uncertainty, and comparing associations across thresholds or alternative models.

associationtesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Traditional genome-wide association studies (GWAS) typically perform separate single-nucleotide polymorphism (SNP) association tests and Hardy–Weinberg equilibrium (HWE) tests, relying on arbitrary thresholds to filter loci—a practice that can introduce false positives or discard valuable information. This work proposes a unified conditional testing framework that integrates HWE information directly into the association analysis by conditioning the Pearson χ² statistic from the 3×2 case–control contingency table on the HWE χ² statistic computed from controls alone. By leveraging asymptotic distribution theory, this approach eliminates the need for a standalone HWE filtering step and yields more accurate p-values. Both simulations and real-data analyses on alopecia demonstrate that the method substantially improves statistical power and SNP ranking accuracy compared to existing retrospective approaches, thereby reducing replication costs and enhancing fine-mapping resolution.

case-control studyfalse positivesGWAS

This study addresses the challenges in testing Hardy–Weinberg equilibrium (HWE) on the X chromosome, where sex-specific genotype structures and sex-differential minor allele frequencies (sdMAF) lead to ambiguous null hypotheses and inconsistent methodologies. The authors propose a unified statistical framework based on allelic regression that reframes HWE testing as an assessment of allelic dependence. This approach explicitly clarifies the null hypotheses, degrees of freedom, and sensitivity to sdMAF across different tests, while naturally accommodating covariate adjustment and correction for population structure. The framework subsumes existing chi-square methods and achieves coherent HWE inference between autosomes and the X chromosome. Simulations and analyses of high-coverage data from the 1000 Genomes Project demonstrate that the proposed method properly controls Type I error and avoids the misleading conclusions often produced by conventional approaches in the presence of sdMAF.

genetic data analysisHardy-Weinberg equilibriumsex differences in minor allele frequency

New Equivalence Tests for Hardy-Weinberg Equilibrium and Multiple Alleles

Jul 14, 2025
VO
Vladimir Ostrovski
🏛️ ERGO Group AG

This paper addresses the underexplored problem of testing Hardy–Weinberg equilibrium (HWE) equivalence in multi-allelic settings. We propose two novel test statistics: one grounded in asymptotic distribution theory and the other employing bootstrap resampling to estimate variance—both avoiding strong assumptions about genotype frequency distributions, thereby markedly improving robustness and applicability in small samples. Our key contribution is the first rigorous, computationally feasible framework for HWE equivalence testing under multi-allelic loci. Through extensive simulations and evaluation on three real-world population genetic datasets, the proposed methods demonstrate superior control of Type I error rates and higher statistical power compared to existing tests—particularly in scenarios involving rare alleles and limited sample sizes.

Assessing finite sample performance via real-data-inspired simulationsProposing two test statistics for asymptotic or bootstrap-based evaluationTesting equivalence to Hardy-Weinberg Equilibrium for multiple alleles

This study addresses a critical gap in genetic research by developing the first closed-form sample size formula for testing the total effect of a single nucleotide polymorphism (SNP)—encompassing both direct effects and indirect effects mediated through longitudinal biomarkers—within joint models of longitudinal and survival data. Existing methods lack appropriate power and sample size calculations for this integrative framework. The proposed approach is grounded in joint modeling theory and rigorously validated via Monte Carlo simulations, demonstrating high accuracy and robustness under finite-sample settings. Simulation studies confirm the reliability of the derived formula, and its practical utility is illustrated through a successful application to data from the Diabetes Control and Complications Trial (DCCT), underscoring its feasibility for real-world genetic study design.

joint modelinglongitudinal and survival datapower calculation

Some Results on Generalized Familywise Error Rate Controlling Procedures under Dependence

Apr 24, 2025
MD
Monitirtha Dey
🏛️ University of Bremen | Indian Statistical Institute

This paper addresses the problem of controlling the generalized family-wise error rate (k-FWER) in multiple hypothesis testing under diverse dependence structures—including positive dependence, negative dependence, and weak dependence commonly encountered in genome-wide association studies (GWAS). Moving beyond the conventional assumption of independence, we derive the first systematic, tight upper bounds on k-FWER for these dependence regimes. Methodologically, we integrate extreme value theory, Gaussian comparison inequalities, and correlation-aware truncation strategies to develop a dependence-adaptive k-FWER control procedure. Theoretically, we establish rigorous finite-sample control guarantees across all considered dependence classes. Numerical experiments demonstrate that our method achieves significantly improved control accuracy and statistical power over existing approaches, particularly in high-dimensional sparse signal detection settings.

Develop improved testing procedures via probability inequalitiesEstablish bounds on error rates for correlated frameworksStudy k-FWER control under varying dependence structures

Latest Papers

What's happening recently
View more

This study addresses the fundamental question of whether a statistical parameter defined through a conditional distribution remains constant across covariates—a problem encompassing treatment effect heterogeneity and conditional association. The authors propose a general nonparametric testing framework based on smooth functionals applied to conditional distributions, yielding functional parameters for which they construct test statistics with tractable asymptotic distributions. Their approach explicitly links to norm-based tests in function spaces and, compared to existing norm-type methods, exhibits superior asymptotic properties under the null hypothesis. Simulation studies demonstrate strong finite-sample performance, and the method is successfully applied to data from a breast cancer clinical trial, effectively identifying key biomarkers predictive of response to adjuvant chemotherapy.

conditional distributionsfunction-valued parametershypothesis testing

Traditional meta-analyses of combined diagnostic tests often yield biased estimates due to the assumption of conditional independence, and existing approaches either require complete joint data or suffer from computational instability. This study proposes a Bayesian hierarchical model that flexibly captures the conditional dependence between two binary diagnostic tests through study-specific log odds ratios, without imputing missing data or assuming a perfect reference standard. The method provides a unified framework for synthesizing heterogeneous study designs—including those without a gold standard or with partial verification—using a stable parameterization that mitigates bias in accuracy estimation. Validation through two real-world meta-analyses demonstrates that ignoring conditional dependence substantially distorts results, whereas the proposed framework yields accurate estimates of joint diagnostic performance while maintaining computational stability.

conditional dependencediagnostic test accuracyimperfect reference standard

This study addresses the limitations of the ASSET method in subset-based meta-analysis, which relies on normality assumptions to compute p-values and whose analytical approximations become inaccurate under extreme tail probabilities or non-normal conditions—such as small sample sizes or low-frequency variants—while conventional Monte Carlo simulations incur prohibitive computational costs. The work presents the first systematic evaluation of ASSET’s accuracy in estimating tail p-values and introduces an efficient importance sampling (IS) algorithm that accurately estimates extremely small p-values in both independent and overlapping study designs. The proposed method maintains high precision even under non-normality and demonstrates substantial gains in computational efficiency. Its practical utility is validated through applications to the OneK1K dataset and a Korean lung cell single-cell eQTL analysis.

ASSETimportance samplingmeta-analysis

Hot Scholars

AC

Aylin Caliskan

Assistant Professor, University of Washington
AI biasAI ethicsmachine learningnatural language processing
MQ

Mengyun Qiao

University College London
Generative AIMulti-agent ReasoningMedical image analysis
BK

Bernhard Kainz

FAU Erlangen-Nürnberg, Imperial College London
human-in-the-loop computingmachine learningmedical image analysis
QL

Qing Lu

Associate Professor, Division of Biostatistics, Department of Epidemiology and Biostatistics
statistical geneticsbioinformaticsgenetic epidemiology
KW

Kyra Wilson

PhD Student, University of Washington