conduct quantitative research

Designs and implements structured numerical studies and measurement instruments (e.g., surveys, experiments, observational protocols) to collect quantitative data from sampled populations. Builds sampling and data-collection plans and performs data cleaning, descriptive and inferential statistical analyses (e.g., hypothesis tests, regression, estimation) to quantify relationships, test hypotheses, and report uncertainty.

conductquantitativeresearch

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

On the role of the design phase in a linear regression

Sep 01, 2025
JC
Junho Choi
🏛️ University of Wisconsin-Madison

This paper investigates how the “design phase”—i.e., subsample selection to achieve covariate balance between treatment and control groups—affects causal inference via linear regression in observational studies. Methodologically, it formalizes subsample selection as an estimator adjustment process centered on covariate balancing, rigorously establishing its theoretical role in mitigating bias from model misspecification. It further introduces a sensitivity analysis framework grounded in imbalance metrics, serving both as a quantitative measure of design quality and a transparency vehicle for results. The key contribution lies in unifying the design and estimation phases within the linear regression framework for the first time, thereby elevating covariate balance from a heuristic practice to a theoretically grounded principle and operational standard for bias control. This integration substantially enhances the robustness and reproducibility of causal inference.

Formalizing estimand adjustment through balanced subsample selectionJustifying design phase utility in linear regression analysisUsing covariate balance as model misspecification sensitivity measure

This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.

experimental designexperimental unitsHasse diagrams

Finite Population Survey Sampling: An Unapologetic Bayesian Perspective.

Jun 18, 2023
SB
Sudipto Banerjee
🏛️ UCLA | University of California Los Angeles

Bayesian inference for finite-population surveys is challenging when sampling units exhibit complex dependencies (e.g., spatial, network, or structural) and nonresponse is nonignorable. Method: We propose a unified hierarchical modeling framework that integrates graphical models and spatial random fields to characterize multivariate dependence; formally adopts the “unapologetic Bayesian” paradigm, embedding design-based weights (e.g., Horvitz–Thompson) naturally into prior and likelihood specifications; incorporates causal ignorability analysis to ensure identifiability under missing-not-at-random (MNAR) mechanisms; and employs MCMC and variational inference for scalable computation. Contribution/Results: The framework achieves improved small-area estimation accuracy and more reliable uncertainty quantification in two empirical spatial finite-population analyses. It rigorously reconciles design-based consistency with model-based flexibility, providing theoretical guarantees for valid Bayesian inference under complex survey designs and nonignorable nonresponse.

Developing Bayesian frameworks for ignorable and nonignorable responsesIncorporating multivariate dependencies using graphical and spatial modelsModeling complex dependencies in finite population sampling

Learning control variables and instruments for causal analysis in observational data

Jul 05, 2024
NA
Nicolas Apfel
🏛️ University of Innsbruck | University of York | University of Fribourg | Heinrich Heine University Düsseldorf

Estimating causal effects from observational data requires selecting appropriate control and instrumental variables that satisfy causal identification conditions—a challenging task often reliant on strong domain knowledge or ad hoc assumptions. Method: This paper proposes the first end-to-end joint learning framework that automatically identifies valid combinations of control and instrumental variables. Grounded in conditional independence testing, the method integrates nonparametric dependence measures with structural search optimization, ensuring statistical consistency in variable selection under mild regularity conditions. Contribution/Results: Unlike conventional approaches requiring prespecified variable sets or strong prior assumptions, our framework is fully data-driven. In simulations, it achieves significantly higher variable identification accuracy. Empirically, applied to the Job Corps study, its estimated treatment effect closely aligns with results from the randomized controlled trial—demonstrating both validity and robustness in real-world causal inference.

Detects control variables and instruments for causal analysis in observational dataLearns partition of instruments and control variables from observed dataTests joint existence of instruments and control variables using machine learning

Planning for Gold: Sample Splitting for Valid Powerful Design of Observational Studies

Jun 02, 2024
WB
William Bekerman
🏛️ University of Pennsylvania | World Bank

To address latent bias arising from unmeasured confounding in observational studies, this paper proposes a novel causal inference paradigm based on sample splitting: data are partitioned into planning and analysis samples, where the former adaptively selects robust design parameters (e.g., matching strategies, covariate sets), and the latter yields unbiased causal estimates. The method innovatively integrates multiple testing correction, heteroskedasticity-robust covariance estimation, and formal sensitivity analysis—extending support to multiple outcome variables, thereby relaxing the conventional single-outcome assumption. We establish theoretical guarantees of statistical validity under latent bias. Simulation studies demonstrate substantially higher statistical power than benchmark methods under strong unmeasured confounding. Empirical application to assessing the multidimensional impacts of floods in Bangladesh confirms practical feasibility and robustness.

Addresses bias from unmeasured covariates in observational studiesDevelops hypothesis screening method using split sample designEnables valid causal inference despite unmeasured confounding

Latest Papers

What's happening recently
View more

This study addresses the coarsening of self-reported numeric variables in surveys—often caused by rounding or heaping—by proposing a novel approach that integrates design-based inference with latent variable modeling. Treating observed values as coarsened manifestations of an underlying continuous latent variable, the method jointly models the coarsening mechanism and the latent distribution via a survey-weighted pseudo-likelihood. It generates posterior predictive replicates to propagate coarsening-induced uncertainty into standard design-based estimators. This framework is the first to explicitly correct for coarsening bias under complex sampling designs, enabling unbiased estimation of means, quantiles, and threshold-based prevalence measures. Simulation studies demonstrate robustness across various model misspecifications and sampling scenarios, and empirical application to Italy’s PASSI behavioral surveillance data shows effective correction of coarsening-related estimation bias.

coarseningdesign-based estimationfinite-population inference

Estimating the functional relationship between a continuous exposure and a binary outcome is challenging when covariates are measured with error. This study presents the first systematic evaluation of Simulation-Extrapolation, Regression Calibration, multiple imputation, and Bayesian correction methods, each coupled with flexible modeling techniques—including B-splines, P-splines, and fractional polynomials—within a multi-team, fully blinded, neutral simulation framework. By generating 155 distinct simulation scenarios and repeated samples, the research quantifies the bias and variance of each approach, revealing their relative strengths and limitations. The findings not only inform method selection under measurement error but also demonstrate the feasibility and value of this neutral comparative paradigm for rigorous methodological assessment.

covariate adjustmentexposure-outcome relationshipfunctional form

This study addresses the lack of a clear operational definition of “statistical purpose” in national statistical practice, which often leads to interpretive discrepancies and compliance risks in data governance. By systematically analyzing relevant laws, statistical standards, and ethical frameworks, the paper integrates legal, statistical, and ethical perspectives—offering a novel, comprehensive definition centered on two core principles: a public-interest orientation focused on large populations and robust confidentiality protections for individual data. The proposed definition provides statistical agencies with both theoretical grounding and practical guidance, while also highlighting key unresolved issues that warrant further investigation to meet emerging challenges in data governance.

confidentialitydata ethicslegal interpretation

Existing data integration methods struggle to accommodate complex survey designs and typically assume that multiple data sources originate from the same population, rendering them unsuitable for non-probability samples. This work proposes a model-assisted calibration framework that extends such integration to multiple probability survey samples—a first in the literature—accepting either individual-level data or aggregated summary statistics as input. The approach guarantees design-consistent estimation without requiring correct specification of the outcome model and naturally accommodates complex sampling designs. Coupled with Taylor linearization for variance estimation, the method substantially enhances the efficiency of regression analysis while preserving validity for finite-population inference. Simulation studies and empirical analyses using NHANES and NHIS data demonstrate consistent efficiency gains across diverse scenarios.

complex sampling designsdata integrationfinite-population inference

Hot Scholars

RD

Ronnie de Souza Santos

Assistant Professor, University of Calgary
Human Aspects of Software EngineeringSoftware TestingSoftware FairnessSoftware Development
GG

George Grispos

Associate Professor of Cybersecurity, University of Nebraska at Omaha
Digital ForensicsCybersecurityCritical Infrastructure ProtectionApplied Computing Science
YH

Yuhan Hu

Research Scientist, Apple Inc.
Human Robot InteractionSoft Robotics
JZ

Jianlong Zhou

University of Technology Sydney (UTS)
AI EthicsAI FairnessAI ExplainabilityHuman Centred AI
JZ

Jingyu Zhang

WNLO Huazhong University of Science and Technology
optical