Score
Designs, specifies, and evaluates statistical or probabilistic models that represent conditional dependence relationships among multiple variables or outcomes, i.e., the multivariate dependence or dependence structure of a joint distribution given covariates. Builds estimation and testing procedures to capture multivariate/conditional dependence, adjust joint estimates for dependence, and reduce bias that arises from incorrect independence assumptions.
This study addresses the challenge of effectively characterizing nonlinear statistical dependence between two random variables while controlling for the influence of covariates. To this end, it proposes partial copula as a theoretical framework for nonlinear partial correlation, extending the classical notion of linear partial correlation to more general nonlinear settings. The work establishes a formal connection between partial copulas and conditional copula dependence structures. Through rigorous theoretical analysis and simulation experiments, the study demonstrates that partial copulas accurately capture dependence relationships after adjusting for covariates. Moreover, it reveals their potential in causal inference for identifying the sign of causal effects, thereby offering a novel tool for nonlinear causal discovery.
Directional data are often constrained to local intervals on the unit circle (e.g., the first quadrant), and existing multivariate models struggle to simultaneously ensure bounded marginal supports and flexible dependence structures. To address this, we propose a copula-based Bayesian multivariate model: marginal variables are defined on subsets of the unit circle, while joint dependence is flexibly modeled via a copula function. We introduce a projected Gamma prior and a two-stage MCMC sampling scheme for posterior inference. This work constitutes the first systematic extension of multivariate modeling frameworks to bounded circular data. Through extensive simulations and real-data applications, we demonstrate the model’s high accuracy in estimating both the joint distribution and underlying parameters. It significantly enhances statistical modeling capability for restricted directional data, offering improved flexibility, interpretability, and inferential precision compared to existing approaches.
This paper addresses the problem of testing multivariate conditional independence between a response variable (Y) and high-dimensional covariates (X) given confounders (Z). We propose the Multivariate Sufficient-statistic-based Conditional Randomization Test (MS-CRT), which constructs conditional exchangeability by leveraging sufficient statistics of (P(X mid Z)), bypassing explicit modeling of (Y) and accommodating arbitrary test statistics. Key contributions include: (i) the first integration of sufficient statistics into the CRT framework; (ii) overcoming the curse of dimensionality in (P(X mid Z)) by reducing dependence on its parametric dimension; (iii) enabling group selection and false discovery rate (FDR) control; and (iv) establishing minimax-optimal detection rates under multivariate normality. Extensive simulations and real-data analyses demonstrate that MS-CRT substantially improves joint signal detection power—particularly when individual components of (X) exert weak effects on (Y)—and consistently outperforms state-of-the-art methods in graphical model learning tasks.
This paper addresses the problem of interpretable detection of dependency structures among random variables. We propose a novel framework integrating a rank-transform-based estimator for the quantile dependence function with a local acceptance region. The method constructs robust quantile dependence measures via rank standardization and employs local hypothesis testing to enable visual diagnostic assessment of dependency patterns and rigorous independence testing under finite samples. Key contributions include: (1) the first nonparametric estimation and theoretical derivation of the quantile dependence function; (2) guaranteed validity of statistical tests at any sample size, balancing high global power with precise localization of heterogeneous dependencies; (3) superior empirical power across diverse alternative models and successful identification of heterogeneous non-independence in real-world data; and (4) a computationally efficient algorithm supporting intuitive, graphical diagnostic interpretation.
In clinical trials—particularly ophthalmology—homogeneity testing for bilateral proportion data is commonly constrained by pre-specified, inflexible dependence structures (e.g., independence, perfect positive/negative dependence), limiting interpretability and adaptability. This paper introduces the Clayton copula—a flexible, interpretable tool for modeling asymmetric lower-tail dependence—into bilateral proportion homogeneity testing for the first time, thereby eliminating reliance on a priori dependence assumptions. We propose three Clayton copula–based test statistics and rigorously evaluate them via Monte Carlo simulation, demonstrating well-controlled Type I error rates and superior statistical power. Furthermore, we validate the robustness and practical utility of our approach on two real-world ophthalmologic datasets. This work establishes a theoretically rigorous, computationally feasible, and clinically meaningful testing paradigm for bilateral proportion data, advancing both methodological foundations and applied biostatistical practice.
This study addresses the challenge of modeling edge effects and dependence structures that evolve with covariates—such as age—in multivariate responses of mixed types. Existing approaches are often hindered by strong assumptions or insufficient flexibility. To overcome these limitations, this work proposes a Bayesian nonparametric framework that integrates adaptive spline-based marginal regression with a covariate-dependent Gaussian copula infinite mixture model. A probit stick-breaking process is introduced to flexibly capture the covariate-driven evolution of dependence patterns, avoiding restrictive global correlation matrix constraints. The method unifies heterogeneous response types and dynamic dependencies through varying-coefficient copula regression and employs Markov chain Monte Carlo algorithms for posterior inference. Simulation studies demonstrate its accuracy and robustness, while empirical analysis of the 2023 Behavioral Risk Factor Surveillance System (BRFSS) data reveals complex age-varying marginal and dependence structures in health outcomes.
This study addresses the challenge of achieving consistent conditional independence testing and association estimation in continuous sample spaces under small-sample, high-dimensional settings. The authors propose a nonparametric multiscale approach that decomposes the continuous space via cascaded 2×2×T contingency tables and conditions on marginal order statistics. This framework extends the Cochran–Mantel–Haenszel (CMH) test and odds ratio estimation to continuous variables for the first time, ensuring statistical consistency without requiring asymptotic layer-wise sample sizes. The method simultaneously supports hypothesis testing and identification of local association strength and direction, with near-linear computational complexity. Empirical evaluations demonstrate its superior or competitive statistical power while properly controlling Type I error, and it successfully uncovers local conditional dependence structures in real-world Uber mobility data.
This study addresses the challenge of modeling electronic health records (EHR), which comprise high-dimensional, mixed-type variables with complex nonlinear dependencies that are poorly captured by traditional statistical methods due to their reliance on strong distributional assumptions such as Gaussianity. To overcome this limitation, the work introduces vine copulas—a flexible probabilistic framework—for the first time in EHR analysis. By decomposing multivariate distributions into a sequence of bivariate conditional dependencies arranged in a tree structure, the approach enables accurate modeling of heterogeneous data types without restrictive parametric assumptions. The proposed method facilitates data-driven variable selection, identification of conditional dependencies among comorbidities, and characterization of patient cohorts. Accompanied by visualization tools and open-source code, this framework promotes reproducible and interpretable probabilistic exploration of healthcare data.
Traditional meta-analyses of combined diagnostic tests often yield biased estimates due to the assumption of conditional independence, and existing approaches either require complete joint data or suffer from computational instability. This study proposes a Bayesian hierarchical model that flexibly captures the conditional dependence between two binary diagnostic tests through study-specific log odds ratios, without imputing missing data or assuming a perfect reference standard. The method provides a unified framework for synthesizing heterogeneous study designs—including those without a gold standard or with partial verification—using a stable parameterization that mitigates bias in accuracy estimation. Validation through two real-world meta-analyses demonstrates that ignoring conditional dependence substantially distorts results, whereas the proposed framework yields accurate estimates of joint diagnostic performance while maintaining computational stability.