Institution profile

Merck & Co., Inc.

Industry researchnorthamerica · us
Official website
Research library37linked papers
Opportunities0open roles
Selected work

Representative Papers

Bayesian Joint Additive Factor Models for Multiview Learning

Jun 02, 2024arXiv.org

To address challenges in multi-omics and other multi-view data—including difficulty modeling cross-view dependencies, strong signal heterogeneity, and insufficient interpretability and uncertainty quantification—this paper proposes JAFAR, a joint Bayesian factor model. Methodologically, JAFAR introduces the Dependency-Cumulative Shrinkage Prior (D-CUSP), which jointly characterizes shared and view-specific latent factor structures while ensuring parameter identifiability. It integrates Bayesian nonparametrics, structured additive designs, partially collapsed Gibbs sampling, and flexible distributional extensions—accommodating non-Gaussian features and survival outcomes. In an application to preterm birth prediction, JAFAR jointly analyzes immunomic, metabolomic, and proteomic data, achieving statistically significant improvements over state-of-the-art methods. The model enables interpretable feature selection and principled uncertainty quantification. An open-source R package implementing JAFAR is publicly available.

1 citationsRead paper

Auditing Pairwise Equivalence Judgments: Self-Critique Effects and Diversity Measurement in Multi-Agent Hypothesis Generation

Oct 02, 2026

This study addresses the interpretability challenges in evaluating the effectiveness of self-critique and diversity within multi-agent hypothesis generation. It systematically investigates how critique rounds and equivalence rules influence mechanistic divergence, employing pairwise equivalence judgment as the core measurement criterion combined with LLM-as-a-Judge, TF-IDF, and dense embedding techniques for hypothesis pair scoring. Results demonstrate that a single round of critique induces 34.5% mechanistic divergence, while varying definitions of equivalence rules cause the identified discrepancy rate to fluctuate dramatically between 38% and 96%, revealing the decisive impact of evaluation criteria selection on outcomes. Overall, this work provides an interpretable analytical framework for understanding and quantifying hypothesis diversity in multi-agent systems.

0 citationsRead paper

Contamination, Prior, or Evidence? Decomposing and Training Evidence Use in Whole-Slide Vision-Language Models

Sep 26, 2026

This study addresses the challenge in evaluating pathology vision-language models, where data contamination and prior knowledge are difficult to disentangle from genuine image evidence. To overcome this, we introduce CleanSlide, a benchmark that employs multiple-choice question-answering auditing to eliminate data leakage. Furthermore, we propose a Pair-DPO loss that leverages authentic counterfactual slide pairs for preference optimization, precisely decoupling non-image signals to strengthen the utilization of visual evidence. Experimental results demonstrate that our approach yields a 15.29% gain in image evidence reliance and improves accuracy on external CPTAC and BCNB cohorts by 9.7% and 3.4%, respectively. Ultimately, this work provides a reliable and unbiased evaluation and training paradigm for multimodal models in computational pathology.

0 citationsRead paper

RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction

Aug 06, 2026

This work addresses the challenge of reaction yield prediction, which is hindered by scarce labeled data, the vast and sparse reaction space, and the inability of existing representations to capture complex chemical transformations. To overcome these limitations, the authors propose RxnCLF, a self-supervised contrastive learning framework that introduces a novel condensed reaction graph (CRG) integrating both reactant and product information. By leveraging graph neural networks, RxnCLF learns explicit and interpretable transformation structures and models chemical reactions within a unified continuous latent space. The method significantly outperforms current graph- and sequence-based models on multiple yield prediction benchmarks, demonstrating substantial improvements in R² scores, and exhibits strong generalization capabilities on downstream tasks such as regioselectivity and enantioselectivity prediction, offering a new paradigm toward a general-purpose foundation model for chemical reactions.

0 citationsRead paper
Recent publications

Latest Papers

Auditing Pairwise Equivalence Judgments: Self-Critique Effects and Diversity Measurement in Multi-Agent Hypothesis Generation

Oct 02, 2026

This study addresses the interpretability challenges in evaluating the effectiveness of self-critique and diversity within multi-agent hypothesis generation. It systematically investigates how critique rounds and equivalence rules influence mechanistic divergence, employing pairwise equivalence judgment as the core measurement criterion combined with LLM-as-a-Judge, TF-IDF, and dense embedding techniques for hypothesis pair scoring. Results demonstrate that a single round of critique induces 34.5% mechanistic divergence, while varying definitions of equivalence rules cause the identified discrepancy rate to fluctuate dramatically between 38% and 96%, revealing the decisive impact of evaluation criteria selection on outcomes. Overall, this work provides an interpretable analytical framework for understanding and quantifying hypothesis diversity in multi-agent systems.

0 citationsRead paper

Contamination, Prior, or Evidence? Decomposing and Training Evidence Use in Whole-Slide Vision-Language Models

Sep 26, 2026

This study addresses the challenge in evaluating pathology vision-language models, where data contamination and prior knowledge are difficult to disentangle from genuine image evidence. To overcome this, we introduce CleanSlide, a benchmark that employs multiple-choice question-answering auditing to eliminate data leakage. Furthermore, we propose a Pair-DPO loss that leverages authentic counterfactual slide pairs for preference optimization, precisely decoupling non-image signals to strengthen the utilization of visual evidence. Experimental results demonstrate that our approach yields a 15.29% gain in image evidence reliance and improves accuracy on external CPTAC and BCNB cohorts by 9.7% and 3.4%, respectively. Ultimately, this work provides a reliable and unbiased evaluation and training paradigm for multimodal models in computational pathology.

0 citationsRead paper

RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction

Aug 06, 2026

This work addresses the challenge of reaction yield prediction, which is hindered by scarce labeled data, the vast and sparse reaction space, and the inability of existing representations to capture complex chemical transformations. To overcome these limitations, the authors propose RxnCLF, a self-supervised contrastive learning framework that introduces a novel condensed reaction graph (CRG) integrating both reactant and product information. By leveraging graph neural networks, RxnCLF learns explicit and interpretable transformation structures and models chemical reactions within a unified continuous latent space. The method significantly outperforms current graph- and sequence-based models on multiple yield prediction benchmarks, demonstrating substantial improvements in R² scores, and exhibits strong generalization capabilities on downstream tasks such as regioselectivity and enantioselectivity prediction, offering a new paradigm toward a general-purpose foundation model for chemical reactions.

0 citationsRead paper

Two-Stage Design with Sample Size Re-estimation Using gsDesign

Aug 04, 2026

This study addresses the challenges in early-phase clinical trials arising from uncertainty in effect size and nuisance parameters, which often lead to inaccurate sample size planning and over-enrollment in group sequential designs. The authors employ a two-stage group sequential framework that dynamically adjusts the sample size at interim analysis based on conditional power, and systematically compare—using the gsDesign R package—the expected sample size and statistical power of conditional power–based designs against more conservative group sequential approaches. Findings indicate that conditional power designs offer only marginal benefits under specific scenarios, whereas conservative designs generally demonstrate superior efficiency, robustness, and the added advantage of avoiding premature disclosure of interim treatment efficacy. These results provide empirical evidence and practical guidance for selecting appropriate group sequential designs in clinical trial settings.

0 citationsRead paper