Institution profile

Wellesley College

Academic institutionnorthamerica · us
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Latent Similarity Gaussian Processes: A Theory-Grounded Approach to Personalized Suicide-Risk Forecasting for Clinical Decision-Support

Oct 05, 2026

This study addresses the challenge that high patient heterogeneity and the low base rate of suicide events constrain the accuracy of personalized risk prediction. To overcome this, we propose a latent similarity Gaussian process that embeds patients into a continuous latent space to jointly model inter-patient similarity and individual risk trajectories. Specifically, an identifiable dual-channel similarity kernel is designed to selectively leverage peer information, while a correction mechanism is introduced to mitigate model degeneration induced by mean-field variational inference. Evaluated on dense longitudinal data, the proposed approach significantly outperforms population-level, individual-level, and hierarchical baselines. Notably, it substantially improves predictive precision for first-time suicide events, thereby establishing a novel paradigm for individualized risk assessment in clinical settings.

0 citationsRead paper

RINI: Seeing the Prior Is Not Enough

Sep 27, 2026

This study addresses the challenge of erroneously claimed novel contributions in research proposals, which are notoriously difficult to correct. To this end, we propose RINI, a method that integrates retrieval-augmented generation with an explicit contribution attribution mechanism. RINI achieves precise corrections through a three-step process: auditing contribution claims, examining residual distinctions, and performing multi-stage local revisions. Experimental results demonstrate that among cases requiring correction, RINI attains a successful repair rate of 72.2%, outperforming a direct revision baseline by 33 percentage points. These findings indicate that RINI effectively mitigates the problem of spurious novelty claims in academic proposals, offering a robust approach for ensuring accurate scholarly attribution.

0 citationsRead paper

Teaching Probabilistic Machine Learning in the Liberal Arts: Empowering Socially and Mathematically Informed AI Discourse

Oct 28, 2025

Underrepresented students—particularly those from humanities backgrounds and non-STEM disciplines—often face significant barriers to engaging with probabilistic machine learning (ML) due to high mathematical thresholds and perceived irrelevance to societal concerns. Method: This study introduces a “framework-focused” pedagogy for an undergraduate probability ML course, anchored in an interdisciplinary narrative—the fictional “Interstellar Hypothetical Hospital”—and integrating probabilistic programming to lower mathematical entry barriers. The curriculum employs open, real-world case studies and counter-narrative discussions to systematically connect foundational ML concepts (e.g., Bayesian modeling) with sociotechnical implications. Contribution/Results: The approach innovatively merges whimsical storytelling with rigorous probabilistic reasoning to enhance accessibility and engagement; uses ethical dilemmas as anchors for developing dialectical AI literacy; and concurrently cultivates modeling competence, critical thinking, and technocivic identity. Empirical evaluation demonstrates significant gains in students’ integrated understanding of ML principles and their social dimensions, alongside increased confidence and capacity to participate diversely in public AI discourse.

0 citationsRead paper

Creating Targeted, Interpretable Topic Models with LLM-Generated Text Augmentation

Apr 24, 2025

Traditional topic modeling methods suffer from poor interpretability and limited capacity to address domain-specific research questions in the social sciences. Method: This paper proposes a large language model (LLM)-based data augmentation framework that integrates controllable semantic text generation into unsupervised topic modeling. Leveraging GPT-4 to synthesize domain-relevant textual data, the approach couples generated corpora with LDA and BERTopic for guided, question-oriented topic discovery—requiring minimal human intervention. A political science–specific corpus and evaluation framework are constructed to support rigorous validation. Contribution/Results: Experiments demonstrate substantial improvements in topic interpretability and task relevance; the method enables direct answering of domain research questions and reduces manual annotation effort by over 70%. By bridging generative AI with social science–driven topic modeling, this work establishes a novel paradigm for theory-informed, question-centered thematic analysis.

0 citationsRead paper
Recent publications

Latest Papers

Latent Similarity Gaussian Processes: A Theory-Grounded Approach to Personalized Suicide-Risk Forecasting for Clinical Decision-Support

Oct 05, 2026

This study addresses the challenge that high patient heterogeneity and the low base rate of suicide events constrain the accuracy of personalized risk prediction. To overcome this, we propose a latent similarity Gaussian process that embeds patients into a continuous latent space to jointly model inter-patient similarity and individual risk trajectories. Specifically, an identifiable dual-channel similarity kernel is designed to selectively leverage peer information, while a correction mechanism is introduced to mitigate model degeneration induced by mean-field variational inference. Evaluated on dense longitudinal data, the proposed approach significantly outperforms population-level, individual-level, and hierarchical baselines. Notably, it substantially improves predictive precision for first-time suicide events, thereby establishing a novel paradigm for individualized risk assessment in clinical settings.

0 citationsRead paper

RINI: Seeing the Prior Is Not Enough

Sep 27, 2026

This study addresses the challenge of erroneously claimed novel contributions in research proposals, which are notoriously difficult to correct. To this end, we propose RINI, a method that integrates retrieval-augmented generation with an explicit contribution attribution mechanism. RINI achieves precise corrections through a three-step process: auditing contribution claims, examining residual distinctions, and performing multi-stage local revisions. Experimental results demonstrate that among cases requiring correction, RINI attains a successful repair rate of 72.2%, outperforming a direct revision baseline by 33 percentage points. These findings indicate that RINI effectively mitigates the problem of spurious novelty claims in academic proposals, offering a robust approach for ensuring accurate scholarly attribution.

0 citationsRead paper

Teaching Probabilistic Machine Learning in the Liberal Arts: Empowering Socially and Mathematically Informed AI Discourse

Oct 28, 2025

Underrepresented students—particularly those from humanities backgrounds and non-STEM disciplines—often face significant barriers to engaging with probabilistic machine learning (ML) due to high mathematical thresholds and perceived irrelevance to societal concerns. Method: This study introduces a “framework-focused” pedagogy for an undergraduate probability ML course, anchored in an interdisciplinary narrative—the fictional “Interstellar Hypothetical Hospital”—and integrating probabilistic programming to lower mathematical entry barriers. The curriculum employs open, real-world case studies and counter-narrative discussions to systematically connect foundational ML concepts (e.g., Bayesian modeling) with sociotechnical implications. Contribution/Results: The approach innovatively merges whimsical storytelling with rigorous probabilistic reasoning to enhance accessibility and engagement; uses ethical dilemmas as anchors for developing dialectical AI literacy; and concurrently cultivates modeling competence, critical thinking, and technocivic identity. Empirical evaluation demonstrates significant gains in students’ integrated understanding of ML principles and their social dimensions, alongside increased confidence and capacity to participate diversely in public AI discourse.

0 citationsRead paper

Creating Targeted, Interpretable Topic Models with LLM-Generated Text Augmentation

Apr 24, 2025

Traditional topic modeling methods suffer from poor interpretability and limited capacity to address domain-specific research questions in the social sciences. Method: This paper proposes a large language model (LLM)-based data augmentation framework that integrates controllable semantic text generation into unsupervised topic modeling. Leveraging GPT-4 to synthesize domain-relevant textual data, the approach couples generated corpora with LDA and BERTopic for guided, question-oriented topic discovery—requiring minimal human intervention. A political science–specific corpus and evaluation framework are constructed to support rigorous validation. Contribution/Results: Experiments demonstrate substantial improvements in topic interpretability and task relevance; the method enables direct answering of domain research questions and reduces manual annotation effort by over 70%. By bridging generative AI with social science–driven topic modeling, this work establishes a novel paradigm for theory-informed, question-centered thematic analysis.

0 citationsRead paper