Institution profile

Bill & Melinda Gates Foundation

Academic institutionnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Sparse probes and murky physics: a case study of interpretability challenges in a foundation model for continuum dynamics

Jun 10, 2026

This study investigates whether the internal mechanisms of the scientific foundation model Walrus align with physical principles when reproducing continuum dynamics, and examines the relationship between its representations and performance. By introducing sparse autoencoders (SAEs) at specific layers, the work pioneers the use of enstrophy—the integral of squared vorticity—for physically grounded filtering and prioritization of large-scale features, complemented by comparative numerical simulations. The findings reveal that while the model’s feature activations exhibit segment-wise consistency, they do not correspond to physically meaningful decompositions. Notably, certain output inaccuracies, such as excessive energy dissipation, can be traced to variations in specific SAE features. The study underscores fundamental challenges in achieving representational fidelity and interpretability in scientific foundation models.

0 citationsRead paper

What a diff makes: automating code migration with large language models

Oct 31, 2025

Semantic version upgrades of software dependency libraries frequently break backward compatibility, necessitating automated code migration solutions. This paper proposes AIMigrate, an LLM-based migration method that innovatively incorporates version-diff information as critical contextual input to the LLM, substantially improving migration accuracy. To support this work, we construct and publicly release the first benchmark dataset specifically designed for code migration tasks, along with the end-to-end tool AIMigrate. Experimental results on real-world migration scenarios show that AIMigrate identifies 65% of necessary changes in a single inference pass and achieves 80% coverage under multiple sampling; among the generated changes, 47% are fully correct. Compared to baseline approaches using source code alone, the diff-augmented strategy consistently outperforms across multiple evaluation metrics.

0 citationsRead paper

A Graph-Based Test-Harness for LLM Evaluation

Aug 28, 2025

This work addresses key limitations of existing LLM evaluation benchmarks for clinical practice guidelines—namely narrow coverage, static design, and heavy reliance on manual annotation—by proposing the first dynamic, systematic framework for assessing LLMs’ guideline-following capabilities. Methodologically, we model the WHO IMCI manual as a directed knowledge graph and automatically generate age-specific, clinically grounded multiple-choice questions via graph traversal, yielding >400 questions and 3.3 trillion answer combinations; contextually plausible distractors further increase difficulty. Our contributions are threefold: (1) enabling automatic benchmark expansion driven by guideline updates; (2) uncovering critical performance gaps—particularly in severity triage, treatment decision-making, and follow-up recommendation—with symptom identification accuracy ranging only from 45% to 67%, and other tasks substantially lower; and (3) generating high-reward, annotation-free samples that effectively support supervised fine-tuning, GRPO, and DPO, ensuring strong scalability and resistance to data contamination.

0 citationsRead paper
Recent publications

Latest Papers

Sparse probes and murky physics: a case study of interpretability challenges in a foundation model for continuum dynamics

Jun 10, 2026

This study investigates whether the internal mechanisms of the scientific foundation model Walrus align with physical principles when reproducing continuum dynamics, and examines the relationship between its representations and performance. By introducing sparse autoencoders (SAEs) at specific layers, the work pioneers the use of enstrophy—the integral of squared vorticity—for physically grounded filtering and prioritization of large-scale features, complemented by comparative numerical simulations. The findings reveal that while the model’s feature activations exhibit segment-wise consistency, they do not correspond to physically meaningful decompositions. Notably, certain output inaccuracies, such as excessive energy dissipation, can be traced to variations in specific SAE features. The study underscores fundamental challenges in achieving representational fidelity and interpretability in scientific foundation models.

0 citationsRead paper

What a diff makes: automating code migration with large language models

Oct 31, 2025

Semantic version upgrades of software dependency libraries frequently break backward compatibility, necessitating automated code migration solutions. This paper proposes AIMigrate, an LLM-based migration method that innovatively incorporates version-diff information as critical contextual input to the LLM, substantially improving migration accuracy. To support this work, we construct and publicly release the first benchmark dataset specifically designed for code migration tasks, along with the end-to-end tool AIMigrate. Experimental results on real-world migration scenarios show that AIMigrate identifies 65% of necessary changes in a single inference pass and achieves 80% coverage under multiple sampling; among the generated changes, 47% are fully correct. Compared to baseline approaches using source code alone, the diff-augmented strategy consistently outperforms across multiple evaluation metrics.

0 citationsRead paper

A Graph-Based Test-Harness for LLM Evaluation

Aug 28, 2025

This work addresses key limitations of existing LLM evaluation benchmarks for clinical practice guidelines—namely narrow coverage, static design, and heavy reliance on manual annotation—by proposing the first dynamic, systematic framework for assessing LLMs’ guideline-following capabilities. Methodologically, we model the WHO IMCI manual as a directed knowledge graph and automatically generate age-specific, clinically grounded multiple-choice questions via graph traversal, yielding >400 questions and 3.3 trillion answer combinations; contextually plausible distractors further increase difficulty. Our contributions are threefold: (1) enabling automatic benchmark expansion driven by guideline updates; (2) uncovering critical performance gaps—particularly in severity triage, treatment decision-making, and follow-up recommendation—with symptom identification accuracy ranging only from 45% to 67%, and other tasks substantially lower; and (3) generating high-reward, annotation-free samples that effectively support supervised fine-tuning, GRPO, and DPO, ensuring strong scalability and resistance to data contamination.

0 citationsRead paper