Institution profile

University of Montana

Academic institutionnorthamerica · us
Official website
Research library11linked papers
Opportunities0open roles
Selected work

Representative Papers

Effective Resistance and Graph Neural Network Reliability in Tissue-Specific Interactomes

Oct 01, 2026

This study addresses the lack of theoretical grounding for assessing the reliability of graph neural networks (GNNs) in protein function prediction, specifically investigating whether tissue-specific network structures can provide node-level prediction confidence. By leveraging effective resistance to quantify GNN reliability within tissue-specific interactomes, this work reveals its degeneration into inverse node degree. Through null model testing and a selective prediction framework, it analyzes the explanatory power of residual signals—obtained after removing degree effects—on prediction errors. The findings demonstrate that these residual signals independently explain node-wise loss, with their explanatory variance increasing monotonically with network depth and reaching 5.6 times that of permutation baselines. This research establishes the first topology-based node-level confidence metric, uncovering both the degree-degeneration limitation of effective resistance and the robustness of residual signals.

0 citationsRead paper

TopoBudget: Persistent-Connectivity-Preserving Web Graph Sparsification for Reusable Community Analytics

Aug 08, 2026

Existing graph sparsification methods struggle to preserve multiscale connectivity structures under edge relevance filtering, often discarding topological evidence critical for community detection. This work proposes a persistence-aware sparsification approach that, given an edge relevance filtration and a proxy partition, selects a budget-constrained subgraph via constrained greedy submodular optimization to exactly maintain connected components across all filtration thresholds—thereby achieving, for the first time, complete preservation of the zeroth-dimensional persistent diagram with a theoretical $(1 - 1/e)$ approximation guarantee. The method integrates a persistent homology skeleton, a skeleton-constrained monotone submodular objective, and a degree-balanced recovery strategy. Experiments on six real-world web and social graphs demonstrate that TopoBudget achieves state-of-the-art community preservation under Louvain, competitive performance under Infomap, zero topological mismatch, and significantly faster runtime than effective-resistance baselines.

0 citationsRead paper

Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models

Aug 05, 2026

This work addresses the challenge of enabling small language models to decide when to defer to human control in privacy-sensitive, offline, and cost-constrained settings while maintaining bounded risk. The authors propose a linguistically grounded confidence-based deferral mechanism and analyze, both theoretically and empirically, how calibration methods influence the trade-off between risk and coverage. Key contributions include establishing three theoretical bounds, introducing a finite-sample risk certification method based on the Clopper–Pearson interval, and correcting an answer-ranking artifact in the multiple-choice formulation of TruthfulQA. Experimental results across 22 model–task pairs show that only three achieve certified autonomy under a 20% risk budget, with none satisfying a stricter 10% threshold; furthermore, Platt scaling reduces expected calibration error (ECE) to as low as 0.02.

0 citationsRead paper

Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index

Aug 05, 2026

This study addresses a critical limitation in current large language models (LLMs) employed as AI tutors: their frequent disregard for learners’ prior knowledge and instructional progression, coupled with a lack of effective evaluation of pedagogical suitability. To bridge this gap, the authors propose the Pedagogical Suitability Index (PSI), a theoretically grounded metric comprising six sub-dimensions that quantifies the alignment between LLM-generated tutoring responses and learners’ readiness as well as curricular pacing. For the first time, PSI is leveraged as a structured feedback signal to guide LLMs in refining their outputs. Experimental results demonstrate that 82.3% of 62 initially low-scoring cases showed significant improvement under PSI guidance, with human evaluators confirming the pedagogical validity of these enhancements—thereby transcending conventional evaluation paradigms that focus solely on answer correctness.

0 citationsRead paper

Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs

Jul 27, 2026

This study addresses the frequent violations of clinical coding standards—such as ICD-10, CPT, and HL7 FHIR—by large language models when generating structured medical data, which impedes integration with electronic health record systems. To mitigate this, the authors propose and validate a closed-loop verification-and-repair framework that automatically detects and iteratively corrects formatting errors. The approach is evaluated using three open-source models—Qwen2.5-7B, Llama3.1-8B, and Gemma2-9B—deployed locally across 320 clinical scenarios. Results demonstrate a substantial improvement in schema compliance across all models, achieving an overall adherence rate of 99.0% and increasing individual model performance by 7.8 to 12.5 percentage points. Notably, 96% of detected errors were attributable to repairable representation-layer issues, with most resolved within one or two correction rounds, effectively compensating for the models’ limited understanding of healthcare IT standards.

0 citationsRead paper
Recent publications

Latest Papers

Effective Resistance and Graph Neural Network Reliability in Tissue-Specific Interactomes

Oct 01, 2026

This study addresses the lack of theoretical grounding for assessing the reliability of graph neural networks (GNNs) in protein function prediction, specifically investigating whether tissue-specific network structures can provide node-level prediction confidence. By leveraging effective resistance to quantify GNN reliability within tissue-specific interactomes, this work reveals its degeneration into inverse node degree. Through null model testing and a selective prediction framework, it analyzes the explanatory power of residual signals—obtained after removing degree effects—on prediction errors. The findings demonstrate that these residual signals independently explain node-wise loss, with their explanatory variance increasing monotonically with network depth and reaching 5.6 times that of permutation baselines. This research establishes the first topology-based node-level confidence metric, uncovering both the degree-degeneration limitation of effective resistance and the robustness of residual signals.

0 citationsRead paper

TopoBudget: Persistent-Connectivity-Preserving Web Graph Sparsification for Reusable Community Analytics

Aug 08, 2026

Existing graph sparsification methods struggle to preserve multiscale connectivity structures under edge relevance filtering, often discarding topological evidence critical for community detection. This work proposes a persistence-aware sparsification approach that, given an edge relevance filtration and a proxy partition, selects a budget-constrained subgraph via constrained greedy submodular optimization to exactly maintain connected components across all filtration thresholds—thereby achieving, for the first time, complete preservation of the zeroth-dimensional persistent diagram with a theoretical $(1 - 1/e)$ approximation guarantee. The method integrates a persistent homology skeleton, a skeleton-constrained monotone submodular objective, and a degree-balanced recovery strategy. Experiments on six real-world web and social graphs demonstrate that TopoBudget achieves state-of-the-art community preservation under Louvain, competitive performance under Infomap, zero topological mismatch, and significantly faster runtime than effective-resistance baselines.

0 citationsRead paper

Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models

Aug 05, 2026

This work addresses the challenge of enabling small language models to decide when to defer to human control in privacy-sensitive, offline, and cost-constrained settings while maintaining bounded risk. The authors propose a linguistically grounded confidence-based deferral mechanism and analyze, both theoretically and empirically, how calibration methods influence the trade-off between risk and coverage. Key contributions include establishing three theoretical bounds, introducing a finite-sample risk certification method based on the Clopper–Pearson interval, and correcting an answer-ranking artifact in the multiple-choice formulation of TruthfulQA. Experimental results across 22 model–task pairs show that only three achieve certified autonomy under a 20% risk budget, with none satisfying a stricter 10% threshold; furthermore, Platt scaling reduces expected calibration error (ECE) to as low as 0.02.

0 citationsRead paper

Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index

Aug 05, 2026

This study addresses a critical limitation in current large language models (LLMs) employed as AI tutors: their frequent disregard for learners’ prior knowledge and instructional progression, coupled with a lack of effective evaluation of pedagogical suitability. To bridge this gap, the authors propose the Pedagogical Suitability Index (PSI), a theoretically grounded metric comprising six sub-dimensions that quantifies the alignment between LLM-generated tutoring responses and learners’ readiness as well as curricular pacing. For the first time, PSI is leveraged as a structured feedback signal to guide LLMs in refining their outputs. Experimental results demonstrate that 82.3% of 62 initially low-scoring cases showed significant improvement under PSI guidance, with human evaluators confirming the pedagogical validity of these enhancements—thereby transcending conventional evaluation paradigms that focus solely on answer correctness.

0 citationsRead paper

Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs

Jul 27, 2026

This study addresses the frequent violations of clinical coding standards—such as ICD-10, CPT, and HL7 FHIR—by large language models when generating structured medical data, which impedes integration with electronic health record systems. To mitigate this, the authors propose and validate a closed-loop verification-and-repair framework that automatically detects and iteratively corrects formatting errors. The approach is evaluated using three open-source models—Qwen2.5-7B, Llama3.1-8B, and Gemma2-9B—deployed locally across 320 clinical scenarios. Results demonstrate a substantial improvement in schema compliance across all models, achieving an overall adherence rate of 99.0% and increasing individual model performance by 7.8 to 12.5 percentage points. Notably, 96% of detected errors were attributable to repairable representation-layer issues, with most resolved within one or two correction rounds, effectively compensating for the models’ limited understanding of healthcare IT standards.

0 citationsRead paper