Institution profile

Lafayette College

Academic institutionnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Align Before You Combine: Reference Space Calibration for Supervision Without Ground Truth

Oct 07, 2026

This study addresses the challenge of multi-source evaluation in the absence of ground-truth labels and a shared annotation space, where incomparable output scales across scorers hinder the construction of effective supervision signals. To overcome this, we propose a calibration-first framework that synthesizes a universal ordinal reference space via ordered calibration features, aligning subset-specific scorers to a unified scale. This approach is further augmented by a low-resolution calibration approximation technique, enabling supervision score fusion independent of training distributions. Evaluated on three benchmark datasets, the proposed method significantly outperforms both uncalibrated averaging and the best individual scorer while substantially reducing computational costs, thereby demonstrating its effectiveness and generalizability in ground-truth-free scenarios.

0 citationsRead paper

"Where Can I Trust You?": Boundary-Aware Evaluation of Surrogate Fidelity

Sep 26, 2026

This study addresses the limitation of existing surrogate model evaluations, which rely on global metrics that overlook local consistency structures and fail to capture distortions near decision boundaries or teacher model boundary instability. We propose a boundary-aware conditional evaluation framework that defines confidence regions and integrates two boundary identification techniques, conducting systematic empirical analyses across multiple datasets and surrogate models. Our findings reveal that surrogate fidelity degrades significantly near teacher decision boundaries; moreover, even when global and boundary fidelities are comparable, different surrogates exhibit markedly divergent responses to perturbations. This work demonstrates the superiority of conditional evaluation over global scoring, emphasizing the need to account for where consistency is distributed and how stable teacher boundaries remain across training runs, thereby offering a new paradigm for enhancing predictive robustness.

0 citationsRead paper

Fidelity Isn't Accuracy: When Linearly Decodable Functions Fail to Match the Ground Truth

Jun 13, 2025

This work investigates whether regression neural networks learn functions that are well-approximated by linear models, and whether linear decodability—i.e., the extent to which network outputs can be recovered via linear projection—is equivalent to predictive accuracy. To this end, we propose λ(f), a formal metric quantifying linear decodability as the R² between the neural network’s predictions and those of its optimal linear proxy. Through systematic experiments on synthetic data (x sin(x) + ε) and real-world regression benchmarks (e.g., Medical Insurance), we empirically establish—and for the first time formally define—that high λ(f) does not imply low true prediction error; thus, linear decodability ≠ high predictive accuracy. This challenges the implicit assumption that linear surrogates suffice for explaining nonlinear regression models, exposing critical limitations of such explanations in high-stakes settings. λ(f) provides a reproducible, quantitative, and comparable framework for assessing intrinsic linear structure within learned regression functions.

0 citationsRead paper
Recent publications

Latest Papers

Align Before You Combine: Reference Space Calibration for Supervision Without Ground Truth

Oct 07, 2026

This study addresses the challenge of multi-source evaluation in the absence of ground-truth labels and a shared annotation space, where incomparable output scales across scorers hinder the construction of effective supervision signals. To overcome this, we propose a calibration-first framework that synthesizes a universal ordinal reference space via ordered calibration features, aligning subset-specific scorers to a unified scale. This approach is further augmented by a low-resolution calibration approximation technique, enabling supervision score fusion independent of training distributions. Evaluated on three benchmark datasets, the proposed method significantly outperforms both uncalibrated averaging and the best individual scorer while substantially reducing computational costs, thereby demonstrating its effectiveness and generalizability in ground-truth-free scenarios.

0 citationsRead paper

"Where Can I Trust You?": Boundary-Aware Evaluation of Surrogate Fidelity

Sep 26, 2026

This study addresses the limitation of existing surrogate model evaluations, which rely on global metrics that overlook local consistency structures and fail to capture distortions near decision boundaries or teacher model boundary instability. We propose a boundary-aware conditional evaluation framework that defines confidence regions and integrates two boundary identification techniques, conducting systematic empirical analyses across multiple datasets and surrogate models. Our findings reveal that surrogate fidelity degrades significantly near teacher decision boundaries; moreover, even when global and boundary fidelities are comparable, different surrogates exhibit markedly divergent responses to perturbations. This work demonstrates the superiority of conditional evaluation over global scoring, emphasizing the need to account for where consistency is distributed and how stable teacher boundaries remain across training runs, thereby offering a new paradigm for enhancing predictive robustness.

0 citationsRead paper

Fidelity Isn't Accuracy: When Linearly Decodable Functions Fail to Match the Ground Truth

Jun 13, 2025

This work investigates whether regression neural networks learn functions that are well-approximated by linear models, and whether linear decodability—i.e., the extent to which network outputs can be recovered via linear projection—is equivalent to predictive accuracy. To this end, we propose λ(f), a formal metric quantifying linear decodability as the R² between the neural network’s predictions and those of its optimal linear proxy. Through systematic experiments on synthetic data (x sin(x) + ε) and real-world regression benchmarks (e.g., Medical Insurance), we empirically establish—and for the first time formally define—that high λ(f) does not imply low true prediction error; thus, linear decodability ≠ high predictive accuracy. This challenges the implicit assumption that linear surrogates suffice for explaining nonlinear regression models, exposing critical limitations of such explanations in high-stakes settings. λ(f) provides a reproducible, quantitative, and comparable framework for assessing intrinsic linear structure within learned regression functions.

0 citationsRead paper