evaluate model accuracy

Designs and implements procedures and analysis pipelines to quantify predictive performance of models by comparing outputs to ground truth across datasets, groups, and spatial partitions. Computes and reports accuracy and error metrics (accuracy, balanced accuracy, worst-group accuracy, RMSE and per-pair RMSE), aggregates challenge/leaderboard scores, estimates confidence intervals, and assesses trade-offs (accuracy–memory–speed), robustness under masking, and geospatial-specific accuracy measures.

evaluatemodelaccuracy

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.

calibrationdistributional validationprobabilistic forecasting

This work addresses the lack of a universal, tunable, and multi-scenario-compatible metric for data quality assessment, which hinders effective comparison of diverse data cleaning pipelines. To overcome this limitation, the authors propose TOMME—a general-purpose data quality measurement framework based on weighted errors—that extends traditional accuracy into a configurable, composite score. By producing a single quantitative metric, TOMME enables flexible adjustment of error weights according to specific use cases, thereby supporting both automated processing and optimization requirements. Experimental results demonstrate that TOMME exhibits strong adaptability, practicality, and comparability across a variety of scenarios, offering an efficient and unified solution for data quality evaluation and decision-making.

accuracydata cleaningdata quality

Traditional cross-validation in spatial prediction suffers from biased risk estimation due to distributional mismatches between validation and deployment tasks, including covariate shift and task difficulty shift. This work proposes Target-Weighted Cross-Validation (TWCV), which for the first time incorporates task distribution alignment into spatial prediction by calibrating weights to match the target-domain distribution and enhancing task difficulty diversity through spatial buffering-based resampling. By integrating importance-weighted risk estimation with task descriptor modeling, TWCV substantially reduces estimation bias. Empirical evaluations on both synthetic data and real-world environmental pollution mapping demonstrate that the method yields more accurate and nearly unbiased estimates of deployment risk compared to existing spatial cross-validation approaches.

covariate shiftcross-validationdistribution shift

Communication barriers between data scientists and domain experts arise from oversimplified, accuracy-centric model performance reporting, hindering shared understanding of model limitations and contextual applicability. Method: We propose a visualization-mediated model explanation framework grounded in human-computer interaction principles, participatory design, and visual narrative techniques. This yields the first domain-expert-oriented model communication guideline—emphasizing risk, trade-offs, and situational appropriateness rather than isolated metrics like accuracy. An iterative empirical study was conducted using regression models, incorporating structured expert feedback for evaluation. Contribution/Results: The framework significantly improves domain experts’ ability to identify model limitations, recognize inherent trade-offs, and proactively make context-driven adoption decisions. Its core innovation lies in repositioning visualization as an interdisciplinary consensus-building medium—shifting the paradigm from “metric reporting” to “collaborative understanding.”

Communication gaps between data scientists and subject matter experts hinder model understanding.Traditional metrics fail to convey model risks, strengths, and limitations effectively.Visualization guidelines improve model performance communication and decision-making confidence.

Consistent Validation for Predictive Methods in Spatial Settings

Feb 05, 2024
DR
David R. Burt
🏛️ Massachusetts Institute for Technology

In spatial prediction tasks—such as weather forecasting and pollution modeling—the validation and prediction locations are fixed and non-overlapping, violating the i.i.d. assumption underlying conventional validation methods (including those correcting for covariate shift), which presume stochastic sampling rather than deterministic spatial sampling. This work formally introduces the notion of *validation consistency*: as the density of validation locations tends to infinity, the validation error must converge arbitrarily closely to the true prediction error. Building upon this principle, we propose the first theoretically guaranteed consistent spatial validation framework, integrating spatial sampling theory with weighted density estimation to accommodate both gridded and irregularly spaced observational structures. We prove its consistency under mild regularity conditions. Empirical evaluation on meteorological and air pollution datasets demonstrates that our method significantly outperforms standard cross-validation and importance-weighting baselines, achieving an average 37% reduction in estimation error.

Addressing failure of classical methods in dense validationProposing adaptive validation for fixed-location spatial dataValidating spatial predictions with mismatched location data

Latest Papers

What's happening recently
View more

This study addresses the error propagation problem in remote sensing agents caused by uncertainty in tool observations. To mitigate this issue, we propose a reliability assessment framework grounded in process evidence verification. The framework integrates Bayesian inference-based reliability modeling with task-tool priors constructed from offline feedback, enabling dynamic evaluation of observation quality through tool invocation monitoring and evidence chain analysis. Based on these assessments, it adaptively executes accept, supplement, or reject decisions. Experimental results demonstrate that the proposed approach significantly suppresses error propagation, effectively improves accuracy across multiple benchmark tasks, and reduces redundant tool invocations.

Error PropagationRemote Sensing AgentsTool Observations

This work addresses the lack of reliable evaluation for large language model (LLM) agents in real-world geographic information system (GIS) multi-step workflow tasks, primarily due to the absence of benchmark datasets with precise ground-truth annotations. To bridge this gap, the authors introduce the first GIS agent benchmark, comprising 349 multi-step tasks derived from GIS Stack Exchange and grounded in publicly available data from six real-world regions. Each task is accompanied by an executable reference trajectory and tolerance-aware ground-truth outputs. This benchmark enables, for the first time, rigorous and reproducible evaluation of LLM agents capable of external tool invocation, moving beyond indirect metrics such as code similarity or model-based judgments. Experimental results reveal that even the best-performing among six state-of-the-art LLM agents achieves only a 32.7% task completion rate under strict scoring, underscoring the significant challenges in automating realistic GIS workflows.

benchmarkGISground truth

This study addresses the limitation that high accuracy on short program outputs often obscures deficiencies in intermediate state tracking during large language model evaluation. To this end, it extends the CRUXEval paradigm by constructing a 400-case benchmark featuring paired short and long execution trajectories alongside multi-dimensional checkpoint tasks. Leveraging Python and C++ static analysis, the authors conduct comparative evaluations across multiple models without requiring code execution environments. The results reveal blind spots in state prediction that conventional single-metric evaluations fail to capture. Under the strongest configuration, accuracy reaches 93.0% for short trajectories but drops to 77.0% for long ones, while reasoning models outperform non-reasoning counterparts by over 33 percentage points, underscoring the persistent challenges of complex state prediction.

benchmark evaluationcheckpoint stateoutput prediction

Hot Scholars

SH

Shaoyi Huang

Assistant Professor, Stevens Institute of Technology
Deep LearningEfficient AISoftware hardware co-design
XA

Xuguang Ai

Biomedical Informatics & Data Science, Yale University
AI in HealthcareData ScienceNLPBiomedical Informatics
HK

Hyunjae Kim

Yale University
Natural Language ProcessingBiomedical InformaticsHealthcare
DL

Dianbo Liu

Assistant professor, National University of Singapore
Push the limits of humanmachine learningbiomedical sciences
QC

Qingyu Chen

Biomedical Informatics & Data Science, Yale University; NCBI-NLM, National Institutes of Health
Text miningMachine learningData curationBioNLP