Score
Designs and implements procedures and analysis pipelines to quantify predictive performance of models by comparing outputs to ground truth across datasets, groups, and spatial partitions. Computes and reports accuracy and error metrics (accuracy, balanced accuracy, worst-group accuracy, RMSE and per-pair RMSE), aggregates challenge/leaderboard scores, estimates confidence intervals, and assesses trade-offs (accuracy–memory–speed), robustness under masking, and geospatial-specific accuracy measures.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
This work addresses the lack of a universal, tunable, and multi-scenario-compatible metric for data quality assessment, which hinders effective comparison of diverse data cleaning pipelines. To overcome this limitation, the authors propose TOMME—a general-purpose data quality measurement framework based on weighted errors—that extends traditional accuracy into a configurable, composite score. By producing a single quantitative metric, TOMME enables flexible adjustment of error weights according to specific use cases, thereby supporting both automated processing and optimization requirements. Experimental results demonstrate that TOMME exhibits strong adaptability, practicality, and comparability across a variety of scenarios, offering an efficient and unified solution for data quality evaluation and decision-making.
Traditional cross-validation in spatial prediction suffers from biased risk estimation due to distributional mismatches between validation and deployment tasks, including covariate shift and task difficulty shift. This work proposes Target-Weighted Cross-Validation (TWCV), which for the first time incorporates task distribution alignment into spatial prediction by calibrating weights to match the target-domain distribution and enhancing task difficulty diversity through spatial buffering-based resampling. By integrating importance-weighted risk estimation with task descriptor modeling, TWCV substantially reduces estimation bias. Empirical evaluations on both synthetic data and real-world environmental pollution mapping demonstrate that the method yields more accurate and nearly unbiased estimates of deployment risk compared to existing spatial cross-validation approaches.
Communication barriers between data scientists and domain experts arise from oversimplified, accuracy-centric model performance reporting, hindering shared understanding of model limitations and contextual applicability. Method: We propose a visualization-mediated model explanation framework grounded in human-computer interaction principles, participatory design, and visual narrative techniques. This yields the first domain-expert-oriented model communication guideline—emphasizing risk, trade-offs, and situational appropriateness rather than isolated metrics like accuracy. An iterative empirical study was conducted using regression models, incorporating structured expert feedback for evaluation. Contribution/Results: The framework significantly improves domain experts’ ability to identify model limitations, recognize inherent trade-offs, and proactively make context-driven adoption decisions. Its core innovation lies in repositioning visualization as an interdisciplinary consensus-building medium—shifting the paradigm from “metric reporting” to “collaborative understanding.”
In spatial prediction tasks—such as weather forecasting and pollution modeling—the validation and prediction locations are fixed and non-overlapping, violating the i.i.d. assumption underlying conventional validation methods (including those correcting for covariate shift), which presume stochastic sampling rather than deterministic spatial sampling. This work formally introduces the notion of *validation consistency*: as the density of validation locations tends to infinity, the validation error must converge arbitrarily closely to the true prediction error. Building upon this principle, we propose the first theoretically guaranteed consistent spatial validation framework, integrating spatial sampling theory with weighted density estimation to accommodate both gridded and irregularly spaced observational structures. We prove its consistency under mild regularity conditions. Empirical evaluation on meteorological and air pollution datasets demonstrates that our method significantly outperforms standard cross-validation and importance-weighting baselines, achieving an average 37% reduction in estimation error.
This study addresses the error propagation problem in remote sensing agents caused by uncertainty in tool observations. To mitigate this issue, we propose a reliability assessment framework grounded in process evidence verification. The framework integrates Bayesian inference-based reliability modeling with task-tool priors constructed from offline feedback, enabling dynamic evaluation of observation quality through tool invocation monitoring and evidence chain analysis. Based on these assessments, it adaptively executes accept, supplement, or reject decisions. Experimental results demonstrate that the proposed approach significantly suppresses error propagation, effectively improves accuracy across multiple benchmark tasks, and reduces redundant tool invocations.
该研究通过多种方法如配置比较、输入压力测试等,解决了表面水分割模型排名稳定性及输入依赖性评估问题。
This work addresses the lack of reliable evaluation for large language model (LLM) agents in real-world geographic information system (GIS) multi-step workflow tasks, primarily due to the absence of benchmark datasets with precise ground-truth annotations. To bridge this gap, the authors introduce the first GIS agent benchmark, comprising 349 multi-step tasks derived from GIS Stack Exchange and grounded in publicly available data from six real-world regions. Each task is accompanied by an executable reference trajectory and tolerance-aware ground-truth outputs. This benchmark enables, for the first time, rigorous and reproducible evaluation of LLM agents capable of external tool invocation, moving beyond indirect metrics such as code similarity or model-based judgments. Experimental results reveal that even the best-performing among six state-of-the-art LLM agents achieves only a 32.7% task completion rate under strict scoring, underscoring the significant challenges in automating realistic GIS workflows.
This study addresses the limitation that high accuracy on short program outputs often obscures deficiencies in intermediate state tracking during large language model evaluation. To this end, it extends the CRUXEval paradigm by constructing a 400-case benchmark featuring paired short and long execution trajectories alongside multi-dimensional checkpoint tasks. Leveraging Python and C++ static analysis, the authors conduct comparative evaluations across multiple models without requiring code execution environments. The results reveal blind spots in state prediction that conventional single-metric evaluations fail to capture. Under the strongest configuration, accuracy reaches 93.0% for short trajectories but drops to 77.0% for long ones, while reasoning models outperform non-reasoning counterparts by over 33 percentage points, underscoring the persistent challenges of complex state prediction.
研究通过审计Praxa AI管道文件及记录,发现代理评估指标与标签含义不一致的问题,并提供了一个可重用的验证包来区分不同类型的性能声明。