Score
Designs and implements scoring methods and evaluation metrics that quantify coherence, fidelity, and overall quality of sequences whose elements are interleaved types or categories; creates analytic tools and benchmarks to compare, rank, and aggregate model outputs on interleaved sequence tasks.
Current evaluations of AI models lack standardized protocols, with institutions selectively employing benchmarks in ways that hinder cross-study comparability and raise concerns about scientific validity. This work introduces Benchmarking-Cultures-25, a dataset encompassing 231 benchmarks from 139 model releases, and combines qualitative content analysis with a unified categorization framework to systematically expose the fragmentation in benchmark selection: 63.2% of benchmarks are used by only a single institution, and 38.5% appear just once. Moreover, many benchmarks marketed as “general-purpose” disproportionately emphasize STEM—particularly mathematics—while often neglecting construct validity. The study further proposes a taxonomy aligning ostensibly disparate terminologies to their underlying measurement signals and develops an interactive tool revealing that benchmarks frequently serve marketing narratives rather than rigorous scientific assessment.
This study identifies systematic inconsistencies in the implementation of machine learning evaluation metrics across mainstream programming languages—Python, R, and MATLAB—spanning ten task categories: classification, regression, clustering, statistical testing, image segmentation, and image-to-image translation. Through the first large-scale, cross-platform empirical analysis, we quantitatively assess consistency across 100+ metrics. Results reveal that 36 metrics—including Accuracy, AUC, and MAE—are robust across implementations, whereas critical metrics such as Precision, F1-score, IoU, and Within-Cluster Sum of Squares (WCSS) exhibit substantial discrepancies. To address this, we propose the first comprehensive, task-agnostic standardization roadmap for ML evaluation, accompanied by a curated recommendation list. This work provides both theoretical foundations and practical guidelines to enhance cross-platform reproducibility and result reliability in ML research and deployment.
To address the misalignment between automatic evaluation metrics and human preferences in generative tasks, this paper proposes MetaMetrics—a calibratable meta-metric that supervisely weights and fuses existing metrics to model fine-grained human preferences across multimodal (language/vision), multilingual, and multi-domain settings. Methodologically, it introduces the first preference-dimension-aware metric calibration framework, enabling cross-modal unified evaluation and plug-and-play integration. The approach combines supervised meta-learning, multi-task joint optimization, and explicit modeling of human preference annotations. Experiments demonstrate that MetaMetrics significantly improves correlation with human judgments across multilingual text and vision generation tasks (average Kendall’s τ increase of +18.7%). Moreover, it exhibits strong generalization to unseen domains and models, maintaining robust alignment with human preferences without task-specific retraining.
Current research on reward modeling and evaluation metrics operates in silos, leading to terminological redundancy, spurious correlations, heightened reward hacking risks, and duplicated efforts in data quality optimization and meta-evaluation. Through a systematic literature review and comparative analysis, we reveal that both reward models and evaluation metrics fundamentally serve the same purpose in language model post-training: preference modeling and performance calibration. Building on this insight, we propose a unified research framework integrating three core directions—preference acquisition, spurious correlation mitigation, and meta-evaluation calibration. Empirical experiments demonstrate that certain evaluation metrics significantly outperform existing reward models on specific tasks. This work clarifies the root causes of conceptual ambiguity in the field and fosters cross-paradigm collaboration, providing both theoretical foundations and practical pathways for developing robust, interpretable, and reusable alignment evaluation systems.
Existing evaluation methods for LLM-generated code comments rely on small-scale datasets and inadequate IR metrics (e.g., BLEU), failing to capture semantic fidelity. Method: We systematically assess GPT-3.5’s Javadoc generation for 23,850 Java code snippets, employing a dual-dimensional evaluation combining quantitative BLEU scoring with qualitative expert human assessment. Contribution/Results: Our study reveals a critical flaw in BLEU: high scores frequently correlate with low-quality, verbatim descriptions, while high-fidelity semantic paraphrasing is systematically penalized. We find that 69.7% of generated Javadocs are semantically equivalent to—or can be refined to match—the original quality, and 22.4% significantly surpass the originals. These results demonstrate that automated metrics alone are unreliable for assessing documentation quality. We advocate human evaluation as the gold standard, with BLEU serving only as a supplementary heuristic—establishing a new, more rigorous paradigm for evaluating code documentation generation.
本文提出一种生命周期感知的框架,结合软件质量评估与大语言模型代码优化,以解决科研软件因开发者缺乏软件工程经验导致的质量问题。
This study addresses the inconsistency in model rankings caused by commonly used ranking metrics—such as MRR, Hits@k, and Mean Rank—in knowledge graph completion (KGC) evaluation, which hinders fair comparison and reproducibility. For the first time, KGC evaluation is framed as a multi-criteria decision-making problem, and seven aggregators are systematically assessed across five dimensions: consistency, cross-dataset stability, metric independence, noise robustness, and generalization capability. Through leave-one-model-out (LOMO) and leave-one-group-out (LOGO) cross-validation, Pareto optimality analysis, and multidimensional sensitivity tests, the Z-score aggregator emerges as the most balanced overall—favoring DualE for tail entity prediction and FMS for relation prediction. The experiments further reveal that consistency and stability are insensitive to removal strategies, whereas generalization and independence exhibit the highest sensitivity.
Existing benchmarks for knowledge work evaluation largely adhere to traditional NLP task paradigms, failing to capture systems’ capabilities in real-world knowledge-intensive settings. This work proposes a three-step framework—explicitly defining work activities, establishing realistic test environments, and focusing evaluation on deliverable outputs—and derives 18 core knowledge work activities from the O*NET database. Innovatively integrating role responsibilities, local tool usage, and downstream usability into benchmark design, the approach establishes a coherent “work activity–test setup–scoring artifact” alignment. Validation through three case studies (GDPval, OfficeQA Pro, and APEX-SWE) exposes critical misalignments in current benchmarks between tasks, environments, and actual work objectives, offering a new paradigm for evaluating knowledge work systems in practical, application-oriented contexts.
Alignment evaluation in machine learning has largely become evaluation of models. Influential benchmarks score model outputs under fixed inputs, such as truthfulness, instruction following, or pairwise preference, and these scores are often used to support claims about deployed alignment. This paper argues that deployment-relevant alignment cannot be inferred from model-level evaluation alone. Alignment claims should instead be indexed to the level at which evidence is collected: model-level, response-level, interaction-level, or deployment-level. Two studies support this position. First, a structured audit of eleven alignment benchmarks, extended to a sixteen-benchmark corpus, dual-coded against an eight-dimension rubric with Cohen's kappa = 0.87, finds that user-facing verification support is absent across every benchmark examined, while process steerability is nearly absent. The few interactional benchmarks identified, including tau-bench, CURATe, Rifts, and Common Ground, remain fragmented in coverage, and benchmark construction rather than data source determines what is measured. Second, a blinded cross-model stress test using 180 transcripts across three frontier models and four scaffolds finds that the same verification scaffold raises one model's verification support to ceiling while leaving another categorically unchanged. This shows that scaffold efficacy is model-dependent and that the gap identified by the audit cannot be closed at the model level alone. We propose a system-level evaluation agenda: alignment profiles instead of single scores, fixed-scaffolding protocols for comparable interactional evaluation, and reporting templates that make the inferential distance between evaluation evidence and deployment claims explicit.
论文针对代理基准测试中的双重测量混淆问题,通过将关键决策转移给模型、使用基于真实值的评分及报告更全面的可靠性指标来解决。