interleaved sequence scoring

Designs and implements scoring methods and evaluation metrics that quantify coherence, fidelity, and overall quality of sequences whose elements are interleaved types or categories; creates analytic tools and benchmarks to compare, rank, and aggregate model outputs on interleaved sequence tasks.

interleavedsequencescoring

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current evaluations of AI models lack standardized protocols, with institutions selectively employing benchmarks in ways that hinder cross-study comparability and raise concerns about scientific validity. This work introduces Benchmarking-Cultures-25, a dataset encompassing 231 benchmarks from 139 model releases, and combines qualitative content analysis with a unified categorization framework to systematically expose the fragmentation in benchmark selection: 63.2% of benchmarks are used by only a single institution, and 38.5% appear just once. Moreover, many benchmarks marketed as “general-purpose” disproportionately emphasize STEM—particularly mathematics—while often neglecting construct validity. The study further proposes a taxonomy aligning ostensibly disparate terminologies to their underlying measurement signals and develops an interactive tool revealing that benchmarks frequently serve marketing narratives rather than rigorous scientific assessment.

AI evaluationbenchmarkingconstruct validity

Machine Learning Evaluation Metric Discrepancies across Programming Languages and Their Components: Need for Standardization

Nov 18, 2024
MR
Mohammad R. Salmanpour
🏛️ University of British Columbia | University of Isfahan | University of Tehran | Shiraz University | TECVICO CORP.

This study identifies systematic inconsistencies in the implementation of machine learning evaluation metrics across mainstream programming languages—Python, R, and MATLAB—spanning ten task categories: classification, regression, clustering, statistical testing, image segmentation, and image-to-image translation. Through the first large-scale, cross-platform empirical analysis, we quantitatively assess consistency across 100+ metrics. Results reveal that 36 metrics—including Accuracy, AUC, and MAE—are robust across implementations, whereas critical metrics such as Precision, F1-score, IoU, and Within-Cluster Sum of Squares (WCSS) exhibit substantial discrepancies. To address this, we propose the first comprehensive, task-agnostic standardization roadmap for ML evaluation, accompanied by a curated recommendation list. This work provides both theoretical foundations and practical guidelines to enhance cross-platform reproducibility and result reliability in ML research and deployment.

Advocates for standardization to ensure reliable ML evaluations.Evaluates discrepancies in ML metrics across Python, R, and Matlab.Highlights inconsistencies in metrics for classification, regression, and clustering.

MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences

Oct 03, 2024
GI
Genta Indra Winata
🏛️ Capital One | University of Toronto | Monash University Indonesia | Boston University

To address the misalignment between automatic evaluation metrics and human preferences in generative tasks, this paper proposes MetaMetrics—a calibratable meta-metric that supervisely weights and fuses existing metrics to model fine-grained human preferences across multimodal (language/vision), multilingual, and multi-domain settings. Methodologically, it introduces the first preference-dimension-aware metric calibration framework, enabling cross-modal unified evaluation and plug-and-play integration. The approach combines supervised meta-learning, multi-task joint optimization, and explicit modeling of human preference annotations. Experiments demonstrate that MetaMetrics significantly improves correlation with human judgments across multilingual text and vision generation tasks (average Kendall’s τ increase of +18.7%). Moreover, it exhibits strong generalization to unseen domains and models, maintaining robust alignment with human preferences without task-specific retraining.

Calibrate metrics to align with human preferences.Evaluate generation tasks across different modalities.Optimize existing metrics for multilingual and multi-domain scenarios.

Reward Models are Metrics in a Trench Coat

Oct 03, 2025
SG
Sebastian Gehrmann
🏛️ Bloomberg

Current research on reward modeling and evaluation metrics operates in silos, leading to terminological redundancy, spurious correlations, heightened reward hacking risks, and duplicated efforts in data quality optimization and meta-evaluation. Through a systematic literature review and comparative analysis, we reveal that both reward models and evaluation metrics fundamentally serve the same purpose in language model post-training: preference modeling and performance calibration. Building on this insight, we propose a unified research framework integrating three core directions—preference acquisition, spurious correlation mitigation, and meta-evaluation calibration. Empirical experiments demonstrate that certain evaluation metrics significantly outperform existing reward models on specific tasks. This work clarifies the root causes of conceptual ambiguity in the field and fosters cross-paradigm collaboration, providing both theoretical foundations and practical pathways for developing robust, interpretable, and reusable alignment evaluation systems.

Both fields struggle with spurious correlations and reward hackingCloser collaboration could improve preference elicitation and meta-evaluation methodsReward models and evaluation metrics face redundant terminology issues

Existing evaluation methods for LLM-generated code comments rely on small-scale datasets and inadequate IR metrics (e.g., BLEU), failing to capture semantic fidelity. Method: We systematically assess GPT-3.5’s Javadoc generation for 23,850 Java code snippets, employing a dual-dimensional evaluation combining quantitative BLEU scoring with qualitative expert human assessment. Contribution/Results: Our study reveals a critical flaw in BLEU: high scores frequently correlate with low-quality, verbatim descriptions, while high-fidelity semantic paraphrasing is systematically penalized. We find that 69.7% of generated Javadocs are semantically equivalent to—or can be refined to match—the original quality, and 22.4% significantly surpass the originals. These results demonstrate that automated metrics alone are unreliable for assessing documentation quality. We advocate human evaluation as the gold standard, with BLEU serving only as a supplementary heuristic—establishing a new, more rigorous paradigm for evaluating code documentation generation.

Evaluates AI-generated code comment quality versus human-written onesExplores relationship between code properties and AI comment effectivenessIdentifies limitations of traditional metrics in assessing documentation quality

Latest Papers

What's happening recently
View more

This study addresses the inconsistency in model rankings caused by commonly used ranking metrics—such as MRR, Hits@k, and Mean Rank—in knowledge graph completion (KGC) evaluation, which hinders fair comparison and reproducibility. For the first time, KGC evaluation is framed as a multi-criteria decision-making problem, and seven aggregators are systematically assessed across five dimensions: consistency, cross-dataset stability, metric independence, noise robustness, and generalization capability. Through leave-one-model-out (LOMO) and leave-one-group-out (LOGO) cross-validation, Pareto optimality analysis, and multidimensional sensitivity tests, the Z-score aggregator emerges as the most balanced overall—favoring DualE for tail entity prediction and FMS for relation prediction. The experiments further reveal that consistency and stability are insensitive to removal strategies, whereas generalization and independence exhibit the highest sensitivity.

Evaluation MetricsKnowledge Graph CompletionMetric Disagreement

Existing benchmarks for knowledge work evaluation largely adhere to traditional NLP task paradigms, failing to capture systems’ capabilities in real-world knowledge-intensive settings. This work proposes a three-step framework—explicitly defining work activities, establishing realistic test environments, and focusing evaluation on deliverable outputs—and derives 18 core knowledge work activities from the O*NET database. Innovatively integrating role responsibilities, local tool usage, and downstream usability into benchmark design, the approach establishes a coherent “work activity–test setup–scoring artifact” alignment. Validation through three case studies (GDPval, OfficeQA Pro, and APEX-SWE) exposes critical misalignments in current benchmarks between tasks, environments, and actual work objectives, offering a new paradigm for evaluating knowledge work systems in practical, application-oriented contexts.

benchmark designevaluationknowledge work

Alignment evaluation in machine learning has largely become evaluation of models. Influential benchmarks score model outputs under fixed inputs, such as truthfulness, instruction following, or pairwise preference, and these scores are often used to support claims about deployed alignment. This paper argues that deployment-relevant alignment cannot be inferred from model-level evaluation alone. Alignment claims should instead be indexed to the level at which evidence is collected: model-level, response-level, interaction-level, or deployment-level. Two studies support this position. First, a structured audit of eleven alignment benchmarks, extended to a sixteen-benchmark corpus, dual-coded against an eight-dimension rubric with Cohen's kappa = 0.87, finds that user-facing verification support is absent across every benchmark examined, while process steerability is nearly absent. The few interactional benchmarks identified, including tau-bench, CURATe, Rifts, and Common Ground, remain fragmented in coverage, and benchmark construction rather than data source determines what is measured. Second, a blinded cross-model stress test using 180 transcripts across three frontier models and four scaffolds finds that the same verification scaffold raises one model's verification support to ceiling while leaving another categorically unchanged. This shows that scaffold efficacy is model-dependent and that the gap identified by the audit cannot be closed at the model level alone. We propose a system-level evaluation agenda: alignment profiles instead of single scores, fixed-scaffolding protocols for comparable interactional evaluation, and reporting templates that make the inferential distance between evaluation evidence and deployment claims explicit.

alignment evaluationbenchmark limitationsdeployment-relevant alignment

Hot Scholars

AG

Arshit Gupta

Sr. Manager, Applied Science, Amazon
LLMNatural Language ProcessingDeep LearningNLU
XE

Xin Eric Wang

Assistant Professor, University of California, Santa Barbara, Simular
NLPCVMLLanguage and Vision
SS

Sailik Sengupta

Amazon Science
DecodingAlignmentRobust NLPMoving Target Defense
JL

Jian Luan

Toshiba, Microsoft, Xiaomi
LLMVLMTTSSinging Synthesis