comparative analysis

Designing and carrying out systematic comparisons between methods, models, or representations to quantify relative performance, robustness, and trade-offs (e.g., accuracy, memory, speed, differentiability). This includes selecting baselines, constructing experiments or case comparisons, and interpreting outcome differences to highlight benefits and limitations.

comparativeanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of systematic and rigorous performance benchmarking methodologies in programming language research, which has undermined the credibility of evaluation results. To remedy this, the paper introduces a closed-loop methodology—Measure-Explain-Test-Improve—that establishes, for the first time, a structured and reproducible workflow for performance assessment in the field. Integrating systematic experimental design, performance metric analysis, result interpretation, and iterative refinement, the approach emphasizes theoretical grounding and practical rigor at every stage. Its key contribution lies in enabling even researchers with limited empirical experience to conduct reliable and methodologically sound performance evaluations, thereby significantly enhancing the scientific validity and reproducibility of performance analysis in programming language research.

benchmarkingperformance evaluationprogramming language research

On the handling of method failure in comparison studies

Aug 21, 2024
MW
Milena Wunsch
🏛️ LMU Munich | Munich Center for Machine Learning | Department of Statistics | MRC Clinical Trials Unit | UCL

In methodological comparative studies, algorithmic failures—such as non-convergence or absence of output—preclude performance evaluation, yet existing literature lacks standardized guidelines for handling such failures, often overlooking or misapplying failure mitigation strategies. Method: We systematically analyze failure causes and risks of improper handling, critically examine prevalent censoring and imputation strategies for their statistical biases, and propose the principle of “context-adapted failure fallback,” establishing a framework grounded in empirically feasible fallback mechanisms. Through statistical modeling, failure root-cause diagnosis, and cross-domain empirical analysis, we identify widespread deficiencies in published studies’ failure handling practices. Contribution/Results: Two representative case studies demonstrate that inappropriate failure handling significantly distorts method rankings and undermines conclusion validity. Our work bridges critical theoretical and practical gaps in the principled treatment of algorithmic failures in empirical methodology research.

Addressing method failure handling in comparison studiesProviding guidance on proper failure interpretation and reportingRecommending realistic fallback strategies for method failures

Aggregating empirical evidence from data strategy studies: a case on model quantization

May 01, 2025
SD
Santiago del Rey
🏛️ Universitat Polit`ecnica de Catalunya | UNIRIO | UFRJ

This study systematically evaluates the impact of model quantization on the correctness and resource efficiency of deep learning systems, while also exploring methodologies for cross-study evidence aggregation in data-driven empirical research. Methodologically, it innovatively applies Structured Synthesis Methods (SSM) for the first time in this domain, integrating findings from six empirical studies covering 19 models through a qualitative-quantitative mixed analysis. Results demonstrate that quantization yields substantial resource gains—average storage compression of ×3.2, inference latency reduction of −41%, and GPU energy consumption decrease of −38%—with only a marginal correctness degradation (−1.7% on average), representing a well-controlled trade-off. The study identifies both consistent patterns and fragmentation bottlenecks in quantization effects, and proposes a refined empirical research framework and methodological guidelines tailored to quantization techniques. These contributions provide foundational methodological support and practical guidance for optimizing trustworthy AI systems.

Assessing model quantization effects on DL correctness and efficiencyEvaluating trade-offs between correctness and resource efficiency in quantizationExploring methodological challenges in aggregating data strategy studies

Manipulation of individual judgments in the quantitative pairwise comparisons method

Nov 01, 2022
MS
M. Strada
🏛️ Aptiv Services Poland S.A. | AGH University of Kraków

In quantitative pairwise comparisons, expert judgments are vulnerable to bribery-based manipulation, leading to distorted global rankings. Method: This paper formally defines the “targeted manipulation” problem for the first time and introduces a unified modeling framework integrating game theory and graph theory to characterize adversarial interventions. It proposes three polynomial-time solvable manipulation algorithms capable of precisely achieving desired rankings. Contribution/Results: Theoretical analysis demonstrates that even minimal bribery costs can significantly distort ranking outcomes. Furthermore, the study uncovers structural properties and inherent vulnerabilities of manipulation strategies, providing a theoretical foundation for detecting anomalous judgments and designing robust aggregation mechanisms. This work bridges a critical gap in the robustness literature on pairwise comparisons by establishing the first formal model of adversarial intervention.

Analyzes defenses against expert judgment manipulationDetects bribery vulnerability in pairwise comparison methodsProposes algorithms to achieve manipulation goals

Latest Papers

What's happening recently
View more

This work addresses the challenge of systematically evaluating concept bottleneck models, whose applicability and failure mechanisms remain poorly understood due to the scarcity of real-world datasets with annotated concept labels. To bridge this gap, we introduce the first controllable synthetic benchmark that leverages parametric generation techniques to precisely modulate data modality, concept selection, annotation quality, and label completeness, thereby simulating diverse real-world relationships between concepts and predictions. This benchmark enables comprehensive evaluation of various concept bottleneck models across both decision-support and fully automated tasks, effectively identifying key performance determinants and characteristic failure modes. Our framework fills a critical void in the current evaluation landscape for concept-based interpretability methods.

concept bottleneck modelsconcept labelsmodel interpretability

This work addresses the challenge of effectively evaluating the reliability of individual predictions made by classifiers under distribution shift. We systematically compare robustness quantification (RQ) and uncertainty quantification (UQ) for this task and, for the first time, explore the potential of integrating both approaches. Through extensive experiments across multiple benchmark datasets—covering both standard training conditions and various distribution shift scenarios—we elucidate the conceptual distinctions between RQ and UQ. Our results demonstrate that RQ consistently matches or outperforms UQ in most settings, and that hybrid strategies combining RQ and UQ significantly enhance the accuracy of reliability assessment for individual predictions. These findings offer a novel perspective toward building more trustworthy AI systems.

classifier predictionsreliability assessmentRobustness Quantification

Current benchmarks for evaluating toxicity in large language models exhibit underappreciated systematic biases that may lead to the deployment of unsafe models. This work systematically investigates how variations in task formulation—such as text completion versus summarization—input data domains, and evaluated models interact with multiple toxicity metrics. It reveals, for the first time, that both task type and data domain significantly influence toxicity scores. Experiments demonstrate that existing benchmarks are prone to misclassifying content as harmful when tasks are altered and show inconsistent performance across domains, highlighting their fragility and dependence on specific model-task configurations. These findings underscore the urgent need for more robust and reliable toxicity evaluation frameworks.

benchmark robustnessevaluation biasLLM evaluation

This work addresses the challenge that state-dependent behaviors in modern computing environments—such as those introduced by adaptive system mechanisms—induce time-dependent biases in traditional software benchmarking, undermining reliable performance comparisons. The paper reframes benchmarking as a decision problem aimed at identifying the fastest program and introduces an experimental paradigm centered on pairwise performance comparisons, thereby avoiding strong assumptions about modeling system dynamics. By leveraging contrastive estimators, consistent statistical inference, and test strategies under finite evaluation budgets, the approach eliminates program-specific biases without relying on absolute performance metrics. The method provides asymptotic guarantees for correct decisions in stateful environments, offering a robust and reliable benchmarking framework for performance-sensitive software development.

biased estimatorsperformance evaluationsoftware benchmarking

Current medical AI evaluation benchmarks predominantly emphasize knowledge acquisition, failing to adequately capture model reliability, safety, and clinical utility in real-world settings. To address this gap, this work proposes the first systematic evaluation framework aligned with clinical workflows, encompassing end-to-end tasks such as clinical documentation, decision support, and administrative processes. The framework integrates authentic multimodal clinical data and introduces task-specific metrics to comprehensively assess generative models, multimodal systems, and AI agents. Empirical results reveal a substantial performance gap between state-of-the-art models on real-world tasks and their scores on medical knowledge exams—scoring 0.74–0.85 in documentation, 0.61–0.76 in clinical decision-making, and 0.53–0.63 in administrative tasks—highlighting the limitations of existing evaluation paradigms and underscoring the critical role of this framework in advancing the clinical deployment of medical AI.

benchmarkingclinical relevancehealthcare AI

Hot Scholars

LC

László Csató

Corvinus University of Budapest
decision theorygame theorymechanism designOR in sports
CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science
AR

Armando Rungi

IMT School for Advanced Studies - Lucca
international economicsindustrial organizationmicroeconometricsmachine learning
MW

Mark Whitmeyer

Arizona State University
Game TheoryMicroeconomic TheoryInformation Economics
NL

Nils Lid Hjort

Professor of Mathematical Statistics, University of Oslo
Theoretical and applied statistics and probability theory