compare model outputs

Designs and implements analyses, metrics, and tools that compare and quantify differences among model outputs across conditions, inputs, or model variants; measures distributional and instance-level output differences, identifies which information sources or inputs dominate outputs, and quantifies and ranks the effect of conflicting inputs on model behavior.

comparemodeloutputs

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Review and Comparison of Different Sensitivity Analysis Techniques in Practice

Apr 11, 2025
DF
Devin Francom
🏛️ Los Alamos National Laboratory

Quantifying how input uncertainty propagates to model outputs remains a fundamental challenge in computational modeling. Method: This study systematically reviews and empirically compares prominent global and local sensitivity analysis (SA) techniques—including Sobol’, FAST, Morris screening, and local derivative-based methods—implemented via standard software packages, supporting both probabilistic modeling and distribution-free settings. Contribution/Results: We propose a practical decision framework that guides method selection based on problem characteristics, analytical objectives, and resource constraints—rejecting the notion of a universally “optimal” SA method and thereby addressing a critical gap in methodological implementation guidance. A reusable, open-source toolkit is developed to enhance the reliability and interpretability of uncertainty attribution. The framework and tools have been validated across multiple engineering and policy modeling applications, demonstrating robustness and scalability in real-world contexts.

Compare sensitivity analysis methods for uncertainty assessment.Guide selection of global vs local sensitivity techniques.Provide practical toolkit for input-output uncertainty analysis.

Communication barriers between data scientists and domain experts arise from oversimplified, accuracy-centric model performance reporting, hindering shared understanding of model limitations and contextual applicability. Method: We propose a visualization-mediated model explanation framework grounded in human-computer interaction principles, participatory design, and visual narrative techniques. This yields the first domain-expert-oriented model communication guideline—emphasizing risk, trade-offs, and situational appropriateness rather than isolated metrics like accuracy. An iterative empirical study was conducted using regression models, incorporating structured expert feedback for evaluation. Contribution/Results: The framework significantly improves domain experts’ ability to identify model limitations, recognize inherent trade-offs, and proactively make context-driven adoption decisions. Its core innovation lies in repositioning visualization as an interdisciplinary consensus-building medium—shifting the paradigm from “metric reporting” to “collaborative understanding.”

Communication gaps between data scientists and subject matter experts hinder model understanding.Traditional metrics fail to convey model risks, strengths, and limitations effectively.Visualization guidelines improve model performance communication and decision-making confidence.

Does the Tool Matter? Exploring Some Causes of Threats to Validity in Mining Software Repositories

Jan 25, 2025
NH
Nicole Hoess
🏛️ Technical University of Applied Sciences Regensburg | University of Hawaii at Mānoa | Siemens AG

Implementation discrepancies across software repository mining tools severely threaten the validity of empirical findings. Method: We conduct a dual-tool comparative analysis of 10 large-scale open-source projects, systematically identifying how minor implementation differences—such as commit parsing logic and author deduplication rules—induce up to 500% deviation in key metrics (e.g., commit count, developer count). We propose a “tool-level configuration + post-hoc normalization” co-optimization framework to mitigate metric divergence and perform multi-tool experiments, quantitative consistency assessment, and code-level root-cause analysis. Contribution/Results: We identify six technical sources undermining data validity and establish the first validity assessment paradigm for Mining Software Projects Research (MSPR) explicitly addressing tool heterogeneity—thereby enabling rigorous, reproducible, and comparable empirical software engineering studies.

Data Analysis VariabilityResearch ReliabilitySoftware Engineering

PRAXA: A Framework for What-If Analysis

Oct 10, 2025
SG
Sneha Gathani
🏛️ University of Maryland, College Park | University of Massachusetts Amherst | MIT CSAIL

Current what-if analysis lacks a unified conceptual framework, leading to terminological inconsistency across domains, structural ambiguity, and divergent interpretations. To address this, we conduct a systematic review of 141 papers in visual analytics and human-computer interaction, proposing Praxa—the first integrative framework that unifies scenario modeling, sensitivity analysis, and counterfactual analysis under a coherent paradigm. Praxa formally defines the underlying motivations, core components (hypothesis generation, intervention modeling, outcome evaluation), and a taxonomy of analytical types. It establishes a standardized terminology and structured model, exposing critical challenges including interpretability, causal modeling fidelity, and alignment with user intent. By clarifying conceptual boundaries and operational relationships among methods, Praxa significantly enhances cross-domain conceptual consistency and application clarity. The framework provides a rigorous foundation for theoretical advancement and the design of next-generation interactive analytical tools.

Establish structural understanding of hypothetical scenario analysisLack unified framework for what-if analysis conceptsNeed standardized vocabulary for cross-domain consistency

Existing methods struggle to align and interpret distribution shifts across heterogeneous, domain-consistent datasets—such as tabular, textual, visual, and time-series data—especially when scale and modality disparities are pronounced, resulting in poor interpretability. This paper introduces the first human-centric, cross-modal distribution discrepancy explanation framework, implemented as an interpretable dataset comparison toolbox. It integrates statistical hypothesis testing, feature importance decomposition, class activation mapping (CAM), contrastive representation learning, and interpretable generative modeling to enable fine-grained, semantically readable attribution and visualization of distributional shifts. Evaluated across diverse real-world scenarios, the framework significantly improves users’ efficiency in understanding shift causes and enhances the accuracy of intervention decisions—thereby overcoming the limitations of conventional black-box shift detection approaches.

Data InterpretationInter-data DifferentiationMulti-type Data Analysis

Latest Papers

What's happening recently
View more

This study addresses the limited diversity and perceived relevance of problem domains in software modeling instruction, which often undermine student motivation and inclusivity. Through parallel surveys of 90 students and 22 instructors, combined with quantitative and qualitative analyses, it reveals a significant mismatch between instructor assumptions and student preferences: learners prioritize socially relevant problem contexts and value autonomy in topic selection. Furthermore, their sense of engagement markedly increases when they perceive their feedback has been explicitly incorporated. The work proposes a learner-centered strategy for selecting problem domains that foregrounds social relevance and autonomy as critical enablers of inclusive learning. It also highlights how seemingly minor instructional design choices can inadvertently foster exclusion, offering empirical insights and actionable guidance for improving pedagogical practice.

Domain DiversityFeedbackInclusion

This study addresses the current lack of human-centered, interpretable, and responsible evaluation criteria for AI in modeling and simulation. The authors propose the first multidimensional benchmark framework specifically designed to assess large language models (LLMs) through a human-centric lens, leveraging an open-source system dynamics AI platform to systematically evaluate performance across qualitative modeling, quantitative modeling, and model discussion tasks—emphasizing human-AI collaboration rather than replacement. The framework incorporates critical capabilities such as causal reasoning, iterative model refinement, and behavioral explanation, while embedding ethical and accountability considerations. Empirical results indicate that existing AI tools perform relatively well in qualitative tasks and model discussions but remain limited in causal reasoning and quantitative error correction; furthermore, different LLMs exhibit distinct strengths, with no single model emerging as universally superior.

AI for Modeling and SimulationBenchmarkingBias in AI

This study investigates the impact of output format on the performance of large language models in single-turn code generation tasks and its interaction with model identity. Through controlled experiments across four open-source projects, three prominent models—Doubao, DeepSeek, and Qwen—were evaluated using three output formats: JSON Patch, unified diff, and full file, resulting in 4,013 test cases. The findings reveal that output format significantly affects success rates, with no universally optimal format; instead, each model exhibits distinct format preferences—for instance, Doubao achieves a 94% success rate with JSON Patch, while DeepSeek performs best (66%) with unified diff. This work is the first to demonstrate a strong interaction effect between output format and model identity, proposing model-specific output strategies and design principles for tools aimed at preventing format misuse.

coding agentformat misuseinteraction effect

Hot Scholars

PN

Ping Nie

Waterloo University
Natural Language ProcessingInformation RetrievalRecommendation SystemsTime Series Forecasting
FP

Fred Prior

Distinguished Professor and Chair, Department of Biomedical Informatics, University of Arkansas for
quantitative imaginginformatics
SQ

Shi Qiu

Peking University
MultimodalityLLM EvaluationNLP
HS

Huan Sun

Endowed CoE Innovation Scholar and Associate Professor, The Ohio State University
AgentsLarge Language ModelsNatural Language ProcessingAI