analyze learning curves

Design, implement, and interpret analyses and visualizations that measure how model performance or loss evolves with training effort (iterations/epochs) or with training set size, including plotting loss and accuracy over time and comparing pooled versus classwise behaviour. Build procedures to estimate learning curves and break-even data volumes, produce decision-curve and training-curve visualizations, and quantify trade-offs revealed by those curves.

analyzelearningcurves

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Decision curve analysis (DCA) and cost curves are widely used for clinical utility assessment of classification models, yet their theoretical relationship and comparative applicability remain unclear. Method: We conduct rigorous mathematical derivation and empirical analysis to compare DCA and the Brier cost curve—specifically examining net benefit and cost under varying threshold probabilities. Contribution/Results: We prove that DCA’s net benefit and the Brier cost curve yield identical optimal model selections across all thresholds and are mathematically equivalent—DCA is a linear transformation of the Brier cost curve. We further introduce the “upper-envelope decision curve” to quantify calibration potential. Crucially, we demonstrate that the Brier curve is more general: its area under the curve equals the Brier score, enabling valid cross-threshold loss comparison, whereas DCA’s net benefit is inherently threshold-dependent and thus not directly comparable across thresholds. This work unifies two dominant evaluation paradigms, providing both theoretical grounding and practical tools for clinical and decision-analytic model selection.

Assessing classification performance across different decision thresholdsComparing decision curve analysis and cost curves for model evaluationDetermining optimal model selection using net benefit and Brier loss

Exploring a Datasets Statistical Effect Size Impact on Model Performance, and Data Sample-Size Sufficiency

Jan 05, 2025
AH
Arya Hatamian
🏛️ University of California, Riverside | University of California, Los Angeles

This study investigates whether effect size measures (e.g., Cohen’s *d*) can serve as prospective proxies for data sufficiency—specifically, to predict model performance (classification accuracy) and training convergence speed. We conduct systematic supervised learning experiments across varying sample sizes, learning rates, and convergence dynamics, quantitatively assessing statistical associations between effect size and model behavior. Our first empirical evaluation reveals no robust correlation between effect size and either accuracy or convergence rate, demonstrating its unreliability for sample-size planning or performance forecasting. These findings expose fundamental limitations of conventional descriptive statistics in assessing data sufficiency, challenge the implicit assumption that effect size serves as a valid proxy for data quality, and underscore the need for a new evaluation framework integrating statistical learning theory with explicit modeling of data-generating mechanisms.

Dataset SufficiencyMachine Learning ModelsPredictive Impact

Communication barriers between data scientists and domain experts arise from oversimplified, accuracy-centric model performance reporting, hindering shared understanding of model limitations and contextual applicability. Method: We propose a visualization-mediated model explanation framework grounded in human-computer interaction principles, participatory design, and visual narrative techniques. This yields the first domain-expert-oriented model communication guideline—emphasizing risk, trade-offs, and situational appropriateness rather than isolated metrics like accuracy. An iterative empirical study was conducted using regression models, incorporating structured expert feedback for evaluation. Contribution/Results: The framework significantly improves domain experts’ ability to identify model limitations, recognize inherent trade-offs, and proactively make context-driven adoption decisions. Its core innovation lies in repositioning visualization as an interdisciplinary consensus-building medium—shifting the paradigm from “metric reporting” to “collaborative understanding.”

Communication gaps between data scientists and subject matter experts hinder model understanding.Traditional metrics fail to convey model risks, strengths, and limitations effectively.Visualization guidelines improve model performance communication and decision-making confidence.

Latest Papers

What's happening recently
View more

Current AI research often treats models as static artifacts, overlooking the fundamental influence of training dynamics on critical properties such as capability, bias, robustness, and safety. This work proposes shifting the focus toward the training process itself to establish a science of AI centered on training dynamics. By analyzing the interactions among data, objectives, architectures, and optimizers, the paper develops a theoretical framework that is predictive, intervenable, and design-oriented. Integrating approaches from mechanistic interpretability, fairness, memory mechanisms, and simplicity biases, it uncovers causal links between early-training signals and final model behavior. The study systematically outlines key challenges and open problems, offering both theoretical pathways and practical foundations for extending scaling laws beyond performance to encompass multidimensional model attributes.

AI sciencemodel behaviorpredictability

Current post-training of language models relies on abstract scalar rewards, which lack transparency regarding the instructional content of preference data and can lead models to learn spurious correlations, resulting in undesirable behaviors such as excessive stylization or sycophancy. This work proposes a data-centric post-training framework that, for the first time, leverages interpretability methods to explicitly model latent conceptual signals within preference data. By analyzing and identifying key features that distinguish preferred from non-preferred responses prior to optimization, the approach integrates interpretability protocols, statistical hypothesis testing, and fine-grained interventions at both feature and data levels. This enables effective diagnosis and suppression of harmful learning signals, significantly reducing off-target behaviors across multiple benchmarks while enhancing model safety and controllable personality.

interpretabilitylearning signalpost-training

This work addresses the lack of intuitive, immediate feedback on fitting errors in existing model-fitting approaches. It proposes an interactive fitting framework that integrates visual and auditory feedback: as users manipulate parametric curves, the system synthesizes audio in real time, with greater model-data discrepancies producing louder and more dissonant sounds. This is the first approach to incorporate auditory cues into model exploration, enabling multisensory assessment of fit quality. Combining interactive visualization, real-time audio synthesis, and Gaussian process regression, the method demonstrates effectiveness and generalizability across four diverse case studies—golf putting, dilution experiments, cosmological parameter estimation, and temperature data fitting—significantly enhancing users’ intuitive perception of model misfit.

curve fittingdata visualizationmodel exploration

The absence of standardized benchmarks hinders systematic evaluation of AI models’ understanding of scatter plots. Method: We introduce ScatterBench, the first comprehensive benchmark for scatter plot comprehension, comprising over 18,000 synthetic samples generated via six data-generation patterns and seventeen chart designs. We propose an N-shot prompting-driven multi-task evaluation framework covering four spatial reasoning tasks: cluster counting, outlier detection, bounding box localization, and centroid coordinate prediction. Results: Experiments reveal that state-of-the-art closed-source models (e.g., GPT-4o, Gemini 2.5 Flash) achieve >90% accuracy on cluster counting but underperform significantly on localization tasks (F1 < 50%). Aspect ratio and color encoding are identified as critical visual factors substantially affecting model performance. This work provides the first systematic characterization of current AI models’ capabilities and design sensitivities in scatter plot spatial reasoning.

Addressing the lack of benchmarks for scatterplot-specific AI model tasksAssessing model performance on cluster counting and outlier localization tasksEvaluating AI models on scatterplot analysis tasks using synthetic datasets

This study addresses the lack of systematic understanding regarding the effectiveness and usage practices of univariate distribution visualizations across diverse tasks and user groups. Through a mixed-methods approach—combining a click-based selection experiment and survey with 215 participants alongside in-depth interviews with five visualization practitioners—the work systematically evaluates the accuracy, user preferences, and common misinterpretations associated with boxplots, violin plots, jittered scatterplots, and histograms in typical analytical tasks. For the first time, it integrates task performance, subjective preference, and real-world practice, revealing a frequent mismatch between chart familiarity and task accuracy, thereby challenging the assumption that commonly used or conventional visualizations are inherently optimal. The findings demonstrate significant performance differences among chart types in low-level tasks, with widely adopted histograms and boxplots not consistently outperforming alternatives.

chart effectivenesstask performanceunivariate distribution

Hot Scholars

AO

Antonio Orvieto

ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems
Deep LearningMachine LearningOptimizationDifferential Equations
FH

Frank Hutter

Prior Labs; ELLIS Institute Tübingen; University of Freiburg
Tabular DataFoundation ModelsAutoMLMeta-Learning
PL

Percy Liang

Associate Professor of Computer Science, Stanford University
machine learningnatural language processing
NZ

Nicolas Zucchet

PhD student, ETH Zurich
machine learningdeep learningneuroscience