analyze training dynamics

Design and build diagnostics, visualizations, metrics, and dynamical models that characterize how model parameters, representations, and behaviors evolve over the course of training — identifying phase transitions, emergent abilities, memorization patterns, and failure modes and quantifying sensitivities to initialization, data, objectives, and hyperparameters. Implement monitoring and trajectory-level analyses that relate trajectory features to final outcomes, detect instabilities or regime changes, and separate conflated failure modes or irreducible sources of opacity.

analyzetrainingdynamics

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$237K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Current AI research often treats models as static artifacts, overlooking the fundamental influence of training dynamics on critical properties such as capability, bias, robustness, and safety. This work proposes shifting the focus toward the training process itself to establish a science of AI centered on training dynamics. By analyzing the interactions among data, objectives, architectures, and optimizers, the paper develops a theoretical framework that is predictive, intervenable, and design-oriented. Integrating approaches from mechanistic interpretability, fairness, memory mechanisms, and simplicity biases, it uncovers causal links between early-training signals and final model behavior. The study systematically outlines key challenges and open problems, offering both theoretical pathways and practical foundations for extending scaling laws beyond performance to encompass multidimensional model attributes.

AI sciencemodel behaviorpredictability

In supervised cybersecurity CTF training, evaluating learning outcomes and identifying flaws in training design remain challenging. To address these issues, this paper proposes an evaluation framework integrating process mining with multidimensional visual analytics. It models participants’ operational behavior sequences and implements an open-source, interactive dashboard supporting temporal pattern recognition, multivariate network visualization, and clustering analysis—rigorously adhering to established visualization design principles. Our key innovation lies in deeply embedding process mining into cybersecurity pedagogical assessment, enabling automated discovery of process deviations, bottlenecks, and organizational anomalies directly from system logs. A case study demonstrates that the framework effectively quantifies learner engagement, pinpoints training deficiencies—including task bottlenecks and imbalanced resource allocation—and substantially enhances the interpretability of evaluation results and their utility for instructional improvement.

Analyzing process-oriented cybersecurity training exercisesEnhancing learning analytics with process mining techniquesVisualizing temporal and clustering data for CTF games

Transition Network Analysis: A Novel Framework for Modeling, Visualizing, and Identifying the Temporal Patterns of Learners and Learning Processes

Nov 23, 2024
MS
Mohammed Saqr
🏛️ University of Eastern Finland | University of Oulu | University of Oslo | FernUniversität in Hagen | University of Jyväskylä

This paper addresses three key challenges in modeling collaborative learning: difficulty in capturing temporal behavioral patterns, weak visualization capabilities, and low robustness in identifying critical events. To tackle these, we propose the Transition Network Analysis (TNA) framework—a novel probabilistic graphical model that unifies relational structure and temporal dynamics. TNA integrates stochastic process mining with multilayer network analysis, enabling centrality computation, community detection, and temporal clustering. A key innovation is the introduction of a Bootstrap-based transition significance test, which effectively filters spurious transitions. Evaluated on real-world collaborative learning data from 191 students, TNA successfully characterizes the dynamic evolution of regulatory processes, accurately identifies critical learning events and behavior clusters, and significantly enhances both the reliability and interpretability of temporal pattern analysis.

Identifying significant learning events and clustersModeling temporal patterns in learning processesVisualizing student and learning dynamics

This work addresses the problem of classifying trajectories generated by distinct nonlinear dynamical systems, where each class corresponds to a unique system. The authors propose Dynafit, a novel method that, for the first time, integrates the Koopman operator framework with kernel methods to achieve global linearization of dynamics in a reproducing kernel Hilbert space. By leveraging the kernel trick, Dynafit efficiently computes dynamical distances between trajectories while allowing incorporation of prior knowledge. The approach demonstrates significant performance gains over baseline methods across three diverse tasks: detecting chaos in logistic maps, recognizing handwritten dynamics, and classifying visual dynamic textures. These results validate Dynafit’s effectiveness and generality in multi-class classification of nonlinear dynamical systems.

chaos detectiondynamical system identificationminimum distance classification

Latest Papers

What's happening recently
View more

This work addresses the inconsistency between the chain-of-thought (CoT) reasoning process and final outputs in large reasoning models (LRMs), which undermines their reliability for safety monitoring. To this end, the authors propose a Probe Trajectory framework that evaluates probes at every generated token to track the continuous evolution of concept probabilities and predict future model behavior. The method innovatively incorporates signal processing features—such as volatility, trend, and steady-state characteristics—to characterize reasoning dynamics. Notably, the study finds that template-based training data can effectively substitute costly dynamically generated data, and reveals that max-pooling is crucial for trajectory stability. Experiments across four datasets and four models demonstrate that the approach substantially improves future state discriminability, achieving up to 95% AUROC with max-pooling, thereby offering a more reliable solution for LRM behavior monitoring.

Chain of ThoughtLarge Reasoning Modelsprobe trajectories

Existing dynamical system reconstruction models exhibit limited out-of-distribution generalization, particularly when extrapolating across critical points. This work identifies three fundamental structural deficiencies underlying this limitation and introduces an improved framework based on topological feature disentanglement and hierarchical modeling. For the first time, the study derives a closed-form theoretical bound characterizing the reliable extrapolation range of such models. The proposed approach enables high-accuracy, zero-shot predictions in unseen dynamical regimes—such as regions straddling bifurcation points—without requiring additional training, thereby substantially enhancing out-of-distribution generalization performance.

dynamical systemsout-of-domain generalizationscientific machine learning

This work proposes a method for automatically mining readable, structured skill libraries from user GUI interaction trajectories to enhance agent policy performance. The approach comprises a three-stage pipeline: trajectory segmentation, unsupervised clustering to generate candidate skills, and skill-aware policy training, integrating trajectory representation learning, offline reward modeling, and the GRPO algorithm. It presents the first systematic validation of the feasibility of unsupervised extraction of interpretable skills from real-world interactions. Experimental results on the InteraSkill Workflows benchmark show that five out of eight clusters achieve purity above 0.95, yet the resulting policy improvement remains marginal—skill-step accuracy in IW tasks increases only from 18.5% to 20.5%—highlighting a disconnect between skill interpretability and effective policy transfer.

computer-using agentsGUI automationinteraction trajectory

This work addresses the performance degradation of time series models in post-training quantization (PTQ), which arises from error propagation and amplification during quantization—particularly challenging in calibration-free or black-box settings where module sensitivity is hard to assess. To tackle this, the paper introduces discrete-time dynamical systems theory into quantization analysis for the first time. By modeling the inference process as a dynamical system, it proposes TQS, a quantizer-agnostic, prior-based sensitivity metric derived from trajectory sensitivity analysis, enabling calibration-free mixed-precision quantization budget allocation. The resulting TQS-PTQ framework significantly outperforms existing PTQ methods without relying on calibration data or second-order approximations, facilitating efficient low-bit deployment.

dynamical systemserror propagationpost-training quantization

Hot Scholars

TS

Tat-Seng Chua

National University of Singapore
Multimedia Information RetrievalLive Social Media Analysis
CT

Christos Thrampoulidis

Assistant Professor, University of British Columbia, ECE Department
data sciencesignal processingoptimizationmachine-learning
MG

Mor Geva

Tel Aviv University, Google Research
Natural Language Processing
OL

Ofir Lindenbaum

Assistant Professor at Bar Ilan University
Computational BiologyMultimodal LearningGenerative ModelsMachine Learning