joint training

Designs and implements training processes that optimize two or more models or model components together, including defining coupled loss functions, optimization schedules, and mechanisms to ensure auxiliary modules provide valid targets. Builds and analyzes strategies to balance competing objectives, stabilize co-training, and coordinate updates across the jointly trained networks.

jointtraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Interactive Training: Feedback-Driven Neural Network Optimization

Oct 02, 2025
WZ
Wentao Zhang
🏛️ University of Waterloo | University of Wisconsin-Madison

Traditional neural network training relies on fixed optimization pipelines, rendering it inflexible in dynamically addressing training instability and anomalies. To address this limitation, we propose the first interactive training framework enabling real-time human–AI collaborative intervention. Our method employs a lightweight control server that integrates expert human directives with AI agent feedback to dynamically adjust hyperparameters, data sampling strategies, and model checkpoints during training. This framework introduces, for the first time, a closed-loop interactive paradigm into neural network training, establishing a scalable human–machine collaboration interface coupled with automated response mechanisms. Experimental results demonstrate significant improvements in training stability, reduced sensitivity to initial hyperparameter configurations, and enhanced real-time responsiveness to user-specified customization requirements. The effectiveness is validated across multiple benchmark tasks.

Allows dynamic adjustment of optimizer hyperparameters and training dataEnables real-time feedback-driven intervention during neural network trainingImproves training stability and adaptability to evolving user needs

MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters

Feb 04, 2024
AS
Arsalan Sharifnassab
🏛️ University of Alberta | Leiden University

To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.

Dynamically adjusting step sizes during model optimizationOptimizing meta-parameters for efficient machine learning trainingReducing regret by considering long-term impact of learning rates

Three Mechanisms of Feature Learning in a Linear Network

Jan 13, 2024
YX
Yizhou Xu
🏛️ Abdus Salam International Center for Theoretical Physics | Massachusetts Institute of Technology | NTT Research

This work investigates how neural network width governs training dynamics. For single-hidden-layer linear networks, we derive the first exact analytical solution of learning dynamics at arbitrary finite width, unifying the characterization of the two-phase evolution—kernel learning and feature learning—and establishing a complete phase diagram parameterized by width, layer-wise learning rates, and initialization scale. Methodologically, we integrate analytical dynamical systems analysis, phase-diagram modeling, and empirical validation on nonlinear networks. Crucially, we identify three novel mechanisms operative during the feature-learning phase: alignment learning, de-alignment learning, and rescaling learning—each transcending the conventional kernel-method paradigm. These theoretical insights are empirically reproduced in realistic deep networks, offering a new conceptual framework for understanding training dynamics and designing adaptive optimization algorithms. (138 words)

Analyzes learning dynamics in neural networksExplores hyperparameter impact on training trajectoriesIdentifies feature learning mechanisms in networks

From Learning to Optimize to Learning Optimization Algorithms

May 28, 2024
CC
Camille Castera
🏛️ University of Tübingen | Saarland University

Learned optimizers (L2Os) suffer from poor out-of-distribution generalization, limiting their applicability beyond the training data distribution. Method: This paper proposes a novel paradigm integrating classical optimization priors with data-driven modeling. It systematically incorporates fundamental optimization principles—specifically scale invariance and affine covariance—into the architecture design. We introduce a parameterized quasi-Newton update module explicitly constrained to preserve BFGS structure, and jointly optimize it via end-to-end training that unifies optimization-theoretic modeling, neural network architecture design, and meta-learning. Contribution/Results: The resulting enhanced BFGS algorithm significantly outperforms both standard L2Os and conventional solvers on unseen problem classes, dimensions, and condition numbers. It achieves over 40% improvement in cross-distribution generalization performance, establishing a new pathway toward more transferable and robust learned optimizers.

Designing learned optimization algorithms usable beyond training settingsDeveloping learning-enhanced BFGS algorithm adaptable to various test settingsSynergy between classical optimization and Learning to Optimize (L2O)

Co-Optimization of Robot Design and Control: Enhancing Performance and Understanding Design Complexity

Sep 13, 2024
EA
Etor Arza
🏛️ Basque Center for Applied Mathematics | University of Oslo

Traditional robot design and control are typically decoupled, leading to morphologies poorly aligned with task requirements. This paper proposes a simulation-driven co-optimization framework for morphology and control, breaking the conventional “design-then-control” paradigm to enable task-oriented, end-to-end joint search. Our method employs gradient-free optimization to simultaneously evolve structural parameters and controller policies within a URDF-based multi-task reinforcement learning simulation environment. Key contributions include: (1) demonstrating that controller retraining significantly improves performance, yielding an average gain of 37%; and (2) revealing an inverse correlation between morphological complexity and controller training budget—providing theoretical justification for structural simplification under resource constraints. We validate the framework across four public simulation benchmarks, showing that co-optimization consistently yields more compact, robust, and task-adapted robot morphologies compared to sequential approaches.

Explores controller training impact on robot performance and designInvestigates computation budget challenges in robot co-optimizationStudies budget allocation effects on design complexity in simulation

Latest Papers

What's happening recently
View more

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

This work addresses the limited understanding of how implicit biases of optimizers arise during training. Departing from prior analyses focused on the geometry of the solution space, it innovatively shifts attention to the trajectory of parameter updates through the lens of dynamic information allocation. The study introduces a preconditioning exponent \( p \) to characterize the relative distribution of training signals between weight and bias pathways. Using a minimal linear model, it reveals that weight update components preserve input-dependent residual structures, while bias updates capture the mean direction of residuals. The relative strength of these two components is shown to govern both learning dynamics and generalization performance. This framework offers a novel mechanistic perspective on optimizer-induced implicit bias, providing a tunable and interpretable viewpoint for its analysis.

information allocationoptimizer implicit biasparameter pathways

This work addresses the challenge of jointly optimizing data and model configurations in large language model training, a task rendered difficult by their high coupling. To this end, we propose JoBS, the first method to enable efficient joint optimization by integrating a scaling law–informed performance predictor into Bayesian optimization and leveraging multi-fidelity evaluation to substantially reduce the cost of full-scale training. JoBS not only yields an optimal budget allocation strategy but also consistently outperforms baselines that optimize only data, only model hyperparameters, or existing multi-fidelity Bayesian optimization approaches, achieving superior performance across diverse large language model tasks under identical computational budgets.

chicken-and-egg dilemmadata configurationjoint optimization

Hot Scholars

ZZ

Zhicheng Zhang

Carnegie Mellon University
Reinforcement LearningExplainable RL
XZ

Xiatian Zhu

University of Surrey
Machine LearningComputer Vision
HY

Hung-yi Lee

National Taiwan University
deep learningspoken language understandingspeech processing
DD

Dwip Dalal

PhD Student, University of Illinois, Urbana-Champaign
VLAsMLLMsAgentsMultimodal Learning