Score
Designs and implements training processes that optimize two or more models or model components together, including defining coupled loss functions, optimization schedules, and mechanisms to ensure auxiliary modules provide valid targets. Builds and analyzes strategies to balance competing objectives, stabilize co-training, and coordinate updates across the jointly trained networks.
Traditional neural network training relies on fixed optimization pipelines, rendering it inflexible in dynamically addressing training instability and anomalies. To address this limitation, we propose the first interactive training framework enabling real-time human–AI collaborative intervention. Our method employs a lightweight control server that integrates expert human directives with AI agent feedback to dynamically adjust hyperparameters, data sampling strategies, and model checkpoints during training. This framework introduces, for the first time, a closed-loop interactive paradigm into neural network training, establishing a scalable human–machine collaboration interface coupled with automated response mechanisms. Experimental results demonstrate significant improvements in training stability, reduced sensitivity to initial hyperparameter configurations, and enhanced real-time responsiveness to user-specified customization requirements. The effectiveness is validated across multiple benchmark tasks.
To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.
This work investigates how neural network width governs training dynamics. For single-hidden-layer linear networks, we derive the first exact analytical solution of learning dynamics at arbitrary finite width, unifying the characterization of the two-phase evolution—kernel learning and feature learning—and establishing a complete phase diagram parameterized by width, layer-wise learning rates, and initialization scale. Methodologically, we integrate analytical dynamical systems analysis, phase-diagram modeling, and empirical validation on nonlinear networks. Crucially, we identify three novel mechanisms operative during the feature-learning phase: alignment learning, de-alignment learning, and rescaling learning—each transcending the conventional kernel-method paradigm. These theoretical insights are empirically reproduced in realistic deep networks, offering a new conceptual framework for understanding training dynamics and designing adaptive optimization algorithms. (138 words)
Learned optimizers (L2Os) suffer from poor out-of-distribution generalization, limiting their applicability beyond the training data distribution. Method: This paper proposes a novel paradigm integrating classical optimization priors with data-driven modeling. It systematically incorporates fundamental optimization principles—specifically scale invariance and affine covariance—into the architecture design. We introduce a parameterized quasi-Newton update module explicitly constrained to preserve BFGS structure, and jointly optimize it via end-to-end training that unifies optimization-theoretic modeling, neural network architecture design, and meta-learning. Contribution/Results: The resulting enhanced BFGS algorithm significantly outperforms both standard L2Os and conventional solvers on unseen problem classes, dimensions, and condition numbers. It achieves over 40% improvement in cross-distribution generalization performance, establishing a new pathway toward more transferable and robust learned optimizers.
Traditional robot design and control are typically decoupled, leading to morphologies poorly aligned with task requirements. This paper proposes a simulation-driven co-optimization framework for morphology and control, breaking the conventional “design-then-control” paradigm to enable task-oriented, end-to-end joint search. Our method employs gradient-free optimization to simultaneously evolve structural parameters and controller policies within a URDF-based multi-task reinforcement learning simulation environment. Key contributions include: (1) demonstrating that controller retraining significantly improves performance, yielding an average gain of 37%; and (2) revealing an inverse correlation between morphological complexity and controller training budget—providing theoretical justification for structural simplification under resource constraints. We validate the framework across four public simulation benchmarks, showing that co-optimization consistently yields more compact, robust, and task-adapted robot morphologies compared to sequential approaches.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
研究通过探索神经网络损失景观中的模式连接现象,以理解其优化过程和特征学习多样性,从而解决对神经网络决策机制理解不足的问题。
本文提出一种多智能体合作机制,通过信息交换减少模型复杂度并保持性能,使用投票、多数和加权平均等策略实现合作,减少了训练参数数量。
This work addresses the limited understanding of how implicit biases of optimizers arise during training. Departing from prior analyses focused on the geometry of the solution space, it innovatively shifts attention to the trajectory of parameter updates through the lens of dynamic information allocation. The study introduces a preconditioning exponent \( p \) to characterize the relative distribution of training signals between weight and bias pathways. Using a minimal linear model, it reveals that weight update components preserve input-dependent residual structures, while bias updates capture the mean direction of residuals. The relative strength of these two components is shown to govern both learning dynamics and generalization performance. This framework offers a novel mechanistic perspective on optimizer-induced implicit bias, providing a tunable and interpretable viewpoint for its analysis.
This work addresses the challenge of jointly optimizing data and model configurations in large language model training, a task rendered difficult by their high coupling. To this end, we propose JoBS, the first method to enable efficient joint optimization by integrating a scaling law–informed performance predictor into Bayesian optimization and leveraging multi-fidelity evaluation to substantially reduce the cost of full-scale training. JoBS not only yields an optimal budget allocation strategy but also consistently outperforms baselines that optimize only data, only model hyperparameters, or existing multi-fidelity Bayesian optimization approaches, achieving superior performance across diverse large language model tasks under identical computational budgets.