Score
Designs parameter-sharing and weighting strategies that specify how model parameters are allocated or shared across network paths or components, including input-aware weighting functions and scale-dependent behaviors; and builds schemes to compress or aggregate residuals or other signals while minimizing added parameter count.
To address policy homogenization induced by parameter sharing in multi-agent reinforcement learning—particularly its inability to accommodate heterogeneous agent identities and task requirements—this paper proposes a zero-overhead, identity-driven adaptive subnet partitioning mechanism. Inspired by neural functional parcellation, the method employs a learnable identity encoder to generate agent-specific binary masks, which dynamically route inputs to localized subnetworks within a shared backbone, thereby enabling differentiated policy representations. Crucially, it introduces no additional parameters and is fully compatible with standard on-policy algorithms such as PPO and A2C. Extensive experiments on StarCraft II, the Multi-Agent Particle Environment (MPE), and custom heterogeneous benchmarks demonstrate an average 12.7% improvement in win rate and a 3.2× increase in inter-agent policy diversity, significantly outperforming both conventional parameter sharing and Hypernetwork-based baselines.
Large-scale deep models incur high memory overhead during inference due to their massive parameter counts. Existing parameter-sharing methods rely on heuristic, adjacent-layer designs and lack systematic scalability across multiple layers. This work pioneers a graph coloring formulation for cross-layer parameter sharing, leveraging structural symmetry in the model’s parameter space for rigorous, system-level modeling. From a group-theoretic perspective, we analyze sharing mechanisms and introduce an analytical criterion grounded in second-order gradient geometry—guiding parameter projection onto low-curvature subspaces. By combining Hessian spectral analysis with Taylor expansion, we formulate optimal parameter grouping as finding the optimal coloring function α: L → C. Evaluated across diverse architectures and tasks, our method consistently outperforms state-of-the-art approaches, achieving superior accuracy at higher compression ratios—demonstrating both theoretical rigor and engineering scalability.
This work addresses the theoretical gap in partial parameter reuse within transfer learning. We investigate the conditions under which transferring only a subset of parameters from a pretrained ReLU convolutional neural network remains effective—and when it fails—for downstream tasks. By establishing a rigorous theoretical link between upstream feature learning capability and downstream performance, we derive a discriminative criterion for parameter transferability, identifying key determinants: feature generality, task alignment, and subspace structure of the transferred parameters. Our analysis reveals that transfer degrades performance below from-scratch training when inherited parameters cannot support discriminative feature representations required by the downstream task. Combining formal derivation with numerical experiments, we validate both the plausibility of our knowledge-transfer-path modeling and the empirical validity of the proposed conditions. This yields the first systematic, theoretically grounded framework for controllable and interpretable partial-parameter transfer.
Existing edge caching mechanisms for AI model delivery in 5G/6G networks overlook parameter-block reuse—e.g., shared knowledge units across CNNs or LLMs—leading to low storage efficiency and limited cache hit rates under stringent latency constraints. Method: We propose a parameter-sharing-aware edge model caching framework that, for the first time, formulates parameter-block reuse as a submodular optimization problem. We design a polynomial-time algorithm with theoretical approximation guarantees and provide a general greedy solution. The framework jointly optimizes storage efficiency and service latency in multi-edge wireless networks. Results: Simulation results demonstrate that our approach significantly improves cache hit rates over conventional content-based caching, validating the effectiveness and practicality of parameter-level sharing for edge AI deployment.
This study addresses the issue that standard evaluation mechanisms—such as win rate—can induce model homogenization in AI markets, thereby undermining consumer utility. To counter this, the authors propose a weighted win rate mechanism that incentivizes model specialization by offering differentiated rewards for high-quality responses. Drawing on game-theoretic and mechanism design frameworks, the work combines theoretical analysis with empirical validation using real-world benchmark data. The results demonstrate that the proposed mechanism effectively promotes model diversity while significantly enhancing consumer welfare, offering a principled approach to aligning model development incentives with user interests in competitive AI ecosystems.
Traditional scaling law estimation suffers from high computational costs due to the absence of efficient budget allocation strategies. This work proposes a novel approach that, for the first time, integrates surrogate-guided pruning into scaling law modeling by combining the Successive Halving algorithm with both parametric and non-parametric surrogate models. This integration enables proactive allocation of computational resources and efficient construction of loss-compute Pareto frontiers. The method substantially improves resource utilization efficiency, achieving relative performance gains of up to 2.84% on real datasets and 5.47% on synthetic datasets, while reducing computational costs by as much as 98.7%.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.
This work addresses the limited understanding of how implicit biases of optimizers arise during training. Departing from prior analyses focused on the geometry of the solution space, it innovatively shifts attention to the trajectory of parameter updates through the lens of dynamic information allocation. The study introduces a preconditioning exponent \( p \) to characterize the relative distribution of training signals between weight and bias pathways. Using a minimal linear model, it reveals that weight update components preserve input-dependent residual structures, while bias updates capture the mean direction of residuals. The relative strength of these two components is shown to govern both learning dynamics and generalization performance. This framework offers a novel mechanistic perspective on optimizer-induced implicit bias, providing a tunable and interpretable viewpoint for its analysis.