Score
Designs, builds, and evaluates neural network architectures and modular interaction patterns specifically intended to generate or improve model explanations, including attention and feature-interaction components. Uses neural architecture search (including bi-level optimization) to discover cross-attention designs, intra-layer interaction modules, inter-layer connection patterns, and feature-interaction functions that optimize explanation quality and fidelity.
Neural architecture design often relies on heuristic rules or expensive search strategies, lacking a principled, differentiable mapping from performance to structure. Method: This paper proposes an automatic architecture optimization framework grounded in structure–performance mapping modeling. We introduce the Architecture Synthesis Neural Network (ASNN), the first model that takes a performance distribution (e.g., accuracy distribution) as input and differentiably synthesizes high-performing architectural parameters—enabling invertible, generalizable mapping from performance to structure. Leveraging a TensorFlow-based multi-layer network performance dataset, ASNN learns this inverse mapping via neural regression and incorporates an iterative prediction mechanism for progressive refinement. Contribution/Results: On both two- and three-layer networks, ASNN discovers novel architectures surpassing the original dataset’s best-performing models, achieving statistically significant average test accuracy improvements. Experiments validate its effectiveness in architecture recommendation, cross-architecture generalization, and iterative optimization.
Neural architecture search (NAS) suffers from high computational cost, manual design of macro-architectural configurations (e.g., depth and width), and unfairness in existing proxy-based performance estimation due to lack of architecture adaptivity. To address these issues, this paper proposes a globally navigable macro-micro joint search framework. We introduce the first macro-micro decoupled search paradigm, enabling fully automated co-optimization of depth and width via a hybrid search space. Furthermore, we design an architecture-aware dynamic training approximation mechanism that delivers low-overhead, differentiated performance prediction. Our method achieves state-of-the-art results on EMNIST and KMNIST, outperforms prior approaches on CIFAR-10, CIFAR-100, and Fashion-MNIST, accelerates search by 2–4× over the fastest global NAS methods, and successfully transfers to face recognition tasks.
This work addresses the limitations of existing neural architecture search (NAS) methods, which are either confined to narrow predefined search spaces or suffer from inefficiency and bias when leveraging large language models (LLMs) for open-ended exploration. To overcome these challenges, the authors propose a semi-automated NAS framework that constructs a prior-informed, open search space by structurally modeling architectural knowledge extracted from scientific literature. The framework integrates the FairNAD algorithm with multiple fairness-aware mutation mechanisms—including fair sampling, Pareto-aware selection, and LLM-driven iterative refinement—to enable efficient, diverse, and high-quality architecture discovery. Empirical evaluations demonstrate consistent improvements over state-of-the-art methods, achieving accuracy gains of 0.84%, 2.17%, and 2.35% on CIFAR-10, CIFAR-100, and ImageNet16-120, respectively.
To address the prohibitively high manual exploration cost arising from the vast design space of hybrid neural architectures under large-scale pretraining, this paper introduces Composer—the first scalable neural architecture search (NAS) framework tailored for hybrid architectures. Its core innovations are: (i) modular modeling of attention and MLP components, enabling fine-grained architectural customization; and (ii) a novel scaling extrapolation strategy that enables efficient transfer from small-scale search to large models (350M–3B parameters). Evaluated on the Llama 3.2 benchmark, architectures discovered by Composer achieve consistently lower validation loss and yield downstream task accuracy gains of 1.1–3.1 percentage points (up to +8.3%), while maintaining competitive training and inference efficiency. This work establishes the first systematic methodology for efficient NAS and cross-scale generalization in hybrid architectures.
Non-modular neural networks suffer from exponential sample complexity growth with input dimensionality in high-dimensional combinatorial tasks—a fundamental bottleneck for generalization. Method: We theoretically and empirically investigate modular neural networks’ generalization mechanisms. We first establish, for the first time, a rigorous theoretical guarantee that modular architectures achieve dimension-independent sample complexity. Methodologically, we propose a task-intrinsic dimensionality–driven modular architecture design and a theory-guided learning rule. Results: Experiments demonstrate significant improvements over baselines in both in-distribution and out-of-distribution generalization, achieving dimension-agnostic efficient learning on high-dimensional combinatorial tasks—thereby breaking conventional scaling laws. Our core contributions are threefold: (i) establishing the first formal theoretical foundation for modular generalization; (ii) devising a provably optimal, theory-grounded learning mechanism; and (iii) empirically validating its fundamental superiority for high-dimensional combinatorial generalization.
This work addresses the diminished understanding of neural network fundamentals caused by the widespread use of high-level deep learning libraries. To bridge this gap, the authors construct a complete neural network framework from scratch, eschewing automatic differentiation and prebuilt modules. The implementation explicitly details forward and backward propagation, incorporates multiple activation functions, L2 regularization, and advanced optimizers such as Adam. Designed to balance pedagogical clarity with engineering scalability, the framework demonstrates numerical stability, correctness, and generalization capability on multiclass classification tasks. It thus provides a reproducible and extensible tool for both research and instruction, fostering deeper insight into the core principles of deep learning.
This work addresses the ongoing debate regarding the ability of deep neural networks (DNNs) to effectively model high-order feature interactions in recommender systems. By investigating the phenomenon of dimensional collapse in embedding representations, we uncover— for the first time—the core mechanism through which DNNs enhance model expressiveness by mitigating such collapse. Through a combination of gradient-based theoretical analysis, ablation studies, and robustness evaluation across embedding dimensions, we systematically compare parallel and stacked DNN architectures. Our findings demonstrate that both structural variants significantly suppress dimensional collapse, thereby improving the modeling of feature interactions and ultimately boosting recommendation performance.
This work proposes a novel paradigm that synergizes large language models (LLMs) with neural architecture search (NAS) to overcome the limited generalizability of conventional NAS methods, which rely on handcrafted search spaces. The approach begins by leveraging an LLM to generate high-quality initial architectures, which are then transformed into “slot-based” architectures containing replaceable modular components. This enables the automatic construction of task-adaptive, structured search spaces that balance open-ended generation with efficient search. Implemented through a modular three-stage pipeline without any human intervention, the method achieves state-of-the-art performance on 11 out of 17 cross-modal tasks, significantly outperforming existing baselines and expert-designed architectures, thereby demonstrating its generality and effectiveness.
This study addresses two key challenges in large language model–mediated creative generation: premature user fixation on suboptimal ideas due to loosely structured outputs, and the lack of fine-grained combinatorial control—termed “combinatorial opacity”—in existing tools. To overcome these limitations, the authors propose a computational pipeline called Cognitive Abstraction, which transforms raw generative outputs into a navigable, transformable design space through functional decomposition, multi-level abstraction, and cross-dimensional recombination. Integrated into the NexusAI system, this approach enables effective human-AI collaborative exploration. The work formally conceptualizes combinatorial opacity as a critical barrier in creative collaboration and introduces a scalable framework of cognitive primitives for creative operations. A user study (N=14) demonstrates that NexusAI significantly enhances exploratory breadth, reduces cognitive load, and facilitates perspective reframing compared to baseline systems.