explanation architecture search

Designs, builds, and evaluates neural network architectures and modular interaction patterns specifically intended to generate or improve model explanations, including attention and feature-interaction components. Uses neural architecture search (including bi-level optimization) to discover cross-attention designs, intra-layer interaction modules, inter-layer connection patterns, and feature-interaction functions that optimize explanation quality and fidelity.

explanationarchitecturesearch

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.54
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Neural architecture design often relies on heuristic rules or expensive search strategies, lacking a principled, differentiable mapping from performance to structure. Method: This paper proposes an automatic architecture optimization framework grounded in structure–performance mapping modeling. We introduce the Architecture Synthesis Neural Network (ASNN), the first model that takes a performance distribution (e.g., accuracy distribution) as input and differentiably synthesizes high-performing architectural parameters—enabling invertible, generalizable mapping from performance to structure. Leveraging a TensorFlow-based multi-layer network performance dataset, ASNN learns this inverse mapping via neural regression and incorporates an iterative prediction mechanism for progressive refinement. Contribution/Results: On both two- and three-layer networks, ASNN discovers novel architectures surpassing the original dataset’s best-performing models, achieving statistically significant average test accuracy improvements. Experiments validate its effectiveness in architecture recommendation, cross-architecture generalization, and iterative optimization.

ASNN automates neural network design efficientlyASNN learns relationship between NN architecture and accuracyASNN suggests improved architectures for better performance

Efficient Global Neural Architecture Search

Feb 05, 2025
SS
Shahid Siddiqui
🏛️ University of Cyprus

Neural architecture search (NAS) suffers from high computational cost, manual design of macro-architectural configurations (e.g., depth and width), and unfairness in existing proxy-based performance estimation due to lack of architecture adaptivity. To address these issues, this paper proposes a globally navigable macro-micro joint search framework. We introduce the first macro-micro decoupled search paradigm, enabling fully automated co-optimization of depth and width via a hybrid search space. Furthermore, we design an architecture-aware dynamic training approximation mechanism that delivers low-overhead, differentiated performance prediction. Our method achieves state-of-the-art results on EMNIST and KMNIST, outperforms prior approaches on CIFAR-10, CIFAR-100, and Fashion-MNIST, accelerates search by 2–4× over the fastest global NAS methods, and successfully transfers to face recognition tasks.

Addresses unfair network comparisons in performance evaluationsAutomates global neural architecture search for efficiencyEnhances NAS with macro-micro search space design

This work addresses the limitations of existing neural architecture search (NAS) methods, which are either confined to narrow predefined search spaces or suffer from inefficiency and bias when leveraging large language models (LLMs) for open-ended exploration. To overcome these challenges, the authors propose a semi-automated NAS framework that constructs a prior-informed, open search space by structurally modeling architectural knowledge extracted from scientific literature. The framework integrates the FairNAD algorithm with multiple fairness-aware mutation mechanisms—including fair sampling, Pareto-aware selection, and LLM-driven iterative refinement—to enable efficient, diverse, and high-quality architecture discovery. Empirical evaluations demonstrate consistent improvements over state-of-the-art methods, achieving accuracy gains of 0.84%, 2.17%, and 2.35% on CIFAR-10, CIFAR-100, and ImageNet16-120, respectively.

Design Knowledge StructuringLarge Language ModelsNeural Architecture Search

Composer: A Search Framework for Hybrid Neural Architecture Design

Sep 30, 2025
BA
Bilge Acun
🏛️ FAIR at Meta | The University of Texas at Austin | Meta

To address the prohibitively high manual exploration cost arising from the vast design space of hybrid neural architectures under large-scale pretraining, this paper introduces Composer—the first scalable neural architecture search (NAS) framework tailored for hybrid architectures. Its core innovations are: (i) modular modeling of attention and MLP components, enabling fine-grained architectural customization; and (ii) a novel scaling extrapolation strategy that enables efficient transfer from small-scale search to large models (350M–3B parameters). Evaluated on the Llama 3.2 benchmark, architectures discovered by Composer achieve consistently lower validation loss and yield downstream task accuracy gains of 1.1–3.1 percentage points (up to +8.3%), while maintaining competitive training and inference efficiency. This work establishes the first systematic methodology for efficient NAS and cross-scale generalization in hybrid architectures.

Automating hybrid neural architecture search for pre-trainingExploring interleaving ratios of computational primitives efficientlyScaling discovered architectures to outperform existing LLM models

Breaking Neural Network Scaling Laws with Modularity

Sep 09, 2024
AB
Akhilan Boopathy
🏛️ Massachusetts Institute of Technology

Non-modular neural networks suffer from exponential sample complexity growth with input dimensionality in high-dimensional combinatorial tasks—a fundamental bottleneck for generalization. Method: We theoretically and empirically investigate modular neural networks’ generalization mechanisms. We first establish, for the first time, a rigorous theoretical guarantee that modular architectures achieve dimension-independent sample complexity. Methodologically, we propose a task-intrinsic dimensionality–driven modular architecture design and a theory-guided learning rule. Results: Experiments demonstrate significant improvements over baselines in both in-distribution and out-of-distribution generalization, achieving dimension-agnostic efficient learning on high-dimensional combinatorial tasks—thereby breaking conventional scaling laws. Our core contributions are threefold: (i) establishing the first formal theoretical foundation for modular generalization; (ii) devising a provably optimal, theory-grounded learning mechanism; and (iii) empirically validating its fundamental superiority for high-dimensional combinatorial generalization.

Demonstrates modular networks' efficiency in high-dimensional tasks.Develops a learning rule for modular networks to enhance generalization.Explains how modularity improves neural network generalizability.

Latest Papers

What's happening recently
View more

This work addresses the diminished understanding of neural network fundamentals caused by the widespread use of high-level deep learning libraries. To bridge this gap, the authors construct a complete neural network framework from scratch, eschewing automatic differentiation and prebuilt modules. The implementation explicitly details forward and backward propagation, incorporates multiple activation functions, L2 regularization, and advanced optimizers such as Adam. Designed to balance pedagogical clarity with engineering scalability, the framework demonstrates numerical stability, correctness, and generalization capability on multiclass classification tasks. It thus provides a reproducible and extensible tool for both research and instruction, fostering deeper insight into the core principles of deep learning.

deep learning librarieseducational gapfundamental understanding

This work addresses the ongoing debate regarding the ability of deep neural networks (DNNs) to effectively model high-order feature interactions in recommender systems. By investigating the phenomenon of dimensional collapse in embedding representations, we uncover— for the first time—the core mechanism through which DNNs enhance model expressiveness by mitigating such collapse. Through a combination of gradient-based theoretical analysis, ablation studies, and robustness evaluation across embedding dimensions, we systematically compare parallel and stacked DNN architectures. Our findings demonstrate that both structural variants significantly suppress dimensional collapse, thereby improving the modeling of feature interactions and ultimately boosting recommendation performance.

dimensional collapseDNNsfeature interaction

This work proposes a novel paradigm that synergizes large language models (LLMs) with neural architecture search (NAS) to overcome the limited generalizability of conventional NAS methods, which rely on handcrafted search spaces. The approach begins by leveraging an LLM to generate high-quality initial architectures, which are then transformed into “slot-based” architectures containing replaceable modular components. This enables the automatic construction of task-adaptive, structured search spaces that balance open-ended generation with efficient search. Implemented through a modular three-stage pipeline without any human intervention, the method achieves state-of-the-art performance on 11 out of 17 cross-modal tasks, significantly outperforming existing baselines and expert-designed architectures, thereby demonstrating its generality and effectiveness.

Architecture GenerationLarge Language ModelsNeural Architecture Search

This study addresses two key challenges in large language model–mediated creative generation: premature user fixation on suboptimal ideas due to loosely structured outputs, and the lack of fine-grained combinatorial control—termed “combinatorial opacity”—in existing tools. To overcome these limitations, the authors propose a computational pipeline called Cognitive Abstraction, which transforms raw generative outputs into a navigable, transformable design space through functional decomposition, multi-level abstraction, and cross-dimensional recombination. Integrated into the NexusAI system, this approach enables effective human-AI collaborative exploration. The work formally conceptualizes combinatorial opacity as a critical barrier in creative collaboration and introduces a scalable framework of cognitive primitives for creative operations. A user study (N=14) demonstrates that NexusAI significantly enhances exploratory breadth, reduces cognitive load, and facilitates perspective reframing compared to baseline systems.

compositional opacitycreative ideationdesign space exploration

Hot Scholars

YZ

Yue Zhao

Assistant Professor of Computer Science, University of Southern California
Anomaly DetectionOut-of-Distribution DetectionTrustworthy AIAI for Science
JL

Jiate Li

University of Southern California
YD

Yushun Dong

Assistant Professor, Department of Computer Science, Florida State University
AI SecurityAI IntegrityGraph Machine LearningLLMs
MV

Michalis Vlachos

Professor, HEC Lausanne, University of Lausanne
Recommender SystemsAI in EducationTime Series MiningDigital Watermarking