Score
Designs and specifies the structure of computational models, including their layers, modules, connectivity patterns, parameterizations, and component choices (e.g., attention, convolution, residual connections, encoders/decoders). Builds and evaluates architectures for capacity, efficiency, generalization, and deployment trade-offs, and analyzes how architectural choices affect learning dynamics, representations, and task performance.
This paper addresses the fundamental problem of how internal neural network architecture choices govern training dynamics. We propose an enhanced transformation layer featuring constrained signal paths and adaptive correction mechanisms, grounded in spectral sensitivity analysis and fixed-point theory to derive interpretable, principled structural design guidelines. These guidelines systematically uncover intrinsic connections among gradient flow characteristics, representation regularity, and training stability. Methodologically, we conduct empirical studies using both synthetic and structured tasks to validate the design. Results demonstrate substantial improvements in optimization smoothness, generalization robustness, and depth scalability. Our approach validates a paradigm shift from heuristic performance tuning toward deliberate learning dynamic control, offering novel theoretical foundations and practical design principles for neural architecture engineering.
To address excessive computational overhead when deploying vision models on resource-constrained devices, this paper proposes an efficient Vision Transformer (ViT) architecture design framework. Methodologically: (1) it optimizes the input-output data pathway to enhance representational capacity of lightweight models; (2) it restructures the context window of computationally constrained attention mechanisms to improve local-global modeling efficiency; and (3) it leverages the invertibility and explicit probabilistic modeling properties of normalizing flows to enable high-fidelity, low-overhead knowledge distillation. Experiments demonstrate that the proposed approach achieves comparable or superior accuracy on benchmarks such as ImageNet, while requiring significantly fewer parameters and FLOPs. It also substantially reduces inference latency and memory footprint. The framework establishes a scalable new paradigm for efficient visual understanding at the edge.
This study addresses the lack of general design principles linking neural network architecture to computational capacity. By systematically evaluating the computational performance of recurrent neural networks with diverse connectivity patterns on Boolean function tasks through large-scale sampling, the work reveals— for the first time—that local 2-cycles and 3-cycles are critical structural motifs for enhancing computational power. It further demonstrates that introducing a small number of sparse connections and biologically inspired interneuron-like units significantly boosts the performance of large-scale networks. The authors construct a comprehensive performance map linking small-network architectures to Boolean function realization, showing that networks containing short cycles achieve optimal performance. Moreover, network performance can be accurately predicted from structural statistics, offering a theoretical foundation and biologically inspired guidance for future neural architecture design.
Machine learning systems lack effective, quantifiable methods to assess how architectural design patterns impact scalability, performance, and cost—leading to subjective, evidence-deficient architecture selection. Method: This paper introduces the first quantitative architectural evaluation framework tailored for ML systems, specifically targeting CPU-based inference. It explicitly models the mappings between architectural patterns and key quality attributes: latency, throughput, and resource overhead. The framework integrates metric-driven modeling, lightweight observability analysis, and standardized benchmarking procedures. Contribution/Results: Evaluated across multiple case studies, the framework enables objective, quantitative ranking of architectural patterns—achieving up to 2.3× higher inference throughput and an average 37% improvement in CPU utilization. It further supports cost-optimized decision-making in production environments. Its core contribution is the establishment of the first rigorous, quantification-oriented paradigm for evaluating ML architectural patterns.
Configuration space explosion complicates performance impact modeling, while gray-box approaches rely on structural knowledge (e.g., module execution graphs) to improve model accuracy—yet the mechanisms by which structural features (e.g., number of modules or configuration options) and structural knowledge influence modeling difficulty and optimization potential remain unclear. Method: We formally define “modeling hardness” and “improvement opportunity,” establishing an analytical framework and matrix to quantify the interplay among system structural complexity, structural knowledge level, and modeling benefit. Controlled experiments on synthetic systems integrate module execution graph analysis with gray-box modeling. Contribution/Results: We identify module count and configuration option count as dominant determinants of modeling hardness. Under high hardness, strong structural knowledge significantly increases improvement opportunity. Structural knowledge primarily enhances ranking accuracy, whereas hardness predominantly degrades prediction accuracy. Our findings provide theoretical foundations and strategic guidance for allocating structural knowledge investment according to specific modeling objectives.
This study investigates differences in neural activation patterns across diverse cognitive tasks among various large language model architectures. Employing a unified framework, the authors systematically analyze final-layer activations, attention entropy, and sparsity across six prominent architectures on twelve task categories, yielding 144 task–model combinations. The work reveals, for the first time, a fundamental distinction between encoder- and decoder-based models in their task-processing mechanisms: mathematical reasoning consistently elicits the highest attention entropy, while decoder-only models exhibit significantly greater activation sparsity. These findings demonstrate the joint influence of architecture type and task category on internal representations, providing empirical guidance for model selection and optimization in large-scale data scenarios.
This study investigates the mechanistic link between scaling laws in large model behavior and the emergence of internal representational structure. By training small Transformers on a controlled sequence modeling task—predicting outputs from a hidden Markov model—and employing residual activation linear encodings alongside probability simplex probing techniques, the work reveals for the first time a predictable correspondence between performance scaling with model size and the geometric evolution of internal belief distributions. This finding provides crucial empirical evidence for understanding the intrinsic mechanisms underlying scaling laws and underscores the central role of internal representational geometry in the qualitative leaps of model capabilities.
This work challenges the prevailing reliance on monolithic architectures—particularly the Transformer—in contemporary AI systems, which overlooks neuroscientific evidence that diverse cognitive functions emerge from heterogeneous, interacting brain regions. The study systematically argues that the Transformer more accurately models hippocampal function rather than serving as a universal cortical analog. Building on this insight and inspired by the structural and functional heterogeneity of the cerebral cortex, the authors propose a modular, heterogeneous network architecture wherein each module embodies a distinct inductive bias and communicates with others through standardized interfaces. Integrating principles from cytoarchitectonics, functionalism, and modular design, this approach establishes a novel paradigm for AI architecture that enhances generalization while reducing data dependency, thereby reclaiming the benefits of efficient, biologically informed inductive biases.
This study addresses the prevailing gap in AI education, which emphasizes model development while neglecting system engineering practices, leaving students ill-equipped to handle real-world challenges such as architectural design, deployment, and monitoring. To bridge this gap, the authors implemented a master’s-level course in which students built a movie recommendation system under realistic constraints, with a focus on integrating AI components into robust software systems, adopting data-driven machine learning practices, and cultivating systems-level thinking. Using a mixed-methods approach—combining analysis of student project artifacts with survey data—the research evaluates learners’ performance in architectural decision-making, integration of heterogeneous models, and adaptation to evolving requirements. Findings reveal common difficulties students encounter in AI system engineering and demonstrate the course’s effectiveness in addressing critical deficiencies in AI engineering education and enhancing systems-aware competencies.