🤖 AI Summary
Neural architecture search (NAS) remains hindered by heavy reliance on human expertise, predefined search spaces, and computationally expensive training-based evaluation. Method: This paper introduces the first end-to-end, Transformer-based neural architecture performance prediction framework. Its core innovation lies in fully tokenizing entire network architectures—integrating edge-free graph encoding with sequential representation—to enable zero-shot, training-free performance regression. Contribution/Results: The approach eliminates conventional NAS constraints on search space design and supports generative architecture exploration. Evaluated on DeepNets-1M, it achieves state-of-the-art performance prediction accuracy, reducing mean absolute error by 18.7% over the best proxy models and graph neural network baselines. This work establishes a new paradigm for fully automated, highly innovative neural network design.
📝 Abstract
In the realm of neural architecture design, achieving high performance is largely reliant on the manual expertise of researchers. Despite the emergence of Neural Architecture Search (NAS) as a promising technique for automating this process, current NAS methods still require human input to expand the search space and cannot generate new architectures. This paper explores the potential of Transformers in comprehending neural architectures and their performance, with the objective of establishing the foundation for utilizing Transformers to generate novel networks. We propose the Token-based Architecture Transformer (TART), which predicts neural network performance without the need to train candidate networks. TART attains state-of-the-art performance on the DeepNets-1M dataset for performance prediction tasks without edge information, indicating the potential of Transformers to aid in discovering novel and high-performing neural architectures.