Neural Architecture Search (NAS) is an automated machine learning technique that searches a defined space of neural network designs to discover architectures that maximise predictive performance or satisfy multi-objective constraints such as latency, parameter count, and energy consumption. NAS algorithms explore the architecture search space using strategies that include reinforcement learning, evolutionary algorithms, differentiable relaxations (DARTS), and predictor-based approaches, each making different trade-offs between search cost and solution quality. The field emerged from the observation that hand-designed architectures require substantial expert knowledge and iterative experimentation, and that automated search can discover non-obvious configurations that outperform human-designed baselines on targeted hardware or task distributions.
Content
- NAS was popularised by the 2017 paper “Neural Architecture Search with Reinforcement Learning” by Zoph and Le at Google Brain, which used an RNN controller trained with REINFORCE to generate CNN cell specifications, discovering architectures competitive with hand-tuned designs on CIFAR-10. The enormous computational cost of early NAS (800 GPUs for 28 days) prompted research into efficiency improvements: weight sharing (ENAS, 2018), one-shot supernets, and differentiable architecture search (DARTS, 2019) reduced search to hours on a single GPU and democratised the approach.
- The NAS pipeline comprises three interdependent components: (1) the search space, which defines the set of possible operations (convolutions, attention heads, pooling) and their wiring patterns (cell-based, graph-based, hierarchical); (2) the search strategy, which determines how candidates are sampled or optimised (RL, evolutionary, gradient-based, Bayesian optimisation, predictor networks); and (3) the performance estimation strategy, which estimates architecture quality without fully training each candidate (weight inheritance, learning curve extrapolation, zero-cost proxies). Hardware-aware NAS adds latency or energy measurements from target devices as constraints or secondary objectives.
- NAS has produced architectures adopted widely in production: EfficientNet (scalable image classifiers), MobileNetV3 (mobile-optimised inference), and ProxylessNAS for embedded deployment. In the transformer era, NAS has extended to searching attention head configurations, feed-forward dimensions, and layer-sharing policies. Multi-task NAS finds single architectures that perform well across several tasks simultaneously, reducing the deployment footprint in resource-constrained systems.
- In 2024–2025 NAS research has shifted toward searching within the space of pre-trained large foundation models—adapting layer granularity, attention sparsity, and mixture-of-experts routing rather than designing from scratch. Zero-cost proxies using gradient statistics at initialisation enable architecture ranking without any training, making NAS practical for rapid hardware-specific co-design. The intersection of NAS with diffusion model architecture and state-space models (Mamba, Hawk) represents an active frontier for discovering efficient sequence modelling configurations.