LP-NAS: Linear Programming-based Neural Architecture Search

2026-08-14Machine Learning

Machine LearningArtificial Intelligence
AI summary

The authors propose a new method called LP-NAS to improve how computers automatically design neural networks. Their approach uses math called linear programming to better guide the search for good network designs, making the process faster and more effective. They tested their method on popular image tasks and showed it finds better architectures quicker than existing methods. They also confirmed that their discovered designs work well on bigger, real-world data like ImageNet.

Neural Architecture SearchDifferentiable NASLinear ProgrammingValidation LossTraining LossGradientHessianDARTSCIFAR-10ImageNet
Authors
Abhishek Shukla, Ankur Sinha, Faiz Hamid
Abstract
Neural Architecture Search (NAS) aims to automate neural network architecture design, reducing reliance on human expertise. Among the various NAS methods, differentiable NAS has gained prominence due to its efficiency and accuracy compared to conventional NAS approaches. Since differentiable NAS relaxes the architecture search space into a continuous domain, it is possible to apply principles from continuous optimization to NAS. In this paper, we propose Linear Programming-based NAS (LP-NAS), a mathematical programming-based framework for differentiable NAS that is applicable to a wide range of continuous search spaces. LP-NAS formulates a linear program (LP) using the validation-loss gradient and the training-loss Hessian to compute an architecture update direction that improves generalization while preserving the optimality of the model parameters. By following this LP-derived descent direction, LP-NAS efficiently navigates the architecture search space, leading to faster and more effective architecture optimization. We introduce two computationally efficient variants of LP-NAS, namely S-LP-NAS and R-LP-NAS. Applying LP-NAS to the Differentiable Architecture Search (DARTS) search space results in two algorithmic variants, S-LP-DARTS and R-LP-DARTS. Both variants achieve faster convergence and significantly higher validation performance during the early search iterations than the standard DARTS algorithm. Extensive experiments on CIFAR-10 and CIFAR-100 show that LP-DARTS outperforms standard DARTS in both the architecture search and evaluation phases. Additionally, we compare our approach with several DARTS variants (P-DARTS, PC-DARTS, and STO-DARTS) on the CIFAR-10 dataset and demonstrate its effectiveness. Furthermore, we validate the transferability of the discovered architectures through experiments on the ImageNet dataset.