DPI Computing SocietyDPI Computing SocietyLearn, build, and grow together
Research

Exploring Neural Architecture Search (NAS) for Edge AI Devices

How automated model search techniques optimize neural network topology for resource-constrained microcontrollers and edge processors without sacrificing classification accuracy.

TTahmid Hasan
1 min read
Exploring Neural Architecture Search (NAS) for Edge AI Devices

The Edge Intelligence Challenge

Deep neural networks (DNNs) have achieved superhuman benchmarks across image classification, speech transcription, and generative modeling. However, state-of-the-art models typically require gigabytes of VRAM and hundreds of Watts of power.

Deploying deep learning models onto edge hardware—such as ARM Cortex-M microcontrollers, ESP32 boards, or mobile NPUs—presents severe constraints:

  • Strict RAM Limits: Frequently under 512KB of SRAM.
  • Flash Storage Constraints: Less than 4MB to 16MB of persistent memory.
  • Thermal and Battery Ceilings: Maximum power draw under 1 to 2 Watts.

What is Neural Architecture Search (NAS)?

Historically, neural network architectures (ResNet, VGG, MobileNet) were handcrafted by human researchers through painstaking trial and error. Neural Architecture Search (NAS) automates this discovery by framing architecture design as an optimization problem:

┌──────────────────┐
│   Search Space   │ (Filter sizes, connections, layers)
└────────┬─────────┘
         ▼
┌──────────────────┐
│ Search Strategy  │ (Reinforcement Learning / Evolutionary / Differentiable)
└────────┬─────────┘
         ▼
┌──────────────────┐
│Performance Eval  │ (Multi-objective: Accuracy + Latency on target hardware)
└──────────────────┘

Hardware-Aware Multi-Objective Optimization

Modern NAS systems do not optimize solely for top-1 test accuracy. Instead, they incorporate real hardware profiling:

$$\mathcal{L}(A) = -\text{Acc}(A) + \lambda \cdot \left(\frac{\text{Latency}(A)}{\text{TargetLatency}}\right)^\beta$$

Where:

  • $\text{Acc}(A)$ is the validation accuracy of candidate architecture $A$.
  • $\text{Latency}(A)$ is the actual measured inference time on the target microcontroller.
  • $\lambda$ and $\beta$ are Pareto weight hyperparameters.

By penalizing candidate layers that violate hardware memory alignment or cache line bounds, the search engine outputs networks that execute with minimum pipeline stalls.


Future Research Directions at DPICS

Our research wing is currently exploring Once-for-All (OFA) networks paired with 4-bit integer quantization (INT4). By decoupling model training from hardware adaptation, we hope to deploy real-time acoustic anomaly detection models for industrial safety monitoring on inexpensive microcontrollers.