YOLO on NPU
Compressing a YOLOv5 detector until it fits the power and memory budget of an embedded NPU, then getting it to actually run there.
Optimisation notebookModel and deployment under NDA
3×
Smaller
Fewer parameters and FLOPs than the YOLOv5s baseline.
−2.8
mAP50-95
Accuracy cost of the full compression chain, on COCO val2017.
5.6×
Faster
On-device throughput, from 25 to 140 FPS.
i.MX 8M+
Target
NXP NPU on a Phytec board, within its power and memory budget.
Overview
Making YOLO fit on an embedded NPU
This was my final-year project at Télécom SudParis, in the HTI specialisation, which I did with Ahmed El Habib Izi for MBDA. The goal was to take a YOLOv5 object detector and run it on the NPU of a Phytec i.MX 8M Plus board, without going over the board’s power and memory budget.
On paper, YOLO is a good fit for an NPU: it is mostly convolutions, batch normalisation and max-pooling, which is exactly what the accelerator is built for. In practice the NPU only runs INT8 operations, and the spec set targets for latency, accuracy and model size all at once. Every optimisation helped one of those and cost us on another.
We split the work into five steps: benchmark a baseline on CPU, compress it, check what the compression cost, deploy it to the NPU, and compare the NPU against the CPU. Along with the model, we delivered our own pipeline tooling and enough documentation for someone else to reproduce the whole thing.
Compression
Pruning and quantisation, measured one at a time
We started from YOLOv5s, the smallest model in the family, pretrained on COCO 2017. It is quick to iterate on and light enough for embedded work. Before changing anything, we benchmarked it on CPU in both PyTorch and TensorFlow Lite so we had a fixed reference for every later result.
For pruning we used structured pruning in PyTorch, which removes whole channels and filters so the tensors actually get smaller. Unstructured pruning only zeroes out individual weights, and that does nothing for speed on hardware without sparsity support, which includes this NPU. We swept the pruning ratio from 20% to 40%. Accuracy held up until 30% and then dropped off quickly, so we stopped there.
For quantisation we compared post-training quantisation, static quantisation and quantisation-aware training, and went with static quantisation in TensorFlow Lite. It calibrates activation ranges on a sample of real data and then converts weights and activations from float32 to INT8, the only format the NPU runs. Together, pruning and quantisation cut parameters and FLOPs by 3× for a loss of 2.8 points of mAP50-95.
Deployment
When hardware stops cooperating
Neither Phytec’s nor NXP’s deployment guides (eIQ, OpenVINO) matched our setup, so we built our own conversion pipeline. The trained PyTorch model is exported to a TensorFlow SavedModel and then converted to an INT8 TensorFlow Lite file.
On the board, a Yocto-built BSP provides the kernel, the NPU drivers and the TFLite runtime, and NXP’s VX delegate hands the model to the accelerator. If the versions of any of these don’t line up, inference quietly falls back to the CPU and nothing tells you. We used TFLite’s benchmark tool to confirm that the whole graph was really running on the NPU.
The hardest part was the output. A quantised model doesn’t return bounding boxes, it returns raw INT8 tensors that have to be dequantised and decoded on the host. Our first runs on the NPU produced piles of spurious detections with broken box coordinates, and fixing the decoding took far more of our time than we had planned for.
Once it was running properly, we measured throughput, power and memory on the device. Throughput went from 25 to 140 FPS, a 5.6× speedup, and the model stayed within the platform’s power and memory budget.