Edge AI accelerator modules on a dedicated server

Edge TPU & AI accelerator modules

A Google Coral Edge TPU or a Hailo-8, fitted as an M.2 module to a dedicated server of your own — for running small, quantised models next to the data they watch, not for training large ones.
See the machines

What a module this size is for — and what it is not:

Vision inference

Object detection, classification and segmentation are what these parts were built for. The networks are small and int8-quantised — the MobileNet and EfficientDet class, compiled for the module ahead of time — and they answer in milliseconds on the machine that holds the frames.

Always-on analytics

Both modules draw single-digit watts, so a camera feed or a sensor stream can be scored around the clock without occupying the host processor. The cores stay free for your application; the module does the arithmetic.

Not language models

A model measured in billions of parameters does not fit either module, at any quantisation, and this page will not pretend otherwise. Sizing a language model is its own arithmetic and lives at what it takes to run a model; the discrete cards that do run one, each priced beside its machine, are at GPU dedicated servers.

Coral M.2 Accelerator

Google’s Coral module carries the Edge TPU, which Google rates at 4 trillion operations per second on 8-bit integers at around two watts. Those are the vendor’s figures, quoted as published — not a measurement of ours.


It runs TensorFlow Lite models that have been quantised to int8 and compiled for the Edge TPU ahead of time. That compile step is the honest cost of the part: a model either compiles for it or it does not, and the classic edge networks — classification, detection, pose — are the ones that do.

Coral Accelerator

Hailo-8 M.2 2280 module

Hailo rates the Hailo-8 at up to 26 tera-operations per second on 8-bit integers — again the maker’s figure, quoted as published. It is the larger of the two modules by some way, and the one to ask for when a Coral-sized network is not enough.


Models reach it from TensorFlow or ONNX through Hailo’s own compiler, and its runtime serves them on your machine. As with the Coral, the compile step decides what runs: this is a neural-network part, not general-purpose compute.

Hailo-8 Module
Its own line on the order form

The module is not a product tier: it is one line on the cartridge’s configurator, added to the machine’s own figure. Neither number is typed here, because a price belongs beside the machine it comes with — the same rule every page in this family follows.

The cartridges that take it

The M.2 slot rides an HPE Moonshot cartridge. Configure the m710x (quad-core Xeon E3-1585L v5) or the m510 (Xeon D, eight or sixteen cores) and add the module as an option.

One tenant, your toolchain

The machine is yours alone, with root. The vendor’s runtime — TensorFlow Lite with the Edge TPU library for the Coral, Hailo’s HailoRT for the Hailo-8 — installs on your operating system and stays at the version you pin. Nothing upgrades underneath you.

When the model outgrows it

The step up is not a bigger module; it is a discrete card in a machine built to take one. GPU dedicated servers is every card we fit, priced beside its machine; multi-GPU and multi-node is more than one.

Fit an accelerator to a cartridge of your own.

Configure a server