AMD Instinct accelerators

AMD Instinct on a machine nobody else is on

Two AMD accelerators go into our dedicated servers: the Instinct MI210 and the Instinct MI50. The card is passed through to your own operating system — your ROCm build, your kernel, your processes, and no hypervisor in between.

MI210: 104 CDNA 2 compute units, 64 GB HBM2e. MI50: 60 GCN 5.1 compute units, 16 GB. One tenant per machine. Not sliced, not time-shared.

Every card we fit — AMD and NVIDIA, current and long out of production — is listed with its own monthly price on the GPU catalogue.

Turn your GPUs into passive monthly revenue

Got idle server or desktop GPU setups? List them on the Primcast marketplace today and earn steady monthly rents from AI teams, developers, and enterprises needing production-grade compute.

Go to Marketplace

What an AMD Instinct server here actually is

A physical server rented to one account, with an AMD accelerator in a PCIe slot and root on the operating system above it. The card is a line on the invoice rather than a product tier, so the machine underneath does not get more expensive because of what is plugged into it.

The card is yours

No partition, no vGPU, no scheduler deciding when your turn is. The accelerator is handed to your kernel directly, so you pin your own ROCm version and load your own modules without asking us.

ROCm, HIP and the usual frameworks

Both parts are supported ROCm targets — gfx90a for the MI210, gfx906 for the MI50 — which is what PyTorch, TensorFlow, JAX and ONNX Runtime build against on AMD. HIP ports CUDA source across rather than requiring a rewrite.

Priced apart from the machine

The accelerator carries its own monthly figure and the server carries another, and you add the two. Swap the card without changing servers. Both figures, for every card in the catalogue, are on the GPU page.

The two AMD Instinct accelerators we fit

Both are orderable today from the same configurator, on different server lines. The figures below are AMD's own published ones.

AMD Instinct MI210 accelerator

AMD Instinct MI210

The CDNA 2 generation in a standard full-height, full-length, dual-slot PCIe card — 104 compute units, 64 GB of HBM2e at up to 1.6 TB/s, and up to three Infinity Fabric links for direct card-to-card traffic. Passively cooled at 300 W over a PCIe Gen 4 host link. This is the one to ask for when the whole question is whether the model fits in memory.

AMD Instinct MI50

The older Vega-generation part: GCN 5.1 silicon, 60 compute units and 16 GB of on-package memory. Nobody rents this card any more, which makes it unobtainable elsewhere rather than merely cheap here — and it is still a supported ROCm target, so it runs the same HIP and PyTorch code as the MI210 for a fraction of the monthly figure. It is the sensible way to get a ROCm toolchain working before committing to the big card.

Matrix throughput on the MI210

AMD publishes 181.0 TFLOPS peak FP16 and BF16 matrix performance for the MI210 at its 1,700 MHz peak boost engine clock, alongside 22.6 TFLOPS peak FP64 and FP32 vector.

Memory is the constraint

64 GB holds a 30-billion-parameter model's weights at 8-bit, or a 13-billion one at 16-bit, with room left over for the KV cache. The GPU page works that arithmetic through for each model size.

What the host gives the card

Host lane generation is the fact almost nobody publishes and the one that decides whether a card runs at its design bandwidth. We publish ours next to every card, on the GPU page.

Nothing counts the bytes

Weights in, checkpoints out, datasets in. The port is unmetered, so a training run that streams a corpus does not arrive as a bandwidth bill at the end of the month.

HPE enterprise grade dedicated servers powered by AMD Instinct™

HPE enterprise

Your AMD INSTINCT™ GPU dedicated server is powered by HPE Enterprise servers, ensuring stable performance for the most demanding workloads.

Hardware upgrades

Easily add resources or additional servers to your server infrastructure. Most upgrades are processed within 24 hours.

24/7 support

Dedicated server experts are available to assist 24/7 via live chat and email.

The MI210 against the NVIDIA cards we also fit

Four accelerators out of the same order catalogue, so this compares parts you can actually put in the same basket. The monthly figures in the last row are read from that catalogue when the page loads.

MI210 L40S A100 H100
GPU Architecture CDNA 2 Ada Lovelace NVIDIA Ampere Hopper
GPU Memory 64GB HBM2e 48GB GDDR6 with ECC 80GB HBM2e 80GB HBM3
GPU Memory Bandwidth Up to 1.6 TB/s 864 GB/s 1935 GB/s 3352 GB/s
FP32 22.6 TFLOPS 91.6 TFLOPS 19.5 TFLOPS 51 TFLOPS
TF32 Tensor Core 366 TFLOPS 312 TFLOPS 756 TFLOPS
FP16/BF16 Matrix 181 TFLOPS 733 TFLOPS 624 TFLOPS 1513 TFLOPS
Power Up to 300W Up to 350W Up to 400W Up to 350W
Loading... Loading... Loading... Loading...

TF32 is an NVIDIA tensor-core format. CDNA 2 does not implement it and AMD publishes no figure for it, so the MI210 column shows nothing there rather than a number borrowed from another vendor. The NVIDIA tensor-core rows are the published figures with sparsity. Every other card we fit, the MI50 included, is priced on the GPU catalogue.

FAQ about AMD Instinct GPU servers

Common questions about deploying and managing AMD Instinct accelerators on bare metal for AI, HPC, and machine learning workloads.

Which AMD Instinct accelerators can I actually order?

Two: the Instinct MI210 and the Instinct MI50. Those are the AMD accelerators in our order catalogue, and the configurator will not offer you anything else from the Instinct line. The MI210 goes on three of our server lines and the MI50 on two. We do not fit the MI250, the MI250X or the MI300A, and never have.

What is the difference between the MI210 and the MI50?

Generation and memory. The MI210 is AMD's CDNA 2 architecture with 104 compute units and 64 GB of HBM2e at up to 1.6 TB/s, in a dual-slot PCIe card drawing up to 300 W over a PCIe Gen 4 host link, with up to three Infinity Fabric links for card-to-card traffic. The MI50 is the earlier GCN 5.1 generation with 60 compute units and 16 GB. Both are supported ROCm targets, so the same HIP and PyTorch code runs on either — the MI210 simply holds four times as much of a model.

How much card memory does the model I want to run need?

The weights alone are parameters multiplied by bytes per parameter: roughly 16 GB for a 7-8 billion parameter model at 16-bit, 28 GB for 13-14 billion, 68 GB for 30-34 billion. The KV cache and the runtime go on top of that, which is why a 16 GB card does not comfortably serve a model whose weights are 16 GB. The MI210's 64 GB covers a 30-billion model quantised to 8-bit with room to work in. The GPU page has the whole table.

How long does it take to deploy an AMD Instinct GPU server?

Where the machine is already racked it is typically activated about 30 minutes after payment clears. Where it is not, it is built to order and the lead time depends on the parts. Ask before you pay if the date matters — instant servers is the page for what is standing in a rack right now. Every server includes instant OS reload, so you can iterate without re-opening a ticket.

What software and frameworks run on these cards?

ROCm, AMD's open-source GPU compute platform. Both parts are supported ROCm targets — gfx90a on the MI210, gfx906 on the MI50 — which covers PyTorch, TensorFlow, JAX and ONNX Runtime, together with the BLAS, FFT, RNG and deep-learning primitive libraries. HIP ports existing CUDA source across rather than requiring a rewrite. Because you have root on bare metal you choose the ROCm and kernel versions yourself, and can run the whole thing in Docker or under Kubernetes.