RTX PRO 6000 Blackwell · 96 GB GDDR7 with ECC

RTX PRO 6000 in a server you rent by the month

Ninety-six gigabytes of error-corrected GDDR7 on one add-in card, in a whole physical machine with nobody else on it. No card we fit holds more — not even the accelerators. What you give up to get that capacity is memory read rate and NVLink, and this page says so before it says anything else.

96 GB, ECC Four encoders, AV1 Passed through, not sliced

The biggest card here, and where it sits on the invoice

Across the 56 cards in the catalogue: RTX PRO 6000 Blackwell 96GB holds 96 GB, no card we fit holds more, and one costs more per month. Read the whole ladder on the whole catalogue, priced.

An H100 is an accelerator: it exists to hold a model and do arithmetic on it, and NVIDIA prices it accordingly. The RTX PRO 6000 is a workstation card that happens to have been given more memory than the accelerator. If the whole of your question is will it fit, read the memory column and the money column in the table below against each other — on this shelf they do not run in the same direction, and that is the reason this page exists.

Where it sits against the other large-memory cards we fit

Cheapest complete machine first. Every figure read from the order catalogue when this page was served.
CardCard memoryThe cardCheapest complete machine
A4048 GB GDDR6$499$572.47
RTX 6000 ADA48 GB GDDR6 ECC$625$698.47
A100 80GB80 GB HBM2e$899$972.47
RTX PRO 6000 Blackwell 96GB96 GB GDDR7 ECC$1,299$1,382.67
H100 80GB80 GB HBM2e$1,599$1,682.67

Every figure above is read from the order catalogue the checkout bills from, at the moment this page was served. Nothing on this page is typed by hand.

The specification, as NVIDIA publishes it

Blackwell, fifth-generation Tensor cores, fourth-generation RT cores. Figures from NVIDIA’s own product page for the Workstation Edition, cited in the source of this page.

RTX PRO 6000 Blackwell 96GB

Architecture

Blackwell

Card memory

96 GB GDDR7 ECC

CUDA cores

24,064

Memory bandwidth

1,792 GB/s

Board power

600 W

Card interface

PCIe Gen 5 x16

Hardware video encoders

4 NVENC, 4 NVDEC, AV1

Error-correcting memory

Yes

What it is good at, and what it is not

The honest version. A card this expensive is the wrong card more often than it is the right one, and the four paragraphs below are the ones that decide it.

Good at: holding something large, whole

Ninety-six gigabytes is enough for a seventy-billion-parameter model at eight-bit precision with head-room left for the context window, and enough for a thirty-billion model at full sixteen-bit. One card, one address space, no sharding and no second machine to keep in step.

Good at: work that is graphics AND arithmetic

Four hardware encoders and four decoders with AV1, plus fourth-generation ray-tracing cores. It renders, it transcodes and it infers on the same board. The accelerators on this shelf do none of the first two — they have no video encoder at all.

Bad at: memory-bound work at scale

GDDR7 is fast, but it is not stacked memory. The A100 and the H100 read their weights faster per second than this card does, and for large-batch token generation that read rate is the ceiling, not the arithmetic. If your bottleneck is bandwidth rather than capacity, the smaller accelerator wins.

Bad at: pooling with a second card

There is no NVLink on this part. Two of them in one machine are two separate 96 GB pools that talk over the slot, not one 192 GB pool. When you genuinely need more than 96 GB in one address space, this is not the card that gets you there.

Which machine it goes in, and what link it gets there

This card is a PCIe Gen 5 part and every platform we run is Gen 3 or Gen 4, so it never gets the full bus it was designed for. That matters to some work and to other work not at all, and the difference is worth more than the silence.

Intel Xeon Silver / Gold

RTX PRO 6000 Blackwell 96GB

Slot: PCIe 3.0, 48 lanes per socket. Card: PCIe Gen 5.

A quarter or less of the bus the card was designed for. Worth avoiding when the work crosses the slot every step, and irrelevant once the weights are resident.

AMD EPYC

RTX PRO 6000 Blackwell 96GB

Slot: PCIe 4.0, 128 lanes per socket. Card: PCIe Gen 5.

About half the bus the card was designed for. That costs a training loop streaming batches off the host; it costs an inference server or a transcoder nothing, because the weights are already on the card.

Once weights are resident on the card, the slot is idle and the generation stops mattering — which is most inference, most rendering and all transcoding. It matters to a training loop that streams a fresh batch across the slot on every step. Ask if you are unsure which of the two you are; the answer decides the machine, not the card.

What actually fits in 96 GB

Weights only, rounded up. A row counts as fitting when it leaves a fifth of the card’s 96 GB free for the KV cache, the activations and the runtime — which is head-room, not luxury.
Model16-bit8-bit4-bit
7-8 billion parameters16 GB — fits8 GB — fits5 GB — fits
13-14 billion parameters28 GB — fits14 GB — fits8 GB — fits
30-34 billion parameters68 GB — fits34 GB — fits19 GB — fits
70 billion parameters140 GB — does not fit70 GB — fits38 GB — fits

Board power is the other ceiling and it is a real one: the Workstation Edition of this card draws up to 600 W on its own, and the Max-Q edition is the same silicon and the same 96 GB at 300 W. Which of the two a given chassis takes depends on the chassis, so ask before ordering rather than after.

What it costs, with the machine it goes in

The card is a separate line on the invoice and the machine is another. Both halves below come from one reading of the order catalogue, so the price and the specification beside it always belong to each other.

Read from the order catalogue when this page was served. The card is a separate line; the machine does not get dearer because of what is plugged into it.
CardThe machine it goes inThe cardComplete, per month
RTX PRO 6000 Blackwell 96GBIntel Xeon Silver / Gold
Intel Xeon Silver 4110 8 Core 2.10 GHz · 16 GB DDR4 · 2 × SATA-SSD 240 GB (RAID 1) · 1 Gbps · PCIe 3.0
$1,299$1,382.67
Configure this build →
RTX PRO 6000 Blackwell 96GBAMD EPYC
AMD EPYC 7413 24 CORE 2.65 GHz 128MB L3 CACHE · 32 GB DDR4 · 1 × SATA-SSD 240 GB · 1 Gbps · PCIe 4.0
$1,299$1,516.59
Configure this build →

Bandwidth is unmetered in both directions with no egress charge, one IPv4 address is included, there is no contract and one month is the default term. For every other card we fit, and the same figures for all of them, read the whole catalogue, priced. For machines racked and ready today rather than built to order, /instant.

Questions people actually ask about this card

Is this the same card as the RTX 6000 Ada?

No, and the names invite the mistake. The RTX 6000 Ada Generation is the previous professional flagship: 48 GB of GDDR6 with ECC on the Ada Lovelace architecture. The RTX PRO 6000 Blackwell has 96 GB of GDDR7 with ECC, twice the memory, on newer silicon. We fit both, and they appear side by side in the ladder above so the step between them is a figure rather than a claim.

96 GB against the H100’s 80 GB — so what is the H100 for?

You are paying for different things. The H100’s money is in stacked memory and the read rate that comes with it, plus NVLink for pooling several cards into one address space. The RTX PRO 6000’s money is in capacity, ray tracing and four video encoders. If the question is will the model fit, this card answers it with more room to spare. If the question is how many tokens a second at a large batch size, the accelerator answers it better and the ladder above shows what that costs.

Is the card shared with anyone else?

No. One tenant per physical machine, and the card is passed straight through to your own operating system — no hypervisor, no MIG partition, no vGPU profile, no time-slicing. You install and pin your own driver and CUDA release, and nothing upgrades underneath you. What you measure on the first afternoon is what you keep.

Can I have two of them in one machine?

It is a build we quote rather than a box you tick, because the answer depends on the chassis, the physical width of the card and the power budget — and at up to 600 W each, the power budget is usually what decides it. Ask, and the reply names the chassis, the count and the lead time. Note that two cards are two separate memory pools here: this part has no NVLink.

How soon can I have one?

It depends on whether the card is already fitted, already on our shelf, or has to be bought in — and you are told which before you pay, not after. /instant lists machines racked and ready right now. Anything else is quoted with its real lead time.