Multi-card and multi-node · Primcast LLC · since 2004

Going wider is not the same trade as going bigger.

One card in one machine is a settled question, and it is priced card by card on the GPU page. This one starts where that stops: a second card in the same chassis, or a second machine beside the first. They sound like the same decision and they are opposites — the platform here that gives a card the most lanes has the lowest network ceiling and sits in the fewest rooms. Nobody publishes that about their own catalogue. It is the first table below.

Compare the platforms What joins two machines

Every machine is one tenant, every card is passed through to your own operating system, and nothing here counts a byte between your nodes.

The register

  • 4Platforms that take a card
  • 100 GbpsFastest port on any of them
  • 2Sockets in one machine, at most
  • 1 TBMemory in one machine, at most
  • 2Rooms holding every platform

Read from the order catalogue when this page was served, not typed in. What is standing in a rack this morning is a different question, and instant servers is where it is answered.

Two different questions

Which one are you actually asking?

People arrive here having decided “I need more” without having decided more of what. The two answers cost differently, are delivered differently, and are limited by different things, so it is worth being explicit before reading another line.

More card in one machine

A second card in the same chassis

You are short of card memory, or short of throughput, and everything else about the machine is fine. This is a build we quote rather than a box you tick in the checkout, because how many cards fit depends on the physical width of the specific card and on the power budget of the specific chassis — two facts that are not in a catalogue and that we will not guess at.

What decides it: the card, not the platform. Ask with the part named and the reply names the chassis, the count and the lead time.

The thing people expect and do not get: two cards are not one larger card. Their memory adds up only if what you are running can split a model across both — tensor or pipeline parallelism, which every serious runtime supports and which costs you some throughput to the traffic between them. For work that is already many independent jobs, two cards are simply twice the machine and nothing has to change.

More machines beside it

A second machine next to the first

You have run out of the machine rather than of the card: out of sockets, out of memory, out of somewhere to put the next job. On the best of the platforms below, that ceiling is 2 sockets and 1 TB, and past either of those the answer is another server rather than a bigger one.

What decides it: the room and the port. Two machines that have to talk to each other have to be in the same building, and what joins them is the port on each of them — bought per machine, so the fabric under four nodes is four ports.

The thing people expect and do not get: there is no cluster product here and no minimum node count. Each machine is ordered, billed, upgraded and replaced on its own, which is why nothing on this page quotes a cluster price — a cluster is n invoices.

The trade

The platform best at one of those is worst at the other.

Every platform we fit a card to, on the axes that decide a cluster rather than a single machine. Read the two middle columns together and the whole argument of this page is in them: they move in opposite directions. The bus generation is the same figure the GPU page publishes, from the same table, so the two pages cannot disagree about a link.

Read from the order catalogue when this page was served. The two middle columns are the trade.
PlatformLanes to the cardNetwork ceilingSocketsMemory ceilingRooms
Intel Xeon Silver / Gold48 lanes per socketPCIe 3.0100 Gbps1 Gbps included21 TB5
Intel Xeon E5-2600 v3/v440 lanes per socketPCIe 3.040 Gbps1 Gbps included21 TB5
Intel Xeon E5-2600 v1/v240 lanes per socketPCIe 3.020 Gbps1 Gbps included2256 GB3
AMD EPYC128 lanes per socketPCIe 4.020 Gbps1 Gbps included21 TB2

Reading that table

  • Lanes to the card decide how fast data crosses between the processor and the accelerator. It matters continuously for a training loop streaming batches, for a render feeding scene data every frame, and for anything spilling out of card memory. It stops mattering almost entirely once a model is resident and staying there.
  • Network ceiling decides how fast one machine can talk to the next one. It is irrelevant to a single server doing its own work and it is the entire question for a job split across several.
  • Rooms decides whether the choice is available where you need it at all, and it is the column people read last and should read first — see below.
  • So: a job that fits inside one machine and hammers the bus wants the widest lanes and does not care about the port. A job spread over several machines wants the port and can live with an older bus. Wanting both at once is the one combination our catalogue does not hold, and we would rather say so here than after you have paid.

Which card goes in any of these, and what each one adds per month, is GPU dedicated servers — every card we fit, priced against the exact machine its figure belongs to. This page prices no card, deliberately: one page owns that money.

What joins them

The interconnect is bought per machine, and that is the arithmetic people get wrong.

A port is a property of one server. Four machines on a given speed is that speed four times, every month, and it is the line item that turns a comfortable cluster budget into an uncomfortable one. Below is the ladder with the multiplication already done, because a page about more than one machine that leaves you to do it yourself is not doing its job.

Per month, on top of the machine. The port is bought per node, so the right-hand columns are what the same speed costs across a cluster.
PortOne machineTwoFourOffered on
1 Gbpsincludedincludedincludedevery platform
2 Gbps$189$378$756every platform
3 Gbps$299$598$1,196every platform
5 Gbps$459$918$1,836every platform
10 Gbps$529$1,058$2,116every platform
20 Gbps$1,629$3,258$6,516every platform
40 Gbps$2,799$5,598$11,1962 of 4
100 Gbps$3,999$7,998$15,9961 of 4

Two things worth noticing before you pick a rung

  • The ladder is not linear, so a wider cluster can be cheaper than a faster one. Compare the rungs against each other rather than against nothing: at several points on that table, two machines a rung lower cost less than one machine a rung higher, and give you two machines. Whether that is a good trade depends entirely on whether your work splits.
  • The top of the ladder is not on every platform. The right-hand column says how many of them reach each speed, and the fastest rungs narrow to a platform that gives a card an older bus. That is the same trade as the table above, arriving from the other direction.

What we do not have, said plainly

Nothing beyond an Ethernet port between machines is in our catalogue. There is no listed InfiniBand fabric and no listed card-to-card bridge, and this page is not going to imply one by writing around it. If a build needs either, ask before you order rather than after: the answer is a quote, and the answer may be no. That is worth more to you than a page that leaves the impression it might be yes.

Every speed above is unmetered in both directions, with no transfer allowance and no egress line on any invoice — so what your nodes send each other is free at any volume, which is the part that is usually metered elsewhere. Unmetered bandwidth is the page for a port considered on its own, on one machine.

Where it can exist

For one machine you choose the platform first. For several, choose the room first.

Machines that have to talk to each other belong in the same building, and our rooms do not all hold the same platforms. Pick the platform first and you can find that the room you need it in does not carry it. This is derived from the same catalogue as everything else, so a line arriving in a new city appears here on its own.

  • New York, US4 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4 · Intel Xeon E5-2600 v1/v2 · AMD EPYC
  • Bucharest, EU4 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4 · Intel Xeon E5-2600 v1/v2 · AMD EPYC
  • Miami, US3 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4 · Intel Xeon E5-2600 v1/v2
  • San Francisco, US2 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4
  • Amsterdam, EU2 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4

Where the building actually is, what it is connected to and who else is in it — data centres describes all five. Latency between two of our cities is a real number rather than a marketing one, and support will measure it for your pair on request rather than have you guess from a map.

How it is quoted, and how soon it exists

Multi-node work is dated before it is paid for, not after.

A build of this kind is several machines and often several parts that are not on a shelf, so the honest thing to publish is not a delivery promise but the ladder the date comes off — the same one build to order and the GPU page use, in the same words.

  1. ~30 minutes

    The machines are already racked as they are

    Nothing is fitted and nothing is moved. Instant servers is the list of what is standing there right now, and for a cluster of ordinary nodes it is worth reading first.

  2. 4–24 hours

    The parts are ones we hold

    The ordinary case. Technicians fit the cards, build the arrays and install the operating systems — per machine, in parallel, not one after another.

  3. 5–10 working days

    Parts have to be bought in

    Something on the specification is not stock we carry. We order it, then build. This is the commonest rung for a multi-card chassis, because the chassis is usually the long pole rather than the card.

  4. 10–15 working days

    A part is scarce

    Current-generation accelerators run on allocation and the date is the distributor’s, not ours. We name the long pole before you commit rather than after. Where hardware is bought in for one customer we ask for three months up front, because it is a machine we cannot re-let if the account closes in week two.

Every rung starts when the order clears payment and fraud review, which is the one step we will not put a promise on — most clear inside a couple of hours, and a large first order from a new account can take longer. For a build of this size a wire or a stablecoin usually clears faster than a card, which is worth knowing before a fraud check delays a rack visit.

Describing a multi-node or multi-card build is faster in a sentence than in a form. — the node count, whether the work splits, and what it has to talk to — and somebody who has quoted one of these before will answer, including when the answer is that one machine would do.

For the record

Everything on this page, as figures.

Read from the order catalogue when this page was served. Card prices are deliberately not among them — GPU dedicated servers holds every one, against the exact machine each figure belongs to.

Cards in one machine
More than one is a build we quote rather than a checkout option. How many fit depends on the physical width of the card and on the power budget, so the honest answer needs the specific card. Ask, and the reply names the chassis, the count and the lead time.
Machines in one cluster
No limit we impose. Each is a whole physical server rented on its own terms, and they are ordered, billed and replaced independently — there is no cluster product and no minimum node count.
Platforms that take a card
4, and they are not interchangeable. The one with the widest bus to the card (AMD EPYC, PCIe 4.0, 128 lanes per socket) has the lowest network ceiling (20 Gbps) and is in 2 of our 5 rooms. The one that reaches 100 Gbps (Intel Xeon Silver / Gold) gives a card an older bus.
What joins two machines
The port on each of them. Speeds run from the included 1 Gbps on every one of them up to 100 Gbps, priced per machine and per month, so a cluster pays for the speed once per node. There is no NVLink fabric and no InfiniBand here; anything beyond an Ethernet port between two of our machines is quoted rather than listed.
Ceilings in one machine before a second one is needed
2 sockets and 1 TB of memory on the best of these platforms. Past either, the answer is another machine.
Where
5 cities: New York, Miami, San Francisco, Amsterdam, Bucharest. Machines that talk to each other should be in one of them together, and New York and Bucharest hold every platform on this page.
Bandwidth
Unmetered, both directions, at every speed on the ladder. No transfer allowance, no egress line on any invoice — which is what makes moving a dataset between nodes free rather than metered.
Tenancy
One tenant per physical machine, on every node. Cards are passed through to your own operating system: no hypervisor, no MIG partition, no vGPU profile, no time-slicing.
How soon
The same ladder as any build here: about four to twenty-four hours when the parts are ones we hold, five to ten working days when they have to be bought in, and ten to fifteen when a part is scarce. A multi-node build is quoted with its long pole named before you commit.
Term
One month, no contract, on each machine separately. Three, six and twelve-month cycles take up to 15% off. Where hardware is bought in for one customer we ask for three months up front.
What this page does not price
Cards. Every graphics card we fit, what each adds per month and what the machine under it costs are on GPU dedicated servers, priced against the exact machine each figure belongs to.