多卡与多节点 · Primcast LLC · 自 2004 年起

横向扩展与纵向扩展不是同一种取舍。

一台机器里装一张卡,这个问题已经尘埃落定,GPU 页面上按卡逐一定价。而这一页从那里结束的地方开始:同一机箱里的第二张卡,或者第一台机器旁边的第二台机器。它们听起来像是同一个决定,实则截然相反 — 这里给一张卡最多通道的平台,网络上限最低,能进的机房也最少。没有哪家会在自己的产品目录里公开这一点。下面的第一张表就是它。

对比各平台 什么把两台机器连起来

每台机器都是单一租户,每张卡都直通给你自己的操作系统,这里没有任何东西会计数你节点之间的一个字节。

登记表

  • 4Platforms that take a card
  • 100 GbpsFastest port on any of them
  • 2Sockets in one machine, at most
  • 1 TBMemory in one machine, at most
  • 2Rooms holding every platform

本页送达时从订单目录中读取,并非手工录入。今天早上机架上有什么是另一个问题,即时服务器会回答它。

把你的 GPU 变成每月被动收入

有闲置的服务器或台式机 GPU 配置吗?今天就上架到 Primcast 市场,从需要生产级算力的 AI 团队、开发者和企业那里赚取稳定的月租。

前往市场

两个不同的问题

你真正在问的是哪一个?

人们来到这里时已经决定了“我需要更多”,却还没决定更多的是什么。两个答案的成本不同、交付方式不同、受限于不同的东西,所以在继续读下去之前,值得先把它说清楚。

一台机器里更多的卡

同一机箱里的第二张卡

你缺的是卡的内存,或者缺吞吐量,而机器的其他一切都很好。这是我们报价的定制方案,而不是你在结账时勾选的选项,因为能装多少张卡取决于具体卡的物理宽度和具体机箱的功率预算 — 这两个事实不在目录里,我们也不会去猜。

决定它的是什么:卡,而不是平台。带着具体型号来问,回复会给出机箱、数量和交付周期。

人们期待却得不到的:两张卡不是一张更大的卡。它们的内存加起来,只有在你运行的东西能把模型拆分到两张卡上时才有意义 — 张量并行或流水线并行,每个严肃的运行时都支持,但你要为它们之间的流量付出一些吞吐量。对于本来就是许多独立作业的工作,两张卡就是两台机器的算力,什么都不用改。

它旁边更多的机器

第一台旁边的第二台机器

你已经用尽了机器本身,而不是卡:插槽用完了、内存用完了、没有地方放下一个作业了。在下面最好的平台上,那个上限是 2 sockets and 1 TB,超过其中任何一个,答案就是另一台服务器,而不是一台更大的。

决定它的是什么:机房和端口。两台必须互相通信的机器必须在同一栋楼里,而连接它们的是每台机器上的端口 — 按机器购买,所以四个节点下的网络就是四个端口。

人们期待却得不到的:这里没有集群产品,也没有最低节点数。每台机器独立订购、计费、升级和更换,这就是为什么本页不报集群价格 — 一个集群就是 n 张发票。

取舍

在其中一项上最好的平台,在另一项上最差。

我们给卡适配的每一个平台,都放在决定集群而非单机的坐标轴上。把中间两列一起读,本页的全部论点就在其中:它们朝相反的方向移动。总线代数是 GPU 页面公布的同一个数字,来自同一张表,所以两页不会在链路上互相矛盾。

Read from the order catalogue when this page was served. The two middle columns are the trade.
PlatformLanes to the cardNetwork ceilingSocketsMemory ceilingRooms
Intel Xeon Silver / Gold48 lanes per socketPCIe 3.0100 Gbps1 Gbps included21 TB5
Intel Xeon E5-2600 v3/v440 lanes per socketPCIe 3.040 Gbps1 Gbps included21 TB5
Intel Xeon E5-2600 v1/v240 lanes per socketPCIe 3.020 Gbps1 Gbps included2256 GB3
AMD EPYC128 lanes per socketPCIe 4.020 Gbps1 Gbps included21 TB2

如何读那张表

  • 到卡的通道决定数据在处理器和加速器之间传输的速度。对于持续流式传输批次的训练循环、每帧都送入场景数据的渲染,以及任何溢出卡内存的工作,它都持续重要。一旦模型常驻且不再移动,它就几乎完全无关紧要了。
  • 网络上限决定一台机器与下一台通信的速度。对于一台做自己工作的服务器,它无关紧要;对于拆分到多台机器上的作业,它就是全部问题。
  • 机房决定这个选择在你需要的地方是否可用,而它是人们最后读、却应该最先读的一列 — 见下文。
  • 所以:一个装在一台机器内、猛打总线的作业,想要最宽的通道,不在乎端口。一个分散在多台机器上的作业,想要端口,可以接受较老的总线。两者同时都要,是我们目录里没有的唯一组合,我们宁愿在这里说清楚,而不是在你付钱之后。

哪张卡装进这些平台中的任何一个,每张卡每月增加多少,都在 GPU 专用服务器 — 我们装的每一张卡,都对着它所属的确切机器定价。本页故意不给任何卡定价:那一页拥有这笔钱。

什么把它们连起来

互连是按机器购买的,而这是人们算错的地方。

端口是一台服务器的属性。四台机器在给定速度上,就是那个速度乘以四,每个月都是,而正是这一行让一个舒适的集群预算变得不舒适。下面是已经算好乘法的阶梯,因为一个关于不止一台机器的页面,如果还让你自己算,就是没尽到本分。

Per month, on top of the machine. The port is bought per node, so the right-hand columns are what the same speed costs across a cluster.
PortOne machineTwoFourOffered on
1 Gbpsincludedincludedincludedevery platform
2 Gbps$189$378$756every platform
3 Gbps$299$598$1,196every platform
5 Gbps$459$918$1,836every platform
10 Gbps$529$1,058$2,116every platform
20 Gbps$1,629$3,258$6,516every platform
40 Gbps$2,799$5,598$11,1962 of 4
100 Gbps$3,999$7,998$15,9961 of 4

在选一档之前,有两件事值得注意

  • 阶梯不是线性的,所以更宽的集群可能比更快的集群便宜。把各档互相比较,而不是跟空气比较:在那张表的几个点上,低一档的两台机器比高一档的一台机器便宜,而且给你两台机器。这是否划算,完全取决于你的工作能否拆分。
  • 阶梯的顶端不是每个平台都有。右列说明了有多少平台能达到每个速度,而最快的几档收窄到一个给卡较老总线的平台。这和上面那张表是同一个取舍,只是从另一个方向到来。

我们没有的,直说

机器之间除了以太网端口之外的东西,都不在我们的目录里。没有列出的 InfiniBand 网络,没有列出的卡间桥接,本页也不会绕着写来暗示有。如果一个方案需要其中任何一个,在下单之前问,而不是之后:答案是一份报价,而答案可能是不行。这对你来说,比一个留下“也许可以”印象的页面更有价值。

上面的每个速度都是双向不限流量,没有传输配额,任何发票上都没有出站流量这一行 — 所以你的节点之间互发什么都是免费的,无论多少量,而这部分在别处通常是计量的。不限流量带宽是单独考虑一台机器上一个端口的页面。

它能在哪里存在

一台机器,你先选平台。多台机器,先选机房。

必须互相通信的机器属于同一栋楼,而我们的机房并不都放着相同的平台。先选平台,你可能会发现你需要的那个机房没有它。这和其他一切一样来自同一份目录,所以一条出现在新城市的线路会自己出现在这里。

  • New York, US4 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4 · Intel Xeon E5-2600 v1/v2 · AMD EPYC
  • Bucharest, EU4 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4 · Intel Xeon E5-2600 v1/v2 · AMD EPYC
  • Miami, US3 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4 · Intel Xeon E5-2600 v1/v2
  • San Francisco, US2 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4
  • Amsterdam, EU2 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4

那栋楼实际在哪里、连接着什么、里面还有谁 — 数据中心描述了全部五个。我们两个城市之间的延迟是一个真实数字,而不是营销数字,支持团队会应要求为你的那一对测量,而不是让你对着地图猜。

如何报价,以及它多快存在

多节点工作在付款之前就定了日期,而不是之后。

这种方案是几台机器,而且往往有几个不在货架上的部件,所以诚实公布的不是一个交付承诺,而是日期所依据的阶梯 — 和 按单定制以及 GPU 页面用的是同一个,措辞也一样。

  1. 约 30 分钟

    机器已经按原样上架

    什么都不装,什么都不搬。即时服务器就是此刻立在那里的清单,对于一个普通节点的集群,值得先读。

  2. 4–24 小时

    部件是我们持有的

    普通情况。技术人员装卡、建阵列、装操作系统 — 按机器并行进行,不是一台接一台。

  3. 5–10 个工作日

    部件需要采购

    规格上的某样东西不是我们常备的库存。我们先订货,再组装。这是多卡机箱最常见的一档,因为机箱通常才是长杆,而不是卡。

  4. 10–15 个工作日

    某个部件稀缺

    当前一代加速器按配额供应,日期是分销商的,不是我们的。我们在你承诺之前就点出长杆,而不是之后。为单个客户采购的硬件,我们要求预付三个月,因为如果账户在第二周关闭,这台机器我们无法再租出去。

每一档都从订单通过付款和反欺诈审核时开始计时,这是唯一我们不会承诺的一步 — 大多数在几个小时内通过,而新账户的大额首单可能需要更久。对于这种规模的方案,电汇或稳定币通常比信用卡更快到账,在反欺诈检查耽误上架参观之前,这一点值得知道。

描述一个多节点或多卡方案,一句话比一张表单更快。 — 节点数、工作是否可拆分、它必须和什么通信 — 一个以前报过这类价格的人会回答你,包括当答案是一台机器就够的时候。

For the record

Everything on this page, as figures.

Read from the order catalogue when this page was served. Card prices are deliberately not among them — GPU dedicated servers holds every one, against the exact machine each figure belongs to.

Cards in one machine
More than one is a build we quote rather than a checkout option. How many fit depends on the physical width of the card and on the power budget, so the honest answer needs the specific card. Ask, and the reply names the chassis, the count and the lead time.
Machines in one cluster
No limit we impose. Each is a whole physical server rented on its own terms, and they are ordered, billed and replaced independently — there is no cluster product and no minimum node count.
Platforms that take a card
4, and they are not interchangeable. The one with the widest bus to the card (AMD EPYC, PCIe 4.0, 128 lanes per socket) has the lowest network ceiling (20 Gbps) and is in 2 of our 5 rooms. The one that reaches 100 Gbps (Intel Xeon Silver / Gold) gives a card an older bus.
What joins two machines
The port on each of them. Speeds run from the included 1 Gbps on every one of them up to 100 Gbps, priced per machine and per month, so a cluster pays for the speed once per node. There is no NVLink fabric and no InfiniBand here; anything beyond an Ethernet port between two of our machines is quoted rather than listed.
Ceilings in one machine before a second one is needed
2 sockets and 1 TB of memory on the best of these platforms. Past either, the answer is another machine.
Where
5 cities: New York, Miami, San Francisco, Amsterdam, Bucharest. Machines that talk to each other should be in one of them together, and New York and Bucharest hold every platform on this page.
Bandwidth
Unmetered, both directions, at every speed on the ladder. No transfer allowance, no egress line on any invoice — which is what makes moving a dataset between nodes free rather than metered.
Tenancy
One tenant per physical machine, on every node. Cards are passed through to your own operating system: no hypervisor, no MIG partition, no vGPU profile, no time-slicing.
How soon
The same ladder as any build here: about four to twenty-four hours when the parts are ones we hold, five to ten working days when they have to be bought in, and ten to fifteen when a part is scarce. A multi-node build is quoted with its long pole named before you commit.
Term
One month, no contract, on each machine separately. Three, six and twelve-month cycles take up to 15% off. Where hardware is bought in for one customer we ask for three months up front.
What this page does not price
Cards. Every graphics card we fit, what each adds per month and what the machine under it costs are on GPU dedicated servers, priced against the exact machine each figure belongs to.