NVIDIA L40S · 48 GB GDDR6 with ECC

L40S — the card that does all three jobs

Forty-eight gigabytes of error-corrected memory, three hardware encoders with AV1, and ray-tracing cores, on one board. It is the only card we fit that will train a model, render a frame and transcode a stream without a compromise on any of the three. The catch is real and this page states it.

48 GB, ECC Three encoders, AV1 No NVLink

The catch, first

Two things this card cannot do, and both of them are the reason it is cheaper than the accelerators it sits beside.

Its memory is fast, not stacked. The L40S reads at 864 GB/s. The A100 alongside it reads at 1,555 GB/s with 40 GB and 1,935 GB/s with 80 GB — roughly twice the rate. For token generation at a large batch size, that read rate is the ceiling and the arithmetic is not, so an A100 will serve more requests a second off the same model even though the L40S has more raw compute on paper. If serving throughput is the whole of your problem, the accelerator is the correct purchase.

There is no NVLink on this part. Two L40S cards in one machine are two separate 48 GB pools that talk over the slot, not one 96 GB pool. The A100 and H100 can pool; this cannot. When you genuinely need one address space larger than 48 GB, this is not the card that gets you there — the ladder below shows what does.

Across the 56 cards in the catalogue: L40S holds 48 GB, 4 cards we fit hold more, and 8 cost more per month. Read the whole ladder on the whole catalogue, priced.

The specification, as NVIDIA publishes it

Ada Lovelace, fourth-generation tensor cores, third-generation ray-tracing cores. Figures from NVIDIA’s own L40S product page, cited in the source of this page.

L40S

Architecture

Ada Lovelace

Card memory

48 GB GDDR6 ECC

CUDA cores

18,176

Memory bandwidth

864 GB/s

Board power

350 W

Card interface

PCIe Gen 4 x16

Hardware video encoders

3 NVENC, 3 NVDEC, AV1

Error-correcting memory

Yes

NVIDIA also publishes 91.6 teraFLOPS of FP32, 212 teraFLOPS of ray-tracing performance and 1,466 teraFLOPS of FP8 tensor throughput with sparsity, from 142 RT cores and 568 tensor cores. Those are peak figures under ideal conditions; the memory bandwidth above is the one that usually decides what you actually get.

What it is good at

Having stated the catch: here is the work this card is genuinely the right answer for.

A pipeline with video in it

Three encoders and three decoders with AV1 in both directions. Ingest, transcode and serve on the same board that runs the model deciding what to do with the frames. Neither the A100 nor the H100 can encode a single frame in hardware — on those cards, video is a CPU job.

Rendering that has to be right

Error-correcting memory and third-generation ray-tracing cores. A render farm on consumer cards costs less per frame and occasionally produces a frame with a flipped bit in it; on a long unattended job that is a re-run. This card gives you forty-eight gigabytes with ECC behind every one of them, which is what a job you cannot babysit actually needs.

Virtual desktops and seats

The L40S supports vGPU and drives four DisplayPort outputs, which is what makes it a card you can slice between virtual workstations rather than a headless accelerator. The A100 and H100 have no display outputs at all.

Inference where the model fits in 48 GB

FP8 tensor support and enough capacity for a thirty-billion-parameter model at eight-bit precision with room to spare. For an endpoint answering real traffic at a modest batch size, the bandwidth ceiling above rarely binds — it binds when you are pushing the card flat out.

What it replaced, and what sits either side of it

From our own shelves rather than from a vendor roadmap. All of these are orderable today.

Cheapest complete machine first. Every figure read from the order catalogue when this page was served.
CardCard memoryThe cardCheapest complete machine
L40S48 GB GDDR6 ECC$449$532.67
A4048 GB GDDR6$499$572.47
RTX 6000 ADA48 GB GDDR6 ECC$625$698.47
A100 40GB40 GB HBM2$699$772.47
A100 80GB80 GB HBM2e$899$972.47

The A40 is the Ampere card the L40S succeeded — the same 48 GB, an older architecture, no AV1 and no FP8. The RTX 6000 Ada is the workstation card with the same capacity and the same generation. The A100s are what you move to when bandwidth or pooling is the binding constraint rather than capacity, and the accelerator page covers those in detail.

Which machine it goes in, and what link it gets there

The L40S is a PCIe Gen 4 card. One of the two platforms that take it is Gen 3, and one is Gen 4.

Intel Xeon Silver / Gold

L40S

Slot: PCIe 3.0, 48 lanes per socket. Card: PCIe Gen 4.

About half the bus the card was designed for. That costs a training loop streaming batches off the host; it costs an inference server or a transcoder nothing, because the weights are already on the card.

AMD EPYC

L40S

Slot: PCIe 4.0, 128 lanes per socket. Card: PCIe Gen 4.

Full width. The slot is at least the generation the card was designed for, so nothing is left on the table.

For an inference endpoint or a transcoder, the slot generation is close to irrelevant — the weights are resident and the frames are small. It matters for a training loop streaming fresh batches off the host on every step, and for that case the Gen 4 platform is worth the difference.

What actually fits in 48 GB

Weights only, rounded up. A row counts as fitting when it leaves a fifth of the card’s 48 GB free for the KV cache, the activations and the runtime — which is head-room, not luxury.
Model16-bit8-bit4-bit
7-8 billion parameters16 GB — fits8 GB — fits5 GB — fits
13-14 billion parameters28 GB — fits14 GB — fits8 GB — fits
30-34 billion parameters68 GB — does not fit34 GB — fits19 GB — fits
70 billion parameters140 GB — does not fit70 GB — does not fit38 GB — fits

The card draws up to 350 W, which is modest for its class and part of why it fits chassis that a larger accelerator does not.

What it costs, with the machine it goes in

The card is one line on the invoice and the machine is another. Both halves below come from one reading of the order catalogue, so the price and the specification beside it always belong to each other.

Read from the order catalogue when this page was served. The card is a separate line; the machine does not get dearer because of what is plugged into it.
CardThe machine it goes inThe cardComplete, per month
L40SIntel Xeon Silver / Gold
Intel Xeon Silver 4110 8 Core 2.10 GHz · 16 GB DDR4 · 2 × SATA-SSD 240 GB (RAID 1) · 1 Gbps · PCIe 3.0
$449$532.67
Configure this build →
L40SAMD EPYC
AMD EPYC 7413 24 CORE 2.65 GHz 128MB L3 CACHE · 32 GB DDR4 · 1 × SATA-SSD 240 GB · 1 Gbps · PCIe 4.0
$449$666.59
Configure this build →

Bandwidth is unmetered in both directions with no egress charge, one IPv4 address is included, there is no contract and one month is the default term. For every other card we fit, priced the same way, read the whole catalogue, priced. For machines racked and ready today, /instant.

Questions people actually ask about this card

L40S or A100, for serving a model?

If the model fits in 48 GB and the traffic is moderate, the L40S is cheaper and does more. If you are pushing the card flat out at a large batch size, the A100 reads its weights about twice as fast and will serve more requests a second despite having less compute on paper. And if the model does not fit in 48 GB, the A100 80 GB is the answer and the L40S is not.

Can I pool two of them to get 96 GB?

No. This part has no NVLink, so two cards are two separate 48 GB pools communicating over the slot. A model larger than 48 GB has to be split across them deliberately by your framework, with the slot as the link. If you want one address space bigger than 48 GB, look at the 80 GB and 96 GB cards in the ladder above instead.

Is the card shared, sliced or virtualised?

Not by us. One tenant per physical machine and the card passed straight through to your own operating system — no hypervisor, no MIG partition, no vGPU profile, no time-slicing. The card supports vGPU, so if you want to slice it between your own virtual desktops you can; that is your decision to make on your own machine, not something we do to it beforehand.

Does it really do AV1?

Yes, encode and decode, on all three of its hardware encoders and decoders. That is the specific thing the RTX 30 series cannot do at all and the accelerators cannot do at all, and it is often the reason this card is on the shortlist.

How soon can I have one?

It depends on whether the card is already fitted, already on our shelf, or has to be bought in — and you are told which before you pay, not after. /instant lists machines racked and ready right now.