Ryzen AI dedicated servers • Inference without a graphics card

Small models run fine on a Ryzen AI machine. That is the point.

Not every workload earns a graphics card. A quantised model of modest size answers from system memory on a whole machine of its own — your runtime, your versions, root on the metal — without paying for silicon the job would leave idle.

Turn your GPUs into passive monthly revenue.

Got idle server or desktop GPU setups? List them on the Primcast marketplace today and earn steady monthly rents from AI teams, developers, and enterprises needing production-grade compute.

Go to Marketplace

Where skipping the card is the right call

The same pattern each time: the model is small, the traffic is modest, and the data has to stay on a machine that is nobody else’s.

Private assistant on internal documents
A private helper on private data

A small quantised model reading your own documents for a handful of colleagues. The corpus never touches a shared endpoint, and the monthly cost is a machine, not a meter.

Batch drafting and summarising
Drafts, summaries, batch runs

Nightly summaries, tagging, rewriting — work that queues. What matters there is the cost per job, not the milliseconds per token, and that is the workload a graphics card is most wasted on.

Internal tools calling a model
Tools that call a model quietly

Ticket routing, log triage, extraction into a database: a small model behind an internal endpoint, on all day, at a flat rate the finance sheet can read.

Evaluating the NPU toolchain
Kicking the NPU’s tyres

Running AMD’s XDNA toolchain against real silicon, rented by the month, before anyone signs for a fleet of it. What refuses to port shows up here first, and cheaply.

The trade, stated plainly

These machines are honest value for the right workload and the wrong purchase for the wrong one. Here is the line between the two.

Memory is the roomy part

The model lives in the machine’s own memory, measured in dozens of gigabytes rather than a card’s allotment. What fits is rarely the problem on this page.

Bandwidth is the narrow part

Every generated token re-reads the weights, and system memory reads slower than card memory. One user on a small model will not mind; a queue of users will. We publish no speed figure — here or anywhere — and say this instead.

Asked before ordering one

Four honest answers — the same ones our engineers give on chat.

Can a server really run a language model without a graphics card?

Yes, within limits worth stating up front: a small quantised model generates from system memory at reading pace, which suits one user, an internal tool or a batch queue — and does not suit a crowd. How much memory a given model needs, and why one that fits can still fail to serve, is exactly what our LLM page works through.

Does the NPU run my model?

It can, through AMD’s own Ryzen AI toolchain — but most of the open-model runtimes people actually deploy use the processor and the integrated Radeon instead, and we would rather say so here than let a spec sheet imply otherwise. The NPU is on the die either way, and nothing stops you targeting it.

When is a dedicated GPU worth it instead?

When the model gets bigger than the memory, when several people need answers at once, or when you are fine-tuning rather than serving. Ryzen AI or a dedicated GPU walks that decision, and the GPU page prices every card against the exact machine it goes in.

Is any of it shared with other customers?

No. You rent the whole machine: processor, graphics, NPU, memory, disk and port belong to one customer at a time, with no hypervisor underneath. That is much of what the monthly figure buys.