Not every workload earns a graphics card. A quantised model of modest size answers from system memory on a whole machine of its own — your runtime, your versions, root on the metal — without paying for silicon the job would leave idle.
The same pattern each time: the model is small, the traffic is modest, and the data has to stay on a machine that is nobody else’s.
These machines are honest value for the right workload and the wrong purchase for the wrong one. Here is the line between the two.
Four honest answers — the same ones our engineers give on chat.