A chatbot that is only yours
Run an open model such as Llama, Qwen or Mistral on a server nobody else uses. Prompts, documents and answers stay on your machine instead of going to somebody else’s API.
NVIDIA CMP 170HX · unlocked when we install it
NVIDIA built the CMP 170HX for mining and locked the rest of it away: the card showed 8 GB of memory and ran floating-point maths at a crawl. The lock turned out to be software. We install a driver that lifts it, so your server arrives with a card that has 64 GB of very fast memory and runs AI models the way its chip was designed to.
Same card, same chip. The grey part is what NVIDIA left on; the rest is what the driver turns on.
What it is good at now
Run an open model such as Llama, Qwen or Mistral on a server nobody else uses. Prompts, documents and answers stay on your machine instead of going to somebody else’s API.
A 70-billion-parameter model fits on one unlocked card at 4-bit. On a card that has not been unlocked, the ceiling is a model of about 8 billion.
Whatever the model does not use goes to its context, so long documents and long chats keep going instead of running out of room halfway through.
A chat model, an embedding model for search and a small model that routes requests can all stay loaded together, rather than taking turns.
One flat monthly price and no meter, on the card or on the bandwidth. Leave a batch running over the weekend and the bill does not move.
Weights only, with a quarter to a half again kept free for the conversation itself. The full sizing guide is on our LLM page.
The same card, before and after
LTT Labs unlocked a CMP 170HX in September 2026 and ran the same models on it both ways. These are their results, not ours. Speeds change with the model and the software, so take them as the size of the jump rather than a number to hold us to.
These numbers come from the LTT Labs write-up of September 2026, which ran llama.cpp on their own test machine. The unlock itself is cmpunlocker, an open-source project by Amogh Munikote.
What happens after you order
Choose Ubuntu 22.04, 24.04 or 26.04 in the configurator. The unlock is a Linux driver, so this is the one choice that matters.
The card goes in, Ubuntu goes on, and the unlocked driver is built into it for the exact kernel your server will run.
The server is switched fully off and back on, which is what the unlock needs, and every card is checked for its full memory before the server is handed over.
When Ubuntu installs a new kernel, the driver rebuilds itself to match. Type cmp-unlock status whenever you want to see how each card is doing.
Worth knowing before you buy
What it costs
The card is its own line on the invoice and the server is another, so you can see what each one costs. The figure shown is one card in the smallest server that takes it. Add memory, bigger disks or more cards in the configurator and it shows the new total as you go.
There is a one-time setup fee because the card is fitted to your order. Bandwidth is unmetered, US sales tax is added at checkout where it applies, and the monthly price is the same when you renew.
$272.67/month
Intel Xeon Silver 4110 8 Core 2.10 GHz · 16 GB DDR4 · 2 × SATA-SSD 240 GB (RAID 1) · 1 Gbps
Questions
NVIDIA sold the CMP 170HX as a mining card and switched most of it off in software: it showed 8 GB of memory and did floating-point maths very slowly. A community driver called cmpunlocker switches those parts back on. An unlocked card has 64 GB of memory and runs at the speed of the chip inside it, which is the same chip as an A100.
64 GB on the 8 GB version of the card, which is what you will see in nvidia-smi. The 10 GB version unlocks to 40 GB, because more than that is not reliable on it. Without the unlocked driver, both show the memory they shipped with.
No. Order the server with Ubuntu 22.04, 24.04 or 26.04 and it arrives unlocked. We install the driver, build the unlock for your kernel and check each card before handing the server over.
Yes. The unlocked driver loads every time the server starts, and when Ubuntu installs a new kernel the unlock rebuilds itself for it. If a card ever shows 8 GB, switch the server off and on again from the control panel, and tell us if that does not fix it.
Anything made for NVIDIA cards: PyTorch, llama.cpp, Ollama, vLLM and the rest. Install the CUDA toolkit with apt install cuda-toolkit. Avoid the packages named cuda and cuda-drivers, which would add a second driver; we set the server to refuse them.
You can, but the unlock only exists for Linux, so on Windows the card stays at 8 GB with its original limits. For anything that needs the memory, choose Ubuntu.
Up to four. Two unlocked cards hold 128 GB between them, enough for a 70-billion-parameter model at 8-bit.
Yes, unlocking takes nothing away. If mining is the main plan, though, our mining page covers the cards and the costs for that.
The card is fitted to your order, which makes the server a custom build, and custom builds carry a one-time setup fee. It is charged once, and the monthly price does not go up when you renew.