Which card holds your model?
Find the model you want to run and read across. Each row is the smallest card we fit that holds that model's files whole on the card, so ComfyUI never has to shuffle pieces in and out of ordinary memory while it works.
| What you want to run | Size of the files | Smallest card we fit that holds it |
|---|---|---|
| SDXL base 1.0 | 6.94 GB | an 8 GB card, such as the RTX 2070 |
| FLUX.1-dev, fp8 single file | 17.2 GB | a 24 GB card, such as the RTX 3090 |
| FLUX.1-dev full model with the smaller fp8 text encoder | about 29 GB | a 32 GB card, such as the Tesla V100 32GB or the RTX 5090 |
| FLUX.1-dev, everything at full 16-bit | about 34 GB | a 40 GB card, such as the A100 40GB, or a 48 GB L40S for more room |
The full Flux model file alone is 23.8 GB; its large text encoder adds 9.79 GB at 16-bit, or 4.89 GB in fp8. A smaller card still works: ComfyUI moves models out to ordinary memory after use and can stream weights onto the card as it goes. That traffic runs over the PCIe bus on every pass and leans on the server's RAM, so a card that holds the model avoids it.
FLUX.1-dev comes under a non-commercial licence; the official download is gated behind accepting it, and repackaged copies such as the fp8 file carry the same licence. If your images are for paying clients, read that licence first.
Will an older card do?
Often, yes, and older cards with lots of memory are some of the best value on the GPU servers page. There is one catch. ComfyUI's install instructions use a PyTorch built for CUDA 13.0, which only supports cards from NVIDIA's Turing generation onward (compute capability 7.5) and needs driver 580 or newer.
Maxwell, Pascal and Volta cards, like the 24 GB Tesla P40 or the 32 GB Tesla V100, need the PyTorch build for CUDA 12.6 instead. It is a one-word change in one install command, shown below. Everything else stays the same.
Memory and disk around the card
The card does the drawing, but the rest of the server matters more than people expect with Flux.
- Ordinary memory (RAM). ComfyUI's own Flux guide says to use the large 16-bit text encoder only when the server has over 32 GB of memory; below that, use the fp8 version. Many of our GPU builds start at 16 or 32 GB, so ask for more if you want the 16-bit encoder.
- Disk space. A complete Flux setup is about 34 GB before you add LoRAs, upscalers and your own outputs. A collection fills a small disk faster than you would think.
- Both can be chosen up front. When you build to order, you choose the memory, up to 1 TB on the lines our cards go in, and the drives too.
Getting ComfyUI running
- Install NVIDIA's driver and check the card shows up with
nvidia-smi. - Log in as a normal user (we call it
comfyhere) and run the commands below. They download ComfyUI, set up a private Python environment for it, and install PyTorch and everything else it needs. - Start it. Out of the box it only listens to the server itself, on port 8188.
- Copy models into the matching folders:
models/checkpointsfor single-file checkpoints,models/diffusion_models,models/text_encodersandmodels/vaefor Flux's separate files.
sudo apt install -y git python3-venv
git clone https://github.com/Comfy-Org/ComfyUI.git
cd ComfyUI
python3 -m venv venv
. venv/bin/activate
# Turing or newer card:
pip install torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu130
# Maxwell, Pascal or Volta card (Tesla M60, P40, P100, V100, TITAN V): use this line instead
# pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu126
pip install -r requirements.txt
python main.py
Keeping it private
ComfyUI has no accounts and no password. Started with --listen and no address, it answers on every network address the server has, and anyone who finds port 8188 can use your card and run any custom node you have installed. So leave it listening only to itself and reach it through SSH from your own computer:
ssh -N -L 8188:127.0.0.1:8188 you@your-server
Then open http://localhost:8188 in your browser and you are looking at the server. If several people need it, put a reverse proxy in front that asks for a login and uses HTTPS. ComfyUI's own --tls-keyfile and --tls-certfile options give you HTTPS, but still no password.
Habits that save trouble later
- Treat custom nodes like software you install, because they are. Each one is Python code running with your user's permissions, so only add ones you trust. The ComfyUI-Manager add-on locks its riskiest features automatically when ComfyUI listens beyond the server itself.
- Keep loaded models on the card when they fit. By default ComfyUI moves a model back to ordinary memory after using it; starting with
--highvramkeeps it on the card, which saves reload time on a card big enough. - Send outputs to the big disk with
--output-directory, and pointextra_model_paths.yamlat a model library stored elsewhere. - Back up your workflows and your list of custom nodes. Models can be downloaded again; the workflows you built cannot.
Why your own server?
Because the card is all yours. Your own operating system drives it directly, in a dedicated server rented to one account, with no virtual layer, no MIG slice and no time-slicing, so a batch that takes ten minutes today takes ten minutes next week. You install and pin your own driver, and nothing changes it underneath you.
The server's own HPE iLO or Dell iDRAC gives you a console, virtual media and power control, so a driver that breaks the network does not lock you out. The port is unmetered, so downloading models and sending back finished images never adds a traffic charge. The card is a separate line on your bill from the server it sits in, so choosing a bigger card adds only what that card costs over the smaller one, not a jump to a new tier.
When a server is not the answer
If you make a handful of images now and then, a service that charges per image costs less and asks nothing of you. A server pays off when you generate every day, when you use your own models and LoRAs, or when your input images should not leave your hands.
And if your workflows need more than the 96 GB of the largest single card we fit, that becomes a multi-GPU build we quote on request. Check that your workflow and its custom nodes can actually use a second card before you order one.
Questions people ask
- What is the smallest GPU that runs Flux in ComfyUI?
- To hold the 17.2 GB fp8 Flux file on the card, a 24 GB card such as the RTX 3090. Smaller cards can still run Flux because ComfyUI streams weights from ordinary memory, which needs more RAM in the server.
- Can I use a cheaper older card like the Tesla P40?
- Yes. Its 24 GB holds Flux in fp8. Being a Pascal-generation card, it needs PyTorch's CUDA 12.6 build rather than the CUDA 13.0 one in ComfyUI's instructions. The newer build will not drive it.
- Is ComfyUI safe to leave open on the internet?
- No. Nothing in ComfyUI asks for a password; whoever finds it can use your GPU and your custom nodes. Keep it listening only to the server and connect through SSH, or place a login page with HTTPS in front of it.
- How much RAM should the server have for Flux?
- ComfyUI's Flux guide keeps the full 16-bit text encoder for servers with over 32 GB of memory and points everyone else to the fp8 one. More RAM also helps when ComfyUI moves models off the card between steps.
- How soon can I have a GPU server?
- About 30 minutes after your order is approved if a racked machine already carries the card you want. Otherwise a card we hold is fitted in 4 to 24 hours, one that has to be bought in takes 5 to 10 working days, and a current-generation accelerator on allocation takes 10 to 15. When we buy a card in for you, we ask for three months up front.
- Is the GPU shared with other customers?
- No. Your operating system gets the card directly, on a server rented to you alone. There is no MIG slice, no vGPU and no time-slicing.