Self-host an LLM: the hardware, and what actually runs on it

The question isn't "is it cheaper than an API?" — on cost, the API almost always wins. It's "what can I run on my own machine, and what does that machine cost?" You self-host for control, privacy, offline use and no rate limits. This page starts with the hardware, shows which open-weight models fit each memory pool, and gives the cost read straight. Every figure is computed from the index and dated.

The hardware

What decides everything is the memory pool a model must sit in — VRAM on a discrete card, unified memory on Apple and DGX. Sticker price is also shown inmonths of Claude Max ($200/mo, the top tier) as a familiar yardstick — a price reference, not a capability match. Prices are a dated snapshot; the 2026 memory shortage is inflating all of them.

32 GBVRAM · +PC
$3,500≈ 18 mo of Claude Max
Runs up to Command R (35B) · 30 of our models fit

Consumer Blackwell, 32 GB GDDR7. Fastest single-GPU option for models that fit under 32 GB; VRAM is the wall. Price excludes the rest of the PC and is volatile under the 2026 memory shortage.

36 GBunified
$2,499≈ 12 mo of Claude Max
Runs up to Command R (35B) · 30 of our models fit

Entry Mac Studio, 36 GB unified memory shared between CPU/GPU. Whole computer, near-silent, low power. Bandwidth below a discrete GPU but the unified pool lets mid-size models fit cheaply.

96 GBunified
$5,499≈ 27 mo of Claude Max
Runs up to Mixtral 8x22B Instruct (141B) · 42 of our models fit

Base M5 Ultra, 96 GB unified memory. Runs 70B-class models in 4-bit that a 32 GB GPU cannot hold, as one quiet box.

96 GBVRAM · +PC
$8,500≈ 43 mo of Claude Max
Runs up to Mixtral 8x22B Instruct (141B) · 42 of our models fit

Workstation Blackwell, 96 GB GDDR7 — same silicon and bandwidth as the RTX 5090 but 3× the memory, so ~123B-class 4-bit models fit. Street price $8–9k; NVIDIA MSRP was pushed to $16k during the shortage. Excludes host PC.

128 GBunified
$4,699≈ 23 mo of Claude Max
Runs up to Mixtral 8x22B Instruct (141B) · 42 of our models fit

Desktop Grace-Blackwell mini-PC, 128 GB unified LPDDR5X. MSRP raised from $3,999 to $4,699 in early 2026 on memory-supply constraints. CUDA-native, so the software path matches cloud NVIDIA.

256 GBunified
$9,499≈ 47 mo of Claude Max
Runs up to Llama 4 Maverick (400B) · 48 of our models fit

M5 Ultra with the 256 GB unified upgrade (+$4,000 over the 96 GB base). Holds very large MoE models in 4-bit that otherwise need a multi-GPU server; a 512 GB tier exists but Apple had not priced it at verification.

Bare workstation GPUs (RTX 5090, RTX PRO 6000) need a host PC on top; Apple and DGX machines are complete. Electricity and idle time aren't included — a mostly-idle box loses to the API on cost every time.

What fits on what

Every listed open-weight model, grouped by the memory it needs in 4-bit. Pick your memory budget above, then see what runs. Sizes and the smallest single GPU per model are computed from the index; click a model for the full run guide with int8/fp16.

Runs on a consumer GPU≤ 32 GB in 4-bit — a single RTX 5090 or an entry Mac Studio · 30 models
ModelParamsVRAM (int4)Smallest single GPUAPI /1M
Llama 3.2 3B Instruct (Parasail)3B1.8RTX 4090 (24 GB)$0.38
Ministral 3 3B 2512 (Mistral)3B1.8RTX 4090 (24 GB)$0.2
Granite 4.2 8B (DeepInfra)8B4.8RTX 4090 (24 GB)$0.25
Llama 3.1 8B (Groq)8B4.8RTX 4090 (24 GB)$0.06
Ministral 3 8B 2512 (Mistral)8B4.8RTX 4090 (24 GB)$0.3
Qwen3 VL 8B Instruct (Alibaba)8B4.8RTX 4090 (24 GB)$0.572
Qwen3.5-9B (SiliconFlow)9B5.4RTX 4090 (24 GB)$0.25
Mistral Nemo (DeepInfra)12B7.2RTX 4090 (24 GB)$0.049
Ministral 3 14B 2512 (Mistral)14B8.4RTX 4090 (24 GB)$0.4
Qwen3 14B (DeepInfra)14B8.4RTX 4090 (24 GB)$0.36
GPT-OSS 20B (Groq)20B12RTX 4090 (24 GB)$0.12
Codestral 2508 (Mistral)22B13.2RTX 4090 (24 GB)$1.2
Mistral Small (Mistral)24B14.4RTX 4090 (24 GB)$0.75
Mistral Small 3.2 24B (DeepInfra)24B14.4RTX 4090 (24 GB)$0.275
Voxtral Small 24B 2507 (Mistral)24B14.4RTX 4090 (24 GB)$0.4
Gemma 4 26B A4B (DeepInfra)26B · 4B act15.6RTX 4090 (24 GB)$0.4
Gemma 3 27B (DeepInfra)27B16.2RTX 4090 (24 GB)$0.24
Qwen3.5-27B (Alibaba)27B16.2RTX 4090 (24 GB)$1.755
Qwen3.6 27B (Chutes)27B16.2RTX 4090 (24 GB)$2.3
Qwen3.8 27B (Groq)27B16.2RTX 4090 (24 GB)$2.44
Nemotron 3 Nano 30B A3B (Crusoe)30B · 3B act18RTX 4090 (24 GB)$0.25
Qwen3 30B A3B Instruct 2507 (SiliconFlow)30B · 3B act18RTX 4090 (24 GB)$0.39
Qwen3 Coder 30B A3B Instruct (Novita)30B · 3B act18RTX 4090 (24 GB)$0.34
Qwen3 VL 30B A3B Instruct (Alibaba)30B · 3B act18RTX 4090 (24 GB)$0.65
Qwen3 VL 30B A3B Thinking (Alibaba)30B · 3B act18RTX 4090 (24 GB)$1.29
Gemma 4 31B (DeepInfra)31B18.6RTX 4090 (24 GB)$0.44
Qwen3 32B (DeepInfra)32B19.2RTX 4090 (24 GB)$0.36
Command R (Cohere)35B21RTX 4090 (24 GB)$0.75
Qwen3.5-35B-A3B (DeepInfra)35B · 3B act21RTX 4090 (24 GB)$1.14
Qwen3.6 35B A3B (AkashML)35B · 3B act21RTX 4090 (24 GB)$1
Needs a workstation or a large unified box33–128 GB — RTX PRO 6000, Mac Studio Ultra, or DGX Spark · 12 models
ModelParamsVRAM (int4)Smallest single GPUAPI /1M
Llama 3.1 70B Instruct (DeepInfra)70B42A100 (80 GB)$0.8
Llama 3.3 70B (Groq)70B42A100 (80 GB)$0.42
Qwen 2.5 72B (DeepInfra)72B43.2A100 (80 GB)$0.76
Qwen3 Next 80B A3B Instruct (DeepInfra)80B · 3B act48A100 (80 GB)$0.878
Qwen3 Next 80B A3B Thinking (Google)80B · 3B act48A100 (80 GB)$1.35
GLM 4.5 Air (Novita)106B · 12B act63.6A100 (80 GB)$0.98
Llama 4 Scout (DeepInfra)109B · 17B act65.4A100 (80 GB)$0.4
Command A (Cohere)111B66.6A100 (80 GB)$12.5
GPT-OSS 120B (Groq)120B72A100 (80 GB)$0.2
Nemotron 3 Super (DeepInfra)120B · 12B act72A100 (80 GB)$0.485
Qwen3.5-122B-A10B (SiliconFlow)122B · 10B act73.2A100 (80 GB)$2.34
Mixtral 8x22B Instruct (Mistral)141B · 39B act84.62× H100 (160 GB)$8
Multi-GPU or datacenter only> 128 GB in 4-bit — beyond a single consumer machine · 16 models
ModelParamsVRAM (int4)Smallest single GPUAPI /1M
Qwen3 235B A22B Instruct 2507 (GMICloud)235B · 22B act1412× H100 (160 GB)$0.438
Qwen3 235B A22B Thinking 2507 (Alibaba)235B · 22B act1412× H100 (160 GB)$2.53
Qwen3 VL 235B A22B Instruct (DeepInfra)235B · 22B act1412× H100 (160 GB)$1.08
Qwen3 VL 235B A22B Thinking (Alibaba)235B · 22B act1412× H100 (160 GB)$4.4
Qwen3.5 397B A17B (Alibaba)397B · 17B act238.24× H100 (320 GB)$2.73
Llama 4 Maverick (DeepInfra)400B · 17B act2404× H100 (320 GB)$0.896
MiniMax M1 (Minimax)456B · 46B act273.64× H100 (320 GB)$2.6
Qwen3 Coder (Google Vertex)480B · 35B act2884× H100 (320 GB)$1.3
Nemotron 3 Ultra (DeepInfra)550B · 55B act3308× H100 (640 GB)$2.7
DeepSeek Chat671B · 37B act402.68× H100 (640 GB)$1.21
DeepSeek V3 0324 (DeepInfra)671B · 37B act402.68× H100 (640 GB)$1.14
DeepSeek V3.1 (DeepInfra)671B · 37B act402.68× H100 (640 GB)$1.2
DeepSeek R1 (DeepInfra)671B · 37B act402.68× H100 (640 GB)$2.65
DeepSeek V3 (DeepInfra)671B · 37B act402.68× H100 (640 GB)$1.21
DeepSeek V3.1 Terminus (SiliconFlow)671B · 37B act402.68× H100 (640 GB)$1.25
Kimi K2 Thinking (Google)1000B · 32B act6008× H100 (640 GB)$3.1

So, is it cheaper than an API?

Straight answer: almost never on cost. Even the smallest model here, Llama 3.2 3B Instruct (Parasail), needs about 653M tokens per month before a GPU rented 24/7 beats the cheapest verified API — and bigger models push that into the billions. Put the other way: Mac Studio M5 Ultra (256 GB unified) costs ≈ 47 months of Claude Max, and it runs open-weight models, not Claude. Self-hosting earns its place on control, privacy, data residency, latency and offline use — treat any cost saving as the exception, at very high sustained volume, not the reason.

Method

VRAM ≈ params(B) × bytes/param × 1.2 (overhead: KV cache + activations + CUDA context). bytes/param: int4=0.5 · int8=1 · fp16=2. Break-even = cost of a GPU running 24/7 at the on-demand rate ÷ API price (blended $/1M): the monthly token volume above which a dedicated GPU is cheaper than the lowest-priced API. Assumes the GPU can be used continuously; real throughput (tokens/s) sets an upper bound not modeled here.

Model sizes verified 2026-09-04 (stated in each model's name/card, or an established published spec). GPU rates verified 2026-09-04 — on-demand, cheapest reliable tier: runpod · lambda · vast. Sizes for 27 newer models aren't sourced yet, so they're left out rather than estimated.

Machine purchase prices verified 2026-09-04: dgx_spark · mac_studio · rtx_5090 · rtx_pro_6000. Months of Claude Max = purchase price ÷ $200/mo (Claude Max 20x, the top tier), a price yardstick only.

FAQ

Is self-hosting an LLM cheaper than using an API?

Almost never, on cost alone. Commodity APIs are so cheap per token that a dedicated GPU only breaks even at very high, sustained volume. As of 2026-09-05, even a small model like Llama 3.2 3B Instruct (Parasail) needs on the order of 653M tokens/month before a 24/7 GPU beats the cheapest API. You self-host for control, privacy, data residency, offline use and no rate limits — not for the bill.

What decides whether a model fits on my machine?

The memory pool, not the compute. A model must sit entirely in memory — VRAM on a discrete GPU, unified memory on Apple or DGX. Rule of thumb: memory needed ≈ parameters (in billions) × bytes per parameter × 1.2 overhead, where bytes per parameter are 0.5 for int4, 1 for int8, 2 for fp16. A 70B model in int4 needs roughly 42 GB.

Does a box worth months of Claude Max give me Claude at home?

No — and that's the honest catch. The hardware price is shown in months of Claude Max only as a familiar yardstick. What you actually run at home is open-weight models (Llama, Qwen, Mistral, Gemma…), not frontier Claude or GPT. Self-hosting buys you control and privacy over those open models, not frontier quality for less.