Self-host an LLM: the hardware, and what actually runs on it
The question isn't "is it cheaper than an API?" — on cost, the API almost always wins. It's "what can I run on my own machine, and what does that machine cost?" You self-host for control, privacy, offline use and no rate limits. This page starts with the hardware, shows which open-weight models fit each memory pool, and gives the cost read straight. Every figure is computed from the index and dated.
The hardware
What decides everything is the memory pool a model must sit in — VRAM on a discrete card, unified memory on Apple and DGX. Sticker price is also shown inmonths of Claude Max ($200/mo, the top tier) as a familiar yardstick — a price reference, not a capability match. Prices are a dated snapshot; the 2026 memory shortage is inflating all of them.
Consumer Blackwell, 32 GB GDDR7. Fastest single-GPU option for models that fit under 32 GB; VRAM is the wall. Price excludes the rest of the PC and is volatile under the 2026 memory shortage.
Entry Mac Studio, 36 GB unified memory shared between CPU/GPU. Whole computer, near-silent, low power. Bandwidth below a discrete GPU but the unified pool lets mid-size models fit cheaply.
Base M5 Ultra, 96 GB unified memory. Runs 70B-class models in 4-bit that a 32 GB GPU cannot hold, as one quiet box.
Workstation Blackwell, 96 GB GDDR7 — same silicon and bandwidth as the RTX 5090 but 3× the memory, so ~123B-class 4-bit models fit. Street price $8–9k; NVIDIA MSRP was pushed to $16k during the shortage. Excludes host PC.
Desktop Grace-Blackwell mini-PC, 128 GB unified LPDDR5X. MSRP raised from $3,999 to $4,699 in early 2026 on memory-supply constraints. CUDA-native, so the software path matches cloud NVIDIA.
M5 Ultra with the 256 GB unified upgrade (+$4,000 over the 96 GB base). Holds very large MoE models in 4-bit that otherwise need a multi-GPU server; a 512 GB tier exists but Apple had not priced it at verification.
Bare workstation GPUs (RTX 5090, RTX PRO 6000) need a host PC on top; Apple and DGX machines are complete. Electricity and idle time aren't included — a mostly-idle box loses to the API on cost every time.
What fits on what
Every listed open-weight model, grouped by the memory it needs in 4-bit. Pick your memory budget above, then see what runs. Sizes and the smallest single GPU per model are computed from the index; click a model for the full run guide with int8/fp16.
Runs on a consumer GPU≤ 32 GB in 4-bit — a single RTX 5090 or an entry Mac Studio · 30 models
Needs a workstation or a large unified box33–128 GB — RTX PRO 6000, Mac Studio Ultra, or DGX Spark · 12 models
| Model | Params | VRAM (int4) | Smallest single GPU | API /1M |
|---|---|---|---|---|
| Llama 3.1 70B Instruct (DeepInfra) | 70B | 42 | A100 (80 GB) | $0.8 |
| Llama 3.3 70B (Groq) | 70B | 42 | A100 (80 GB) | $0.42 |
| Qwen 2.5 72B (DeepInfra) | 72B | 43.2 | A100 (80 GB) | $0.76 |
| Qwen3 Next 80B A3B Instruct (DeepInfra) | 80B · 3B act | 48 | A100 (80 GB) | $0.878 |
| Qwen3 Next 80B A3B Thinking (Google) | 80B · 3B act | 48 | A100 (80 GB) | $1.35 |
| GLM 4.5 Air (Novita) | 106B · 12B act | 63.6 | A100 (80 GB) | $0.98 |
| Llama 4 Scout (DeepInfra) | 109B · 17B act | 65.4 | A100 (80 GB) | $0.4 |
| Command A (Cohere) | 111B | 66.6 | A100 (80 GB) | $12.5 |
| GPT-OSS 120B (Groq) | 120B | 72 | A100 (80 GB) | $0.2 |
| Nemotron 3 Super (DeepInfra) | 120B · 12B act | 72 | A100 (80 GB) | $0.485 |
| Qwen3.5-122B-A10B (SiliconFlow) | 122B · 10B act | 73.2 | A100 (80 GB) | $2.34 |
| Mixtral 8x22B Instruct (Mistral) | 141B · 39B act | 84.6 | 2× H100 (160 GB) | $8 |
Multi-GPU or datacenter only> 128 GB in 4-bit — beyond a single consumer machine · 16 models
| Model | Params | VRAM (int4) | Smallest single GPU | API /1M |
|---|---|---|---|---|
| Qwen3 235B A22B Instruct 2507 (GMICloud) | 235B · 22B act | 141 | 2× H100 (160 GB) | $0.438 |
| Qwen3 235B A22B Thinking 2507 (Alibaba) | 235B · 22B act | 141 | 2× H100 (160 GB) | $2.53 |
| Qwen3 VL 235B A22B Instruct (DeepInfra) | 235B · 22B act | 141 | 2× H100 (160 GB) | $1.08 |
| Qwen3 VL 235B A22B Thinking (Alibaba) | 235B · 22B act | 141 | 2× H100 (160 GB) | $4.4 |
| Qwen3.5 397B A17B (Alibaba) | 397B · 17B act | 238.2 | 4× H100 (320 GB) | $2.73 |
| Llama 4 Maverick (DeepInfra) | 400B · 17B act | 240 | 4× H100 (320 GB) | $0.896 |
| MiniMax M1 (Minimax) | 456B · 46B act | 273.6 | 4× H100 (320 GB) | $2.6 |
| Qwen3 Coder (Google Vertex) | 480B · 35B act | 288 | 4× H100 (320 GB) | $1.3 |
| Nemotron 3 Ultra (DeepInfra) | 550B · 55B act | 330 | 8× H100 (640 GB) | $2.7 |
| DeepSeek Chat | 671B · 37B act | 402.6 | 8× H100 (640 GB) | $1.21 |
| DeepSeek V3 0324 (DeepInfra) | 671B · 37B act | 402.6 | 8× H100 (640 GB) | $1.14 |
| DeepSeek V3.1 (DeepInfra) | 671B · 37B act | 402.6 | 8× H100 (640 GB) | $1.2 |
| DeepSeek R1 (DeepInfra) | 671B · 37B act | 402.6 | 8× H100 (640 GB) | $2.65 |
| DeepSeek V3 (DeepInfra) | 671B · 37B act | 402.6 | 8× H100 (640 GB) | $1.21 |
| DeepSeek V3.1 Terminus (SiliconFlow) | 671B · 37B act | 402.6 | 8× H100 (640 GB) | $1.25 |
| Kimi K2 Thinking (Google) | 1000B · 32B act | 600 | 8× H100 (640 GB) | $3.1 |
So, is it cheaper than an API?
Straight answer: almost never on cost. Even the smallest model here, Llama 3.2 3B Instruct (Parasail), needs about 653M tokens per month before a GPU rented 24/7 beats the cheapest verified API — and bigger models push that into the billions. Put the other way: Mac Studio M5 Ultra (256 GB unified) costs ≈ 47 months of Claude Max, and it runs open-weight models, not Claude. Self-hosting earns its place on control, privacy, data residency, latency and offline use — treat any cost saving as the exception, at very high sustained volume, not the reason.
Method
VRAM ≈ params(B) × bytes/param × 1.2 (overhead: KV cache + activations + CUDA context). bytes/param: int4=0.5 · int8=1 · fp16=2. Break-even = cost of a GPU running 24/7 at the on-demand rate ÷ API price (blended $/1M): the monthly token volume above which a dedicated GPU is cheaper than the lowest-priced API. Assumes the GPU can be used continuously; real throughput (tokens/s) sets an upper bound not modeled here.
Model sizes verified 2026-09-04 (stated in each model's name/card, or an established published spec). GPU rates verified 2026-09-04 — on-demand, cheapest reliable tier: runpod · lambda · vast. Sizes for 27 newer models aren't sourced yet, so they're left out rather than estimated.
Machine purchase prices verified 2026-09-04: dgx_spark · mac_studio · rtx_5090 · rtx_pro_6000. Months of Claude Max = purchase price ÷ $200/mo (Claude Max 20x, the top tier), a price yardstick only.
FAQ
Is self-hosting an LLM cheaper than using an API?
Almost never, on cost alone. Commodity APIs are so cheap per token that a dedicated GPU only breaks even at very high, sustained volume. As of 2026-09-05, even a small model like Llama 3.2 3B Instruct (Parasail) needs on the order of 653M tokens/month before a 24/7 GPU beats the cheapest API. You self-host for control, privacy, data residency, offline use and no rate limits — not for the bill.
What decides whether a model fits on my machine?
The memory pool, not the compute. A model must sit entirely in memory — VRAM on a discrete GPU, unified memory on Apple or DGX. Rule of thumb: memory needed ≈ parameters (in billions) × bytes per parameter × 1.2 overhead, where bytes per parameter are 0.5 for int4, 1 for int8, 2 for fp16. A 70B model in int4 needs roughly 42 GB.
Does a box worth months of Claude Max give me Claude at home?
No — and that's the honest catch. The hardware price is shown in months of Claude Max only as a familiar yardstick. What you actually run at home is open-weight models (Llama, Qwen, Mistral, Gemma…), not frontier Claude or GPT. Self-hosting buys you control and privacy over those open models, not frontier quality for less.