Run Qwen3 Coder locally

Qwen3 Coder is 480B parameters (35B active per token — a mixture-of-experts, so it still needs room for all 480B in memory). Here's the VRAM it needs at each quantization, the smallest GPU that fits, and how a self-hosted box compares with the cheapest verified API.

VRAM by quantization

QuantizationVRAM neededFits onNotes
int4 (Q4)288 GB4× H100 (320 GB) ($7.96/hr)smallest footprint, minor quality loss
int8 (Q8)576 GB8× H100 (640 GB) ($15.92/hr)near-lossless
fp16 (full)1152 GBmulti-node (>640 GB)reference quality

VRAM ≈ params × bytes/param × 1.2 (overhead). See the fullmethod and break-even calculator.

Self-host vs API

Cheapest API for Qwen3 Coder is $1.3 / 1M tokens(blended) via DeepInfra. Running it yourself in int4 fits a4× H100 (320 GB) at $7.96/hr — about $5,810.8/month at 24/7. Those two lines cross at roughly 4.5B tokens/month: below that the API wins on cost, above it the dedicated GPU does (assuming you keep it busy). Tune your own volume in the break-even calculator.

Run it at home

In int4, Qwen3 Coder needs 288 GB — more than any consumer desktop or Apple machine we track holds, so at home it means a multi-GPU server. Renting cloud GPUs by the hour, or the API, is the practical route. See the machines table for the ceiling.

Get the weights

Quantized builds on the model hubs — links search live, so they track new builds as they appear:

License: open weights under Apache 2.0 · terms. Compare all providers on the qwen3-coder API page.