Run Qwen3.5-35B-A3B locally
Qwen3.5-35B-A3B is 35B parameters (3B active per token — a mixture-of-experts, so it still needs room for all 35B in memory). Here's the VRAM it needs at each quantization, the smallest GPU that fits, and how a self-hosted box compares with the cheapest verified API.
VRAM by quantization
| Quantization | VRAM needed | Fits on | Notes |
|---|---|---|---|
| int4 (Q4) | 21 GB | RTX 4090 (24 GB) ($0.34/hr) | smallest footprint, minor quality loss |
| int8 (Q8) | 42 GB | A100 (80 GB) ($1.19/hr) | near-lossless |
| fp16 (full) | 84 GB | 2× H100 (160 GB) ($3.98/hr) | reference quality |
VRAM ≈ params × bytes/param × 1.2 (overhead). See the fullmethod and break-even calculator.
Self-host vs API
Cheapest API for Qwen3.5-35B-A3B is $1.14 / 1M tokens(blended) via DeepInfra. Running it yourself in int4 fits aRTX 4090 (24 GB) at $0.34/hr — about $248.2/month at 24/7. Those two lines cross at roughly 218M tokens/month: below that the API wins on cost, above it the dedicated GPU does (assuming you keep it busy). Tune your own volume in the break-even calculator.
Run it at home
No cloud account needed: Qwen3.5-35B-A3B in int4 (21 GB) fits a RTX 5090 (32 GB) + host PC (plus a host PC), a $3,500 one-off buy. Amortized over 3 years that's about $0.1332/hr — hardware only, electricity aside. Against the DeepInfra API at $1.14/1M, buying it pays for itself after roughly 3.1B tokens total. Below that the API is cheaper; a machine that mostly sits idle rarely earns back its price.
Get the weights
Quantized builds on the model hubs — links search live, so they track new builds as they appear:
- GGUF builds llama.cpp · Ollama · LM Studio
- AWQ builds GPU serving (vLLM)
- GPTQ builds GPU serving
- MLX builds Apple silicon
- Ollama library one-command local run
License: open weights under Apache 2.0 · terms. Compare all providers on the qwen3.5-35b-a3b API page.