Run Ministral 3 3B 2512 locally

Ministral 3 3B 2512 is 3B parameters (dense). Here's the VRAM it needs at each quantization, the smallest GPU that fits, and how a self-hosted box compares with the cheapest verified API.

VRAM by quantization

QuantizationVRAM neededFits onNotes
int4 (Q4)1.8 GBRTX 4090 (24 GB) ($0.34/hr)smallest footprint, minor quality loss
int8 (Q8)3.6 GBRTX 4090 (24 GB) ($0.34/hr)near-lossless
fp16 (full)7.2 GBRTX 4090 (24 GB) ($0.34/hr)reference quality

VRAM ≈ params × bytes/param × 1.2 (overhead). See the fullmethod and break-even calculator.

Self-host vs API

Cheapest API for Ministral 3 3B 2512 is $0.2 / 1M tokens(blended) via Mistral AI. Running it yourself in int4 fits aRTX 4090 (24 GB) at $0.34/hr — about $248.2/month at 24/7. Those two lines cross at roughly 1.2B tokens/month: below that the API wins on cost, above it the dedicated GPU does (assuming you keep it busy). Tune your own volume in the break-even calculator.

Run it at home

No cloud account needed: Ministral 3 3B 2512 in int4 (1.8 GB) fits a RTX 5090 (32 GB) + host PC (plus a host PC), a $3,500 one-off buy. Amortized over 3 years that's about $0.1332/hr — hardware only, electricity aside. Against the Mistral AI API at $0.2/1M, buying it pays for itself after roughly 17.5B tokens total. Below that the API is cheaper; a machine that mostly sits idle rarely earns back its price.

Get the weights

Quantized builds on the model hubs — links search live, so they track new builds as they appear: