Cheapest box to run Qwen3-4B at BF16
Alibaba’s Qwen3-4B is 4B parameters, dense. The number below is computed from that and each vendor’s published specs, not asserted. Switch the quant or the context and it recomputes in place.
Qwen3-4B at BF16 needs 10GB.
Cheapest Config that fits: Mac mini, M6, 24GB. Fits, 11GB headroom. Est. 13.8 tok/s.
Best offer today: $1,299 at Expercom — in stock, checked 1 hour ago.
Collecting price history — 3 of 90 days. No timing verdict yet.
Offers — Mac mini, M6, 24GB, 512GB
- Mac mini, M6, 16GB$1,099Tight3GB headroomest. 13.8 tok/s
- Mac mini, M6, 24GB$1,299Fits11GB headroomest. 13.8 tok/s
- Mac mini, M6, 32GB$1,499Fits19GB headroomest. 13.8 tok/s
- Mac mini, M5 Pro, 24GB$1,699Fits11GB headroomest. 24.9 tok/s
- Mac mini, M5 Pro, 48GB$2,499Fits34GB headroomest. 24.9 tok/s
- Mac Studio, M5 Max, 36GB$2,499Fits23GB headroomest. 49.9 tok/s
- Mac mini, M5 Pro, 64GB$2,899Fits49GB headroomest. 24.9 tok/s
- GeForce RTX 5090, 32GB$2,999.99+hostFits20GB headroomest. 146 tok/s
- Fits49GB headroomest. 44.4 tok/s
- Mac Studio, M5 Max, 48GB$3,099Fits34GB headroomest. 49.9 tok/s
- Mac Studio, M5 Max, 64GB$3,499Fits49GB headroomest. 49.9 tok/s
- Fits110GB headroomest. 22.2 tok/s
- Fits78GB headroomest. 66.5 tok/s
- Fits110GB headroomest. 44.4 tok/s
- Fits110GB headroomest. 22.2 tok/s
- Fits110GB headroomest. 49.9 tok/s
- Fits78GB headroomest. 97.5 tok/s
- Fits238GB headroomest. 66.5 tok/s
- Fits238GB headroomest. 97.5 tok/s
- Fits494GB headroomest. 97.5 tok/s
| Config | Memory | Price | $/GB | Fit | Headroom | Est. speed |
|---|---|---|---|---|---|---|
| Mac mini, M6 | 16GB | $1,099 | $69 | Tight | 3GB | 13.8 tok/s |
| Mac mini, M6 | 24GB | $1,299 | $54 | Fits | 11GB | 13.8 tok/s |
| Mac mini, M6 | 32GB | $1,499 | $47 | Fits | 19GB | 13.8 tok/s |
| Mac mini, M5 Pro | 24GB | $1,699 | $71 | Fits | 11GB | 24.9 tok/s |
| Mac mini, M5 Pro | 48GB | $2,499 | $52 | Fits | 34GB | 24.9 tok/s |
| Mac Studio, M5 Max | 36GB | $2,499 | $69 | Fits | 23GB | 49.9 tok/s |
| Mac mini, M5 Pro | 64GB | $2,899 | $45 | Fits | 49GB | 24.9 tok/s |
| GeForce RTX 5090 | 32GB | $2,999.99+host | $94 | Fits | 20GB | 146 tok/s |
| Mac Studio, M4 Max, refurb | 64GB | $3,049 | $48 | Fits | 49GB | 44.4 tok/s |
| Mac Studio, M5 Max | 48GB | $3,099 | $65 | Fits | 34GB | 49.9 tok/s |
| Mac Studio, M5 Max | 64GB | $3,499 | $55 | Fits | 49GB | 49.9 tok/s |
| Ascent GX10, GB10 Grace Blackwell | 128GB | $3,999.99 | $31 | Fits | 110GB | 22.2 tok/s |
| Mac Studio, M3 Ultra, refurb | 96GB | $4,329 | $45 | Fits | 78GB | 66.5 tok/s |
| Mac Studio, M4 Max, refurb | 128GB | $4,409 | $34 | Fits | 110GB | 44.4 tok/s |
| DGX Spark, GB10 Grace Blackwell | 128GB | $4,699 | $37 | Fits | 110GB | 22.2 tok/s |
| Mac Studio, M5 Max | 128GB | $5,099 | $40 | Fits | 110GB | 49.9 tok/s |
| Mac Studio, M5 Ultra | 96GB | $5,499 | $57 | Fits | 78GB | 97.5 tok/s |
| Mac Studio, M3 Ultra, refurb | 256GB | $7,729 | $30 | Fits | 238GB | 66.5 tok/s |
| Mac Studio, M5 Ultra | 256GB | $9,499 | $37 | Fits | 238GB | 97.5 tok/s |
| Mac Studio, M5 Ultra | 512GB | — | — | Fits | 494GB | 97.5 tok/s |
GeForce RTX 5090 is a card, not a computer: every price shown for it is the cheapest listing we found plus $1,000 for the minimum host machine.
memory = 4B × 2.000 B/param (16 bits) + 0.147 GB/1K × 8K K/V + 1.2 GB runtime
speed = bandwidth × 0.65 ÷ (4B active × 2.000 B/param) = 8.0 GB read per token
usable = memory − operating system reserve (8% on a Mac, 6% on a DGX Spark, 4% on a card)
Bandwidth is the vendor’s published figure. Every speed is an estimate and says so. Methodology.
What else Mac mini M6 24GB runs
- Qwen3-32BQ3_K_M, est. 7.1 tok/s
- Qwen3-30B-A3BQ4_K_M, est. 42.7 tok/s
- Gemma-3-27BQ4_K_M, est. 6.8 tok/s
- Devstral-Small-2507Q5_K_M, est. 6.5 tok/s
- gpt-oss-20bQ6_K, est. 28.8 tok/s
- Phi-4Q8_0, est. 7.4 tok/s
Best quant that still fits at 8K context, largest model first. Full list on the Config page.
Bandwidth 170GB/s, usable memory 21GB of 24GB.