Cheapest box to run Qwen3-4B at Q6_K
Alibaba’s Qwen3-4B is 4B parameters, dense. The number below is computed from that and each vendor’s published specs, not asserted. Switch the quant or the context and it recomputes in place.
Qwen3-4B at Q6_K needs 6GB.
Cheapest Config that fits: Mac mini, M6, 16GB. Fits, 7GB headroom. Est. 33.7 tok/s.
Best offer today: $1,099 at Apple — in stock, checked 42 hours ago.
Collecting price history — 3 of 90 days. No timing verdict yet.
Offers — Mac mini, M6, 16GB, 512GB
- Applelist pricein stockchecked 42 hours ago$1,099
- Mac mini, M6, 16GB$1,099Fits7GB headroomest. 33.7 tok/s
- Mac mini, M6, 24GB$1,299Fits15GB headroomest. 33.7 tok/s
- Mac mini, M6, 32GB$1,499Fits23GB headroomest. 33.7 tok/s
- Mac mini, M5 Pro, 24GB$1,699Fits15GB headroomest. 60.8 tok/s
- Mac mini, M5 Pro, 48GB$2,499Fits39GB headroomest. 60.8 tok/s
- Mac Studio, M5 Max, 36GB$2,499Fits27GB headroomest. 122 tok/s
- Mac mini, M5 Pro, 64GB$2,899Fits53GB headroomest. 60.8 tok/s
- GeForce RTX 5090, 32GB$2,999.99+hostFits25GB headroomest. 355 tok/s
- Fits53GB headroomest. 108 tok/s
- Mac Studio, M5 Max, 48GB$3,099Fits39GB headroomest. 122 tok/s
- Mac Studio, M5 Max, 64GB$3,499Fits53GB headroomest. 122 tok/s
- Fits115GB headroomest. 54.1 tok/s
- Fits83GB headroomest. 162 tok/s
- Fits114GB headroomest. 108 tok/s
- Fits115GB headroomest. 54.1 tok/s
- Fits114GB headroomest. 122 tok/s
- Fits83GB headroomest. 238 tok/s
- Fits242GB headroomest. 162 tok/s
- Fits242GB headroomest. 238 tok/s
- Fits498GB headroomest. 238 tok/s
| Config | Memory | Price | $/GB | Fit | Headroom | Est. speed |
|---|---|---|---|---|---|---|
| Mac mini, M6 | 16GB | $1,099 | $69 | Fits | 7GB | 33.7 tok/s |
| Mac mini, M6 | 24GB | $1,299 | $54 | Fits | 15GB | 33.7 tok/s |
| Mac mini, M6 | 32GB | $1,499 | $47 | Fits | 23GB | 33.7 tok/s |
| Mac mini, M5 Pro | 24GB | $1,699 | $71 | Fits | 15GB | 60.8 tok/s |
| Mac mini, M5 Pro | 48GB | $2,499 | $52 | Fits | 39GB | 60.8 tok/s |
| Mac Studio, M5 Max | 36GB | $2,499 | $69 | Fits | 27GB | 122 tok/s |
| Mac mini, M5 Pro | 64GB | $2,899 | $45 | Fits | 53GB | 60.8 tok/s |
| GeForce RTX 5090 | 32GB | $2,999.99+host | $94 | Fits | 25GB | 355 tok/s |
| Mac Studio, M4 Max, refurb | 64GB | $3,049 | $48 | Fits | 53GB | 108 tok/s |
| Mac Studio, M5 Max | 48GB | $3,099 | $65 | Fits | 39GB | 122 tok/s |
| Mac Studio, M5 Max | 64GB | $3,499 | $55 | Fits | 53GB | 122 tok/s |
| Ascent GX10, GB10 Grace Blackwell | 128GB | $3,999.99 | $31 | Fits | 115GB | 54.1 tok/s |
| Mac Studio, M3 Ultra, refurb | 96GB | $4,329 | $45 | Fits | 83GB | 162 tok/s |
| Mac Studio, M4 Max, refurb | 128GB | $4,409 | $34 | Fits | 114GB | 108 tok/s |
| DGX Spark, GB10 Grace Blackwell | 128GB | $4,699 | $37 | Fits | 115GB | 54.1 tok/s |
| Mac Studio, M5 Max | 128GB | $5,099 | $40 | Fits | 114GB | 122 tok/s |
| Mac Studio, M5 Ultra | 96GB | $5,499 | $57 | Fits | 83GB | 238 tok/s |
| Mac Studio, M3 Ultra, refurb | 256GB | $7,729 | $30 | Fits | 242GB | 162 tok/s |
| Mac Studio, M5 Ultra | 256GB | $9,499 | $37 | Fits | 242GB | 238 tok/s |
| Mac Studio, M5 Ultra | 512GB | — | — | Fits | 498GB | 238 tok/s |
GeForce RTX 5090 is a card, not a computer: every price shown for it is the cheapest listing we found plus $1,000 for the minimum host machine.
memory = 4B × 0.820 B/param (6.56 bits) + 0.147 GB/1K × 8K K/V + 1.2 GB runtime
speed = bandwidth × 0.65 ÷ (4B active × 0.820 B/param) = 3.3 GB read per token
usable = memory − operating system reserve (8% on a Mac, 6% on a DGX Spark, 4% on a card)
Bandwidth is the vendor’s published figure. Every speed is an estimate and says so. Methodology.
What else Mac mini M6 16GB runs
- gpt-oss-20bQ3_K_M, est. 48.3 tok/s
- Phi-4Q5_K_M, est. 11.1 tok/s
- Qwen3-VL-8B-InstructQ8_0, est. 13.0 tok/s
- Gemma-3n-E4BQ8_0, est. 26.0 tok/s
- Qwen3-4BBF16, est. 13.8 tok/s
Best quant that still fits at 8K context, largest model first. Full list on the Config page.
Bandwidth 170GB/s, usable memory 13GB of 16GB.