Pre-launch · Fit Finder live, offers pending
Indexed by model, not by SKU
Which Mac runs the model you want, and what it costs today.
You decided to run a 235B-parameter model on your own desk, with nobody metering the tokens. Then came the arithmetic. Which chip, and how much unified memory? Is the 256GB machine enough, or do you need the one that passes $10,000? The memory is soldered, so a wrong answer cannot be fixed later.
Qwen3-235B-A22B at Q4_K_M needs 145GB.
Smallest Config that holds it: Mac Studio, M3 Ultra, 256GB. Fits, 103GB headroom. Est. 30.8 tok/s.
- Mac mini, M432GBWon't fit−116GB headroom
- Mac mini, M4 Pro64GBWon't fit−86GB headroom
- Mac Studio, M4 Max128GBWon't fit−25GB headroom
- Mac Studio, M3 Ultra256GBFits103GB headroomest. 30.8 tok/s
- Mac Studio, M3 Ultra512GBFits359GB headroomest. 30.8 tok/s
| Config | Memory | Bandwidth | Fit | Headroom | Est. speed |
|---|---|---|---|---|---|
| Mac mini, M4 | 32GB | 120GB/s | Won't fit | −116GB | — |
| Mac mini, M4 Pro | 64GB | 273GB/s | Won't fit | −86GB | — |
| Mac Studio, M4 Max | 128GB | 546GB/s | Won't fit | −25GB | — |
| Mac Studio, M3 Ultra | 256GB | 819GB/s | Fits | 103GB | 30.8 tok/s |
| Mac Studio, M3 Ultra | 512GB | 819GB/s | Fits | 359GB | 30.8 tok/s |
memory = 235B × 0.604 B/param (4.83 bits) + 0.193 GB/1K × 8K K/V + 1.2 GB runtime
speed = bandwidth × 0.5 ÷ (22B active × 0.604 B/param) = 13.3 GB read per token
usable = unified memory − macOS reserve (8GB above 100GB, 8% below)
Bandwidth is Apple’s published figure. Efficiency is 0.65 dense and 0.5 MoE, calibrated against llama.cpp runs on Apple silicon. Every speed on this page is an estimate and says so. Reseller prices are not live yet. Fit and Speed are computed from specs, so they are true today.
Every other site is organized by product. We are organized by the model. Pick one and the answer is already on the screen.
- Models in the library
- 9
- Configs in the Grid
- 5
- Reseller offers live
- 0
The gap
The question has two halves, and nobody joins them.
So the buyer overspends by thousands on memory they will never fill, or buys short and finds out the model does not load. Both mistakes are permanent, because Apple solders the memory.
01
Deals sites
AppleInsider, 9to5Toys, MacRumors, Slickdeals
Every price, per SKU, across every reseller.
Nothing about models. A table of Mac Studio prices cannot tell you whether the box loads the weights you downloaded.
02
Memory calculators
Model cards, memory tables, spreadsheets
How many gigabytes a model needs at a given quant.
Nothing about money. They stop at a number, and the number is not a machine you can order tonight.
03
Forum threads
r/LocalLLaMA, llama.cpp benchmark issues
Real measured numbers from people who actually bought one.
Freshness. The good comment is nine months old, three chips ago, and the prices in it are fiction.
The principle
We print the formula instead of asserting the number.
Decode on Apple silicon is bound by memory bandwidth, not by compute. That makes speed computable: bandwidth divided by the bytes read per token, corrected by one efficiency figure calibrated on published llama.cpp runs. The inputs sit on every page, so you can check the estimate rather than trust it.
It matters most where no independent benchmark exists yet. A 512GB M5 Ultra has no measured tok/s in public. It has an arithmetic one, and we will show the arithmetic.
params × bytes-per-weight + K/V per 1K × context + runtimeWeights at the real effective bit width of the quant, not the nominal one. K/V from the model's published layer and head counts.
bandwidth × efficiency ÷ (active params × bytes-per-weight)Active parameters, so a mixture-of-experts model is not charged for weights it never reads. Efficiency is 0.65 dense, 0.50 MoE.
unified memory − macOS reserve − requirement = headroomThree states only: Fits, Tight, Won't fit — with the headroom in GB next to it, so Tight is a number and not an adjective.
Send us a benchmark that contradicts an estimate and we check it. If you are right, the calibration changes and the correction is logged with a date. We never quietly overwrite a wrong number.
What ships
What is live now, and what lands next.
Free, no accounts, no paywall, and nothing to buy from us. We earn a commission when you buy through a reseller link, which is why the ranking is the cheapest that fits and never the best-paying.
Fit Finder
Above, working. Pick a model and a quant; read the smallest Config that holds it and the estimated decode rate.
liveThe Grid
Every buyable Mac — chip, unified memory, bandwidth, storage — ranked by dollars per GB of unified memory and by est. tok/s on your model.
nextOffers
Amazon, B&H, Adorama, Best Buy, Expercom and OWC, each with a price, a stock state and a checked-at time in minutes, not days.
nextModel pages
One page per model and quant: the memory arithmetic, the offers table, and the configs that come up short marked Tight or Won't fit.
nextPrice alert
One email when a Config you chose drops below a target price. One email, then we stop. No account.
laterMethodology
The formula, the calibration sources, and a dated changelog of every correction we make. When a number is wrong we log it, not overwrite it.
later
A 235B-parameter model, on a desk, answering to nobody.
That is the part that is worth the money. Our job is the boring half: which box, how much headroom, how many tokens per second, and which reseller has it in stock this afternoon.
Or write to hello@local-ai-deals.nanocorp.app and name the model you want in the library first.
