local ai deals

Pre-launch

Which Mac runs the model you want, and what it costs today.

You decided to run a 235B-parameter model on your own desk, with nobody metering the tokens. Then came the arithmetic. Which chip, and how much unified memory? Is the 256GB machine enough, or do you need the one that passes $10,000? The memory is soldered, so a wrong answer cannot be fixed later.

Fit finderspecs, not prices
Quant
Context

Qwen3-235B-A22B at Q4_K_M needs 145GB.

Smallest Config that holds it: Mac Studio, M3 Ultra, 256GB. Fits, 103GB headroom. Est. 30.8 tok/s.

  • Mac mini, M432GB
    Won't fit116GB headroom
  • Mac mini, M4 Pro64GB
    Won't fit86GB headroom
  • Mac Studio, M4 Max128GB
    Won't fit25GB headroom
  • Mac Studio, M3 Ultra256GB
    Fits103GB headroomest. 30.8 tok/s
  • Mac Studio, M3 Ultra512GB
    Fits359GB headroomest. 30.8 tok/s

memory = 235B × 0.604 B/param (4.83 bits) + 0.193 GB/1K × 8K K/V + 1.2 GB runtime
speed = bandwidth × 0.5 ÷ (22B active × 0.604 B/param) = 13.3 GB read per token
usable = unified memory − macOS reserve (8GB above 100GB, 8% below)

Bandwidth is Apple’s published figure. Efficiency is 0.65 dense and 0.5 MoE, calibrated against llama.cpp runs on Apple silicon. Every speed on this page is an estimate and says so. Reseller prices are not live yet. Fit and Speed are computed from specs, so they are true today.

Every other site is organized by product. We are organized by the model. Pick one and the answer is already on the screen.

Models in the library
9
Configs in the Grid
5
Reseller offers live
0
Nothing to sign up for. The link is the whole ask.

The gap

The question has two halves, and nobody joins them.

So the buyer overspends by thousands on memory they will never fill, or buys short and finds out the model does not load. Both mistakes are permanent, because Apple solders the memory.

01

Deals sites

AppleInsider, 9to5Toys, MacRumors, Slickdeals

Every price, per SKU, across every reseller.

Nothing about models. A table of Mac Studio prices cannot tell you whether the box loads the weights you downloaded.

02

Memory calculators

Model cards, memory tables, spreadsheets

How many gigabytes a model needs at a given quant.

Nothing about money. They stop at a number, and the number is not a machine you can order tonight.

03

Forum threads

r/LocalLLaMA, llama.cpp benchmark issues

Real measured numbers from people who actually bought one.

Freshness. The good comment is nine months old, three chips ago, and the prices in it are fiction.

The principle

We print the formula instead of asserting the number.

Decode on Apple silicon is bound by memory bandwidth, not by compute. That makes speed computable: bandwidth divided by the bytes read per token, corrected by one efficiency figure calibrated on published llama.cpp runs. The inputs sit on every page, so you can check the estimate rather than trust it.

It matters most where no independent benchmark exists yet. A 512GB M5 Ultra has no measured tok/s in public. It has an arithmetic one, and we will show the arithmetic.

Methodology, in three lines
Memoryparams × bytes-per-weight + K/V per 1K × context + runtime

Weights at the real effective bit width of the quant, not the nominal one. K/V from the model's published layer and head counts.

Speedbandwidth × efficiency ÷ (active params × bytes-per-weight)

Active parameters, so a mixture-of-experts model is not charged for weights it never reads. Efficiency is 0.65 dense, 0.50 MoE.

Fitunified memory − macOS reserve − requirement = headroom

Three states only: Fits, Tight, Won't fit — with the headroom in GB next to it, so Tight is a number and not an adjective.

Send us a benchmark that contradicts an estimate and we check it. If you are right, the calibration changes and the correction is logged with a date. We never quietly overwrite a wrong number.

What ships

What is live now, and what lands next.

Free, no accounts, no paywall, and nothing to buy from us. We earn a commission when you buy through a reseller link, which is why the ranking is the cheapest that fits and never the best-paying.

  1. Fit Finder

    Above, working. Pick a model and a quant; read the smallest Config that holds it and the estimated decode rate.

    live
  2. The Grid

    Every buyable Mac — chip, unified memory, bandwidth, storage — ranked by dollars per GB of unified memory and by est. tok/s on your model.

    next
  3. Offers

    Amazon, B&H, Adorama, Best Buy, Expercom and OWC, each with a price, a stock state and a checked-at time in minutes, not days.

    next
  4. Model pages

    One page per model and quant: the memory arithmetic, the offers table, and the configs that come up short marked Tight or Won't fit.

    next
  5. Price alert

    One email when a Config you chose drops below a target price. One email, then we stop. No account.

    later
  6. Methodology

    The formula, the calibration sources, and a dated changelog of every correction we make. When a number is wrong we log it, not overwrite it.

    later

A 235B-parameter model, on a desk, answering to nobody.

That is the part that is worth the money. Our job is the boring half: which box, how much headroom, how many tokens per second, and which reseller has it in stock this afternoon.

Nothing to sign up for. The link is the whole ask.

Or write to hello@local-ai-deals.nanocorp.app and name the model you want in the library first.