Which box runs the model you want, and what it costs today.

You decided to run a 235B-parameter model on your own desk, with nobody metering the tokens. Then came the arithmetic. Which chip, and how much memory? Is a 32GB card enough, or do you need the 256GB machine that passes $9,000? The memory is soldered, so a wrong answer cannot be fixed later.

Fit finderprices checked 2 hours ago
Quant
Context

Qwen3-235B-A22B at Q4_K_M needs 145GB.

Cheapest Config that fits: Mac Studio, M3 Ultra, 256GB, refurbished. Fits, 103GB headroom. Est. 30.8 tok/s.

Best offer today: $7,729 at Apple Certified Refurbished — in stock, checked 2 hours ago.

Price history: 45 days recorded.

GeForce RTX 5090 is a card, not a computer: every price shown for it is the cheapest listing we found plus $1,000 for the minimum host machine.

memory = 235B × 0.604 B/param (4.83 bits) + 0.193 GB/1K × 8K K/V + 1.2 GB runtime
speed = bandwidth × 0.5 ÷ (22B active × 0.604 B/param) = 13.3 GB read per token
usable = memory − operating system reserve (8% on a Mac, 6% on a DGX Spark, 4% on a card)

Bandwidth is the vendor’s published figure. Every speed is an estimate and says so. Methodology.

Every other site is organized by product. We are organized by the model. Pick one and the answer is already on the screen.

Models in the library
30
Configs in the Grid
20
Prices live
38
Cheapest Config
$859

The gap

The question has two halves, and nobody joins them.

So the buyer overspends by thousands on memory they will never fill, or buys short and finds out the model does not load. Both mistakes are permanent, because nobody solders in more later.

01

Deals sites

AppleInsider, 9to5Toys, MacRumors, Slickdeals

Every price, per SKU, across every reseller.

Nothing about models. A table of Mac Studio and RTX 5090 prices cannot tell you whether the box loads the weights you downloaded.

02

Memory calculators

Model cards, memory tables, spreadsheets

How many gigabytes a model needs at a given quant.

Nothing about money. They stop at a number, and the number is not a machine you can order tonight.

03

Forum threads

r/LocalLLaMA, llama.cpp benchmark issues

Real measured numbers from people who actually bought one.

Freshness. The good comment is nine months old, three chips ago, and the prices in it are fiction.

The principle

We print the formula instead of asserting the number.

Decode is bound by memory bandwidth, not by compute, on Apple silicon and on a Blackwell card alike. That makes speed computable: bandwidth divided by the bytes read per token, corrected by one efficiency figure calibrated on published llama.cpp runs. The inputs sit on every page, so you can check the estimate rather than trust it.

It matters most where no independent benchmark exists yet. A 512GB M5 Ultra has no measured tok/s in public. It has an arithmetic one, and we will show the arithmetic.

Methodology, in three lines
Memoryparams × bytes-per-weight + K/V per 1K × context + runtime

Weights at the real effective bit width of the quant, not the nominal one. K/V from the model's published layer and head counts.

Speedbandwidth × efficiency ÷ (active params × bytes-per-weight)

Active parameters, so a mixture-of-experts model is not charged for weights it never reads. Efficiency is 0.65 dense, 0.50 MoE.

Fitmemory − operating system reserve − requirement = headroom

Three states only: Fits, Tight, Won't fit — with the headroom in GB next to it, so Tight is a number and not an adjective.

Send us a benchmark that contradicts an estimate and we check it. If you are right, the calibration changes and the correction is logged with a date. We never quietly overwrite a wrong number.

A 235B-parameter model, on a desk, answering to nobody.

That is the part that is worth the money. Our job is the boring half: which box, how much headroom, how many tokens per second, and which reseller has it in stock this afternoon.

Nothing to sign up for. The link is the whole ask.

Or write to hello@local-ai-deals.nanocorp.app and name the model you want in the library first.