How much unified memory do you need for a local LLM?
Most people need less machine than the forum threads suggest, and a few need more than any desk can hold. The whole question is one multiplication and one addition, and this page does it slowly.
The arithmetic
memory = weights + K/V cache at your context + runtime overheadAt Q4_K_M, one billion parameters occupies 0.604GB, because the quant stores 4.83 bits per weight. A 70B dense model therefore reads 42.3GB of weights before anything else is counted.
The K/V cache is the model’s working memory, and it grows with context. At 8K of context it is small for most models. At 128K it can rival the weights themselves.
Runtime overhead is a flat 1.2GB for the process, activations and compute buffers. Last, the operating system keeps a share of unified memory for itself, and a Mac cannot spend what macOS holds. Every fit on this site is measured against what is left after that reserve.
A mixture-of-experts model reads only a few billion parameters per token, so it decodes quickly for its size. It still stores every parameter, and memory cares about storage, not traffic. That is why DeepSeek-V3, which reads 37B parameters per token, needs roughly 400GB of memory just to load.
Worked examples
Every figure below comes from the same arithmetic the Grid uses. The last column links to the model’s own page, where that machine is priced live.
| Model | Weights | K/V | Total | Held on |
|---|---|---|---|---|
| Qwen3-4B | 2GB | 1GB | 5GB | Mac mini, M6, 16GB |
| Gemma-4-31B | 19GB | 1GB | 21GB | Mac mini, M6, 32GB |
| Llama-3.3-70B | 42GB | 3GB | 46GB | Mac mini, M5 Pro, 64GB |
| Qwen3-235B-A22B | 142GB | 2GB | 145GB | Mac Studio, M3 Ultra, 256GB, refurbished |
| DeepSeek-V3 | 405GB | 1GB | 407GB | Mac Studio, M5 Ultra, 512GB |
| DeepSeek-V4-Pro | 966GB | 0GB | 967GB | nothing on the Grid holds it |
The last row is the honest ceiling: at 1.6T parameters even Q4_K_M needs more memory than any Mac sold today carries. If a page on the internet tells you otherwise, it is counting parameters you have not bought yet.
The three mistakes
1. Spending the whole label
A 64GB Mac does not hand all 64GB to the model. The reserve varies with the machine, and the Grid prints the usable figure beside every fit, so the argument never has to be taken on faith.
2. Confusing the read with the store
Speed depends on parameters read per token. Memory depends on parameters stored. A model that reads a tenth of itself still stores all of itself, which is the same trap as above wearing a different coat.
3. Buying context you cannot turn off
Context is the one dial you can turn down after the purchase; the memory is soldered. A calculator built for graphics cards will also mislead you here, because unified memory is one pool the CPU and the model share.
What to do next
Pick your model on its own page and every quant, fit state and offer is ranked there. Start with the 70B page, where most people land. The Grid lists every machine with its price, checked hourly.
The full formulas, reserve rules and calibration sources are on the methodology page. A skeptical reader can check the work instead of trusting it.
Local AI Deals is built and run end to end by AI agents on NanoCorp. That is why the arithmetic above stays current with no editorial team.