The best local LLM for a 512GB Mac
M3 Ultra 512GB โ about 496.0 GiB usable once macOS has taken its share.
Unified memory is the whole game on Apple Silicon. Your 512GB is shared between macOS, your apps and the model, so the honest budget is closer to 496.0 GiB than to 512. Everything below is sized against that number, at Q4_K_M with an 8K context.
The picks
Best all-round: DeepSeek R1
The largest general-purpose model that still leaves room to work. It loads in 411.3 GiB and generates around 18 tokens/sec on an M3 Ultra 512GB. Full breakdown โ
Best for coding: Qwen2.5-Coder 32B
Trained specifically on code, and worth the swap if that is your workload. It loads in 22.4 GiB and generates around 33 tokens/sec on an M3 Ultra 512GB. Full breakdown โ
Best for reasoning: DeepSeek-R1-Distill-Llama-70B
Thinks before answering; slower per question, better on hard ones. It loads in 45.6 GiB and generates around 15 tokens/sec on an M3 Ultra 512GB. Full breakdown โ
Fastest usable: Llama 3.2 1B
When latency matters more than depth โ voice assistants, autocomplete, agents. It loads in 1.8 GiB and generates around 863 tokens/sec on an M3 Ultra 512GB. Full breakdown โ
Everything that fits in 512GB
| Model | Params | Loaded | Tok/s | Max ctx |
|---|---|---|---|---|
| DeepSeek R1 | 684.49B | 411.3 GiB | 18 | 128K |
| Llama 3.1 405B | 405B | 247.3 GiB | 3 | 128K |
| DeepSeek V4 Flash Vision Exp | 304.65B | 183.9 GiB | 4 | 128K |
| DeepSeek V4 Flash 0731 | 304.18B | 183.7 GiB | 4 | 128K |
| DeepSeek V4 Flash | 290.94B | 175.7 GiB | 4 | 128K |
| gpt-oss-120b | 116.8B | 71.3 GiB | 131 | 128K |
| Qwen2.5 72B | 72.7B | 46.8 GiB | 15 | 128K |
| DeepSeek-R1-Distill-Llama-70B | 70.6B | 45.6 GiB | 15 | 128K |
| Llama 3.3 70B | 70.6B | 45.6 GiB | 15 | 128K |
| Mixtral 8x7B | 46.7B | 29.8 GiB | 52 | 32K |
| DeepSeek-R1-Distill-Qwen-32B | 32.8B | 22.4 GiB | 33 | 128K |
| Qwen2.5 32B | 32.8B | 22.4 GiB | 33 | 128K |
| Qwen2.5-Coder 32B | 32.8B | 22.4 GiB | 33 | 128K |
| Qwen3 32B | 32.8B | 22.4 GiB | 33 | 128K |
| Qwen3 30B-A3B | 30.5B | 19.8 GiB | 203 | 128K |
| Gemma 3 27B | 27.4B | 21.1 GiB | 39 | 128K |
| Gemma 2 27B | 27.2B | 20.0 GiB | 39 | 8K |
| Mistral Small 3 24B | 23.6B | 16.2 GiB | 45 | 32K |
| gpt-oss-20b | 20.9B | 13.7 GiB | 186 | 128K |
| DeepSeek-R1-Distill-Qwen-14B | 14.8B | 11.2 GiB | 72 | 128K |
| Qwen2.5 14B | 14.8B | 11.2 GiB | 72 | 128K |
| Qwen2.5-Coder 14B | 14.8B | 11.2 GiB | 72 | 128K |
| Qwen3 14B | 14.8B | 10.9 GiB | 72 | 128K |
| Phi-4 14B | 14.7B | 11.2 GiB | 73 | 16K |
| Gemma 3 12B | 12.2B | 11.1 GiB | 88 | 128K |
| Gemma 2 9B | 9.24B | 9.0 GiB | 116 | 8K |
| Qwen3 8B | 8.2B | 6.8 GiB | 130 | 128K |
| Llama 3.1 8B | 8.03B | 6.6 GiB | 133 | 128K |
| DeepSeek-R1-Distill-Qwen-7B | 7.62B | 5.8 GiB | 140 | 128K |
| Qwen2.5 7B | 7.62B | 5.8 GiB | 140 | 128K |
| Qwen2.5-Coder 7B | 7.62B | 5.8 GiB | 140 | 128K |
| Mistral 7B v0.3 | 7.25B | 6.1 GiB | 148 | 32K |
| Gemma 3 4B | 4.3B | 4.4 GiB | 249 | 128K |
| Llama 3.2 3B | 3.21B | 3.6 GiB | 333 | 128K |
| Qwen2.5 3B | 3.09B | 2.9 GiB | 346 | 32K |
| Gemma 2 2B | 2.61B | 3.0 GiB | 410 | 8K |
| Qwen2.5 1.5B | 1.54B | 1.9 GiB | 695 | 32K |
| Llama 3.2 1B | 1.24B | 1.8 GiB | 863 | 128K |
What does not fit
| Model | Needs | Short by |
|---|
A model that is a gigabyte or two over can often be rescued by dropping to Q3_K_M or quantising the KV cache. Anything further over than that is better solved by picking a smaller model โ a 14B at Q4 beats a 32B at Q2 on almost every task.