Can It Run?
Home โ€บ Guides โ€บ 96GB

The best local LLM for a 96GB Mac

M3 Ultra 96GB, M2 Max 96GB โ€” about 81.6 GiB usable once macOS has taken its share.

Unified memory is the whole game on Apple Silicon. Your 96GB is shared between macOS, your apps and the model, so the honest budget is closer to 81.6 GiB than to 96. Everything below is sized against that number, at Q4_K_M with an 8K context.

The picks

Best all-round: Qwen2.5 72B

The largest general-purpose model that still leaves room to work. It loads in 46.8 GiB and generates around 15 tokens/sec on an M3 Ultra 96GB. Full breakdown โ†’

Best for coding: Qwen2.5-Coder 32B

Trained specifically on code, and worth the swap if that is your workload. It loads in 22.4 GiB and generates around 33 tokens/sec on an M3 Ultra 96GB. Full breakdown โ†’

Best for reasoning: DeepSeek-R1-Distill-Llama-70B

Thinks before answering; slower per question, better on hard ones. It loads in 45.6 GiB and generates around 15 tokens/sec on an M3 Ultra 96GB. Full breakdown โ†’

Fastest usable: Llama 3.2 1B

When latency matters more than depth โ€” voice assistants, autocomplete, agents. It loads in 1.8 GiB and generates around 863 tokens/sec on an M3 Ultra 96GB. Full breakdown โ†’

Everything that fits in 96GB

ModelParamsLoadedTok/sMax ctx
Qwen2.5 72B72.7B46.8 GiB1564K
DeepSeek-R1-Distill-Llama-70B70.6B45.6 GiB1564K
Llama 3.3 70B70.6B45.6 GiB1564K
Mixtral 8x7B46.7B29.8 GiB5232K
DeepSeek-R1-Distill-Qwen-32B32.8B22.4 GiB33128K
Qwen2.5 32B32.8B22.4 GiB33128K
Qwen2.5-Coder 32B32.8B22.4 GiB33128K
Qwen3 32B32.8B22.4 GiB33128K
Qwen3 30B-A3B30.5B19.8 GiB203128K
Gemma 3 27B27.4B21.1 GiB3964K
Gemma 2 27B27.2B20.0 GiB398K
Mistral Small 3 24B23.6B16.2 GiB4532K
gpt-oss-20b20.9B13.7 GiB186128K
DeepSeek-R1-Distill-Qwen-14B14.8B11.2 GiB72128K
Qwen2.5 14B14.8B11.2 GiB72128K
Qwen2.5-Coder 14B14.8B11.2 GiB72128K
Qwen3 14B14.8B10.9 GiB72128K
Phi-4 14B14.7B11.2 GiB7316K
Gemma 3 12B12.2B11.1 GiB88128K
Gemma 2 9B9.24B9.0 GiB1168K
Qwen3 8B8.2B6.8 GiB130128K
Llama 3.1 8B8.03B6.6 GiB133128K
DeepSeek-R1-Distill-Qwen-7B7.62B5.8 GiB140128K
Qwen2.5 7B7.62B5.8 GiB140128K
Qwen2.5-Coder 7B7.62B5.8 GiB140128K
Mistral 7B v0.37.25B6.1 GiB14832K
Gemma 3 4B4.3B4.4 GiB249128K
Llama 3.2 3B3.21B3.6 GiB333128K
Qwen2.5 3B3.09B2.9 GiB34632K
Gemma 2 2B2.61B3.0 GiB4108K
Qwen2.5 1.5B1.54B1.9 GiB69532K
Llama 3.2 1B1.24B1.8 GiB863128K

What does not fit

ModelNeedsShort by
DeepSeek V4 Flash175.7 GiB94.1 GiB
DeepSeek V4 Flash 0731183.7 GiB102.1 GiB
DeepSeek V4 Flash Vision Exp183.9 GiB102.3 GiB
Llama 3.1 405B247.3 GiB165.7 GiB
DeepSeek R1411.3 GiB329.7 GiB
DeepSeek V4.1 Flash458.6 GiB377.0 GiB

A model that is a gigabyte or two over can often be rescued by dropping to Q3_K_M or quantising the KV cache. Anything further over than that is better solved by picking a smaller model โ€” a 14B at Q4 beats a 32B at Q2 on almost every task.

Machines in this tier

M3 Ultra 96GB
819 GB/s ยท Mac Studio (M3 Ultra)
M2 Max 96GB
400 GB/s ยท MacBook Pro 16" (M2 Max)