Can It Run?
Home โ€บ Guides โ€บ 18GB

The best local LLM for a 18GB Mac

M3 Pro 18GB โ€” about 15.0 GiB usable once macOS has taken its share.

Unified memory is the whole game on Apple Silicon. Your 18GB is shared between macOS, your apps and the model, so the honest budget is closer to 15.0 GiB than to 18. Everything below is sized against that number, at Q4_K_M with an 8K context.

The picks

Best all-round: DeepSeek-R1-Distill-Qwen-14B

The largest general-purpose model that still leaves room to work. It loads in 11.2 GiB and generates around 13 tokens/sec on an M3 Pro 18GB. Full breakdown โ†’

Best for coding: Qwen2.5-Coder 14B

Trained specifically on code, and worth the swap if that is your workload. It loads in 11.2 GiB and generates around 13 tokens/sec on an M3 Pro 18GB. Full breakdown โ†’

Best for reasoning: DeepSeek-R1-Distill-Qwen-14B

Thinks before answering; slower per question, better on hard ones. It loads in 11.2 GiB and generates around 13 tokens/sec on an M3 Pro 18GB. Full breakdown โ†’

Fastest usable: Llama 3.2 1B

When latency matters more than depth โ€” voice assistants, autocomplete, agents. It loads in 1.8 GiB and generates around 158 tokens/sec on an M3 Pro 18GB. Full breakdown โ†’

Everything that fits in 18GB

ModelParamsLoadedTok/sMax ctx
DeepSeek-R1-Distill-Qwen-14B14.8B11.2 GiB1316K
Qwen2.5 14B14.8B11.2 GiB1316K
Qwen2.5-Coder 14B14.8B11.2 GiB1316K
Qwen3 14B14.8B10.9 GiB1316K
Phi-4 14B14.7B11.2 GiB1316K
Gemma 3 12B12.2B11.1 GiB1616K
Gemma 2 9B9.24B9.0 GiB218K
Qwen3 8B8.2B6.8 GiB2432K
Llama 3.1 8B8.03B6.6 GiB2432K
DeepSeek-R1-Distill-Qwen-7B7.62B5.8 GiB2664K
Qwen2.5 7B7.62B5.8 GiB2664K
Qwen2.5-Coder 7B7.62B5.8 GiB2664K
Mistral 7B v0.37.25B6.1 GiB2732K
Gemma 3 4B4.3B4.4 GiB4664K
Llama 3.2 3B3.21B3.6 GiB6164K
Qwen2.5 3B3.09B2.9 GiB6332K
Gemma 2 2B2.61B3.0 GiB758K
Qwen2.5 1.5B1.54B1.9 GiB12732K
Llama 3.2 1B1.24B1.8 GiB158128K

What does not fit

ModelNeedsShort by
Mistral Small 3 24B16.2 GiB1.2 GiB
Qwen3 30B-A3B19.8 GiB4.8 GiB
Gemma 2 27B20.0 GiB5.0 GiB
Gemma 3 27B21.1 GiB6.1 GiB
DeepSeek-R1-Distill-Qwen-32B22.4 GiB7.4 GiB
Qwen2.5 32B22.4 GiB7.4 GiB
Qwen2.5-Coder 32B22.4 GiB7.4 GiB
Qwen3 32B22.4 GiB7.4 GiB
Mixtral 8x7B29.8 GiB14.8 GiB
DeepSeek-R1-Distill-Llama-70B45.6 GiB30.6 GiB

A model that is a gigabyte or two over can often be rescued by dropping to Q3_K_M or quantising the KV cache. Anything further over than that is better solved by picking a smaller model โ€” a 14B at Q4 beats a 32B at Q2 on almost every task.

Machines in this tier

M3 Pro 18GB
150 GB/s ยท MacBook Pro 14" (M3 Pro)