Can It Run?
HomeGuides › 24GB

The best local LLM for a 24GB Mac

M4 Pro 24GB, M4 24GB, M2 24GB — about 20.4 GiB usable once macOS has taken its share.

Unified memory is the whole game on Apple Silicon. Your 24GB is shared between macOS, your apps and the model, so the honest budget is closer to 20.4 GiB than to 24. Everything below is sized against that number, at Q4_K_M with an 8K context.

The picks

Best all-round: Mistral Small 3 24B

The largest general-purpose model that still leaves room to work. It loads in 16.2 GiB and generates around 15 tokens/sec on an M4 Pro 24GB. Full breakdown →

Best for coding: Qwen2.5-Coder 14B

Trained specifically on code, and worth the swap if that is your workload. It loads in 11.2 GiB and generates around 24 tokens/sec on an M4 Pro 24GB. Full breakdown →

Best for reasoning: DeepSeek-R1-Distill-Qwen-14B

Thinks before answering; slower per question, better on hard ones. It loads in 11.2 GiB and generates around 24 tokens/sec on an M4 Pro 24GB. Full breakdown →

Fastest usable: Llama 3.2 1B

When latency matters more than depth — voice assistants, autocomplete, agents. It loads in 1.8 GiB and generates around 288 tokens/sec on an M4 Pro 24GB. Full breakdown →

Everything that fits in 24GB

ModelParamsLoadedTok/sMax ctx
Mistral Small 3 24B23.6B16.2 GiB1516K
gpt-oss-20b20.9B13.7 GiB6264K
DeepSeek-R1-Distill-Qwen-14B14.8B11.2 GiB2432K
Qwen2.5 14B14.8B11.2 GiB2432K
Qwen2.5-Coder 14B14.8B11.2 GiB2432K
Qwen3 14B14.8B10.9 GiB2432K
Phi-4 14B14.7B11.2 GiB2416K
Gemma 3 12B12.2B11.1 GiB2916K
Gemma 2 9B9.24B9.0 GiB398K
Qwen3 8B8.2B6.8 GiB4364K
Llama 3.1 8B8.03B6.6 GiB4464K
DeepSeek-R1-Distill-Qwen-7B7.62B5.8 GiB47128K
Qwen2.5 7B7.62B5.8 GiB47128K
Qwen2.5-Coder 7B7.62B5.8 GiB47128K
Mistral 7B v0.37.25B6.1 GiB4932K
Gemma 3 4B4.3B4.4 GiB8364K
Llama 3.2 3B3.21B3.6 GiB11164K
Qwen2.5 3B3.09B2.9 GiB11532K
Gemma 2 2B2.61B3.0 GiB1378K
Qwen2.5 1.5B1.54B1.9 GiB23232K
Llama 3.2 1B1.24B1.8 GiB288128K

What does not fit

ModelNeedsShort by
Gemma 3 27B21.1 GiB0.7 GiB
DeepSeek-R1-Distill-Qwen-32B22.4 GiB2.0 GiB
Qwen2.5 32B22.4 GiB2.0 GiB
Qwen2.5-Coder 32B22.4 GiB2.0 GiB
Qwen3 32B22.4 GiB2.0 GiB
Mixtral 8x7B29.8 GiB9.4 GiB
DeepSeek-R1-Distill-Llama-70B45.6 GiB25.2 GiB
Llama 3.3 70B45.6 GiB25.2 GiB
Qwen2.5 72B46.8 GiB26.4 GiB
gpt-oss-120b71.3 GiB50.9 GiB

A model that is a gigabyte or two over can often be rescued by dropping to Q3_K_M or quantising the KV cache. Anything further over than that is better solved by picking a smaller model — a 14B at Q4 beats a 32B at Q2 on almost every task.

Machines in this tier

M4 Pro 24GB
273 GB/s · Mac mini (M4 Pro)
M4 24GB
120 GB/s · MacBook Air 15" (M4)
M2 24GB
100 GB/s · MacBook Air 15" (M2)