The best local LLM for a 8GB Mac
M1 8GB โ about 5.0 GiB usable once macOS has taken its share.
Unified memory is the whole game on Apple Silicon. Your 8GB is shared between macOS, your apps and the model, so the honest budget is closer to 5.0 GiB than to 8. Everything below is sized against that number, at Q4_K_M with an 8K context.
The picks
Best all-round: Llama 3.2 3B
The largest general-purpose model that still leaves room to work. It loads in 3.6 GiB and generates around 28 tokens/sec on an M1 8GB. Full breakdown โ
Fastest usable: Llama 3.2 1B
When latency matters more than depth โ voice assistants, autocomplete, agents. It loads in 1.8 GiB and generates around 72 tokens/sec on an M1 8GB. Full breakdown โ
Everything that fits in 8GB
| Model | Params | Loaded | Tok/s | Max ctx |
|---|---|---|---|---|
| Llama 3.2 3B | 3.21B | 3.6 GiB | 28 | 16K |
| Qwen2.5 3B | 3.09B | 2.9 GiB | 29 | 32K |
| Gemma 2 2B | 2.61B | 3.0 GiB | 34 | 8K |
| Qwen2.5 1.5B | 1.54B | 1.9 GiB | 58 | 32K |
| Llama 3.2 1B | 1.24B | 1.8 GiB | 72 | 32K |
What does not fit
| Model | Needs | Short by |
|---|---|---|
| DeepSeek-R1-Distill-Qwen-7B | 5.8 GiB | 0.8 GiB |
| Qwen2.5 7B | 5.8 GiB | 0.8 GiB |
| Qwen2.5-Coder 7B | 5.8 GiB | 0.8 GiB |
| Mistral 7B v0.3 | 6.1 GiB | 1.1 GiB |
| Llama 3.1 8B | 6.6 GiB | 1.6 GiB |
| Qwen3 8B | 6.8 GiB | 1.8 GiB |
| Gemma 2 9B | 9.0 GiB | 4.0 GiB |
| Qwen3 14B | 10.9 GiB | 5.9 GiB |
| Gemma 3 12B | 11.1 GiB | 6.1 GiB |
| DeepSeek-R1-Distill-Qwen-14B | 11.2 GiB | 6.2 GiB |
A model that is a gigabyte or two over can often be rescued by dropping to Q3_K_M or quantising the KV cache. Anything further over than that is better solved by picking a smaller model โ a 14B at Q4 beats a 32B at Q2 on almost every task.