Which AI models run on an M1 Pro 16GB Mac?
M1 Pro · 16-core GPU · 200 GB/s memory bandwidth · MacBook Pro 14" (M1 Pro), MacBook Pro 16" (M1 Pro)
Unified memory
16 GB
soldered, not upgradeable
Usable for a model
13.0 GiB
after macOS takes its share
Bandwidth
200 GB/s
this sets your token speed
Models that fit
14 of 39
comfortably, at Q4_K_M
Short answerThe largest model this machine runs comfortably is Qwen3 14B — 10.9 GiB loaded, about 18 tokens/sec.
Every model on an M1 Pro 16GB
Assumes Q4_K_M weights and an 8K context with an FP16 KV cache. Click any row for the full breakdown and what to do if it does not fit.
| Model | Params | Needs | Tok/s | Verdict |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | 763.21B | 458.6 GiB | 0 | No |
| DeepSeek R1 | 684.49B | 411.3 GiB | 4 | No |
| Llama 3.1 405B | 405B | 247.3 GiB | 1 | No |
| DeepSeek V4 Flash Vision Exp | 304.65B | 183.9 GiB | 1 | No |
| DeepSeek V4 Flash 0731 | 304.18B | 183.7 GiB | 1 | No |
| DeepSeek V4 Flash | 290.94B | 175.7 GiB | 1 | No |
| gpt-oss-120b | 116.8B | 71.3 GiB | 32 | No |
| Qwen2.5 72B | 72.7B | 46.8 GiB | 4 | No |
| DeepSeek-R1-Distill-Llama-70B | 70.6B | 45.6 GiB | 4 | No |
| Llama 3.3 70B | 70.6B | 45.6 GiB | 4 | No |
| Mixtral 8x7B | 46.7B | 29.8 GiB | 13 | No |
| DeepSeek-R1-Distill-Qwen-32B | 32.8B | 22.4 GiB | 8 | No |
| Qwen2.5 32B | 32.8B | 22.4 GiB | 8 | No |
| Qwen2.5-Coder 32B | 32.8B | 22.4 GiB | 8 | No |
| Qwen3 32B | 32.8B | 22.4 GiB | 8 | No |
| Qwen3 30B-A3B | 30.5B | 19.8 GiB | 49 | No |
| Gemma 3 27B | 27.4B | 21.1 GiB | 10 | No |
| Gemma 2 27B | 27.2B | 20.0 GiB | 10 | No |
| Mistral Small 3 24B | 23.6B | 16.2 GiB | 11 | No |
| gpt-oss-20b | 20.9B | 13.7 GiB | 45 | No |
| DeepSeek-R1-Distill-Qwen-14B | 14.8B | 11.2 GiB | 18 | Yes, but tight |
| Qwen2.5 14B | 14.8B | 11.2 GiB | 18 | Yes, but tight |
| Qwen2.5-Coder 14B | 14.8B | 11.2 GiB | 18 | Yes, but tight |
| Qwen3 14B | 14.8B | 10.9 GiB | 18 | Yes, comfortably |
| Phi-4 14B | 14.7B | 11.2 GiB | 18 | Yes, but tight |
| Gemma 3 12B | 12.2B | 11.1 GiB | 21 | Yes, but tight |
| Gemma 2 9B | 9.24B | 9.0 GiB | 28 | Yes, comfortably |
| Qwen3 8B | 8.2B | 6.8 GiB | 32 | Yes, comfortably |
| Llama 3.1 8B | 8.03B | 6.6 GiB | 33 | Yes, comfortably |
| DeepSeek-R1-Distill-Qwen-7B | 7.62B | 5.8 GiB | 34 | Yes, comfortably |
| Qwen2.5 7B | 7.62B | 5.8 GiB | 34 | Yes, comfortably |
| Qwen2.5-Coder 7B | 7.62B | 5.8 GiB | 34 | Yes, comfortably |
| Mistral 7B v0.3 | 7.25B | 6.1 GiB | 36 | Yes, comfortably |
| Gemma 3 4B | 4.3B | 4.4 GiB | 61 | Yes, comfortably |
| Llama 3.2 3B | 3.21B | 3.6 GiB | 81 | Yes, comfortably |
| Qwen2.5 3B | 3.09B | 2.9 GiB | 85 | Yes, comfortably |
| Gemma 2 2B | 2.61B | 3.0 GiB | 100 | Yes, comfortably |
| Qwen2.5 1.5B | 1.54B | 1.9 GiB | 170 | Yes, comfortably |
| Llama 3.2 1B | 1.24B | 1.8 GiB | 211 | Yes, comfortably |
Working with 16 GBQuit your browser before loading a model — Chrome and Safari can hold several gigabytes each. Raising the Metal limit with
sudo sysctl iogpu.wired_limit_mb=13926 gives a model more room, at the cost of leaving macOS less.