Which AI models run on an M3 Pro 18GB Mac?
M3 Pro · 14-core GPU · 150 GB/s memory bandwidth · MacBook Pro 14" (M3 Pro), MacBook Pro 16" (M3 Pro)
Unified memory
18 GB
soldered, not upgradeable
Usable for a model
15.0 GiB
after macOS takes its share
Bandwidth
150 GB/s
this sets your token speed
Models that fit
19 of 39
comfortably, at Q4_K_M
Short answerThe largest model this machine runs comfortably is DeepSeek-R1-Distill-Qwen-14B — 11.2 GiB loaded, about 13 tokens/sec.
Every model on an M3 Pro 18GB
Assumes Q4_K_M weights and an 8K context with an FP16 KV cache. Click any row for the full breakdown and what to do if it does not fit.
| Model | Params | Needs | Tok/s | Verdict |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | 763.21B | 458.6 GiB | 0 | No |
| DeepSeek R1 | 684.49B | 411.3 GiB | 3 | No |
| Llama 3.1 405B | 405B | 247.3 GiB | 0 | No |
| DeepSeek V4 Flash Vision Exp | 304.65B | 183.9 GiB | 1 | No |
| DeepSeek V4 Flash 0731 | 304.18B | 183.7 GiB | 1 | No |
| DeepSeek V4 Flash | 290.94B | 175.7 GiB | 1 | No |
| gpt-oss-120b | 116.8B | 71.3 GiB | 24 | No |
| Qwen2.5 72B | 72.7B | 46.8 GiB | 3 | No |
| DeepSeek-R1-Distill-Llama-70B | 70.6B | 45.6 GiB | 3 | No |
| Llama 3.3 70B | 70.6B | 45.6 GiB | 3 | No |
| Mixtral 8x7B | 46.7B | 29.8 GiB | 9 | No |
| DeepSeek-R1-Distill-Qwen-32B | 32.8B | 22.4 GiB | 6 | No |
| Qwen2.5 32B | 32.8B | 22.4 GiB | 6 | No |
| Qwen2.5-Coder 32B | 32.8B | 22.4 GiB | 6 | No |
| Qwen3 32B | 32.8B | 22.4 GiB | 6 | No |
| Qwen3 30B-A3B | 30.5B | 19.8 GiB | 37 | No |
| Gemma 3 27B | 27.4B | 21.1 GiB | 7 | No |
| Gemma 2 27B | 27.2B | 20.0 GiB | 7 | No |
| Mistral Small 3 24B | 23.6B | 16.2 GiB | 8 | No |
| gpt-oss-20b | 20.9B | 13.7 GiB | 34 | Yes, but tight |
| DeepSeek-R1-Distill-Qwen-14B | 14.8B | 11.2 GiB | 13 | Yes, comfortably |
| Qwen2.5 14B | 14.8B | 11.2 GiB | 13 | Yes, comfortably |
| Qwen2.5-Coder 14B | 14.8B | 11.2 GiB | 13 | Yes, comfortably |
| Qwen3 14B | 14.8B | 10.9 GiB | 13 | Yes, comfortably |
| Phi-4 14B | 14.7B | 11.2 GiB | 13 | Yes, comfortably |
| Gemma 3 12B | 12.2B | 11.1 GiB | 16 | Yes, comfortably |
| Gemma 2 9B | 9.24B | 9.0 GiB | 21 | Yes, comfortably |
| Qwen3 8B | 8.2B | 6.8 GiB | 24 | Yes, comfortably |
| Llama 3.1 8B | 8.03B | 6.6 GiB | 24 | Yes, comfortably |
| DeepSeek-R1-Distill-Qwen-7B | 7.62B | 5.8 GiB | 26 | Yes, comfortably |
| Qwen2.5 7B | 7.62B | 5.8 GiB | 26 | Yes, comfortably |
| Qwen2.5-Coder 7B | 7.62B | 5.8 GiB | 26 | Yes, comfortably |
| Mistral 7B v0.3 | 7.25B | 6.1 GiB | 27 | Yes, comfortably |
| Gemma 3 4B | 4.3B | 4.4 GiB | 46 | Yes, comfortably |
| Llama 3.2 3B | 3.21B | 3.6 GiB | 61 | Yes, comfortably |
| Qwen2.5 3B | 3.09B | 2.9 GiB | 63 | Yes, comfortably |
| Gemma 2 2B | 2.61B | 3.0 GiB | 75 | Yes, comfortably |
| Qwen2.5 1.5B | 1.54B | 1.9 GiB | 127 | Yes, comfortably |
| Llama 3.2 1B | 1.24B | 1.8 GiB | 158 | Yes, comfortably |