Can It Run?
HomeMacs › M1 Pro 16GB

Which AI models run on an M1 Pro 16GB Mac?

M1 Pro · 16-core GPU · 200 GB/s memory bandwidth · MacBook Pro 14" (M1 Pro), MacBook Pro 16" (M1 Pro)

Unified memory
16 GB
soldered, not upgradeable
Usable for a model
13.0 GiB
after macOS takes its share
Bandwidth
200 GB/s
this sets your token speed
Models that fit
14 of 39
comfortably, at Q4_K_M
Short answerThe largest model this machine runs comfortably is Qwen3 14B — 10.9 GiB loaded, about 18 tokens/sec.

Every model on an M1 Pro 16GB

Assumes Q4_K_M weights and an 8K context with an FP16 KV cache. Click any row for the full breakdown and what to do if it does not fit.

ModelParamsNeedsTok/sVerdict
DeepSeek V4.1 Flash763.21B458.6 GiB0No
DeepSeek R1684.49B411.3 GiB4No
Llama 3.1 405B405B247.3 GiB1No
DeepSeek V4 Flash Vision Exp304.65B183.9 GiB1No
DeepSeek V4 Flash 0731304.18B183.7 GiB1No
DeepSeek V4 Flash290.94B175.7 GiB1No
gpt-oss-120b116.8B71.3 GiB32No
Qwen2.5 72B72.7B46.8 GiB4No
DeepSeek-R1-Distill-Llama-70B70.6B45.6 GiB4No
Llama 3.3 70B70.6B45.6 GiB4No
Mixtral 8x7B46.7B29.8 GiB13No
DeepSeek-R1-Distill-Qwen-32B32.8B22.4 GiB8No
Qwen2.5 32B32.8B22.4 GiB8No
Qwen2.5-Coder 32B32.8B22.4 GiB8No
Qwen3 32B32.8B22.4 GiB8No
Qwen3 30B-A3B30.5B19.8 GiB49No
Gemma 3 27B27.4B21.1 GiB10No
Gemma 2 27B27.2B20.0 GiB10No
Mistral Small 3 24B23.6B16.2 GiB11No
gpt-oss-20b20.9B13.7 GiB45No
DeepSeek-R1-Distill-Qwen-14B14.8B11.2 GiB18Yes, but tight
Qwen2.5 14B14.8B11.2 GiB18Yes, but tight
Qwen2.5-Coder 14B14.8B11.2 GiB18Yes, but tight
Qwen3 14B14.8B10.9 GiB18Yes, comfortably
Phi-4 14B14.7B11.2 GiB18Yes, but tight
Gemma 3 12B12.2B11.1 GiB21Yes, but tight
Gemma 2 9B9.24B9.0 GiB28Yes, comfortably
Qwen3 8B8.2B6.8 GiB32Yes, comfortably
Llama 3.1 8B8.03B6.6 GiB33Yes, comfortably
DeepSeek-R1-Distill-Qwen-7B7.62B5.8 GiB34Yes, comfortably
Qwen2.5 7B7.62B5.8 GiB34Yes, comfortably
Qwen2.5-Coder 7B7.62B5.8 GiB34Yes, comfortably
Mistral 7B v0.37.25B6.1 GiB36Yes, comfortably
Gemma 3 4B4.3B4.4 GiB61Yes, comfortably
Llama 3.2 3B3.21B3.6 GiB81Yes, comfortably
Qwen2.5 3B3.09B2.9 GiB85Yes, comfortably
Gemma 2 2B2.61B3.0 GiB100Yes, comfortably
Qwen2.5 1.5B1.54B1.9 GiB170Yes, comfortably
Llama 3.2 1B1.24B1.8 GiB211Yes, comfortably
Working with 16 GBQuit your browser before loading a model — Chrome and Safari can hold several gigabytes each. Raising the Metal limit with sudo sysctl iogpu.wired_limit_mb=13926 gives a model more room, at the cost of leaving macOS less.