Can It Run?
HomeMacs › M3 Ultra 256GB

Which AI models run on an M3 Ultra 256GB Mac?

M3 Ultra · 80-core GPU · 819 GB/s memory bandwidth · Mac Studio (M3 Ultra)

Unified memory
256 GB
soldered, not upgradeable
Usable for a model
240.0 GiB
after macOS takes its share
Bandwidth
819 GB/s
this sets your token speed
Models that fit
36 of 39
comfortably, at Q4_K_M
Short answerThe largest model this machine runs comfortably is DeepSeek V4 Flash Vision Exp — 183.9 GiB loaded, about 4 tokens/sec.

Every model on an M3 Ultra 256GB

Assumes Q4_K_M weights and an 8K context with an FP16 KV cache. Click any row for the full breakdown and what to do if it does not fit.

ModelParamsNeedsTok/sVerdict
DeepSeek V4.1 Flash763.21B458.6 GiB1No
DeepSeek R1684.49B411.3 GiB18No
Llama 3.1 405B405B247.3 GiB3No
DeepSeek V4 Flash Vision Exp304.65B183.9 GiB4Yes, comfortably
DeepSeek V4 Flash 0731304.18B183.7 GiB4Yes, comfortably
DeepSeek V4 Flash290.94B175.7 GiB4Yes, comfortably
gpt-oss-120b116.8B71.3 GiB131Yes, comfortably
Qwen2.5 72B72.7B46.8 GiB15Yes, comfortably
DeepSeek-R1-Distill-Llama-70B70.6B45.6 GiB15Yes, comfortably
Llama 3.3 70B70.6B45.6 GiB15Yes, comfortably
Mixtral 8x7B46.7B29.8 GiB52Yes, comfortably
DeepSeek-R1-Distill-Qwen-32B32.8B22.4 GiB33Yes, comfortably
Qwen2.5 32B32.8B22.4 GiB33Yes, comfortably
Qwen2.5-Coder 32B32.8B22.4 GiB33Yes, comfortably
Qwen3 32B32.8B22.4 GiB33Yes, comfortably
Qwen3 30B-A3B30.5B19.8 GiB203Yes, comfortably
Gemma 3 27B27.4B21.1 GiB39Yes, comfortably
Gemma 2 27B27.2B20.0 GiB39Yes, comfortably
Mistral Small 3 24B23.6B16.2 GiB45Yes, comfortably
gpt-oss-20b20.9B13.7 GiB186Yes, comfortably
DeepSeek-R1-Distill-Qwen-14B14.8B11.2 GiB72Yes, comfortably
Qwen2.5 14B14.8B11.2 GiB72Yes, comfortably
Qwen2.5-Coder 14B14.8B11.2 GiB72Yes, comfortably
Qwen3 14B14.8B10.9 GiB72Yes, comfortably
Phi-4 14B14.7B11.2 GiB73Yes, comfortably
Gemma 3 12B12.2B11.1 GiB88Yes, comfortably
Gemma 2 9B9.24B9.0 GiB116Yes, comfortably
Qwen3 8B8.2B6.8 GiB130Yes, comfortably
Llama 3.1 8B8.03B6.6 GiB133Yes, comfortably
DeepSeek-R1-Distill-Qwen-7B7.62B5.8 GiB140Yes, comfortably
Qwen2.5 7B7.62B5.8 GiB140Yes, comfortably
Qwen2.5-Coder 7B7.62B5.8 GiB140Yes, comfortably
Mistral 7B v0.37.25B6.1 GiB148Yes, comfortably
Gemma 3 4B4.3B4.4 GiB249Yes, comfortably
Llama 3.2 3B3.21B3.6 GiB333Yes, comfortably
Qwen2.5 3B3.09B2.9 GiB346Yes, comfortably
Gemma 2 2B2.61B3.0 GiB410Yes, comfortably
Qwen2.5 1.5B1.54B1.9 GiB695Yes, comfortably
Llama 3.2 1B1.24B1.8 GiB863Yes, comfortably