Which AI models actually run on your Mac?
Unified memory decides everything. Pick your Mac, see what fits — with the arithmetic shown, not guessed.
Start with your memory
8 GB
5.0 GiB usable for a model
16 GB
13.0 GiB usable for a model
18 GB
15.0 GiB usable for a model
24 GB
20.4 GiB usable for a model
32 GB
27.2 GiB usable for a model
36 GB
30.6 GiB usable for a model
48 GB
40.8 GiB usable for a model
64 GB
54.4 GiB usable for a model
96 GB
81.6 GiB usable for a model
128 GB
112.0 GiB usable for a model
192 GB
176.0 GiB usable for a model
256 GB
240.0 GiB usable for a model
512 GB
496.0 GiB usable for a model
Frequently checked
| Question | Answer | Tok/s |
|---|---|---|
| Llama 3.1 8B on M4 16GB | Yes, comfortably | 20 |
| Llama 3.3 70B on M4 Max 64GB | Yes, comfortably | 10 |
| Qwen2.5-Coder 32B on M4 Pro 48GB | Yes, comfortably | 11 |
| gpt-oss-20b on M4 24GB | Yes, comfortably | 27 |
| Gemma 3 27B on M4 Max 36GB | Yes, comfortably | 20 |
| Qwen3 32B on M4 Pro 48GB | Yes, comfortably | 11 |
| DeepSeek-R1-Distill-Qwen-14B on M4 24GB | Yes, comfortably | 11 |
| Llama 3.2 3B on M1 8GB | Yes, comfortably | 28 |
| Mistral Small 3 24B on M4 Pro 24GB | Yes, comfortably | 15 |
Why memory, and not the GPU
Generating one token means reading every active parameter out of memory. That makes local inference bandwidth bound, so the two numbers that decide your experience are how much unified memory you have (what fits) and how fast it is (how quickly it runs). GPU core count barely moves the needle for chat.
A Mac mini M4 with 24GB will run an 8B model at reading speed. A MacBook Pro M4 Max with 128GB will run a 70B model at the same speed, because its memory is four and a half times faster. Neither can be upgraded after purchase, which is why the configuration you buy is the decision that matters.