gpt-oss-20b on a Mac
OpenAI · 20.9B parameters (3.6B active — mixture of experts) · Apache 2.0 · released 2025-08
Weights @ Q4_K_M
11.9 GiB
the usual download size
KV cache
48 KB
per token, FP16
Max context
128K
tokens
Smallest Mac
18 GB
M3 Pro
Short answerA M3 Pro 18GB is the smallest Apple Silicon machine that loads gpt-oss-20b at Q4_K_M with an 8K context, at roughly 34 tokens/sec. MacBook Pro 14" (M3 Pro) is the cheapest way to get that configuration.
Every Mac, ranked
| Mac | Memory | Needs | Tok/s | Verdict |
|---|---|---|---|---|
| M1 8GB | 8 GB | 13.7 GiB | 15 | No |
| M1 16GB | 16 GB | 13.7 GiB | 15 | No |
| M2 16GB | 16 GB | 13.7 GiB | 23 | No |
| M4 16GB | 16 GB | 13.7 GiB | 27 | No |
| M1 Pro 16GB | 16 GB | 13.7 GiB | 45 | No |
| M3 Pro 18GB | 18 GB | 13.7 GiB | 34 | Yes, but tight |
| M2 24GB | 24 GB | 13.7 GiB | 23 | Yes, comfortably |
| M4 24GB | 24 GB | 13.7 GiB | 27 | Yes, comfortably |
| M4 Pro 24GB | 24 GB | 13.7 GiB | 62 | Yes, comfortably |
| M4 32GB | 32 GB | 13.7 GiB | 27 | Yes, comfortably |
| M1 Max 32GB | 32 GB | 13.7 GiB | 91 | Yes, comfortably |
| M2 Max 32GB | 32 GB | 13.7 GiB | 91 | Yes, comfortably |
| M3 Max 36GB | 36 GB | 13.7 GiB | 68 | Yes, comfortably |
| M4 Max 36GB | 36 GB | 13.7 GiB | 93 | Yes, comfortably |
| M4 Pro 48GB | 48 GB | 13.7 GiB | 62 | Yes, comfortably |
| M4 Max 48GB | 48 GB | 13.7 GiB | 124 | Yes, comfortably |
| M4 Pro 64GB | 64 GB | 13.7 GiB | 62 | Yes, comfortably |
| M1 Max 64GB | 64 GB | 13.7 GiB | 91 | Yes, comfortably |
| M4 Max 64GB | 64 GB | 13.7 GiB | 124 | Yes, comfortably |
| M2 Ultra 64GB | 64 GB | 13.7 GiB | 181 | Yes, comfortably |
| M2 Max 96GB | 96 GB | 13.7 GiB | 91 | Yes, comfortably |
| M3 Ultra 96GB | 96 GB | 13.7 GiB | 186 | Yes, comfortably |
| M3 Max 128GB | 128 GB | 13.7 GiB | 91 | Yes, comfortably |
| M4 Max 128GB | 128 GB | 13.7 GiB | 124 | Yes, comfortably |
| M2 Ultra 192GB | 192 GB | 13.7 GiB | 181 | Yes, comfortably |
| M3 Ultra 256GB | 256 GB | 13.7 GiB | 186 | Yes, comfortably |
| M3 Ultra 512GB | 512 GB | 13.7 GiB | 186 | Yes, comfortably |
Download size by quantisation
Weight memory scales linearly with bits per parameter. Everything below assumes an 8K context on top.
| Quantisation | Weights | Total @ 8K | Notes |
|---|---|---|---|
| FP16 | 38.9 GiB | 42.1 GiB | Full precision. Reference quality, twice the memory of Q8. |
| Q8_0 | 20.7 GiB | 22.9 GiB | Indistinguishable from FP16 in practice, at half the size. |
| Q6_K | 16.1 GiB | 18.0 GiB | Near-lossless. A good stop when you have memory to spare. |
| Q5_K_M | 13.9 GiB | 15.7 GiB | Slightly better than Q4_K_M, noticeably bigger. |
| Q4_K_M | 11.9 GiB | 13.7 GiB | The default. Best quality-per-gigabyte for most people. |
| Q3_K_M | 9.5 GiB | 11.1 GiB | Visible quality loss. Use to squeeze one size class up. |
| Q2_K | 7.3 GiB | 8.8 GiB | Last resort. Often worse than a smaller model at Q4. |
Mixture of expertsgpt-oss-20b holds 20.9B parameters in memory but only reads about 3.6B per token. You pay the full memory cost of a 20.9B model and get roughly the speed of a 3.6B one — an excellent trade if you have the RAM.