Can It Run?
HomeModels › gpt-oss-120b

gpt-oss-120b on a Mac

OpenAI · 116.8B parameters (5.1B active — mixture of experts) · Apache 2.0 · released 2025-08

Weights @ Q4_K_M
66.6 GiB
the usual download size
KV cache
72 KB
per token, FP16
Max context
128K
tokens
Smallest Mac
96 GB
M2 Max
Short answerA M2 Max 96GB is the smallest Apple Silicon machine that loads gpt-oss-120b at Q4_K_M with an 8K context, at roughly 64 tokens/sec. MacBook Pro 16" (M2 Max) is the cheapest way to get that configuration.

Every Mac, ranked

MacMemoryNeedsTok/sVerdict
M1 8GB8 GB71.3 GiB11No
M1 16GB16 GB71.3 GiB11No
M2 16GB16 GB71.3 GiB16No
M4 16GB16 GB71.3 GiB19No
M1 Pro 16GB16 GB71.3 GiB32No
M3 Pro 18GB18 GB71.3 GiB24No
M2 24GB24 GB71.3 GiB16No
M4 24GB24 GB71.3 GiB19No
M4 Pro 24GB24 GB71.3 GiB44No
M4 32GB32 GB71.3 GiB19No
M1 Max 32GB32 GB71.3 GiB64No
M2 Max 32GB32 GB71.3 GiB64No
M3 Max 36GB36 GB71.3 GiB48No
M4 Max 36GB36 GB71.3 GiB66No
M4 Pro 48GB48 GB71.3 GiB44No
M4 Max 48GB48 GB71.3 GiB87No
M4 Pro 64GB64 GB71.3 GiB44No
M1 Max 64GB64 GB71.3 GiB64No
M4 Max 64GB64 GB71.3 GiB87No
M2 Ultra 64GB64 GB71.3 GiB128No
M2 Max 96GB96 GB71.3 GiB64Yes, but tight
M3 Ultra 96GB96 GB71.3 GiB131Yes, but tight
M3 Max 128GB128 GB71.3 GiB64Yes, comfortably
M4 Max 128GB128 GB71.3 GiB87Yes, comfortably
M2 Ultra 192GB192 GB71.3 GiB128Yes, comfortably
M3 Ultra 256GB256 GB71.3 GiB131Yes, comfortably
M3 Ultra 512GB512 GB71.3 GiB131Yes, comfortably

Download size by quantisation

Weight memory scales linearly with bits per parameter. Everything below assumes an 8K context on top.

QuantisationWeightsTotal @ 8KNotes
FP16217.6 GiB229.8 GiBFull precision. Reference quality, twice the memory of Q8.
Q8_0115.6 GiB122.7 GiBIndistinguishable from FP16 in practice, at half the size.
Q6_K89.7 GiB95.6 GiBNear-lossless. A good stop when you have memory to spare.
Q5_K_M77.5 GiB82.7 GiBSlightly better than Q4_K_M, noticeably bigger.
Q4_K_M66.6 GiB71.3 GiBThe default. Best quality-per-gigabyte for most people.
Q3_K_M53.0 GiB57.0 GiBVisible quality loss. Use to squeeze one size class up.
Q2_K40.8 GiB44.2 GiBLast resort. Often worse than a smaller model at Q4.
Mixture of expertsgpt-oss-120b holds 116.8B parameters in memory but only reads about 5.1B per token. You pay the full memory cost of a 116.8B model and get roughly the speed of a 5.1B one — an excellent trade if you have the RAM.