Can It Run?
HomeModels › Mixtral 8x7B

Mixtral 8x7B on a Mac

Mistral AI · 46.7B parameters (12.9B active — mixture of experts) · Apache 2.0 · released 2023-12

Weights @ Q4_K_M
26.6 GiB
the usual download size
KV cache
128 KB
per token, FP16
Max context
32K
tokens
Smallest Mac
36 GB
M3 Max
Short answerA M3 Max 36GB is the smallest Apple Silicon machine that loads Mixtral 8x7B at Q4_K_M with an 8K context, at roughly 19 tokens/sec. MacBook Pro 14" (M3 Max) is the cheapest way to get that configuration.

Every Mac, ranked

MacMemoryNeedsTok/sVerdict
M1 8GB8 GB29.8 GiB4No
M1 16GB16 GB29.8 GiB4No
M2 16GB16 GB29.8 GiB6No
M4 16GB16 GB29.8 GiB8No
M1 Pro 16GB16 GB29.8 GiB13No
M3 Pro 18GB18 GB29.8 GiB9No
M2 24GB24 GB29.8 GiB6No
M4 24GB24 GB29.8 GiB8No
M4 Pro 24GB24 GB29.8 GiB17No
M4 32GB32 GB29.8 GiB8No
M1 Max 32GB32 GB29.8 GiB25No
M2 Max 32GB32 GB29.8 GiB25No
M3 Max 36GB36 GB29.8 GiB19Yes, but tight
M4 Max 36GB36 GB29.8 GiB26Yes, but tight
M4 Pro 48GB48 GB29.8 GiB17Yes, comfortably
M4 Max 48GB48 GB29.8 GiB35Yes, comfortably
M4 Pro 64GB64 GB29.8 GiB17Yes, comfortably
M1 Max 64GB64 GB29.8 GiB25Yes, comfortably
M4 Max 64GB64 GB29.8 GiB35Yes, comfortably
M2 Ultra 64GB64 GB29.8 GiB51Yes, comfortably
M2 Max 96GB96 GB29.8 GiB25Yes, comfortably
M3 Ultra 96GB96 GB29.8 GiB52Yes, comfortably
M3 Max 128GB128 GB29.8 GiB25Yes, comfortably
M4 Max 128GB128 GB29.8 GiB35Yes, comfortably
M2 Ultra 192GB192 GB29.8 GiB51Yes, comfortably
M3 Ultra 256GB256 GB29.8 GiB52Yes, comfortably
M3 Ultra 512GB512 GB29.8 GiB52Yes, comfortably

Download size by quantisation

Weight memory scales linearly with bits per parameter. Everything below assumes an 8K context on top.

QuantisationWeightsTotal @ 8KNotes
FP1687.0 GiB93.1 GiBFull precision. Reference quality, twice the memory of Q8.
Q8_046.2 GiB50.3 GiBIndistinguishable from FP16 in practice, at half the size.
Q6_K35.9 GiB39.5 GiBNear-lossless. A good stop when you have memory to spare.
Q5_K_M31.0 GiB34.3 GiBSlightly better than Q4_K_M, noticeably bigger.
Q4_K_M26.6 GiB29.8 GiBThe default. Best quality-per-gigabyte for most people.
Q3_K_M21.2 GiB24.1 GiBVisible quality loss. Use to squeeze one size class up.
Q2_K16.3 GiB18.9 GiBLast resort. Often worse than a smaller model at Q4.
Mixture of expertsMixtral 8x7B holds 46.7B parameters in memory but only reads about 12.9B per token. You pay the full memory cost of a 46.7B model and get roughly the speed of a 12.9B one — an excellent trade if you have the RAM.