Qwen2.5 3B on a Mac
Alibaba · 3.09B parameters · Qwen Research · released 2024-09
Weights @ Q4_K_M
1.8 GiB
the usual download size
KV cache
36 KB
per token, FP16
Max context
32K
tokens
Smallest Mac
8 GB
M1
Short answerA M1 8GB is the smallest Apple Silicon machine that loads Qwen2.5 3B at Q4_K_M with an 8K context, at roughly 29 tokens/sec. MacBook Air (M1) is the cheapest way to get that configuration.
Every Mac, ranked
| Mac | Memory | Needs | Tok/s | Verdict |
|---|---|---|---|---|
| M1 8GB | 8 GB | 2.9 GiB | 29 | Yes, comfortably |
| M1 16GB | 16 GB | 2.9 GiB | 29 | Yes, comfortably |
| M2 16GB | 16 GB | 2.9 GiB | 42 | Yes, comfortably |
| M4 16GB | 16 GB | 2.9 GiB | 51 | Yes, comfortably |
| M1 Pro 16GB | 16 GB | 2.9 GiB | 85 | Yes, comfortably |
| M3 Pro 18GB | 18 GB | 2.9 GiB | 63 | Yes, comfortably |
| M2 24GB | 24 GB | 2.9 GiB | 42 | Yes, comfortably |
| M4 24GB | 24 GB | 2.9 GiB | 51 | Yes, comfortably |
| M4 Pro 24GB | 24 GB | 2.9 GiB | 115 | Yes, comfortably |
| M4 32GB | 32 GB | 2.9 GiB | 51 | Yes, comfortably |
| M1 Max 32GB | 32 GB | 2.9 GiB | 169 | Yes, comfortably |
| M2 Max 32GB | 32 GB | 2.9 GiB | 169 | Yes, comfortably |
| M3 Max 36GB | 36 GB | 2.9 GiB | 127 | Yes, comfortably |
| M4 Max 36GB | 36 GB | 2.9 GiB | 173 | Yes, comfortably |
| M4 Pro 48GB | 48 GB | 2.9 GiB | 115 | Yes, comfortably |
| M4 Max 48GB | 48 GB | 2.9 GiB | 231 | Yes, comfortably |
| M4 Pro 64GB | 64 GB | 2.9 GiB | 115 | Yes, comfortably |
| M1 Max 64GB | 64 GB | 2.9 GiB | 169 | Yes, comfortably |
| M4 Max 64GB | 64 GB | 2.9 GiB | 231 | Yes, comfortably |
| M2 Ultra 64GB | 64 GB | 2.9 GiB | 338 | Yes, comfortably |
| M2 Max 96GB | 96 GB | 2.9 GiB | 169 | Yes, comfortably |
| M3 Ultra 96GB | 96 GB | 2.9 GiB | 346 | Yes, comfortably |
| M3 Max 128GB | 128 GB | 2.9 GiB | 169 | Yes, comfortably |
| M4 Max 128GB | 128 GB | 2.9 GiB | 231 | Yes, comfortably |
| M2 Ultra 192GB | 192 GB | 2.9 GiB | 338 | Yes, comfortably |
| M3 Ultra 256GB | 256 GB | 2.9 GiB | 346 | Yes, comfortably |
| M3 Ultra 512GB | 512 GB | 2.9 GiB | 346 | Yes, comfortably |
Download size by quantisation
Weight memory scales linearly with bits per parameter. Everything below assumes an 8K context on top.
| Quantisation | Weights | Total @ 8K | Notes |
|---|---|---|---|
| FP16 | 5.8 GiB | 7.1 GiB | Full precision. Reference quality, twice the memory of Q8. |
| Q8_0 | 3.1 GiB | 4.3 GiB | Indistinguishable from FP16 in practice, at half the size. |
| Q6_K | 2.4 GiB | 3.6 GiB | Near-lossless. A good stop when you have memory to spare. |
| Q5_K_M | 2.1 GiB | 3.2 GiB | Slightly better than Q4_K_M, noticeably bigger. |
| Q4_K_M | 1.8 GiB | 2.9 GiB | The default. Best quality-per-gigabyte for most people. |
| Q3_K_M | 1.4 GiB | 2.6 GiB | Visible quality loss. Use to squeeze one size class up. |
| Q2_K | 1.1 GiB | 2.2 GiB | Last resort. Often worse than a smaller model at Q4. |