Llama 3.2 3B on a Mac
Meta · 3.21B parameters · Llama 3.2 Community · released 2024-09
Weights @ Q4_K_M
1.8 GiB
the usual download size
KV cache
112 KB
per token, FP16
Max context
128K
tokens
Smallest Mac
8 GB
M1
Short answerA M1 8GB is the smallest Apple Silicon machine that loads Llama 3.2 3B at Q4_K_M with an 8K context, at roughly 28 tokens/sec. MacBook Air (M1) is the cheapest way to get that configuration.
Every Mac, ranked
| Mac | Memory | Needs | Tok/s | Verdict |
|---|---|---|---|---|
| M1 8GB | 8 GB | 3.6 GiB | 28 | Yes, comfortably |
| M1 16GB | 16 GB | 3.6 GiB | 28 | Yes, comfortably |
| M2 16GB | 16 GB | 3.6 GiB | 41 | Yes, comfortably |
| M4 16GB | 16 GB | 3.6 GiB | 49 | Yes, comfortably |
| M1 Pro 16GB | 16 GB | 3.6 GiB | 81 | Yes, comfortably |
| M3 Pro 18GB | 18 GB | 3.6 GiB | 61 | Yes, comfortably |
| M2 24GB | 24 GB | 3.6 GiB | 41 | Yes, comfortably |
| M4 24GB | 24 GB | 3.6 GiB | 49 | Yes, comfortably |
| M4 Pro 24GB | 24 GB | 3.6 GiB | 111 | Yes, comfortably |
| M4 32GB | 32 GB | 3.6 GiB | 49 | Yes, comfortably |
| M1 Max 32GB | 32 GB | 3.6 GiB | 163 | Yes, comfortably |
| M2 Max 32GB | 32 GB | 3.6 GiB | 163 | Yes, comfortably |
| M3 Max 36GB | 36 GB | 3.6 GiB | 122 | Yes, comfortably |
| M4 Max 36GB | 36 GB | 3.6 GiB | 167 | Yes, comfortably |
| M4 Pro 48GB | 48 GB | 3.6 GiB | 111 | Yes, comfortably |
| M4 Max 48GB | 48 GB | 3.6 GiB | 222 | Yes, comfortably |
| M4 Pro 64GB | 64 GB | 3.6 GiB | 111 | Yes, comfortably |
| M1 Max 64GB | 64 GB | 3.6 GiB | 163 | Yes, comfortably |
| M4 Max 64GB | 64 GB | 3.6 GiB | 222 | Yes, comfortably |
| M2 Ultra 64GB | 64 GB | 3.6 GiB | 326 | Yes, comfortably |
| M2 Max 96GB | 96 GB | 3.6 GiB | 163 | Yes, comfortably |
| M3 Ultra 96GB | 96 GB | 3.6 GiB | 333 | Yes, comfortably |
| M3 Max 128GB | 128 GB | 3.6 GiB | 163 | Yes, comfortably |
| M4 Max 128GB | 128 GB | 3.6 GiB | 222 | Yes, comfortably |
| M2 Ultra 192GB | 192 GB | 3.6 GiB | 326 | Yes, comfortably |
| M3 Ultra 256GB | 256 GB | 3.6 GiB | 333 | Yes, comfortably |
| M3 Ultra 512GB | 512 GB | 3.6 GiB | 333 | Yes, comfortably |
Download size by quantisation
Weight memory scales linearly with bits per parameter. Everything below assumes an 8K context on top.
| Quantisation | Weights | Total @ 8K | Notes |
|---|---|---|---|
| FP16 | 6.0 GiB | 8.0 GiB | Full precision. Reference quality, twice the memory of Q8. |
| Q8_0 | 3.2 GiB | 5.0 GiB | Indistinguishable from FP16 in practice, at half the size. |
| Q6_K | 2.5 GiB | 4.3 GiB | Near-lossless. A good stop when you have memory to spare. |
| Q5_K_M | 2.1 GiB | 3.9 GiB | Slightly better than Q4_K_M, noticeably bigger. |
| Q4_K_M | 1.8 GiB | 3.6 GiB | The default. Best quality-per-gigabyte for most people. |
| Q3_K_M | 1.5 GiB | 3.2 GiB | Visible quality loss. Use to squeeze one size class up. |
| Q2_K | 1.1 GiB | 2.9 GiB | Last resort. Often worse than a smaller model at Q4. |