Llama 3.2 1B on a Mac
Meta · 1.24B parameters · Llama 3.2 Community · released 2024-09
Weights @ Q4_K_M
0.7 GiB
the usual download size
KV cache
32 KB
per token, FP16
Max context
128K
tokens
Smallest Mac
8 GB
M1
Short answerA M1 8GB is the smallest Apple Silicon machine that loads Llama 3.2 1B at Q4_K_M with an 8K context, at roughly 72 tokens/sec. MacBook Air (M1) is the cheapest way to get that configuration.
Every Mac, ranked
| Mac | Memory | Needs | Tok/s | Verdict |
|---|---|---|---|---|
| M1 8GB | 8 GB | 1.8 GiB | 72 | Yes, comfortably |
| M1 16GB | 16 GB | 1.8 GiB | 72 | Yes, comfortably |
| M2 16GB | 16 GB | 1.8 GiB | 105 | Yes, comfortably |
| M4 16GB | 16 GB | 1.8 GiB | 126 | Yes, comfortably |
| M1 Pro 16GB | 16 GB | 1.8 GiB | 211 | Yes, comfortably |
| M3 Pro 18GB | 18 GB | 1.8 GiB | 158 | Yes, comfortably |
| M2 24GB | 24 GB | 1.8 GiB | 105 | Yes, comfortably |
| M4 24GB | 24 GB | 1.8 GiB | 126 | Yes, comfortably |
| M4 Pro 24GB | 24 GB | 1.8 GiB | 288 | Yes, comfortably |
| M4 32GB | 32 GB | 1.8 GiB | 126 | Yes, comfortably |
| M1 Max 32GB | 32 GB | 1.8 GiB | 421 | Yes, comfortably |
| M2 Max 32GB | 32 GB | 1.8 GiB | 421 | Yes, comfortably |
| M3 Max 36GB | 36 GB | 1.8 GiB | 316 | Yes, comfortably |
| M4 Max 36GB | 36 GB | 1.8 GiB | 432 | Yes, comfortably |
| M4 Pro 48GB | 48 GB | 1.8 GiB | 288 | Yes, comfortably |
| M4 Max 48GB | 48 GB | 1.8 GiB | 575 | Yes, comfortably |
| M4 Pro 64GB | 64 GB | 1.8 GiB | 288 | Yes, comfortably |
| M1 Max 64GB | 64 GB | 1.8 GiB | 421 | Yes, comfortably |
| M4 Max 64GB | 64 GB | 1.8 GiB | 575 | Yes, comfortably |
| M2 Ultra 64GB | 64 GB | 1.8 GiB | 843 | Yes, comfortably |
| M2 Max 96GB | 96 GB | 1.8 GiB | 421 | Yes, comfortably |
| M3 Ultra 96GB | 96 GB | 1.8 GiB | 863 | Yes, comfortably |
| M3 Max 128GB | 128 GB | 1.8 GiB | 421 | Yes, comfortably |
| M4 Max 128GB | 128 GB | 1.8 GiB | 575 | Yes, comfortably |
| M2 Ultra 192GB | 192 GB | 1.8 GiB | 843 | Yes, comfortably |
| M3 Ultra 256GB | 256 GB | 1.8 GiB | 863 | Yes, comfortably |
| M3 Ultra 512GB | 512 GB | 1.8 GiB | 863 | Yes, comfortably |
Download size by quantisation
Weight memory scales linearly with bits per parameter. Everything below assumes an 8K context on top.
| Quantisation | Weights | Total @ 8K | Notes |
|---|---|---|---|
| FP16 | 2.3 GiB | 3.5 GiB | Full precision. Reference quality, twice the memory of Q8. |
| Q8_0 | 1.2 GiB | 2.3 GiB | Indistinguishable from FP16 in practice, at half the size. |
| Q6_K | 1.0 GiB | 2.1 GiB | Near-lossless. A good stop when you have memory to spare. |
| Q5_K_M | 0.8 GiB | 1.9 GiB | Slightly better than Q4_K_M, noticeably bigger. |
| Q4_K_M | 0.7 GiB | 1.8 GiB | The default. Best quality-per-gigabyte for most people. |
| Q3_K_M | 0.6 GiB | 1.6 GiB | Visible quality loss. Use to squeeze one size class up. |
| Q2_K | 0.4 GiB | 1.5 GiB | Last resort. Often worse than a smaller model at Q4. |