Gemma 2 2B on a Mac
Google · 2.61B parameters · Gemma Terms · released 2024-07
Weights @ Q4_K_M
1.5 GiB
the usual download size
KV cache
84 KB
per token, FP16
Max context
8K
tokens
Smallest Mac
8 GB
M1
Short answerA M1 8GB is the smallest Apple Silicon machine that loads Gemma 2 2B at Q4_K_M with an 8K context, at roughly 34 tokens/sec. MacBook Air (M1) is the cheapest way to get that configuration.
Every Mac, ranked
| Mac | Memory | Needs | Tok/s | Verdict |
|---|---|---|---|---|
| M1 8GB | 8 GB | 3.0 GiB | 34 | Yes, comfortably |
| M1 16GB | 16 GB | 3.0 GiB | 34 | Yes, comfortably |
| M2 16GB | 16 GB | 3.0 GiB | 50 | Yes, comfortably |
| M4 16GB | 16 GB | 3.0 GiB | 60 | Yes, comfortably |
| M1 Pro 16GB | 16 GB | 3.0 GiB | 100 | Yes, comfortably |
| M3 Pro 18GB | 18 GB | 3.0 GiB | 75 | Yes, comfortably |
| M2 24GB | 24 GB | 3.0 GiB | 50 | Yes, comfortably |
| M4 24GB | 24 GB | 3.0 GiB | 60 | Yes, comfortably |
| M4 Pro 24GB | 24 GB | 3.0 GiB | 137 | Yes, comfortably |
| M4 32GB | 32 GB | 3.0 GiB | 60 | Yes, comfortably |
| M1 Max 32GB | 32 GB | 3.0 GiB | 200 | Yes, comfortably |
| M2 Max 32GB | 32 GB | 3.0 GiB | 200 | Yes, comfortably |
| M3 Max 36GB | 36 GB | 3.0 GiB | 150 | Yes, comfortably |
| M4 Max 36GB | 36 GB | 3.0 GiB | 205 | Yes, comfortably |
| M4 Pro 48GB | 48 GB | 3.0 GiB | 137 | Yes, comfortably |
| M4 Max 48GB | 48 GB | 3.0 GiB | 273 | Yes, comfortably |
| M4 Pro 64GB | 64 GB | 3.0 GiB | 137 | Yes, comfortably |
| M1 Max 64GB | 64 GB | 3.0 GiB | 200 | Yes, comfortably |
| M4 Max 64GB | 64 GB | 3.0 GiB | 273 | Yes, comfortably |
| M2 Ultra 64GB | 64 GB | 3.0 GiB | 400 | Yes, comfortably |
| M2 Max 96GB | 96 GB | 3.0 GiB | 200 | Yes, comfortably |
| M3 Ultra 96GB | 96 GB | 3.0 GiB | 410 | Yes, comfortably |
| M3 Max 128GB | 128 GB | 3.0 GiB | 200 | Yes, comfortably |
| M4 Max 128GB | 128 GB | 3.0 GiB | 273 | Yes, comfortably |
| M2 Ultra 192GB | 192 GB | 3.0 GiB | 400 | Yes, comfortably |
| M3 Ultra 256GB | 256 GB | 3.0 GiB | 410 | Yes, comfortably |
| M3 Ultra 512GB | 512 GB | 3.0 GiB | 410 | Yes, comfortably |
Download size by quantisation
Weight memory scales linearly with bits per parameter. Everything below assumes an 8K context on top.
| Quantisation | Weights | Total @ 8K | Notes |
|---|---|---|---|
| FP16 | 4.9 GiB | 6.6 GiB | Full precision. Reference quality, twice the memory of Q8. |
| Q8_0 | 2.6 GiB | 4.2 GiB | Indistinguishable from FP16 in practice, at half the size. |
| Q6_K | 2.0 GiB | 3.6 GiB | Near-lossless. A good stop when you have memory to spare. |
| Q5_K_M | 1.7 GiB | 3.3 GiB | Slightly better than Q4_K_M, noticeably bigger. |
| Q4_K_M | 1.5 GiB | 3.0 GiB | The default. Best quality-per-gigabyte for most people. |
| Q3_K_M | 1.2 GiB | 2.7 GiB | Visible quality loss. Use to squeeze one size class up. |
| Q2_K | 0.9 GiB | 2.4 GiB | Last resort. Often worse than a smaller model at Q4. |