Gemma 3 4B on a Mac
Google · 4.3B parameters · Gemma Terms · released 2025-03
Weights @ Q4_K_M
2.5 GiB
the usual download size
KV cache
136 KB
per token, FP16
Max context
128K
tokens
Smallest Mac
8 GB
M1
Short answerA M1 8GB is the smallest Apple Silicon machine that loads Gemma 3 4B at Q4_K_M with an 8K context, at roughly 21 tokens/sec. MacBook Air (M1) is the cheapest way to get that configuration.
Every Mac, ranked
| Mac | Memory | Needs | Tok/s | Verdict |
|---|---|---|---|---|
| M1 8GB | 8 GB | 4.4 GiB | 21 | Yes, but tight |
| M1 16GB | 16 GB | 4.4 GiB | 21 | Yes, comfortably |
| M2 16GB | 16 GB | 4.4 GiB | 30 | Yes, comfortably |
| M4 16GB | 16 GB | 4.4 GiB | 36 | Yes, comfortably |
| M1 Pro 16GB | 16 GB | 4.4 GiB | 61 | Yes, comfortably |
| M3 Pro 18GB | 18 GB | 4.4 GiB | 46 | Yes, comfortably |
| M2 24GB | 24 GB | 4.4 GiB | 30 | Yes, comfortably |
| M4 24GB | 24 GB | 4.4 GiB | 36 | Yes, comfortably |
| M4 Pro 24GB | 24 GB | 4.4 GiB | 83 | Yes, comfortably |
| M4 32GB | 32 GB | 4.4 GiB | 36 | Yes, comfortably |
| M1 Max 32GB | 32 GB | 4.4 GiB | 121 | Yes, comfortably |
| M2 Max 32GB | 32 GB | 4.4 GiB | 121 | Yes, comfortably |
| M3 Max 36GB | 36 GB | 4.4 GiB | 91 | Yes, comfortably |
| M4 Max 36GB | 36 GB | 4.4 GiB | 125 | Yes, comfortably |
| M4 Pro 48GB | 48 GB | 4.4 GiB | 83 | Yes, comfortably |
| M4 Max 48GB | 48 GB | 4.4 GiB | 166 | Yes, comfortably |
| M4 Pro 64GB | 64 GB | 4.4 GiB | 83 | Yes, comfortably |
| M1 Max 64GB | 64 GB | 4.4 GiB | 121 | Yes, comfortably |
| M4 Max 64GB | 64 GB | 4.4 GiB | 166 | Yes, comfortably |
| M2 Ultra 64GB | 64 GB | 4.4 GiB | 243 | Yes, comfortably |
| M2 Max 96GB | 96 GB | 4.4 GiB | 121 | Yes, comfortably |
| M3 Ultra 96GB | 96 GB | 4.4 GiB | 249 | Yes, comfortably |
| M3 Max 128GB | 128 GB | 4.4 GiB | 121 | Yes, comfortably |
| M4 Max 128GB | 128 GB | 4.4 GiB | 166 | Yes, comfortably |
| M2 Ultra 192GB | 192 GB | 4.4 GiB | 243 | Yes, comfortably |
| M3 Ultra 256GB | 256 GB | 4.4 GiB | 249 | Yes, comfortably |
| M3 Ultra 512GB | 512 GB | 4.4 GiB | 249 | Yes, comfortably |
Download size by quantisation
Weight memory scales linearly with bits per parameter. Everything below assumes an 8K context on top.
| Quantisation | Weights | Total @ 8K | Notes |
|---|---|---|---|
| FP16 | 8.0 GiB | 10.3 GiB | Full precision. Reference quality, twice the memory of Q8. |
| Q8_0 | 4.3 GiB | 6.3 GiB | Indistinguishable from FP16 in practice, at half the size. |
| Q6_K | 3.3 GiB | 5.3 GiB | Near-lossless. A good stop when you have memory to spare. |
| Q5_K_M | 2.9 GiB | 4.9 GiB | Slightly better than Q4_K_M, noticeably bigger. |
| Q4_K_M | 2.5 GiB | 4.4 GiB | The default. Best quality-per-gigabyte for most people. |
| Q3_K_M | 2.0 GiB | 3.9 GiB | Visible quality loss. Use to squeeze one size class up. |
| Q2_K | 1.5 GiB | 3.4 GiB | Last resort. Often worse than a smaller model at Q4. |