Gemma 4 hardware requirements: 12B, 26B-A4B and 31B
- Page updated
- Sources checked
- File sizes from Hugging Face, read October 9, 2026
- Capacity estimates, labeled
Gemma 4 is Google's family of open-weight models with image input and a 262K context, released under Apache 2.0. This page covers the three sizes for PCs and Macs: the dense 12B, the mixture-of-experts 26B-A4B (26B parameters, about 4B active per token) and the dense 31B. The phone-sized E2B and E4B are not covered.
Short answer: at a 32K context with llama.cpp (RigCheck estimates), Gemma 4 12B at 4-bit (UD-Q4_K_XL, 7.4 GB) needs about 10 GB, so it fits a 12 GB card; an 8 GB card holds the 3-bit UD-IQ3_XXS. Gemma 4 26B-A4B at 4-bit (UD-Q4_K_XL, 17.0 GB) needs about 20 GB and 31B at 4-bit about 22 GB with IQ4_XS (16.4 GB): both fit a 24 GB card (RTX 3090, 4090, RX 7900 XTX) or a Mac with 32 GB, while the 31B's Q4_K_M (about 24 GB in total) needs a 32 GB card or Mac. 16 GB cards take their 3-bit or 2-bit files. Most of Gemma 4's layers use a 1,024-token sliding window, so its KV cache stays small: about 1 GB for the 12B and 26B-A4B and 3.9 GB for the 31B at 32K.
Can my hardware run Gemma 4?
Terms such as GGUF, KV cache or offload: glossary.
RigCheck estimate for llama.cpp with Unsloth's GGUF files. Memory needed = the file + KV cache for your context + 1.5 GB of compute buffers. The KV cache comes from Google's config.json: on the sliding-window layers K and V for 1,536 tokens at most (llama.cpp caches the 1,024-token window plus one 512-token batch, unless --swa-full), on the full-attention layers for the whole context (head dim 512); q8_0 is 34/64 of f16. The 1.5 GB of buffers is our assumption: about 1 GB plus the 0.54 GB logits buffer for a 512-token batch over the 262,144-token vocabulary. Usable memory keeps 8% of each GPU for the driver and runtime, 8 GB of system RAM for the OS, and 10% (at least 8 GB) of unified memory for the OS. When the file does not fit on the GPU, llama.cpp keeps whole layers on the GPU and runs the rest from system RAM. Radeon cards are the ones Ollama lists for its ROCm support; for them the result is capacity only (we found no source for these models with llama.cpp's own AMD builds), and Ollama's own packages are different files. No speed is estimated.
How much VRAM does Gemma 4 need?
Unsloth's GGUF files as Hugging Face lists them (decimal GB, read October 9, 2026), and the total memory each needs at a 32K context with an f16 KV cache and no image input (RigCheck estimate). Official BF16 weights: 23.9 GB (12B), 51.6 GB (26B-A4B), 62.5 GB (31B).
Gemma 4 31B
| File | File size | Needs at 32K | Source |
|---|---|---|---|
| BF16 | 61.4 GB | 66.8 GB | Unsloth |
| Q8_0 | 32.6 GB | 38.0 GB | Unsloth |
| Q6_K | 25.2 GB | 30.6 GB | Unsloth |
| Q5_K_M | 21.7 GB | 27.1 GB | Unsloth |
| UD-Q4_K_XL | 18.8 GB | 24.2 GB | Unsloth |
| Q4_K_M | 18.3 GB | 23.7 GB | Unsloth |
| IQ4_XS | 16.4 GB | 21.8 GB | Unsloth |
| Q3_K_M | 14.7 GB | 20.1 GB | Unsloth |
| UD-IQ3_XXS | 11.8 GB | 17.2 GB | Unsloth |
| UD-IQ2_XXS | 8.5 GB | 13.9 GB | Unsloth |
Gemma 4 26B-A4B
| File | File size | Needs at 32K | Source |
|---|---|---|---|
| BF16 | 50.5 GB | 53.0 GB | Unsloth |
| Q8_0 | 26.9 GB | 29.4 GB | Unsloth |
| UD-Q6_K_XL | 23.3 GB | 25.8 GB | Unsloth |
| UD-Q5_K_XL | 21.2 GB | 23.7 GB | Unsloth |
| UD-Q4_K_XL | 17.0 GB | 19.5 GB | Unsloth |
| MXFP4_MOE | 16.6 GB | 19.1 GB | Unsloth |
| UD-IQ4_XS | 13.6 GB | 16.1 GB | Unsloth |
| UD-Q3_K_XL | 12.9 GB | 15.4 GB | Unsloth |
| UD-IQ3_XXS | 11.4 GB | 13.9 GB | Unsloth |
| UD-IQ2_XXS | 9.9 GB | 12.4 GB | Unsloth |
Gemma 4 12B
| File | File size | Needs at 32K | Source |
|---|---|---|---|
| BF16 | 23.8 GB | 26.3 GB | Unsloth |
| Q8_0 | 12.7 GB | 15.2 GB | Unsloth |
| Q6_K | 9.8 GB | 12.3 GB | Unsloth |
| Q5_K_M | 8.4 GB | 10.9 GB | Unsloth |
| UD-Q4_K_XL | 7.4 GB | 9.9 GB | Unsloth |
| Q4_K_M | 7.1 GB | 9.6 GB | Unsloth |
| IQ4_XS | 6.4 GB | 8.9 GB | Unsloth |
| Q3_K_M | 5.7 GB | 8.2 GB | Unsloth |
| UD-IQ3_XXS | 4.6 GB | 7.1 GB | Unsloth |
| UD-IQ2_M | 4.2 GB | 6.7 GB | Unsloth |
Unsloth's table gives recommended total memory (RAM + VRAM, or unified memory) for 4-bit, 8-bit and BF16: 7–8, 13–14 and 25 GB for 12B; 16–18, 28–30 and 52 GB for 26B-A4B; 17–20, 34–38 and 62 GB for 31B, and says llama.cpp still runs below that with RAM or disk offload, more slowly. Ollama's default packages, as its tags page lists them: gemma4:12b 7.7–8.0 GB, gemma4:26b 16–19 GB, gemma4:31b 19–20 GB.
What hardware can run Gemma 4?
RigCheck estimate at a 32K context with an f16 KV cache and no image input, from the file sizes above and the assumptions under the checker. "Largest" means the largest of the checker's ten files for that model.
Gemma 4 31B
| Hardware | Largest file that fits fully | With layers in system RAM (slower) |
|---|---|---|
| No GPU + 32 GB RAM | Q4_K_M (all in RAM, CPU only) | no larger file |
| RTX 4060 (8 GB) + 32 GB RAM | none | Q6_K (48 of 60 layers, 23 GB in RAM) |
| RTX 3060 (12 GB) + 32 GB RAM | none | Q6_K (41 of 60 layers, 20 GB in RAM) |
| RTX 5060 Ti (16 GB) + 32 GB RAM | UD-IQ2_XXS | Q8_0 (39 of 60 layers, 24 GB in RAM) |
| RX 9070 XT (16 GB) + 32 GB RAM | UD-IQ2_XXS | Q8_0 (39 of 60 layers, 24 GB in RAM) |
| RTX 4090 (24 GB) + 32 GB RAM | IQ4_XS | Q8_0 (27 of 60 layers, 16 GB in RAM) |
| RTX 5090 (32 GB) + 32 GB RAM | Q5_K_M | Q8_0 (15 of 60 layers, 9 GB in RAM) |
| Mac, 16 GB unified | none | not applicable (unified memory) |
| Mac, 24 GB unified | UD-IQ2_XXS | not applicable (unified memory) |
| Mac, 32 GB unified | Q4_K_M | not applicable (unified memory) |
Gemma 4 26B-A4B
| Hardware | Largest file that fits fully | With layers in system RAM (slower) |
|---|---|---|
| No GPU + 32 GB RAM | UD-Q5_K_XL (all in RAM, CPU only) | no larger file |
| RTX 4060 (8 GB) + 32 GB RAM | none | Q8_0 (24 of 30 layers, 22 GB in RAM) |
| RTX 3060 (12 GB) + 32 GB RAM | none | Q8_0 (20 of 30 layers, 19 GB in RAM) |
| RTX 5060 Ti (16 GB) + 32 GB RAM | UD-IQ3_XXS | Q8_0 (16 of 30 layers, 15 GB in RAM) |
| RX 9070 XT (16 GB) + 32 GB RAM | UD-IQ3_XXS | Q8_0 (16 of 30 layers, 15 GB in RAM) |
| RTX 4090 (24 GB) + 32 GB RAM | UD-Q4_K_XL | Q8_0 (8 of 30 layers, 7 GB in RAM) |
| RTX 5090 (32 GB) + 32 GB RAM | Q8_0 | BF16 (14 of 30 layers, 24 GB in RAM) |
| Mac, 16 GB unified | none | not applicable (unified memory) |
| Mac, 24 GB unified | UD-Q3_K_XL | not applicable (unified memory) |
| Mac, 32 GB unified | UD-Q5_K_XL | not applicable (unified memory) |
Gemma 4 12B
| Hardware | Largest file that fits fully | With layers in system RAM (slower) |
|---|---|---|
| No GPU + 32 GB RAM | Q8_0 (all in RAM, CPU only) | no larger file |
| RTX 4060 (8 GB) + 32 GB RAM | UD-IQ3_XXS | BF16 (37 of 48 layers, 19 GB in RAM) |
| RTX 3060 (12 GB) + 32 GB RAM | Q5_K_M | BF16 (30 of 48 layers, 16 GB in RAM) |
| RTX 5060 Ti (16 GB) + 32 GB RAM | Q6_K | BF16 (23 of 48 layers, 12 GB in RAM) |
| RX 9070 XT (16 GB) + 32 GB RAM | Q6_K | BF16 (23 of 48 layers, 12 GB in RAM) |
| RTX 4090 (24 GB) + 32 GB RAM | Q8_0 | BF16 (9 of 48 layers, 5 GB in RAM) |
| RTX 5090 (32 GB) + 32 GB RAM | BF16 | no larger file |
| Mac, 16 GB unified | UD-IQ3_XXS | not applicable (unified memory) |
| Mac, 24 GB unified | Q8_0 | not applicable (unified memory) |
| Mac, 32 GB unified | Q8_0 | not applicable (unified memory) |
Capacity only, not tested setups. With layers in system RAM, any of them runs slower; we found no published speeds to say how much slower for each.
Measured Gemma 4 on your own hardware? Send us the numbers with the log. After checking it we add it to this table, labeled as a community report.
Gemma 4 specs
| 12B | 26B-A4B | 31B | |
|---|---|---|---|
| Type | dense | mixture of experts: 128 experts, 8 per token | dense |
| Layers | 48 (40 sliding-window, 8 full attention) | 30 (25 + 5) | 60 (50 + 10) |
| KV cache at 32K, f16 | 1.0 GB | 1.0 GB | 3.9 GB |
| Input | text, image, audio | text, image | text, image |
Layers, experts and inputs from Google's config.json files; KV cache is a RigCheck calculation. All three: 262,144-token context, Apache 2.0 license. Runtimes: llama.cpp supports the architecture, Unsloth publishes GGUF files and Ollama has gemma4 packages (checked October 9, 2026).
FAQ
Which Gemma 4 model should I run on my GPU?
By our estimate at a 32K context: on an 8 GB card the 12B at 3-bit; on a 12 GB card the 12B at up to 5-bit; on a 16 GB card the 12B at 6-bit, or the 26B-A4B and 31B at 2–3 bits; on a 24 GB card the 26B-A4B or 31B at 4-bit. Which of those suits you best depends on the task; this page only answers what fits.
How much VRAM does Gemma 4 26B-A4B need?
About 20 GB at a 32K context for the 4-bit UD-Q4_K_XL (17.0 GB file), by our estimate, so a 24 GB card holds it. A 16 GB card holds UD-IQ3_XXS (11.4 GB) fully, or the 4-bit file with part of it in system RAM. Unsloth recommends 16–18 GB of total memory for 4-bit.
How much VRAM does Gemma 4 31B need?
About 24 GB at a 32K context for the 4-bit Q4_K_M (18.3 GB), so on a 24 GB card the slightly smaller IQ4_XS (16.4 GB) is the largest that fits fully; a 32 GB card holds Q5_K_M. Unsloth recommends 17–20 GB of total memory for 4-bit.
Can Gemma 4 run on a Mac?
Yes. By our conservative estimate a Mac with 16 GB holds the 12B at 3-bit, 24 GB the 12B at 8-bit or the 26B-A4B at 3-bit, and 32 GB the 26B-A4B or 31B at 4–5 bits, at a 32K context. Use llama.cpp (Metal) or Ollama, which also lists MLX builds.
Can I run Gemma 4 without a GPU?
Yes, with llama.cpp's CPU build: by our estimate 32 GB of RAM holds the 12B at 8-bit, the 26B-A4B at 5-bit or the 31B at 4-bit at a 32K context. The 26B-A4B, with about 4B parameters active per token, is the one best suited to a CPU; we found no published CPU speed.
Is Gemma 4 on Ollama?
Yes: ollama run gemma4:12b (7.7–8.0 GB), gemma4:26b (16–19 GB) and gemma4:31b (19–20 GB), sizes as its tags page lists them, plus q8_0, bf16, QAT and MLX variants (checked October 9, 2026).
Sources
- Hugging Face: google/gemma-4-12B-it, google/gemma-4-26B-A4B-it, google/gemma-4-31B-it (config.json, license, file sizes)
- Unsloth GGUF: 12B, 26B-A4B, 31B, and Unsloth's Gemma 4 guide (memory table, llama.cpp commands)
- Ollama library: gemma4 tags and Ollama's GPU support list (checked October 9, 2026)
- llama.cpp: src/models/gemma4.cpp, src/llama-kv-cache-iswa.cpp (sliding-window cache size) and tools/server/README.md
- Usable-memory reserves, buffers and the offload model are RigCheck assumptions, stated under the checker.