Gemma 4 hardware requirements: 12B, 26B-A4B and 31B

Gemma 4 is Google's family of open-weight models with image input and a 262K context, released under Apache 2.0. This page covers the three sizes for PCs and Macs: the dense 12B, the mixture-of-experts 26B-A4B (26B parameters, about 4B active per token) and the dense 31B. The phone-sized E2B and E4B are not covered.

Short answer: at a 32K context with llama.cpp (RigCheck estimates), Gemma 4 12B at 4-bit (UD-Q4_K_XL, 7.4 GB) needs about 10 GB, so it fits a 12 GB card; an 8 GB card holds the 3-bit UD-IQ3_XXS. Gemma 4 26B-A4B at 4-bit (UD-Q4_K_XL, 17.0 GB) needs about 20 GB and 31B at 4-bit about 22 GB with IQ4_XS (16.4 GB): both fit a 24 GB card (RTX 3090, 4090, RX 7900 XTX) or a Mac with 32 GB, while the 31B's Q4_K_M (about 24 GB in total) needs a 32 GB card or Mac. 16 GB cards take their 3-bit or 2-bit files. Most of Gemma 4's layers use a 1,024-token sliding window, so its KV cache stays small: about 1 GB for the 12B and 26B-A4B and 3.9 GB for the 31B at 32K.

Can my hardware run Gemma 4?

Setup

Terms such as GGUF, KV cache or offload: glossary.

RigCheck estimate for llama.cpp with Unsloth's GGUF files. Memory needed = the file + KV cache for your context + 1.5 GB of compute buffers. The KV cache comes from Google's config.json: on the sliding-window layers K and V for 1,536 tokens at most (llama.cpp caches the 1,024-token window plus one 512-token batch, unless --swa-full), on the full-attention layers for the whole context (head dim 512); q8_0 is 34/64 of f16. The 1.5 GB of buffers is our assumption: about 1 GB plus the 0.54 GB logits buffer for a 512-token batch over the 262,144-token vocabulary. Usable memory keeps 8% of each GPU for the driver and runtime, 8 GB of system RAM for the OS, and 10% (at least 8 GB) of unified memory for the OS. When the file does not fit on the GPU, llama.cpp keeps whole layers on the GPU and runs the rest from system RAM. Radeon cards are the ones Ollama lists for its ROCm support; for them the result is capacity only (we found no source for these models with llama.cpp's own AMD builds), and Ollama's own packages are different files. No speed is estimated.

How much VRAM does Gemma 4 need?

Unsloth's GGUF files as Hugging Face lists them (decimal GB, read October 9, 2026), and the total memory each needs at a 32K context with an f16 KV cache and no image input (RigCheck estimate). Official BF16 weights: 23.9 GB (12B), 51.6 GB (26B-A4B), 62.5 GB (31B).

Gemma 4 31B

FileFile sizeNeeds at 32KSource
BF1661.4 GB66.8 GBUnsloth
Q8_032.6 GB38.0 GBUnsloth
Q6_K25.2 GB30.6 GBUnsloth
Q5_K_M21.7 GB27.1 GBUnsloth
UD-Q4_K_XL18.8 GB24.2 GBUnsloth
Q4_K_M18.3 GB23.7 GBUnsloth
IQ4_XS16.4 GB21.8 GBUnsloth
Q3_K_M14.7 GB20.1 GBUnsloth
UD-IQ3_XXS11.8 GB17.2 GBUnsloth
UD-IQ2_XXS8.5 GB13.9 GBUnsloth

Gemma 4 26B-A4B

FileFile sizeNeeds at 32KSource
BF1650.5 GB53.0 GBUnsloth
Q8_026.9 GB29.4 GBUnsloth
UD-Q6_K_XL23.3 GB25.8 GBUnsloth
UD-Q5_K_XL21.2 GB23.7 GBUnsloth
UD-Q4_K_XL17.0 GB19.5 GBUnsloth
MXFP4_MOE16.6 GB19.1 GBUnsloth
UD-IQ4_XS13.6 GB16.1 GBUnsloth
UD-Q3_K_XL12.9 GB15.4 GBUnsloth
UD-IQ3_XXS11.4 GB13.9 GBUnsloth
UD-IQ2_XXS9.9 GB12.4 GBUnsloth

Gemma 4 12B

FileFile sizeNeeds at 32KSource
BF1623.8 GB26.3 GBUnsloth
Q8_012.7 GB15.2 GBUnsloth
Q6_K9.8 GB12.3 GBUnsloth
Q5_K_M8.4 GB10.9 GBUnsloth
UD-Q4_K_XL7.4 GB9.9 GBUnsloth
Q4_K_M7.1 GB9.6 GBUnsloth
IQ4_XS6.4 GB8.9 GBUnsloth
Q3_K_M5.7 GB8.2 GBUnsloth
UD-IQ3_XXS4.6 GB7.1 GBUnsloth
UD-IQ2_M4.2 GB6.7 GBUnsloth

Unsloth's table gives recommended total memory (RAM + VRAM, or unified memory) for 4-bit, 8-bit and BF16: 7–8, 13–14 and 25 GB for 12B; 16–18, 28–30 and 52 GB for 26B-A4B; 17–20, 34–38 and 62 GB for 31B, and says llama.cpp still runs below that with RAM or disk offload, more slowly. Ollama's default packages, as its tags page lists them: gemma4:12b 7.7–8.0 GB, gemma4:26b 16–19 GB, gemma4:31b 19–20 GB.

What hardware can run Gemma 4?

RigCheck estimate at a 32K context with an f16 KV cache and no image input, from the file sizes above and the assumptions under the checker. "Largest" means the largest of the checker's ten files for that model.

Gemma 4 31B

HardwareLargest file that fits fullyWith layers in system RAM (slower)
No GPU + 32 GB RAMQ4_K_M (all in RAM, CPU only)no larger file
RTX 4060 (8 GB) + 32 GB RAMnoneQ6_K (48 of 60 layers, 23 GB in RAM)
RTX 3060 (12 GB) + 32 GB RAMnoneQ6_K (41 of 60 layers, 20 GB in RAM)
RTX 5060 Ti (16 GB) + 32 GB RAMUD-IQ2_XXSQ8_0 (39 of 60 layers, 24 GB in RAM)
RX 9070 XT (16 GB) + 32 GB RAMUD-IQ2_XXSQ8_0 (39 of 60 layers, 24 GB in RAM)
RTX 4090 (24 GB) + 32 GB RAMIQ4_XSQ8_0 (27 of 60 layers, 16 GB in RAM)
RTX 5090 (32 GB) + 32 GB RAMQ5_K_MQ8_0 (15 of 60 layers, 9 GB in RAM)
Mac, 16 GB unifiednonenot applicable (unified memory)
Mac, 24 GB unifiedUD-IQ2_XXSnot applicable (unified memory)
Mac, 32 GB unifiedQ4_K_Mnot applicable (unified memory)

Gemma 4 26B-A4B

HardwareLargest file that fits fullyWith layers in system RAM (slower)
No GPU + 32 GB RAMUD-Q5_K_XL (all in RAM, CPU only)no larger file
RTX 4060 (8 GB) + 32 GB RAMnoneQ8_0 (24 of 30 layers, 22 GB in RAM)
RTX 3060 (12 GB) + 32 GB RAMnoneQ8_0 (20 of 30 layers, 19 GB in RAM)
RTX 5060 Ti (16 GB) + 32 GB RAMUD-IQ3_XXSQ8_0 (16 of 30 layers, 15 GB in RAM)
RX 9070 XT (16 GB) + 32 GB RAMUD-IQ3_XXSQ8_0 (16 of 30 layers, 15 GB in RAM)
RTX 4090 (24 GB) + 32 GB RAMUD-Q4_K_XLQ8_0 (8 of 30 layers, 7 GB in RAM)
RTX 5090 (32 GB) + 32 GB RAMQ8_0BF16 (14 of 30 layers, 24 GB in RAM)
Mac, 16 GB unifiednonenot applicable (unified memory)
Mac, 24 GB unifiedUD-Q3_K_XLnot applicable (unified memory)
Mac, 32 GB unifiedUD-Q5_K_XLnot applicable (unified memory)

Gemma 4 12B

HardwareLargest file that fits fullyWith layers in system RAM (slower)
No GPU + 32 GB RAMQ8_0 (all in RAM, CPU only)no larger file
RTX 4060 (8 GB) + 32 GB RAMUD-IQ3_XXSBF16 (37 of 48 layers, 19 GB in RAM)
RTX 3060 (12 GB) + 32 GB RAMQ5_K_MBF16 (30 of 48 layers, 16 GB in RAM)
RTX 5060 Ti (16 GB) + 32 GB RAMQ6_KBF16 (23 of 48 layers, 12 GB in RAM)
RX 9070 XT (16 GB) + 32 GB RAMQ6_KBF16 (23 of 48 layers, 12 GB in RAM)
RTX 4090 (24 GB) + 32 GB RAMQ8_0BF16 (9 of 48 layers, 5 GB in RAM)
RTX 5090 (32 GB) + 32 GB RAMBF16no larger file
Mac, 16 GB unifiedUD-IQ3_XXSnot applicable (unified memory)
Mac, 24 GB unifiedQ8_0not applicable (unified memory)
Mac, 32 GB unifiedQ8_0not applicable (unified memory)

Capacity only, not tested setups. With layers in system RAM, any of them runs slower; we found no published speeds to say how much slower for each.

Measured Gemma 4 on your own hardware? Send us the numbers with the log. After checking it we add it to this table, labeled as a community report.

Gemma 4 specs

12B26B-A4B31B
Typedensemixture of experts: 128 experts, 8 per tokendense
Layers48 (40 sliding-window, 8 full attention)30 (25 + 5)60 (50 + 10)
KV cache at 32K, f161.0 GB1.0 GB3.9 GB
Inputtext, image, audiotext, imagetext, image

Layers, experts and inputs from Google's config.json files; KV cache is a RigCheck calculation. All three: 262,144-token context, Apache 2.0 license. Runtimes: llama.cpp supports the architecture, Unsloth publishes GGUF files and Ollama has gemma4 packages (checked October 9, 2026).

FAQ

Which Gemma 4 model should I run on my GPU?

By our estimate at a 32K context: on an 8 GB card the 12B at 3-bit; on a 12 GB card the 12B at up to 5-bit; on a 16 GB card the 12B at 6-bit, or the 26B-A4B and 31B at 2–3 bits; on a 24 GB card the 26B-A4B or 31B at 4-bit. Which of those suits you best depends on the task; this page only answers what fits.

How much VRAM does Gemma 4 26B-A4B need?

About 20 GB at a 32K context for the 4-bit UD-Q4_K_XL (17.0 GB file), by our estimate, so a 24 GB card holds it. A 16 GB card holds UD-IQ3_XXS (11.4 GB) fully, or the 4-bit file with part of it in system RAM. Unsloth recommends 16–18 GB of total memory for 4-bit.

How much VRAM does Gemma 4 31B need?

About 24 GB at a 32K context for the 4-bit Q4_K_M (18.3 GB), so on a 24 GB card the slightly smaller IQ4_XS (16.4 GB) is the largest that fits fully; a 32 GB card holds Q5_K_M. Unsloth recommends 17–20 GB of total memory for 4-bit.

Can Gemma 4 run on a Mac?

Yes. By our conservative estimate a Mac with 16 GB holds the 12B at 3-bit, 24 GB the 12B at 8-bit or the 26B-A4B at 3-bit, and 32 GB the 26B-A4B or 31B at 4–5 bits, at a 32K context. Use llama.cpp (Metal) or Ollama, which also lists MLX builds.

Can I run Gemma 4 without a GPU?

Yes, with llama.cpp's CPU build: by our estimate 32 GB of RAM holds the 12B at 8-bit, the 26B-A4B at 5-bit or the 31B at 4-bit at a 32K context. The 26B-A4B, with about 4B parameters active per token, is the one best suited to a CPU; we found no published CPU speed.

Is Gemma 4 on Ollama?

Yes: ollama run gemma4:12b (7.7–8.0 GB), gemma4:26b (16–19 GB) and gemma4:31b (19–20 GB), sizes as its tags page lists them, plus q8_0, bf16, QAT and MLX variants (checked October 9, 2026).

Sources