Qwen3.8-27B hardware requirements: VRAM, GPU and RAM
- Page updated
- Sources checked
- File sizes from Hugging Face, read October 9, 2026
- Capacity estimates, labeled
Qwen3.8-27B is Qwen's 27-billion-parameter open-weight model with a vision encoder and a 262K context, released under Apache 2.0. It is a dense model, so unlike the mixture-of-experts Qwen3.8-Flash-Next it fits a single gaming GPU once quantized.
Short answer: with llama.cpp the 4-bit UD-Q4_K_XL file is 17.6 GB and needs about 21 GB in total at a 32K context (RigCheck estimate: file + 2.1 GB KV cache + 0.16 GB DeltaNet state + 1.5 GB buffers). That fits a 24 GB card (RTX 3090, 4090, RX 7900 XTX) or a Mac with 32 GB. A 16 GB card holds the 3-bit UD-IQ3_XXS (10.9 GB) at 32K, or the 4-bit file with part of it in system RAM, which is slower. An 8 GB card holds most layers of the 1-bit UD-IQ1_S (6.2 GB) but needs system RAM for the rest; with the 4-bit file most layers run from system RAM. Unsloth's own table, for total memory (RAM + VRAM), is 16–19 GB for 4-bit and 12–14 GB for 3-bit.
Can my hardware run Qwen3.8-27B?
Terms such as GGUF, KV cache or offload: glossary.
RigCheck estimate for llama.cpp with Unsloth's GGUF files. Memory needed = the file + KV cache for your context + 0.16 GB of DeltaNet state (one server slot) + 1.5 GB of compute buffers. The KV cache is 64 KiB per token at f16, only on the 16 attention layers (4 KV heads × 256 dims, from Qwen's config.json); q8_0 is 34/64 of that. The 48 Gated DeltaNet layers keep a fixed F32 state instead (128 × 6,144 values plus the convolution state per layer, as llama.cpp sizes it). The 1.5 GB of buffers is our assumption: about 1 GB plus the 0.51 GB logits buffer for a 512-token batch over the 248,320-token vocabulary. Usable memory keeps 8% of each GPU for the driver and runtime, 8 GB of system RAM for the OS, and 10% (at least 8 GB) of unified memory for the OS. When the file does not fit on the GPU, llama.cpp keeps whole layers on the GPU and runs the rest from system RAM. Radeon cards are the ones Ollama lists for its ROCm support; for them the result is capacity only (we found no source for this model with llama.cpp's own AMD builds), and Ollama's own package is a different file. No speed is estimated.
How much VRAM does Qwen3.8-27B need?
Unsloth's GGUF files as Hugging Face lists them (decimal GB, read October 9, 2026), and the total memory each needs at a 32K context with an f16 KV cache and no image input (RigCheck estimate). The official weights are BF16, 55.6 GB in safetensors.
| File | File size | Needs at 32K | Source |
|---|---|---|---|
| BF16 | 54.7 GB | 58.5 GB | Unsloth |
| Q8_0 | 29.0 GB | 32.8 GB | Unsloth |
| UD-Q6_K_XL | 25.3 GB | 29.1 GB | Unsloth |
| UD-Q5_K_XL | 20.9 GB | 24.7 GB | Unsloth |
| UD-Q4_K_XL | 17.6 GB | 21.4 GB | Unsloth |
| UD-IQ4_XS | 14.3 GB | 18.1 GB | Unsloth |
| UD-Q3_K_XL | 13.1 GB | 16.9 GB | Unsloth |
| UD-IQ3_XXS | 10.9 GB | 14.7 GB | Unsloth |
| UD-Q2_K_XL | 9.8 GB | 13.6 GB | Unsloth |
| UD-IQ2_XXS | 7.3 GB | 11.1 GB | Unsloth |
| UD-IQ1_S | 6.2 GB | 10.0 GB | Unsloth |
Unsloth's table gives total memory (RAM + VRAM, or unified memory) of 7–8 GB for 1-bit, 9–11 GB for 2-bit, 12–14 GB for 3-bit, 16–19 GB for 4-bit, 23–26 GB for 6-bit, 31 GB for 8-bit and 56 GB for BF16, and says 4-bit "works on 16-19GB VRAM like RTX 5080, 4090 or a Mac with 24GB RAM". Our estimates are higher because they include a 32K KV cache, buffers and an OS reserve. Ollama's default qwen3.8:27b package is 18 GB (q4_K_M).
What hardware can run Qwen3.8-27B?
RigCheck estimate at a 32K context with an f16 KV cache and no image input, from the file sizes above and the assumptions under the checker. "Largest" means the largest of the checker's eleven files.
| Hardware | Largest file that fits fully | With layers in system RAM (slower) |
|---|---|---|
| No GPU + 32 GB RAM | UD-Q4_K_XL (all in RAM, CPU only) | no larger file |
| RTX 4060 (8 GB) + 32 GB RAM | none | UD-Q6_K_XL (51 of 64 layers, 22 GB in RAM) |
| RTX 3060 (12 GB) + 32 GB RAM | UD-IQ1_S | Q8_0 (45 of 64 layers, 22 GB in RAM) |
| RTX 5060 Ti (16 GB) + 32 GB RAM | UD-IQ3_XXS | Q8_0 (38 of 64 layers, 18 GB in RAM) |
| RX 9070 XT (16 GB) + 32 GB RAM | UD-IQ3_XXS | Q8_0 (38 of 64 layers, 18 GB in RAM) |
| RTX 4090 (24 GB) + 32 GB RAM | UD-Q4_K_XL | Q8_0 (23 of 64 layers, 11 GB in RAM) |
| RX 7900 XTX (24 GB) + 32 GB RAM | UD-Q4_K_XL | Q8_0 (23 of 64 layers, 11 GB in RAM) |
| RTX 5090 (32 GB) + 32 GB RAM | UD-Q6_K_XL | Q8_0 (8 of 64 layers, 4 GB in RAM) |
| Mac, 16 GB unified | none | not applicable (unified memory) |
| Mac, 24 GB unified | UD-IQ3_XXS | not applicable (unified memory) |
| Mac, 32 GB unified | UD-Q4_K_XL | not applicable (unified memory) |
| Mac, 64 GB unified | Q8_0 | not applicable (unified memory) |
Capacity only, not tested setups. A shorter context or a q8_0 KV cache leaves room for a larger file: at 8K, for example, the KV cache is 0.5 GB instead of 2.1 GB.
Measured Qwen3.8-27B on your own hardware? Send us the numbers with the log. After checking it we add it to this table, labeled as a community report.
Qwen3.8-27B specs
| Parameters | 27B, dense Model card |
|---|---|
| Architecture | 64 layers: 16 × (3 Gated DeltaNet + 1 Gated Attention), each followed by a feed-forward block; vision encoder for image and video input Model card |
| Context length | 262,144 tokens natively, up to 1,000,000 with YaRN Model card |
| KV cache | 64 KiB per token at f16 (16 attention layers × 4 KV heads × 256 dims): 2.1 GB at 32K, 8.6 GB at 128K, 17.2 GB at 256K; plus a fixed 0.16 GB DeltaNet state RigCheck calculation |
| License | Apache 2.0 Model card |
| Runtimes | Qwen names Transformers, vLLM, SGLang and TokenSpeed; llama.cpp supports the architecture, Unsloth publishes GGUF files and Ollama has qwen3.8:27b (checked October 9, 2026) |
FAQ
How much VRAM does Qwen3.8-27B need?
By our estimate about 21 GB at a 32K context for the 4-bit UD-Q4_K_XL (17.6 GB file), so a 24 GB card holds it, and about 15 GB for the 3-bit UD-IQ3_XXS (10.9 GB), which fits a 16 GB card. Unsloth's table lists 16–19 GB of total memory (RAM + VRAM) for 4-bit and does not state the context or cache it assumes.
Can Qwen3.8-27B run on a 16 GB GPU like the RTX 5060 Ti, 5080 or RX 9070 XT?
Yes, by our estimate: UD-IQ3_XXS fits fully at a 32K context, and the 4-bit file runs with part of its layers in system RAM, which is slower.
Can Qwen3.8-27B run on an 8 GB or 12 GB GPU?
With part of it in system RAM. On an 8 GB card with 32 GB of RAM the 1-bit UD-IQ1_S keeps 43 of its 64 layers on the GPU, the 4-bit UD-Q4_K_XL 18 (our estimate at 32K). On a 12 GB card only the 1-bit UD-IQ1_S (6.2 GB) fits fully at 32K; larger files run with some layers in system RAM. For these cards, the Gemma 4 12B page may suit better.
Can I run Qwen3.8-27B on a Mac?
Yes. By our conservative estimate a Mac with 24 GB holds UD-IQ3_XXS and one with 32 GB holds UD-Q4_K_XL at a 32K context; Unsloth says the 4-bit files work on a Mac with 24 GB. Use llama.cpp (Metal) or Ollama.
Can I run Qwen3.8-27B without a GPU?
Yes, with llama.cpp's CPU build: by our estimate 32 GB of RAM holds UD-Q4_K_XL at a 32K context. A dense 27B model is slow on a CPU; we found no published CPU speed.
Is Qwen3.8-27B on Ollama?
Yes: ollama run qwen3.8:27b downloads an 18 GB q4_K_M package. Ollama also lists q8_0 (30 GB), bf16 (56 GB), nvfp4, mxfp8 and MLX builds (checked October 9, 2026).
Qwen3.8-27B or Qwen3.8-Flash-Next?
They need very different hardware. The 27B fits one 24 GB card at 4-bit. Flash-Next is a 125B mixture-of-experts model whose smallest file is 72.5 GB, so it needs a lot of system RAM, or the Strata engine on a gaming PC.
Sources
- Hugging Face: Qwen/Qwen3.8-27B (model card, config.json, license, file sizes)
- Hugging Face: unsloth/Qwen3.8-27B-GGUF and Unsloth's Qwen3.8 guide (GGUF sizes, memory table, llama.cpp commands)
- Ollama library: qwen3.8 tags and Ollama's GPU support list (checked October 9, 2026)
- llama.cpp: src/models/qwen35.cpp, src/llama-hparams.cpp (recurrent state size) and tools/server/README.md
- Usable-memory reserves, buffers and the offload model are RigCheck assumptions, stated under the checker.