Qwen3.8-Flash-Next hardware requirements: VRAM, GPU and RAM
- Updated
- File sizes from Hugging Face, read October 8, 2026
- Speeds only as published, with their setup
Qwen3.8-Flash-Next is Qwen's open-weight mixture-of-experts model: 125B parameters with 6B active per token, plus a 51B-parameter n-gram embedding table, a vision encoder and a 262K context. It is a preview of the architecture behind Qwen4. This page covers the general runtimes (llama.cpp, Ollama, LM Studio, vLLM, SGLang); the Strata engine, which runs it on a gaming PC, has its own checker.
Short answer: with llama.cpp you need about 76 GB of memory in total (system RAM plus VRAM, or unified memory) for the smallest 1-bit file, about 85 GB for 3-bit (UD-IQ3_XXS) and 114 GB for the 4-bit UD-Q4_K_XL at a 32K context (RigCheck estimate; Unsloth lists 75 GB for 1-bit and 96–114 GB for 4-bit). Even 1-bit is 72.5 GB, because the 51B-parameter n-gram table stays at 4 bits or more (28.8 GB), and llama.cpp keeps that table on the CPU side, never on the GPU. A GPU is optional: by our estimate a PC with no GPU and 96 GB of RAM holds UD-IQ3_XXS; an RTX 4090 or 5090 with 128 GB of RAM holds UD-Q4_K_XL with most of it in RAM; one 96 GB RTX PRO 6000 with 64 GB of RAM holds UD-Q4_K_XL with only the table in RAM. A 128 GB Mac or DGX Spark just holds UD-Q4_K_XL at 32K without image input (a Ryzen AI Max+ 395 too, if about 86 GB is GPU-accessible and the rest stays with the CPU; untested); the 8-bit file (188.2 GB) needs a 256 GB Mac. The KV cache is small: 1.0 GB at 32K. With 12–16 GB of VRAM and 32–64 GB of RAM, use the Strata engine instead.
Can my hardware run Qwen3.8-Flash-Next?
RigCheck estimate for llama.cpp with Unsloth's GGUF files. Memory needed = the file + KV cache for your context + 0.12 GB of DeltaNet state (one server slot) + 2 GB of compute buffers. The KV cache is 30 KiB per token at f16 (24 KiB for the 12 attention layers, 6 KiB for their sparse-attention indexer; q8_0 is 34/64 of that), calculated from Qwen's config.json and llama.cpp's source. Usable memory keeps 8% of each GPU for the driver and runtime, 8 GB of system RAM for the OS, and 10% (at least 8 GB) of unified memory for the OS. On a PC the n-gram table (28.8 GB in UD-Q4_K_XL and smaller files, 54.4 GB in UD-Q5_K_XL and larger, read from the GGUF headers) never goes on the GPU, because llama.cpp keeps input layers (this table and the token embedding) on the CPU; everything else goes on the GPU as far as it fits (llama.cpp fits the model to your devices by default) and the rest stays in RAM. By default (--lazy-mode auto) llama.cpp reads the table's rows from the file on demand instead of keeping it resident, which can need less RAM; we count the whole table in RAM, as with --lazy-mode off: this estimate is the worst case, with the whole file in memory. Memory sizes count as decimal GB, like the files (128 GB of RAM is really 137 decimal GB, so the estimate is conservative). llama-server downloads the 0.9 GB vision file automatically with -hf (--no-mmproj skips it), and uses the model's full 262K context unless you set -c or its fitting shrinks it. No speed is estimated: see the published measurements below.
How much VRAM does Qwen3.8-Flash-Next need?
File sizes as Hugging Face lists them (all parts of each file, decimal GB, read October 8, 2026). Top-1 is how often the file picks the same next token as BF16, as measured by Unsloth. Add the KV cache for your context (see Specs) and a few GB for buffers.
| Format | File size | Top-1 vs BF16 | Runs with | Source |
|---|---|---|---|---|
| BF16 (official weights) | 360.0 GB | — | Transformers, vLLM, SGLang | Qwen |
| BF16 (GGUF) | 354.0 GB | — | llama.cpp | Unsloth |
| Q8_0 | 188.2 GB | 94.1% | llama.cpp | Unsloth |
| FP8 (official weights) | 185.5 GB | — | vLLM, SGLang | Qwen |
| UD-Q6_K_XL | 169.2 GB | 94.1% | llama.cpp | Unsloth |
| UD-Q5_K_XL | 158.3 GB | 93.7% | llama.cpp | Unsloth |
| NVFP4, RadixArk export | 135.2 GB | — | SGLang on NVIDIA Blackwell (B200, B300, GB300, RTX PRO 6000, DGX Spark) | RadixArk |
| NVFP4 (experts 4-bit, n-gram table FP8) | 132.7 GB | — | vLLM on NVIDIA B200 / B300, Linux (NVIDIA); SGLang on RTX PRO 6000 and DGX Spark | NVIDIA |
| UD-Q4_K_XL (dynamic 4-bit) | 111.3 GB | 92.3% | llama.cpp | Unsloth |
| UD-IQ4_XS | 93.7 GB | 89.6% | llama.cpp | Unsloth |
| UD-Q3_K_XL | 90.0 GB | 88.3% | llama.cpp | Unsloth |
| UD-IQ3_XXS | 82.0 GB | 85.4% | llama.cpp | Unsloth |
| UD-Q2_K_XL | 78.9 GB | 82.7% | llama.cpp | Unsloth |
| UD-IQ1_M | 74.5 GB | 79.7% | llama.cpp | Unsloth |
| UD-IQ1_S | 72.5 GB | 77.3% | llama.cpp | Unsloth |
Unsloth keeps the n-gram (PLE) table at 4 bits or more in every file, because heavier quantization of a randomly accessed table damages the model; that is why even 1-bit is over 70 GB. In the GGUF headers the table is 51.2B values: 28.80 GB (IQ4_NL) in UD-Q4_K_XL and every smaller file, 54.40 GB (Q8_0) in UD-Q5_K_XL, UD-Q6_K_XL and Q8_0. Unsloth's guide lists total memory (RAM + VRAM, or unified) of 75 GB for 1-bit, 79 GB for 2-bit, 90 GB for 3-bit, 96–114 GB for 4-bit, 163 GB for 5-bit, 200 GB for 8-bit and 355 GB for BF16, and advises 1–2 GB of extra headroom for MTP (its MTP files are 1.9–7.8 GB). Other uploaders' files differ by a few GB (bartowski's Q4_K_M is 119.6 GB). The vision file (mmproj) is 0.9 GB at BF16. NVIDIA's NVFP4 build is a 79.0 GB model file plus a 53.7 GB FP8 file holding the n-gram table and MTP layers. Ollama packages its own builds of qwen3.8-flash-next: q4_K_M 120 GB, q8_0 189 GB, bf16 355 GB, and MLX builds of 105 GB (nvfp4) and 360 GB (bf16), all with text and image input (Ollama library, checked October 8, 2026).
What hardware can run Qwen3.8-Flash-Next?
RigCheck estimate at a 32K context with an f16 KV cache and no image input, from the file sizes above and the assumptions under the checker. "Largest" means the largest of the checker's ten Unsloth files. On a PC we count the n-gram table in system RAM (worst case; llama.cpp never puts it on the GPU).
| Hardware | Largest file with all but the n-gram table on the GPU, or in unified memory | Largest with more in system RAM (PC) |
|---|---|---|
| No GPU + 96 GB RAM | — | UD-IQ3_XXS (all in RAM) |
| No GPU + 128 GB RAM | — | UD-Q4_K_XL (all in RAM) |
| 1× RTX 3060 (12 GB) + 96 GB RAM | none | UD-IQ4_XS (86 GB in RAM) |
| 1× RTX 4090 (24 GB) + 64 GB RAM | none | UD-IQ1_M (56 GB in RAM) |
| 1× RTX 5090 (32 GB) + 64 GB RAM | none | UD-IQ3_XXS (56 GB in RAM) |
| 1× RTX 4090 (24 GB) + 128 GB RAM | none | UD-Q4_K_XL (92 GB in RAM) |
| 1× RTX 5090 (32 GB) + 128 GB RAM | none | UD-Q4_K_XL (85 GB in RAM) |
| 2× RTX 3090 (48 GB) + 64 GB RAM | none | UD-IQ4_XS (53 GB in RAM) |
| 4× RTX 3090 (96 GB) + 64 GB RAM | UD-Q4_K_XL | — |
| 1× RTX PRO 6000 (96 GB) + 64 GB RAM | UD-Q4_K_XL | — |
| 2× RTX PRO 6000 (192 GB) + 64 GB RAM | Q8_0 | — |
| 1× H200 (141 GB) + 128 GB RAM | UD-Q6_K_XL | Q8_0 (62 GB in RAM) |
| NVIDIA DGX Spark (128 GB unified) | UD-Q4_K_XL | — |
| AMD Ryzen AI Max+ 395 (128 GB) | UD-Q4_K_XL (if ~86 GB is GPU-accessible and the rest stays with the CPU) | — |
| Mac with M5 Max, 40-core GPU (128 GB unified) | UD-Q4_K_XL | — |
| Mac Studio M3 Ultra (256 GB unified) | Q8_0 | — |
| Mac Studio M5 Ultra (256 GB unified) | Q8_0 | — |
Capacity only, not tested setups. On a 128 GB machine UD-Q4_K_XL just fits at 32K by our conservative estimate, and only without the vision file: llama-server's -hf downloads and loads it unless you pass --no-mmproj. With image input, or at 64K or more, UD-IQ4_XS is the largest that fits. Unsloth's guide builds llama.cpp with CUDA, or without it for CPU-only PCs, with Metal on Macs, and says that running from system RAM or unified memory costs this model relatively little against VRAM; it gives no numbers for that (checked October 8, 2026). llama.cpp has other GPU backends, but we found no report of this model on Ryzen AI Max hardware.
How fast is Qwen3.8-Flash-Next locally?
Each row is a published measurement with its own setup; we do not convert one setup's speed into another's. Single-request generation speed:
| Hardware | Setup | Generation rate | Source |
|---|---|---|---|
| 1× RTX PRO 6000 (96 GB) | SGLang, RadixArk's NVFP4 export, n-gram table in pinned system RAM, 1,024 tokens in / 256 out | 83 tok/s; 148 tok/s with MTP | SGLang cookbook |
| 1× RTX PRO 6000 | Unsloth; runtime and file not stated | about 100 tok/s; 170 tok/s with MTP | Unsloth guide |
| 1× NVIDIA DGX Spark | SGLang, RadixArk's NVFP4 export, n-gram table file-backed on NVMe | 15.9 tok/s; 27.5 tok/s with MTP | SGLang cookbook |
| 2× NVIDIA DGX Spark (TP=2) | SGLang, NVIDIA's NVFP4 export with RadixArk's MTP draft, 200GbE link (RadixArk's export measured the same) | 47.4 tok/s with MTP | SGLang cookbook |
We found no published llama.cpp speed for a consumer NVIDIA GPU with the experts in system RAM, or for a Mac. For gaming GPUs with 12–16 GB, the Strata page lists the Strata engine's measured speeds per card (a different engine, so they are not llama.cpp numbers). MTP (multi-token prediction) is a built-in draft layer. Unsloth advises 1–2 GB of extra headroom for it and its guide uses its own llama.cpp branch; llama.cpp master also has an MTP draft mode (--spec-type draft-mtp, read October 8, 2026).
Qwen3.8-Flash-Next specs
| Parameters | 125B in the language model, 6B active per token; plus 51B n-gram embedding and 4B MTP Model card |
|---|---|
| Architecture | 48 layers: 12 × (3 Gated DeltaNet + 1 Qwen Sparse Attention), each followed by a mixture of experts (512 experts, 10 routed + 1 shared per token); vision encoder for image and video input Model card |
| Context length | 262,144 tokens natively, up to 1,000,000 with YaRN Model card |
| KV cache | 30 KiB per token at f16 (12 attention layers × 2 KV heads × 256 dims, plus the sparse-attention indexer): 1.0 GB at 32K, 4.0 GB at 128K, 8.1 GB at 256K; q8_0 is 34/64 of that. The 36 DeltaNet layers keep a fixed 0.12 GB state per conversation (per llama-server slot) instead RigCheck calculation from config.json and llama.cpp's source |
| Official weights | BF16, 360.0 GB; FP8, 185.5 GB Hugging Face |
| License | Qwen Community License 1.0: free use, but very large products must display the model name, and a company running a hosted-model or AI coding / office assistant business needs a separate license from Qwen for any commercial use (internal use excepted) LICENSE |
| Runtimes | Qwen names Transformers, vLLM, SGLang, TokenSpeed and KTransformers; llama.cpp supports the architecture (qwen4exp), and Ollama and LM Studio's community catalog carry builds (checked October 8, 2026) |
FAQ
How much RAM or VRAM does Qwen3.8-Flash-Next need?
About 76 GB of total memory (system RAM plus VRAM, or unified memory) for the smallest 1-bit GGUF, about 85 GB for the 3-bit UD-IQ3_XXS and 114 GB for the 4-bit UD-Q4_K_XL at a 32K context, by our estimate; Unsloth lists 75 GB for 1-bit and 96–114 GB for 4-bit. On a PC, 28.8 GB of that (the n-gram table) never goes on the GPU. The files themselves are 72.5 GB (UD-IQ1_S) to 188.2 GB (Q8_0); BF16 is 354–360 GB. A GPU is not required.
Can Qwen3.8-Flash-Next run on an RTX 4090 or RTX 5090?
Not on the card alone: even the 72.5 GB 1-bit file is far larger than 24 or 32 GB, and llama.cpp keeps its 28.8 GB n-gram table on the CPU side anyway. With enough RAM, llama.cpp keeps the rest there too: by our estimate, with 64 GB of RAM an RTX 4090 holds UD-IQ1_M and an RTX 5090 UD-IQ3_XXS, each with about 56 GB in RAM; with 128 GB of RAM either card holds the 4-bit UD-Q4_K_XL. For 12–16 GB cards and 32–64 GB of RAM, the Strata engine is built for exactly this model.
Can I run Qwen3.8-Flash-Next without a GPU?
Yes. Unsloth's guide includes a CPU-only llama.cpp build and lists 75 GB of RAM as the minimum. By our estimate 96 GB of RAM holds UD-IQ3_XXS and 128 GB holds UD-Q4_K_XL at a 32K context. We found no published CPU-only speed.
Can I run Qwen3.8-Flash-Next on a Mac?
Yes, with 128 GB of unified memory or more, using llama.cpp (Metal) or Ollama (its GGUF or MLX builds). By our conservative estimate a 128 GB Mac just holds UD-Q4_K_XL at 32K without image input (UD-IQ4_XS with image input or at 64K or more), and a 256 GB Mac holds Q8_0. Ollama's MLX build is 105 GB. Raise the GPU wired-memory limit first. The Strata engine does not run on Macs.
Will Qwen3.8-Flash-Next run on a DGX Spark?
Yes. In llama.cpp, UD-Q4_K_XL just fits its 128 GB at 32K without image input, by our estimate. SGLang's cookbook measured RadixArk's NVFP4 export on one Spark with the n-gram table read from NVMe: 15.9 tokens per second, 27.5 with MTP; on two Sparks linked at 200GbE, NVIDIA's NVFP4 export (with RadixArk's MTP draft) reached 47.4 with MTP, the same as RadixArk's export.
Why is even the 1-bit Qwen3.8-Flash-Next so big?
The 51B-parameter n-gram embedding table is looked up at random, and Unsloth found that quantizing it below 4 bits damages the model, so every Unsloth file keeps it at 4 bits or more. That puts the 1-bit files at 72.5–74.5 GB. llama.cpp keeps the table on the CPU side (never on the GPU) and by default reads its rows from the file on demand (--lazy-mode auto); Unsloth notes it can stay on SSD through mmap, which uses less RAM.
Is Qwen3.8-Flash-Next on Ollama?
Yes. Ollama's library lists qwen3.8-flash-next builds: q4_K_M (120 GB), q8_0 (189 GB) and bf16 (355 GB), plus MLX builds of 105 GB (nvfp4) and 360 GB (bf16), with text and image input and a 256K context (checked October 8, 2026).
Does LM Studio support Qwen3.8-Flash-Next?
LM Studio's community catalog has GGUF builds made with llama.cpp release b10662: Q4_K_M (119.2 GB), Q6_K (167.6 GB) and Q8_0 (188.2 GB) (checked October 8, 2026). The memory needs are the same as for llama.cpp.
Which GPUs does Qwen3.8-Flash-Next need with vLLM or SGLang?
Server GPUs, for the official weights. vLLM's recipe puts the FP8 checkpoint at 172.78 GiB and runs it on 4× H100 only with the n-gram table offloaded to at least 51 GB of system RAM, on 8× H200 with expert parallelism, or on 2–4 GB300s. SGLang's cookbook runs an NVFP4 checkpoint on a single 96 GB RTX PRO 6000 with the n-gram table in pinned system RAM, which needs at least 64 GB of free RAM. On AMD, vLLM's recipe has a 4× MI355X FP8 configuration and SGLang lists MI350X and MI355X for BF16 and FP8.
Strata or llama.cpp for Qwen3.8-Flash-Next?
They answer different hardware. For llama.cpp, Ollama and LM Studio our estimate assumes the whole file in RAM plus VRAM (72.5 GB or more, plus the context); mapping the n-gram table from SSD can lower that, by an amount nobody has published. Strata, an engine built for this model, keeps the most-used experts on the GPU, all experts in RAM and the n-gram table on SSD, so it runs with 12 GB of VRAM and 32 GB of RAM for its code-focused Coder cut (64 GB of RAM for every size), on Windows and Linux with NVIDIA or supported AMD cards. Check it on the Strata page.
Is Qwen3.8-Flash-Next open source?
The weights are open under the Qwen Community License 1.0. It allows free use, modification and redistribution; products with more than 100 million monthly users or US$20 million in monthly revenue must display the model name, and a company running a hosted-model or AI coding / office assistant business needs a separate license from Qwen for any commercial use (internal use that exposes nothing to third parties is exempt).
Sources
- Hugging Face: Qwen/Qwen3.8-Flash-Next (model card, config.json, LICENSE, file sizes) and Qwen/Qwen3.8-Flash-Next-FP8 (read October 8, 2026)
- Hugging Face: unsloth/Qwen3.8-Flash-Next-GGUF and Unsloth's guide (GGUF sizes, top-1 accuracy, total-memory table, CPU / Metal builds, MTP, n-gram on SSD, RTX PRO 6000 speed)
- bartowski GGUF, lmstudio-community GGUF nvidia/Qwen3.8-Flash-Next-NVFP4 and RadixArk/Qwen3.8-Flash-Next-NVFP4 (file sizes, NVFP4 runtime and hardware); n-gram table size per file from the GGUF headers of Unsloth's files (read October 8, 2026)
- Ollama library: qwen3.8-flash-next tags (checked October 8, 2026)
- vLLM recipe and SGLang cookbook (server GPUs, n-gram offload, measured speeds)
- llama.cpp: qwen4exp.cpp and llama-memory-hybrid-idx.cpp (KV and indexer cache layout), llama-model.cpp (input layer kept on the CPU, lazy mapping), server README (--fit, --n-cpu-moe, --lazy-mode, KV cache types, -hf mmproj download, MTP draft mode, default context), ggml-common.h (q8_0 block size)
- Memory: NVIDIA DGX Spark, Apple Mac Studio (M5), Apple Mac Studio 2025 (M3 Ultra), AMD Ryzen AI Max+ 395; MLX LM README (raising the macOS GPU memory limit)
- Memory reserves, compute buffers and the GPU + RAM pool are RigCheck assumptions, stated under the checker.