Strata LLM requirements & compatibility checker
- Updated
- Numbers from the official Strata repo
- Calculated in your browser
Strata runs the 125B-class Qwen3.8-Flash-Next model on a gaming PC by splitting it across your graphics card, RAM and SSD. Pick your hardware below to see whether it runs, which model size to install and what published benchmarks say about its speed.
Minimum to run Strata: an NVIDIA GeForce RTX 20–50 series or a supported AMD Radeon card with 12 GB of VRAM (NVIDIA 8 GB runs, slowly), 32 GB of RAM (64 GB runs every size), ~80 GB free on an SSD, and Windows 10/11 or Linux. Macs are not supported.
Check your PC
Not affiliated with the Strata project. Requirements, sizes and speeds come from Strata's official repository and its published community benchmark reports, linked next to each figure. The checker adds a few labeled RigCheck assumptions (extra RAM headroom for "tight" fits, which benchmark is the closest card). Last checked (Strata engine 0.1.40).
Strata system requirements
| Part | Minimum | Recommended |
|---|---|---|
| Graphics card | NVIDIA RTX 20 series or newer with 8 GB (slow) · AMD RX 6800 / 7700 XT / 9060 XT or newer with 12 GB+ | 12 GB+ VRAM; 16–24 GB is noticeably faster |
| RAM | 32 GB (Coder only) | 64 GB (every size fits) |
| Disk | ~80 GB free (standard sizes) | ~90 GB+ free on an NVMe SSD; more for UD-IQ4_XS, AVX-512 Q2_0 or ROCm |
| Processor | x86-64 with AVX2 | AVX-512 (Ryzen 7000/9000) is a bit faster |
| System | Windows 10/11 or Linux | Ubuntu 22.04/24.04 installs everything automatically |
| Driver | NVIDIA 580 or newer · AMD: current Adrenalin (Windows) or the kernel amdgpu driver (Linux) | — |
Supported AMD cards: Radeon RX 7900 XT / XTX, RX 9070 / 9070 XT and Radeon AI PRO R9700 (validated by the maintainers); RX 7800 XT / 7700 XT and RX 9060 XT 16 GB (validated by their owners); RX 6800 / 6900 series (community-reported). Older NVIDIA cards (GTX 10, Tesla P40 / V100) are experimental and use Strata's separate CUDA 12 engine on Windows or Linux. Radeon VII / MI50, RX 6700 XT, Intel Arc and AMD Strix Halo run through experimental community builds on Linux (Strix Halo is also in the Windows engine from 0.1.40, unvalidated). Two or three cards can share the model: on NVIDIA under Windows and Linux, on AMD only under Linux.
Which model size for your RAM?
Strata keeps the model's experts in your RAM, so RAM, not VRAM, decides which size fits. A size fits when your RAM is at least its experts plus about 10 GB for Windows and other programs. More VRAM makes it faster but does not lower the RAM needed, except in low-RAM mode.
| Your RAM | Take | Why |
|---|---|---|
| 32 GB | Coder | The only size that fits. Best at code, weaker at other tasks and non-English text. With a 24 GB card, Q2_0 and IQ2_XS also run in low-RAM mode. |
| 48 GB | IQ2_XS (or Q2_0, fastest) | The IQ3 sizes do not fit. |
| 64 GB | IQ2_XS, or IQ3_XXS / IQ3_S | Every size fits; IQ3_S is the best and the slowest. |
| 96 GB+ | IQ3_S, or Unsloth UD-IQ4_XS (~4-bit) | Room for the largest sizes with everything else open. UD-IQ4_XS needs an NVIDIA card. |
| Size | Experts in RAM | RAM + VRAM total | Download | Quality |
|---|---|---|---|---|
| Coder (IQ1_M) | 23 GB | — | — | 91% of the full model on SWE-bench Verified (per its authors); half the experts |
| Q2_0 | 34 GB | 37.6 GB | 66 GB | good, fastest |
| IQ2_XS | 35.5 GB | 39.2 GB | 68 GB | better (recommended) |
| IQ3_XXS | 43 GB | 47.0 GB | 76 GB | great |
| IQ3_S | 50 GB | 54.8 GB | — | best: matches the full model on the published tests |
| Unsloth UD-IQ4_XS | 59.5 GB | — | 94 GB | ~4-bit; part read from SSD below ~80 GB RAM |
The first start also downloads the MTP draft layer (~6 GB, +1 GB with images). Source: Strata docs/MODELS.md.
How fast is Strata? Measured speeds by GPU
"Writes answers" is output speed in a short chat; a token is about ¾ of a word, so 60 tokens/s is faster than you can read.
| Card | Size | Writes answers | At 128K context | Source |
|---|---|---|---|---|
| RTX 5090 32 GB | IQ2_XS | 179 tok/s | 165 tok/s | Community |
| RTX 3090 24 GB | Q2_0 | ~140 tok/s | ~100 tok/s | Official estimate ±20% |
| RTX 5070 12 GB | Q2_0 | 94 tok/s | 76 tok/s | Official |
| RTX 5070 12 GB | IQ2_XS | 79 tok/s | 63 tok/s | Official |
| RTX 5070 12 GB | IQ3_S | 53 tok/s | 46 tok/s | Official |
| RTX 5060 Ti 16 GB | Q2_0 | ~87 tok/s | ~62 tok/s | Official estimate ±20% |
| RX 9070 XT 16 GB | Q2_0 | 60 tok/s | 48 tok/s | Official |
| RX 9070 XT 16 GB | Coder | 44 tok/s | 33 tok/s | Official |
| RX 7800 XT 16 GB | Coder | 38–44 tok/s | — | Community |
| RX 9060 XT 16 GB | Coder | 27–31 tok/s | — | Community |
| 2× Intel Arc Pro B60 (layer split) | IQ2_XS | 55–61 tok/s short / 2K code · 39 prose | — | Community SYCL port, Linux |
RTX 5070 and RX 9070 XT: MODELS.md. Estimates: DETAILS.md. Community: COMMUNITY_BENCHMARKS.md and AMD_HIP.md. More VRAM matters more than a faster GPU: every extra GB holds ~700 more experts that the CPU no longer has to compute. Slow RAM (EXPO/XMP off) and a monitor plugged into the card also lower speed.
Connect Strata to Claude Code, Codex and other apps
Strata serves OpenAI- and Anthropic-compatible APIs on http://127.0.0.1:8080. Any API key and any model name work unless you set a key yourself.
Claude Code
export ANTHROPIC_BASE_URL=http://127.0.0.1:8080
claude
Any OpenAI-compatible app (Continue, Open WebUI, scripts)
Base URL: http://127.0.0.1:8080/v1
API key: anything
Model: anything
Python (OpenAI SDK)
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="local")
reply = client.chat.completions.create(
model="strata",
messages=[{"role": "user", "content": "Hello from my own GPU"}],
)
print(reply.choices[0].message.content)
Codex CLI and other Responses API apps
Point them at http://127.0.0.1:8080/v1/responses; Strata's DETAILS.md has a ready config.toml.
Strata answers one request at a time by default; set "parallel": 2 to serve two (slower per answer on a 12 GB card). To use it from another device, start it with --host 0.0.0.0 --api-key <secret> and always set a key.
Common Strata errors and fixes
| What you see | Fix |
|---|---|
| PC freezes for minutes on the first start | Normal: it loads 35–55 GB into RAM. Wait. Still frozen after 10 minutes: restart, close other programs, or pick Q2_0 / IQ2_XS. |
| Very slow, disk light blinking, or "the engine stopped unexpectedly" | Out of free RAM. Close the browser and other programs, or pick a smaller size. |
| "NVIDIA driver is too old" | Install driver 580 or newer, restart, run START-HERE.bat again. |
| "Port 8080 is already in use" | Strata is already running, or use START-HERE.bat --port 8081. |
| Windows blocked the engine (Smart App Control) | Strata's executables are unsigned. Turning off Smart App Control lets them run, but it cannot be turned back on without resetting Windows. If only the image encoder is blocked, run setup with --vision no instead. |
| Linux: engine does not compile (gcc 15) | Install g++-14 and run CXX=g++-14 CUDAHOSTCXX=g++-14 ./setup.sh. |
| Slower than the tables | Unplug the monitor from the card, close GPU apps, enable EXPO/XMP for your RAM, or run START-HERE.bat --calibrate (NVIDIA). |
Full list: TROUBLESHOOTING.md.
FAQ
What is Strata LLM?
Strata is a free, MIT-licensed inference engine that runs Alibaba's Qwen3.8-Flash-Next, a large mixture-of-experts model that normally needs a server, on a consumer PC. It keeps the most-used experts on the graphics card, all experts in RAM and a large lookup table on the SSD, and uses a small draft model to speed up output 1.6–1.8×.
Can Strata run on a GPU with 8 GB of VRAM?
On NVIDIA RTX 20–50 series cards, yes, but slowly: fewer experts fit on the card, so the CPU does more of the work. 12 GB is the recommended minimum. AMD cards need 12 GB or more.
Can I run Strata with 16 GB or 32 GB of RAM?
16 GB is below the minimum. 32 GB runs the Coder version (best for code). With 32 GB of RAM and a 24 GB graphics card, Q2_0 and IQ2_XS also run in low-RAM mode. 64 GB runs every size.
Does Strata work on a Mac?
No. Strata supports Windows 10/11 and Linux with NVIDIA or AMD graphics cards. Apple Silicon is not supported.
Does Strata support AMD graphics cards?
Yes. RX 7900 XT / XTX, RX 9070 / 9070 XT and Radeon AI PRO R9700 are validated; RX 7800 XT / 7700 XT and RX 9060 XT 16 GB were validated by owners; RX 6800 / 6900 series are community-reported. On Windows, AMD cards cannot read pictures yet.
How big is the Strata download?
The model is 66–76 GB for Q2_0, IQ2_XS and IQ3_XXS, plus ~6 GB for the MTP draft layer (+1 GB with images). Unsloth's UD-IQ4_XS is a 94 GB download. Plan for about 80–90 GB of free SSD space for the standard sizes, ~100 GB for UD-IQ4_XS, plus ~40 GB for Q2_0 on an AVX-512 CPU and ~10 GB if setup installs ROCm for an AMD card on Linux.
Which Strata model size should I choose?
IQ2_XS is the official default for 48 GB of RAM or more. Q2_0 is the fastest, IQ3_S the highest quality (64 GB+), and the Coder is the only size that fits 32 GB, but it is weaker outside code and in non-English text.
Is Strata faster than Ollama or llama.cpp?
Strata builds on parts of llama.cpp / ggml but adds its own expert caching across GPU, RAM and SSD and speculative decoding for this model. On an RX 9060 XT the community report measured 27–31 tokens/s with Strata against 20 tokens/s for llama.cpp's HIP build on the same card.
Is this checker official?
No. This is an independent tool that applies the rules and numbers published in Strata's repository. The installer itself also checks your card, RAM and disk and recommends a size.
Sources
- Niko1221/Strata on GitHub: README, requirements and install guide
- docs/MODELS.md: sizes, RAM rules and speeds
- docs/DETAILS.md: full measurements and estimates for other GPUs
- docs/AMD_HIP.md: AMD per-card results
- Qwen/Qwen3.8-Flash-Next on Hugging Face