Mistral Medium 3.5 hardware requirements: VRAM, GPU and RAM

Mistral Medium 3.5 is Mistral AI's dense 128-billion-parameter model with a 256K context, released as open weights under a Modified MIT license on April 28, 2026. Unlike a mixture-of-experts model, it uses every parameter for every token, so memory decides whether it fits and memory bandwidth decides how fast it can be.

Short answer: the 4-bit GGUF files are about 75 GB (Q4_K_M 74.9 GB, Unsloth's dynamic UD-Q4_K_XL 75.7 GB), 8-bit is 133 GB, BF16 250 GB and the smallest 2-bit file 35 GB. Add the KV cache: 11.8 GB at a 32K context (16-bit). By our estimate at 32K, UD-Q4_K_XL fits fully on a 128 GB Mac or DGX Spark, and one 96 GB RTX PRO 6000 holds IQ4_XS. With 64 GB of RAM, a single RTX 4090 runs Q3_K_M and a single RTX 5090 runs IQ4_XS only with offload: in each case about 53 GB of layers and their KV cache sit in system RAM. With two RTX 3090s, UD-Q4_K_XL puts about 48 GB of layers and KV cache in RAM. That caps speed at roughly 2 tokens per second (RigCheck ceiling at 90 GB/s dual-channel DDR5, short context).

Can my hardware run Mistral Medium 3.5?

Setup

RigCheck estimate for llama.cpp with Unsloth's GGUF files. Memory needed = the file + KV cache for your context (f16, or q8_0 at 34/64 of that) + 5% of the file for compute buffers. Usable memory keeps 8% of each GPU for the driver and runtime, 8 GB of system RAM for the OS, and 10% (at least 8 GB) of unified memory for the OS. With too little VRAM, whole layers and their KV cache run from system RAM (llama.cpp's -ngl); layers and KV split across several GPUs by default. The speed ceiling divides the weights read per token by memory bandwidth: short context, batch of one, no speculative decoding; real speed is lower. Unified-machine bandwidths are vendor specs (the Ryzen AI Max+ 395 figure is calculated from AMD's 256-bit LPDDR5x-8000). The 90 GB/s RAM default is an illustrative dual-channel DDR5 figure; use your own.

How much VRAM does Mistral Medium 3.5 need?

File sizes as Hugging Face lists them (all parts of each file, decimal GB, read October 8, 2026). Add the KV cache for your context (see Specs) and a few GB for buffers.

FormatFile sizeRuns withSource
BF16 (GGUF)250.1 GBllama.cppUnsloth
FP8 (official weights)133.6 GBvLLM, SGLangMistral
Q8_0132.9 GBllama.cppUnsloth
Q6_K102.6 GBllama.cppUnsloth
NVFP495.2 GBvLLM on NVIDIA Blackwell, LinuxNVIDIA
Q5_K_M88.3 GBllama.cppUnsloth
UD-Q4_K_XL (dynamic 4-bit)75.7 GBllama.cppUnsloth
Q4_K_M74.9 GBllama.cppUnsloth
IQ4_XS67.1 GBllama.cppUnsloth
Q3_K_M60.6 GBllama.cppUnsloth
UD-IQ3_XXS49.3 GBllama.cppUnsloth
UD-IQ2_M44.1 GBllama.cppUnsloth
UD-IQ2_XXS34.9 GBllama.cppUnsloth

Unsloth's GGUFs are the ones Mistral's model card links for llama.cpp; other uploaders' files differ by a few GB (bartowski's Q4_K_M is 78.4 GB). Image input needs the separate vision file (mmproj, 5.4 GB at BF16), which Unsloth says does not work in Ollama. NVIDIA's card lists vLLM, NVIDIA Blackwell and Linux for its NVFP4 build and keeps part of the model above 4 bits (MLP layers 4–86 in NVFP4), so it is larger than the bit width suggests. Ollama packages its own builds: mistral-medium-3.5 is 80 GB (q4_K_M, the default), with q8_0 at 138 GB and bf16 at 255 GB, text and image input (Ollama library).

What hardware can run Mistral Medium 3.5?

RigCheck estimate at a 32K context with an f16 KV cache, from the file sizes above and the assumptions under the checker. "Largest" means the largest of the checker's ten files.

HardwareLargest file that fits fullyWith layers in system RAM (slow)
1× RTX 4090 (24 GB) + 64 GB RAMnoneQ3_K_M (65 of 88 layers, 53 GB in RAM)
1× RTX 5090 (32 GB) + 64 GB RAMnoneIQ4_XS (59 of 88 layers, 53 GB in RAM)
2× RTX 3090 (48 GB) + 64 GB RAMnoneUD-Q4_K_XL (48 of 88 layers, 48 GB in RAM)
2× RTX 5090 (64 GB) + 64 GB RAMUD-IQ2_MQ5_K_M (41 of 88 layers, 47 GB in RAM)
4× RTX 3090 (96 GB) + 64 GB RAMIQ4_XSQ6_K (25 of 88 layers, 33 GB in RAM)
1× RTX PRO 6000 (96 GB) + 64 GB RAMIQ4_XSQ6_K (25 of 88 layers, 33 GB in RAM)
2× RTX PRO 6000 (192 GB) + 64 GB RAMQ8_0—
1× H100 (80 GB)UD-IQ3_XXS—
1× H200 (141 GB)Q6_K—
NVIDIA DGX Spark (128 GB unified)Q5_K_M—
AMD Ryzen AI Max+ 395 (128 GB)Q5_K_M (if ~105 GB is GPU-accessible)—
Mac with M5 Max, 40-core GPU (128 GB unified)Q5_K_M—
Mac Studio M3 Ultra (256 GB unified)Q8_0—
Mac Studio M5 Ultra (256 GB unified)Q8_0—

Capacity only, not tested setups. Mistral's model card lists llama.cpp (via Unsloth's GGUFs), Ollama, vLLM, SGLang and transformers; Unsloth's guide builds llama.cpp with CUDA, or without it for CPU and unified-memory machines, and warns against CUDA 13.2, which may produce gibberish (checked October 8, 2026). llama.cpp also has AMD backends, but we found no report of this model on Ryzen AI Max hardware. A longer context needs more memory: at 128K the KV cache alone is 47.2 GB (f16) or 25.1 GB (q8_0).

How fast is Mistral Medium 3.5 locally?

A dense model reads all of its weights for every token, so generation speed is bounded by memory bandwidth ÷ file size. Measured numbers are scarce; this is everything we found:

HardwareSetupReported generation rateSource
NVIDIA DGX SparkvLLM, community NVFP4 build zdy1995love/Mistral-Medium-3.5-128B-NVFP4 + Mistral's EAGLE draft model (3 speculative tokens), FP8 KV cache, text only, 16,384-token maximum contextQ&A 6.9, code 8.9, JSON 8.8, long code 9.1, short math 2.8–2.9 tok/s (output tokens ÷ elapsed time)Community report, NVIDIA forum, May 4, 2026
NVIDIA DGX Sparkllama.cpp without EAGLE (file not stated)"about 2 tokens per second"Same report

Ceilings for the 75.7 GB UD-Q4_K_XL file from vendor bandwidth (RigCheck estimate, not a measurement; short context, no speculative decoding): DGX Spark ≤ 3.6 tok/s (273 GB/s), Ryzen AI Max+ 395 ≤ 3.4 (256 GB/s), M5 Max ≤ 8.1 (614 GB/s), M3 Ultra ≤ 10.8 (819 GB/s), M5 Ultra ≤ 15.9 (1.2 TB/s). Speculative decoding (Mistral's EAGLE model for vLLM and SGLang) can beat these ceilings, as the DGX Spark report shows with a different, NVFP4 file.

Mistral Medium 3.5 specs

Parameters128B, dense Model card
Context length256K (262,144 tokens) Model card, params.json
Architecture88 layers, 96 attention heads, 8 KV heads × 128 dims, plus a vision encoder for image input params.json
KV cache352 KiB per token at f16: 11.8 GB at 32K, 47.2 GB at 128K, 94.5 GB at 256K; q8_0 is 34/64 of that RigCheck calculation from params.json
Official weightsFP8 (e4m3), 133.6 GB Hugging Face
LicenseModified MIT: commercial and non-commercial use, with exceptions for companies with large revenue Model card
ReleaseApril 28, 2026, as open weights Mistral changelog
RuntimesvLLM (recommended by Mistral), llama.cpp, Ollama, SGLang, transformers; LM Studio "WIP" (checked October 8, 2026) Model card
API$1.5 per million input tokens, $7.5 per million output tokens Mistral announcement

FAQ

How much VRAM does Mistral Medium 3.5 need?

About 75 GB for a 4-bit GGUF file, 133 GB at 8-bit and 250 GB at BF16, plus the KV cache (11.8 GB at 32K context, f16) and a few GB of buffers. The smallest Unsloth file, UD-IQ2_XXS, is 34.9 GB.

How big is Mistral Medium 3.5?

128 billion parameters, all used for every token (a dense model). Mistral's official FP8 weights are 133.6 GB on Hugging Face; Unsloth's GGUF files run from 34.9 GB (UD-IQ2_XXS) to 250.1 GB (BF16).

Can Mistral Medium 3.5 run on an RTX 4090 or RTX 5090?

Not fully on the card: even the 34.9 GB 2-bit file plus its KV cache needs more than 24 or 32 GB. With 64 GB of system RAM, llama.cpp can run the remaining layers from RAM: Q3_K_M on a 4090 or IQ4_XS on a 5090, each with about 53 GB of layers and their KV cache in RAM at 32K context (RigCheck estimate). A dense model reads those layers for every token, which caps speed at roughly 2 tokens per second with 90 GB/s dual-channel DDR5.

Will two RTX 3090s run Mistral Medium 3.5?

A 4-bit file only with offloading: 48 GB of VRAM leaves 48 of UD-Q4_K_XL's 88 layers, about 48 GB with their KV cache, in system RAM at 32K context (RigCheck estimate), roughly 2 tokens per second at best with 90 GB/s RAM. The 2-bit UD-IQ2_XXS fits fully at 8K context. Four RTX 3090s (96 GB) hold IQ4_XS fully at 32K.

Can I run Mistral Medium 3.5 on a Mac?

Yes, with enough unified memory. By our estimate a 128 GB Mac holds up to Q5_K_M at 32K context and 256 GB holds Q8_0; a 64 GB Mac only the 2-bit UD-IQ2_XXS (Unsloth's guide lists 64 GB of total memory for 3-bit, which leaves almost nothing for macOS and the context). Raise the GPU wired-memory limit first. Speed is bounded by memory bandwidth: at most about 8 tok/s on an M5 Max and 11 on an M3 Ultra for a 4-bit file.

How fast is Mistral Medium 3.5 on a DGX Spark?

One community report (NVIDIA forum, May 4, 2026) measured 6.9–9.1 tokens per second on Q&A, code, JSON and long-code prompts, and 2.8–2.9 on a 9-token math answer, with vLLM, a community NVFP4 build and Mistral's EAGLE draft model. The same user saw about 2 tokens per second with llama.cpp without EAGLE.

Can I run the NVFP4 version of Mistral Medium 3.5?

NVIDIA's NVFP4 build (95.2 GB) lists vLLM on NVIDIA Blackwell GPUs and Linux. On older GPUs, Macs or with llama.cpp, use a GGUF file instead.

Is Mistral Medium 3.5 on Ollama?

Yes. ollama run mistral-medium-3.5 downloads Ollama's 80 GB q4_K_M build; q8_0 (138 GB) and bf16 (255 GB) tags are also listed, with text and image input (checked October 8, 2026). Unsloth's separate vision files do not work in Ollama.

When was Mistral Medium 3.5 released?

April 28, 2026, as open weights under a Modified MIT license (Mistral changelog). The weights are on Hugging Face as mistralai/Mistral-Medium-3.5-128B.

Do I need to re-download my Mistral Medium 3.5 GGUF?

If you downloaded it before May 1, 2026, yes. Mistral's model card warns that GGUFs built from the original Transformers config lose long-context quality, and Unsloth released updated GGUFs with the fix on May 1, 2026.

Mistral Medium 3.5 or Mistral Large 4 for a home setup?

Medium 3.5. Large 4 has 1.05T parameters and needs about 210 GB even at the smallest quant, by our estimate (see Mistral Large 4 hardware requirements), while a 4-bit Medium 3.5 fits a 128 GB machine.

Sources