Mistral Large 4 hardware requirements: VRAM, GPUs & quantization

Mistral Large 4 is a ~1-trillion-parameter mixture-of-experts model with 49B parameters active per token. Mistral released it as a public preview on October 6, 2026; the open weights are announced for the end of October. Here is how much memory it takes to run it yourself.

Short answer: the weights alone take about 2.0 TB at BF16, 1.0 TB at FP8, ~0.53–0.6 TB at 4-bit and ~0.26–0.38 TB at 2–3-bit GGUF quants. Only 49B parameters are active per token, but all ~1T must be held in memory. By capacity alone that is an 8-GPU server class machine (e.g. 8× B200 or 8× MI300X at FP8), or one 512 GB Mac Studio for the smallest quants. It will not fit on one RTX 4090 or 5090. These are memory estimates: which runtimes will support the model is not known until the weights ship.

Weights: announced for the end of October 2026. Every figure below is an estimate from the published parameter count until an official model card lists file sizes and context length.

How much VRAM does Mistral Large 4 need?

Weights = parameters × bits per weight ÷ 8. "Block" is the storage rate of the format itself; GGUF "_K"/"_M" files keep some tensors at higher precision, so we assume a slightly higher whole-model average (the number used for the size). Device counts add 10% for KV cache and buffers and assume a usable budget per device (H200 131 GB, B200 168 GB, H100 74 GB, 512 GB Mac 460 GB). They are theoretical capacity minimums, not validated deployments.

FormatBlock bitsAssumed avgWeightsH200B200H100Mac 512 GB
BF161616~2.0 TB1714305
FP888~1.0 TB97153
Q8_08.58.5~1.06 TB97163
Q6_K6.56~6.6~825 GB76132
Q4_K_M4.5~4.8~600 GB6492
NVFP44.54.5~563 GB5492
MXFP44.254.25~531 GB5482
Q3_K_M3.44~3.9~488 GB5482
Q2_K2.63~3.0~375 GB4361
IQ2_XXS2.06~2.1~263 GB3241
IQ1_S1.56~1.6~200 GB2231

Block bits: llama.cpp ggml-common.h; NVFP4 stores an 8-bit scale per 16 values, MXFP4 per 32. Whole-model averages for this architecture are unknown until real files exist; set your own in the calculator. Below ~3 bits, quality on a model this size is untested.

Mistral Large 4 VRAM & memory calculator

CPU-RAM estimate: an idealized batch-of-one figure that assumes every token reads all 49B active weights from system RAM. Keeping layers on a GPU, batching and speculative decoding change it. Illustrative bandwidths: dual-channel DDR5 ≈ 90 GB/s, 8-channel server DDR5 ≈ 300 GB/s, 12-channel ≈ 450–600 GB/s, Mac Studio M3 Ultra 819 GB/s (Apple spec). Sustained bandwidth is lower.

Can you run Mistral Large 4 on a consumer GPU?

Not on GPUs alone: even IQ1_S (~200 GB) is more than six RTX 5090s. If runtimes add support once the weights ship, the home-scale options by memory capacity are:

Mixture-of-experts helps speed, not memory: only 49B parameters work per token, but the router can pick any expert, so all of them must be loaded.

Self-host vs. the Mistral API

The API costs $1.36 per million input tokens and $4.18 per million output tokens (model id mistral-large-4-0). Compare with renting the hardware from the calculator above:

The default $2.50/hour is an example; use your provider's price. Self-hosting usually only makes sense at very high, steady volume or when data must stay on your own hardware.

Mistral Large 4 specs

Total parameters~1 trillion
Active parameters49B per token (mixture of experts)
TypeHybrid instruct + reasoning, natively multimodal (text and images)
Languages160+
ReleasePublic preview October 6, 2026; open weights announced for the end of October 2026
APImistral-large-4-0, $1.36 / $4.18 per million input / output tokens
Context length, license, file formatsNot stated in the announcement (checked October 7, 2026)

FAQ

How much VRAM does Mistral Large 4 need?

By our estimate about 2 TB at BF16, 1 TB at FP8 and roughly 530–600 GB at 4-bit for the weights alone, plus memory for the KV cache. At FP8 that is the capacity of an 8-GPU server such as 8× B200 or 8× MI300X. Official file sizes are not published yet.

Can Mistral Large 4 run on an RTX 4090 or RTX 5090?

No. Even a ~1.6-bit quant is about 200 GB, far beyond one 24–32 GB card. A server with enough system RAM could offload the experts to the CPU if a runtime supports the model, with speed limited by RAM bandwidth.

Can I run Mistral Large 4 on a Mac Studio?

By memory capacity, a 512 GB M3 Ultra can hold Q2_K, IQ2_XXS or IQ1_S quants; Q3_K_M and 4-bit need two clustered machines. That also requires MLX or llama.cpp support for the model, which is not confirmed yet.

Why does a 49B-active model need memory for 1T parameters?

In a mixture-of-experts model a router picks different experts for every token, so all experts must be loaded. Active parameters decide speed (bytes read per token); total parameters decide memory.

When will Mistral Large 4 weights be released?

Mistral announced the weights for the end of October 2026. Community GGUF, MLX or Ollama builds need both the weights and runtime support; no date for those is confirmed.

Is Mistral Large 4 on Ollama?

Not yet. Ollama and llama.cpp builds need the open weights, announced for the end of October 2026, and support for the architecture. Until then it is available through Mistral's API.

Sources