Local LLM glossary
- Page updated
- Sources checked
- Model figures link to their sources
The words the checkers use, in plain English. Figures about a specific model, file or tool link to their source or to the RigCheck page that sources it; the checkers' own settings, such as the 32K default context, are explained under each checker.
In one paragraph: a model runs at full speed when its file and a working cache fit in fast memory. VRAM is the graphics card's own memory and the fastest place for it; system RAM is slower; a Mac or DGX Spark has one pool of unified memory. Quantization shrinks the file by storing each weight in fewer bits (Q8_0 ≈ 8.5 bits per weight, Q4_K_M ≈ 4.9, IQ2_XS ≈ 2.6 in llama.cpp's own measurements), trading some quality for size. On top of the file, the KV cache grows with the length of the conversation. What does not fit on the GPU can be offloaded to system RAM, which works but is slower; some engines, such as Strata, can even read part of the model from the SSD while it answers, slower still.
Memory
- VRAM
- The memory on the graphics card. The model runs fastest when it fits here, because the GPU reads every weight it needs from VRAM for each token.
- System RAM
- The computer's main memory. Runtimes such as llama.cpp keep what does not fit in VRAM here; the CPU then works on that part, which is slower. How much must stay free for the operating system differs by checker: the llama.cpp checkers (Qwen3.8-Flash-Next, Mistral Medium 3.5) keep 8 GB of a PC's system RAM free, and 10% (at least 8 GB) of a unified-memory machine's, RigCheck assumptions stated under each; the Strata checker follows Strata's own rule of the experts plus about 10 GB.
- Unified memory
- One memory pool shared by the CPU and GPU, as on Apple Silicon Macs, NVIDIA DGX Spark and AMD Ryzen AI Max. A Mac can give most of its memory to the model, but that needs the GPU memory limit raised (the checkers give the command), and on a Ryzen AI Max how much the GPU may use depends on the OS and BIOS.
- KV cache
- Memory the model uses to remember the conversation so far: it grows with every token of context. llama.cpp stores it at 16 bits (f16) by default and can store it at about 8 bits (q8_0), which roughly halves it (llama-server options
-ctk/-ctv). Its size per token depends on the model: the checkers calculate it from each model's configuration. - Context length
- How many tokens the model can keep in view at once: your messages, its answers and any files you paste. The checkers use 32K (32,768 tokens) by default; a longer context needs a bigger KV cache. Without
-c, llama-server starts from the model's full context and its automatic fitting may shrink it to what your memory holds (llama-server README); the commands on this site set-cto the context the checker assumed.
Files and quantization
- Weights and parameters
- The numbers that make up a model. "125B" means 125 billion of them. At 16 bits (2 bytes) each, a billion parameters take 2 GB, so a 125B model is about 250 GB before quantization (arithmetic, not a file size: the real files are on each model's page).
- BF16, FP8, NVFP4
- Number formats for the weights, named after their bits per value: 16 (BF16, usually the original release), 8 (FP8) and 4 (NVFP4, for recent NVIDIA GPUs; block scales add a little, see the Mistral Large 4 storage rates). Servers such as vLLM and SGLang run these directly.
- GGUF
- The model file format of llama.cpp, also used by Ollama and LM Studio. A GGUF holds the weights, quantized or not (BF16 GGUFs exist too), together with the model's settings and tokenizer. Image input needs a separate vision file (see mmproj below). Large models are split into numbered parts (
-00001-of-00003.gguf); the checkers count all parts. - Quantization
- Storing weights in fewer bits to make the file smaller, at some cost in quality. The fewer the bits, the larger the cost; the checkers show quality measurements where the file's maker published them.
- Q4_K_M, Q8_0, IQ2_XS and the other names
- llama.cpp's quantization types. The number is roughly the bits per weight; "K" types and "IQ" (importance-based) types are newer schemes that keep more quality at the same size; S, M, L and XS, XXS mark smaller or larger variants. In llama.cpp's own measurements on one 8B model the averages were Q8_0 8.50, Q6_K 6.56, Q4_K_M 4.89, IQ4_XS 4.46, Q3_K_M 4.00, IQ3_XXS 3.25, IQ2_XS 2.59 and IQ1_M 2.15 bits per weight (llama-quantize README). Real files mix types, so file sizes, not names, are what the checkers use.
- UD- (Unsloth Dynamic)
- Unsloth's GGUFs whose names start with UD- choose the quantization type layer by layer for each model instead of using one type everywhere (Unsloth Dynamic docs). The rest of the name (UD-Q4_K_XL) gives the general size class.
- mmproj (vision file)
- A separate file that lets a model read images. llama-server downloads it automatically with
-hfunless you pass--no-mmproj; the checkers add its size only when you ask for image input.
Model types
- Dense model
- A model that uses all of its weights for every token, such as Mistral Medium 3.5 (128B, per its model card). Its speed is limited by how fast the hardware can read the whole file for each token, so the part in slower memory sets the pace.
- MoE (mixture of experts) and active parameters
- A model split into many "experts", of which only a few are used for each token. Qwen3.8-Flash-Next has 125B parameters, 6B of them active per token (per its model card). The experts usually all sit in memory, but each token reads only a small part, which is why MoE models can run usefully with most of the experts in system RAM; Strata can even read some of them from the SSD, more slowly.
- Open weights
- A model whose weight files you can download and run yourself. Until the weights are out (as with Mistral Large 4 before its release), RigCheck only estimates and gives no verdict.
Running a model
- llama.cpp, Ollama, vLLM, SGLang
- Programs that run a model. llama.cpp runs GGUF files on GPUs, CPUs and Macs and can split a model between VRAM and RAM; Ollama and LM Studio are built on it. vLLM and SGLang are servers aimed at data-center GPUs. Strata is an engine made for one model family on gaming PCs.
- Offload
- Putting part of a model on the GPU and the rest in system RAM. In llama.cpp,
-nglsets how many layers go to the GPU, and--n-cpu-moekeeps MoE expert weights in RAM; by default it fits the model to your devices itself (llama-server README). - tok/s
- Tokens per second, the speed of writing an answer. A token is a word or a piece of one. Reading a long prompt ("prompt processing") is a separate, much faster speed.
- MTP and speculative decoding
- A small draft model or extra layer (MTP, multi-token prediction) guesses the next few tokens, and the main model checks them in one step, which speeds up writing when the guesses are right. Strata downloads an MTP draft layer on its first start (~6 GB, Strata MODELS.md); llama.cpp can use one with
--spec-type draft-mtp(llama-server README). - Low-RAM mode (Strata)
- Strata's way to run a size larger than your RAM allows: experts the graphics card does not hold are read from the SSD instead, which is much slower with a 12–16 GB card (Strata MODELS.md).
Missing a term, or found a mistake? Email hello@localrigcheck.com.