skip to content

In llama.cpp GGUF filenames, what does a quant level like Q4_K_M mean?

level: middleimportance: must knowfreq 70%

answer

  1. Bits, block format, size variant
  2. K means per-block scales
  3. S, M, L are tensor mixes
  4. Effective bpw exceeds the nominal number
  5. Q4_K_M is the usual default

basics

~20 s

Q4 is the nominal bits per weight, _K marks llama.cpp's k-quant block format with per-block scales, and _S/_M/_L choose how much extra precision goes to the most sensitive tensors. Q4_K_M averages roughly 4.8 bits per weight.

solid answer

~50 s

A GGUF quant name encodes three things. The number after **Q** is the nominal bit width of the stored weights — Q2 through Q8. The **_K** suffix means the k-quant family, where weights are grouped into small blocks inside super-blocks, and each block carries its own scale (and for some types a minimum), so the quantization error is local rather than per-tensor. The trailing **_S/_M/_L** is the size variant: the medium and large mixes keep certain sensitive tensors — notably attention value projections and feed-forward down projections — at a higher bit width than the rest, which is why Q4_K_M lands around 4.8 effective bits per weight rather than 4.0. Older names like **Q8_0** and **Q4_0** are the legacy block formats with one scale per 32-weight block and no min. The practical guidance for a Llama build is that Q4_K_M is the usual quality/size sweet spot, Q8_0 is effectively lossless but twice the size, and anything below Q3 degrades noticeably.

code

bash · 1 line
bash
llama-quantize models/llama-3.1-8b-f16.gguf models/llama-3.1-8b-Q4_K_M.gguf Q4_K_M

go deeper

for a junior

Be able to read the name: the digit is the bit width, bigger digit means bigger file and better quality, and Q4_K_M is the common default people download.

for a middle

Explain the mechanics — per-block scales in the k-quant format, the S/M/L mixes promoting sensitive tensors, and why effective bits per weight exceed the nominal number.

for a senior

Show you pick a level from measurements, not folklore: know that Q8_0 is your near-lossless reference, that small models degrade far faster than large ones at the same level, and that file size predicts loaded weight memory.

for a principal

Own the policy question — which quant levels you standardise on across a fleet, whether the imatrix calibration set is representative of production traffic, and how you keep quantized builds reproducible and re-verified when weights are updated.

## What GGUF is GGUF is the single-file model format used by llama.cpp (it replaced the older GGML format). One GGUF file holds the tensors, the tokenizer, the architecture metadata and the chat template. When people say "a Q4_K_M of Llama 3.1 8B", they mean a GGUF file whose tensors were converted from the original FP16/BF16 weights down to a 4-bit block format. ## Reading the name A name such as `Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf` decomposes as: - **Q4** — the nominal number of bits used to store each weight. Types run from Q2 (very small, very lossy) to Q8 (near-lossless). - **_K** — the k-quant family. Instead of a single scale for a whole tensor, weights are split into blocks (typically 16 or 32 values) that are grouped into super-blocks. Each block stores its own quantized scale, and some types also store a per-block minimum so the representable range can be shifted, not just stretched. Localising the scale is what lets 4 bits stay usable: an outlier weight only ruins its own block's resolution instead of the whole tensor's. - **_M** — the size variant within that bit width, one of `_S` (small), `_M` (medium), `_L` (large). These are *mixes*: llama.cpp does not use one type for every tensor. The larger variants promote the tensors that empirically hurt most when crushed — attention value projections and feed-forward down projections are the classic ones — to a higher-precision type such as Q6_K, while everything else stays at the nominal width. Token embedding and output tensors are also commonly kept at a higher precision. Because of the block scales and the promoted tensors, the *effective* bits per weight is always above the nominal number. Q4_K_M is roughly 4.8 bpw; Q5_K_M roughly 5.7; Q8_0 roughly 8.5. ## Legacy and importance-matrix types Names without `_K` — `Q4_0`, `Q4_1`, `Q5_0`, `Q8_0` — are the older, simpler block formats. `_0` means one fp16 scale per block; `_1` adds a minimum. They are simpler and still common for Q8_0, which is used precisely because it is dumb, fast and effectively lossless. The `IQ` family (`IQ2_XXS`, `IQ3_M`, `IQ4_XS`, `IQ4_NL`, …) is newer and uses non-linear codebooks plus an *importance matrix* — statistics collected by running calibration text through the model with `--imatrix` — to decide which weights deserve resolution. IQ types reach lower bits per weight at a given quality than the equivalent K type, at the cost of a calibration pass and somewhat slower kernels on some hardware. ## Producing one You convert the Hugging Face weights to a full-precision GGUF, then run the quantizer: `llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_M` No GPU and no calibration data are needed for plain K-quants; the process is a straightforward per-block rounding of the stored tensors. That is a real difference from GPTQ and AWQ, which run calibration text through the model. ## Choosing a level The practical ladder for a Llama-family model: - **Q8_0** — treat as a reference; differences from FP16 are usually within measurement noise. Use it when you have the memory and want to rule quantization out as a suspect. - **Q6_K / Q5_K_M** — very small quality loss, meaningful size saving. - **Q4_K_M** — the default recommendation, and the level most published GGUF builds lead with. Small but measurable loss. - **Q4_K_S / Q3_K_M** — noticeably lossier; acceptable when memory is the binding constraint. - **Q2 / IQ2** — only worth it to fit a much larger model that would otherwise not run at all, and even then the small models in a family suffer far more than the large ones. The size on disk is a good proxy for the memory the weights will occupy when loaded, which is why the file size is often the fastest way to decide whether a build fits your hardware. ## Common misreadings The `_M` is not a mixed-precision *calibration* setting, the `K` does not mean "thousands" or a group size of K, and `Q4_0` is not simply an older name for `Q4_K_S` — they are different block layouts with different error characteristics. Nor does a 4-bit quant mean the arithmetic happens in 4 bits: llama.cpp dequantizes blocks on the fly into the compute type, so the win is memory footprint and memory bandwidth, not integer math.

  • What does the trailing _0 in Q8_0 mean, and why is that type still used?
    The `_0` marks the legacy block layout: each block of 32 weights stores a single fp16 scale and no minimum. It is less clever than the k-quants, but at 8 bits the extra cleverness buys nothing — Q8_0 is effectively indistinguishable from FP16 on most tasks. It is the type people reach for as a quality reference point, or when memory is plentiful and they want quantization ruled out as a source of regressions.
  • What do the IQ quant types add over the K types, and what do they cost?
    IQ types use non-linear codebooks plus an importance matrix — per-weight statistics gathered by running calibration text through the model with llama.cpp's `--imatrix` — so bits are spent where activations say they matter. That lets IQ2/IQ3 builds stay usable at bit widths where plain K-quants collapse. The costs are a calibration pass before quantizing, dependence on how representative that text was, and kernels that can be slower than the equivalent K type on some backends.
  • Why is a Q4_K_M file bigger than four bits times the parameter count?
    Two reasons. Each block stores its own scale, and for some types a minimum, which is real overhead spread across every block. And the `_M` mix deliberately promotes sensitive tensors — attention value and feed-forward down projections, plus usually embeddings and the output layer — to a higher-precision type. Together those push the average to about 4.8 bits per weight, so a 8B model lands near 4.9 GB rather than 4.0 GB.

saying these in an interview costs you the question

  • Thinking the K in Q4_K_M means thousands or group size
  • Claiming 4-bit quants do the matrix math in 4 bits
  • Treating Q4_0 as just an older name for Q4_K_S
  • Assuming every tensor in Q4_K_M is stored at 4 bits
  • Believing GGUF K-quants need a calibration dataset

context