skip to content

Meta Llama

Meta's open-weight family — the models you can download, quantize, fine-tune and serve yourself. Interviews use Llama as the vehicle for every self-hosting question: sizing, serving throughput, and the license fine print.

on this pageshow

explore

questions

page 1 of 2

What does tokenizer.apply_chat_template() do when prompting a Llama model?

level: juniorimportance: must knowfreq 60%

answer

  1. the format ships with the model
  2. a Jinja template in the tokenizer files
  3. messages in, exact prompt string out
  4. one flag opens the assistant turn
  5. add_generation_prompt=True

basics

~20 s

apply_chat_template() turns a list of {role, content} messages into the exact prompt string the checkpoint was fine-tuned on, using the Jinja template shipped alongside that model's tokenizer — so the special tokens come from the model, not from your code.

solid answer

~40 s

In Hugging Face transformers, every instruct checkpoint ships a Jinja chat template with its tokenizer files. `tokenizer.apply_chat_template(messages, tokenize=False)` renders your messages through it and returns the prompt string; with `tokenize=True` you get token IDs directly. The argument that matters most is `add_generation_prompt=True`: it appends the opening assistant header (for Llama 3, `<|start_header_id|>assistant<|end_header_id|>` plus the blank line) so the model continues *as the assistant*. Leave it off and the model is just as likely to invent another user turn. The practical value is portability: swap a Llama 2 checkpoint for a Llama 3 one and the same call emits `[INST]`-style or header-style markup automatically. Hand-written format strings do not survive that swap, and they fail silently rather than loudly.

code

python · 15 lines
python
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")

messages = [
    {"role": "system", "content": "You are terse."},
    {"role": "user", "content": "Name one benefit of GQA."},
]

prompt = tok.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
print(repr(prompt))

go deeper

for a junior

Be able to write the call from memory: pass a list of role/content dicts, set add_generation_prompt=True, and know the result is the prompt string the model expects.

for a middle

Explain where the template lives (a Jinja template in the model repo's tokenizer files) and why that makes checkpoint swaps safe, plus what tokenize=False versus tokenize=True returns.

for a senior

Talk about the failure modes you have hit: duplicated special tokens, forks shipping edited templates, and double rendering when a server also applies the template. Say how you verify the rendered prompt in CI.

for a principal

Frame the template as part of the model's public interface and version it with the weights. Decide organizationally whether rendering happens in the client or the inference server, and make that boundary explicit so two teams never apply it twice.

## The problem it solves An instruction-tuned open-weight model only behaves well when the prompt matches the layout it was tuned on, down to the marker tokens and the blank lines. Hardcoding that layout couples your code to one checkpoint. The chat template exists so the layout ships **with the model** instead: a Jinja template stored in the model repo's tokenizer files (historically in the `chat_template` field of `tokenizer_config.json`, and in recent transformers releases as a standalone `chat_template.jinja` file in the repo). ## The call ``` messages = [ {"role": "system", "content": "You are terse."}, {"role": "user", "content": "Explain RoPE in one line."}, ] prompt = tokenizer.apply_chat_template( messages, tokenize=False, add_generation_prompt=True ) ``` - `messages` — a list of dicts with `role` and `content`. Llama templates accept `system`, `user` and `assistant`; Llama 3.1+ adds `ipython` for tool results. - `tokenize=False` returns the rendered **string**, which is what you want when inspecting or when the string goes to a separate serving process. `tokenize=True` (the default) returns token IDs and skips the round-trip. - `add_generation_prompt=True` appends the assistant turn opener. This is the single most common omission: without it the sequence ends at a closed turn and the model may write `<|start_header_id|>user<|end_header_id|>` and continue the conversation on your behalf. - `continue_final_message=True` does the opposite for prefill: instead of opening a new assistant turn, it leaves the last assistant message unterminated so the model continues that text. Use it to force a reply to start with a given prefix; do not combine it with `add_generation_prompt`. ## What it does *not* do It does not validate your role sequence for semantics; most Llama templates do enforce strict alternation of user and assistant and will raise if you interleave them wrongly. It does not add a system prompt if you omit one — Llama 3 is happy without a system turn. And it does not exist for every checkpoint: base (non-instruct) models often ship no chat template, and modern transformers raises an explicit error rather than guessing a default. Seeing that error on a `-base` model is the signal that you picked the wrong artifact, not that the library is broken. ## The pitfalls worth naming in an interview 1. **Double special tokens.** The rendered string already contains BOS. Passing it back through `tokenizer(prompt)` with default settings prepends another one; use `add_special_tokens=False`, or let `apply_chat_template` tokenize for you. 2. **Stopping.** The template renders the prompt, not the stop criteria. You still have to stop generation on the model's end-of-turn token. 3. **Template drift.** Community re-uploads and quantized conversions sometimes carry a hand-edited or missing template. When a fork behaves worse than the official weights, diff the rendered prompt against the reference repo's before blaming the quantization. 4. **Server-side rendering.** OpenAI-compatible servers apply the template themselves when you post a `messages` array; if you post an already-rendered string to a completion endpoint you can get the template applied twice. Pick one layer and be explicit about which. ## Why interviewers like it It separates candidates who have actually served open weights from those who have only called a hosted API. The hosted path hides formatting entirely; the moment you download weights, this call is the boundary between "works" and "quietly mediocre". The right answer names the function, names `add_generation_prompt`, and says where the template comes from.

  • What happens if you forget add_generation_prompt=True?
    The rendered prompt ends at a completed turn, so the model has no cue that it should speak next. It commonly opens a fresh user header and writes both sides of the conversation, or answers with a stray leading newline. With the flag, the sequence ends inside an open assistant turn and the very next token is the start of the reply.
  • You load a checkpoint and apply_chat_template raises because there is no template. What does that tell you?
    Almost always that you pulled a base (pre-trained) model rather than an instruct/chat variant — base weights ship no chat template because they were never tuned on a conversation layout. Modern transformers raises instead of falling back to a guessed default. Switch to the instruct checkpoint, or supply a template explicitly if you fine-tuned your own conversational format.
  • How would you pin a reply to start with a fixed prefix such as "```json"?
    Append an assistant message containing the prefix and render with `continue_final_message=True` instead of `add_generation_prompt=True`. That leaves the assistant turn unterminated so generation continues the prefix rather than starting a new turn. Remember to prepend the prefix back onto the decoded output, since the model only returns the continuation.

saying these in an interview costs you the question

  • Thinks the template is the same for every open model
  • Hardcodes [INST] markers instead of rendering from the tokenizer
  • Omits add_generation_prompt and blames the model for rambling
  • Believes a base checkpoint has a chat template
  • Applies the template and then posts to a chat endpoint that applies it again

context

open as a page

Why is Meta's Llama called open-weight rather than open-source?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Llama weights are downloadable and usable commercially, but the Llama Community License imposes field-of-use limits through an acceptable-use policy and singles out very large operators for extra permission. Those restrictions fail the Open Source Definition, so the accurate term is open-weight.

open as a page

How much memory do Llama 70B weights need at FP16 versus 4-bit quantization?

level: juniorimportance: must knowfreq 65%

basics

~20 s

Weights cost bytes-per-parameter times parameter count: about 140 GB for 70B at FP16, about 70 GB at 8-bit, and roughly 35-40 GB at 4-bit once block scales are counted. KV cache and activations are extra and are not shrunk by weight quantization.

open as a page

When should you serve Llama with Ollama versus vLLM?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Ollama wraps llama.cpp for local, single-user use: one command pulls a quantized model and runs it on a laptop, CPU included. vLLM is a GPU server built for concurrent production traffic, using continuous batching and a paged KV cache.

open as a page

How does Llama 3's chat template differ from Llama 2's [INST] and <<SYS>> format?

level: middleimportance: must knowfreq 68%

basics

~20 s

Llama 2 wraps each user turn in [INST] ... [/INST] plain-text markers and nests the system prompt inside the first one with <<SYS>>. Llama 3 drops both, using real special tokens: a role header per turn, closed by <|eot_id|>.

open as a page

In LoRA fine-tuning of Llama, what do rank r and lora_alpha control?

level: middleimportance: must knowfreq 72%

basics

~20 s

Rank r sets the width of LoRA's low-rank update matrices, fixing adapter capacity and the trainable-parameter count. lora_alpha scales that update — PEFT applies it as alpha/r — so raising alpha strengthens the adapter's effect without adding parameters.

open as a page

What does QLoRA change compared with plain LoRA when fine-tuning Llama?

level: middleimportance: must knowfreq 60%

basics

~20 s

QLoRA loads the frozen Llama base in 4-bit NF4 instead of 16-bit, then trains ordinary LoRA adapters in bfloat16 on top. Weights are dequantized block by block during the forward and backward passes, cutting base-model memory roughly fourfold.

open as a page

What is the 700 million monthly-active-user clause in Meta's Llama Community License?

level: middleimportance: must knowfreq 70%

basics

~20 s

Meta's Llama Community License grants free commercial use unless the licensee's products had more than 700 million monthly active users in the calendar month before that Llama version's release date. Above that line you must request a separate licence from Meta, which Meta may refuse.

open as a page

How do Llama 2, Llama 3 and Llama 3.1 differ in sizes and context length?

level: middleimportance: must knowfreq 72%

basics

~20 s

Llama 2 shipped 7B, 13B and 70B with a 4K context. Llama 3 narrowed to 8B and 70B at 8K context with a much larger tokenizer vocabulary. Llama 3.1 kept those sizes, added a 405B flagship, and extended all of them to 128K context.

open as a page

Why does every Llama 3 size use grouped-query attention rather than multi-head attention?

level: middleimportance: must knowfreq 58%

basics

~20 s

Grouped-query attention lets several query heads share one key/value head, so the KV cache shrinks by that ratio. Llama 3 uses 8 key/value heads at every size, cutting per-token cache memory about fourfold on 8B and eightfold on 70B, which is what makes long contexts and large batches affordable.

open as a page

In llama.cpp GGUF filenames, what does a quant level like Q4_K_M mean?

level: middleimportance: must knowfreq 70%

basics

~20 s

Q4 is the nominal bits per weight, _K marks llama.cpp's k-quant block format with per-block scales, and _S/_M/_L choose how much extra precision goes to the most sensitive tensors. Q4_K_M averages roughly 4.8 bits per weight.

open as a page

How does continuous batching in vLLM raise Llama serving throughput?

level: middleimportance: must knowfreq 64%

basics

~20 s

Continuous batching schedules per decode iteration instead of per batch: a finished sequence leaves the running batch immediately and a queued request takes its place at the next step, so the GPU never idles waiting for the slowest generation in a fixed batch.

open as a page

Why do self-hosted Llama servers expose an OpenAI-compatible endpoint?

level: middleimportance: must knowfreq 58%

basics

~20 s

Because the OpenAI chat-completions shape became the de facto wire protocol: exposing /v1/chat/completions lets existing SDKs, gateways and frameworks talk to a local model by changing only the base URL. Compatibility covers the common fields, not every parameter or feature.

open as a page

When is fine-tuning Llama the wrong tool compared with prompting plus RAG?

level: principalimportance: must knowfreq 50%

basics

~20 s

Fine-tuning teaches form — output shape, tone, task convention, refusal behaviour. It is a poor way to install facts that change, because updating them means retraining. Retrieval handles changing knowledge; fine-tuning handles how the answer is produced.

open as a page

What distinguishes Llama 3.2's 1B and 3B models from its 11B and 90B ones?

level: juniorimportance: should knowfreq 42%

basics

~20 s

Llama 3.2 shipped two different tiers. The 1B and 3B are small text-only models built for on-device and edge use, produced by pruning and distilling larger Llama 3.1 checkpoints. The 11B and 90B are vision models that accept images alongside text.

open as a page

Why must Llama 3 Instruct generation stop on <|eot_id|>, not just EOS?

level: middleimportance: should knowfreq 48%

basics

~20 s

Llama 3 Instruct closes each turn with <|eot_id|>; the base end-of-sequence token <|end_of_text|> almost never appears in chat output. Stop only on the base EOS and the model runs past its answer and writes the next user turn itself.

open as a page

When should you merge a LoRA adapter into Llama's base weights?

level: middleimportance: should knowfreq 38%

basics

~20 s

Merge when one adapter will serve all traffic and you want zero adapter overhead — merging folds the low-rank update into the weights, producing a plain checkpoint. Keep it unmerged when you need to hot-swap several adapters over one shared base.

open as a page

What attribution and naming rules apply when redistributing a fine-tuned Llama model?

level: middleimportance: should knowfreq 45%

basics

~20 s

Redistribution under the Llama Community License requires shipping a copy of the agreement, a NOTICE file with Meta's copyright text, prominent "Built with Llama" attribution on any product using it, and a derivative model name that begins with "Llama".

open as a page

What does bitsandbytes NF4 loading give up versus a prequantized GPTQ Llama?

level: middleimportance: should knowfreq 45%

basics

~20 s

bitsandbytes quantizes weights to NF4 in memory as the checkpoint loads, with no calibration data and no separate build artifact. You give up serving throughput — its dequantize-and-matmul path is generally slower than tuned GPTQ or AWQ kernels — and you re-pay the quantization cost on every load.

open as a page

How do GPTQ and AWQ differ when quantizing Llama weights to 4 bits?

level: middleimportance: should knowfreq 55%

basics

~20 s

Both are calibration-based post-training weight quantizers. GPTQ quantizes column by column and adjusts the not-yet-quantized weights to cancel the error it just introduced. AWQ instead measures activations, finds the few salient weight channels, and rescales them so rounding hurts them less.

open as a page

What problem does PagedAttention solve for the KV cache in vLLM?

level: middleimportance: should knowfreq 50%

basics

~20 s

It stores each sequence's key/value cache in small fixed-size blocks tracked by a block table instead of one contiguous reservation sized for the maximum output length. That removes the huge reserved-but-unused waste, so far more sequences fit in the same VRAM.

open as a page

Why can a Llama prompt end up with two BOS tokens, and how do you avoid it?

level: seniorimportance: should knowfreq 38%

basics

~10 s

The rendered chat template already contains the model's beginning-of-sequence token. Tokenizing that string again with default settings prepends a second one. Pass add_special_tokens=False when encoding pre-rendered text, or let the template tokenize directly.

open as a page

When aligning a fine-tuned Llama, how does DPO differ from PPO-based RLHF?

level: seniorimportance: should knowfreq 44%

basics

~20 s

DPO trains directly on chosen/rejected preference pairs with a classification-style loss against a frozen reference model. PPO-based RLHF first trains a separate reward model, then optimizes the policy by sampling completions and scoring them — far more moving parts.

open as a page

Why does full fine-tuning of an 8B Llama need far more VRAM than its weights?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Training stores far more than weights: a gradient per parameter, two AdamW moment tensors, usually an fp32 master copy, plus activations. That is roughly 16 bytes per parameter, so an 8B model needs well over 100 GB before activation memory is counted.

open as a page

What does Meta's Llama Acceptable Use Policy restrict, and what happens if you breach it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The Acceptable Use Policy is incorporated into the Llama Community License and bars categories such as illegal activity, weapons development, malware, critical-infrastructure interference, exploitation of minors, and deceptive impersonation. Breach is a licence breach: Meta may terminate, after which you must stop using and delete the materials.

open as a page

How did the Llama 3.1 licence change the rules on training other models with Llama outputs?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Llama 2 and Llama 3 forbade using the models' outputs to improve any other large language model. The Llama 3.1 Community License dropped that prohibition and expressly permits using outputs to train other models, provided any resulting distributed model's name begins with "Llama".

open as a page

How does Llama 3.1 reach a 128K context when Llama 3 was trained at 8K?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Llama 3.1 rescales the rotary position embeddings so positions past the original 8,192 fall back into a range the model has seen, then continues pretraining on progressively longer sequences up to 131,072 tokens. The config records this as a rope_scaling block with rope_type llama3 and factor 8.

open as a page

After quantizing Llama to 4-bit, how do you verify the accuracy loss is acceptable?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Compare the quantized build against the same unquantized weights on your own task, not just on perplexity. Perplexity averages over tokens and hides narrow failures, so add task evals for structured output, tool-call validity, code and long context, plus a side-by-side sample review.

open as a page

How do you size GPU VRAM to serve Llama 3.1 8B at long context?

level: seniorimportance: should knowfreq 56%

basics

~20 s

Budget three things: weights (about 16 GB for 8B at fp16), KV cache, and activation/framework overhead. Llama 3.1 8B costs roughly 128 KiB of fp16 KV per token, so a single 128k-token context needs about 16 GB of cache on its own.

open as a page

How do you choose between Llama 8B, 70B and 405B for a production workload?

level: principalimportance: should knowfreq 50%

basics

~20 s

Start from a task evaluation, not a leaderboard. Pick the smallest size that clears your quality bar, then check the footprint: at bf16 the weights alone run roughly 16 GB at 8B, 140 GB at 70B and over 800 GB at 405B. Most teams land on 70B.

open as a page

showing 1–30 of 34