skip to content

Large Language Models

What is actually inside the model you call: transformer layers and attention, tokenization, the pretraining-then-alignment pipeline, decoding at inference, and the context window and knowledge cutoff that bound it. Interviews use this layer to check that your application knowledge sits on real mechanics.

on this pageshow

explore

questions

100 · 8 sections

Why does a transformer need positional encoding, given that attention sees every token?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Self-attention treats its input as a set — reorder the tokens and each one gets the same representation back, just moved. Order has to be injected separately, or "the auditor approved the invoice" and "the invoice approved the auditor" are indistinguishable.

open as a page

In transformer self-attention, what roles do the query, key and value vectors play?

level: juniorimportance: must knowfreq 86%
basics
~20 s

Each token is projected into three vectors: a query saying what it is looking for, a key advertising what it offers, and a value carrying the content it contributes. Query-key dot products score matches, and those scores weight a sum of values.

open as a page

Does FlashAttention change attention's output, and where does its speedup come from?

level: middleimportance: must knowfreq 58%
basics
~20 s

No — FlashAttention computes exactly the same attention, bit-for-bit equivalent up to floating-point reordering. It is faster because it tiles the computation in fast on-chip memory and never writes the full score matrix out to the GPU's main memory, so it moves far less data.

open as a page

In transformer LLMs, how do GQA and MQA differ from multi-head attention?

level: middleimportance: must knowfreq 72%
basics
~20 s

Multi-head attention gives every query head its own key and value projections. Multi-query attention makes all query heads share a single key/value head. Grouped-query attention sits between them: query heads are split into groups, and each group shares one key/value head.

open as a page

An MoE model has 670B total and 37B active parameters — which number sizes GPU memory?

level: middleimportance: must knowfreq 66%
basics
~20 s

Total parameters size memory. Every expert's weights must be resident because any token may route to any expert. Active parameters size compute — the arithmetic and latency per token. A sparse model is therefore memory-hungry like a huge model but computes like a small one.

open as a page

How do you count the tokens an LLM prompt will use before sending it?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Run the target model's own tokenizer over the fully assembled request, or call a provider token-counting endpoint when no offline tokenizer exists. Character heuristics such as roughly four characters per token are planning guardrails, not exact counts.

open as a page

Why do LLMs tokenize text into subwords instead of whole words or characters?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Subword tokenization keeps a fixed vocabulary that can still encode any input. Word-level vocabularies fail on words they never saw; character-level ones make sequences several times longer. Subword pieces keep common words whole and split rare ones.

open as a page

Why do LLMs miscount letters and judge 9.11 as larger than 9.9?

level: middleimportance: must knowfreq 58%
basics
~10 s

The model never sees characters. Text arrives as subword token IDs, so spelling is a memorized association rather than something readable, and long numbers split into chunks whose boundaries differ between similar-looking values.

open as a page

How does byte-pair encoding learn its merge table when a tokenizer is trained?

level: middleimportance: must knowfreq 66%
basics
~20 s

BPE starts with every word split into single symbols, counts every adjacent symbol pair across the corpus, merges the single most frequent pair into a new token, and repeats until the vocabulary target is reached. The ordered list of merges is the tokenizer.

open as a page

Which hidden tokens does an LLM chat request add beyond your message text?

level: middleimportance: should knowfreq 42%
basics
~20 s

Chat models are fed one rendered sequence, not a list of strings. A template wraps every message in role and turn-delimiter tokens and prepends the system prompt and tool schemas, all billed as input and invisible to a plain character count.

open as a page

Why does a base LLM checkpoint continue your prompt instead of answering it?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A base checkpoint is trained only to continue documents, so a question is most plausibly followed by more document text rather than an answer. Answering, and stopping when finished, come from the post-training stages that produce an instruct checkpoint.

open as a page

What is an LLM's knowledge cutoff, and why is it fuzzy rather than a hard date?

level: middleimportance: must knowfreq 52%
basics
~20 s

A knowledge cutoff is the point after which no training text was collected. It is fuzzy because crawls trail real events, coverage of recent months is thin, later training stages can add newer data, and the model has no reliable sense of its own horizon.

open as a page

What does next-token cross-entropy actually optimise during LLM pretraining?

level: middleimportance: must knowfreq 72%
basics
~20 s

Next-token cross-entropy maximises the probability the model assigns to each real next token in the corpus. It optimises fit to the data distribution, including that data's errors and style, never truthfulness, helpfulness or task success.

open as a page

When would you choose offline DPO over an online RL run like GRPO?

level: middleimportance: must knowfreq 58%
basics
~20 s

Choose DPO when good preference pairs already exist and the goal is bounded - tone, format, a known bad habit. It needs no sampling loop, no reward model and no rollout infrastructure. Online RL wins when the reward is verifiable or the model must be pushed well past its current behaviour.

open as a page

Why does GRPO drop the value critic that PPO-style RLHF requires?

level: middleimportance: must knowfreq 62%
basics
~20 s

PPO-style RLHF trains a second network to predict expected return as a baseline. GRPO samples a group of responses per prompt instead and uses the group's own mean reward as that baseline, removing a whole model from memory and from the failure surface.

open as a page

What does temperature do to an LLM's next-token distribution during sampling?

level: juniorimportance: must knowfreq 82%
basics
~20 s

Temperature divides the model's raw scores (logits) before they are turned into probabilities. Below 1 it sharpens the distribution toward the top-scoring tokens; above 1 it flattens it, so rarer tokens get picked. At 0 it collapses to always taking the highest-scoring token.

open as a page

In LLM serving, what do time-to-first-token and inter-token latency each measure?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Time-to-first-token is the wait from sending a request until the first output token arrives, and it grows with prompt length. Inter-token latency is the gap between successive tokens after that, and it sets how fast the answer streams.

open as a page

How do top-k, top-p and min-p differ as token truncation strategies?

level: middleimportance: must knowfreq 62%
basics
~20 s

Top-k keeps a fixed number of highest-probability tokens. Top-p (nucleus) keeps the smallest set whose probabilities sum past a threshold, so the set size shrinks when the model is confident. Min-p keeps every token above a fraction of the top token's probability.

open as a page

Why can an LLM cache keys and values across decoding steps but not queries?

level: middleimportance: must knowfreq 62%
basics
~20 s

Under a causal mask, a past token's key and value vectors never change as new tokens arrive, so every later step reuses them. That token's query was consumed once, to produce its own output, and is never read again.

open as a page

Why is LLM prefill compute-bound while token-by-token decode is bandwidth-bound?

level: middleimportance: must knowfreq 70%
basics
~20 s

Prefill pushes every prompt token through the model in one pass, so each weight read is amortized over many tokens and the math units saturate. Decode produces one token per pass, re-reading all weights and the cache for that single token, so memory bandwidth is the ceiling.

open as a page

How do you extend a shipped 8K-context LLM to 128K without pretraining from scratch?

level: middleimportance: must knowfreq 60%
basics
~20 s

Rescale the model's rotary position encodings so 128K positions fall inside the range it was trained on — uniform position interpolation, or frequency-aware variants such as NTK-aware scaling, YaRN and LongRoPE — then continue training briefly on long sequences so the model adapts.

open as a page

Why does a model's effective context length fall short of its advertised token limit?

level: middleimportance: must knowfreq 72%
basics
~20 s

Advertised length is the maximum input the model accepts without erroring; effective length is how far accuracy actually holds. Attention spreads thinner, long-range dependencies are rare in training data, and more text means more plausible-but-wrong material to attend to.

open as a page

How does lost-in-the-middle position bias affect where you place key facts in a prompt?

level: middleimportance: must knowfreq 66%
basics
~20 s

Models attend most reliably to the beginning and end of a long prompt and least reliably to its middle, producing a U-shaped recall curve. Put instructions and the most critical evidence at the edges, not buried mid-context.

open as a page

In a sliding-window model, why do attention sinks keep an endless stream coherent?

level: seniorimportance: should knowfreq 28%
basics
~20 s

Attention distributions dump surplus probability mass onto the first few tokens regardless of their meaning. Evict those tokens as the window slides and the mass redistributes onto real content, destabilising the model. Pinning them keeps generation coherent indefinitely.

open as a page

Why can extending a model's context window regress its short-prompt quality?

level: seniorimportance: should knowfreq 42%
basics
~20 s

Extension changes the model everywhere, not only past the old limit. Rescaled positions blunt fine-grained local distance signals, and the continued training on long data shifts the model away from short-prompt behaviour, so short-input accuracy typically drops a few points.

open as a page

What is hallucination in an LLM, and why doesn't telling it "don't make things up" fix it?

level: juniorimportance: must knowfreq 85%
basics
~20 s

Hallucination is a language model stating false or unsupported claims in the same fluent, confident register as correct ones. A prompt instruction cannot fix it because the model has no internal signal separating what it reliably knows from what it is inventing.

open as a page

What does in-context learning change in an LLM if no weights are updated?

level: juniorimportance: must knowfreq 76%
basics
~20 s

In-context learning changes only the next-token probabilities for that one request. Examples and instructions in the prompt condition the output distribution during the forward pass; the parameters stay frozen, so nothing carries over to the next call.

open as a page

How do factual and faithfulness hallucinations differ, and why does the distinction change your fix?

level: middleimportance: must knowfreq 68%
basics
~20 s

A factual hallucination is wrong about the world; a faithfulness hallucination contradicts or exceeds the source text the model was given, even if the claim happens to be true. The first is fixed by supplying sources, the second by constraining and checking generation against them.

open as a page

How does a reasoning effort level differ from setting a fixed thinking-token budget?

level: middleimportance: must knowfreq 55%
basics
~20 s

An effort level sets a policy — think shallowly, moderately or deeply — and lets the model decide per request how far to go. A fixed token budget makes you guess a number in advance for a difficulty you cannot see yet.

open as a page

How did training on verifiable rewards produce models that think before answering?

level: middleimportance: must knowfreq 60%
basics
~20 s

Training rewarded only whether the final answer checked out — an exact string match, a passing test suite, a proof checker. With the outcome graded and the path free, models learned to spend tokens exploring, verifying and backtracking, because that reliably produced correct answers.

open as a page

How much GPU memory do a 70B model's weights need at bf16, int8 and int4?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Multiply parameters by bytes per parameter. A 70-billion-parameter dense model needs roughly 140 GB of weights at bf16 (2 bytes), 70 GB at int8 (1 byte) and 35 GB at int4 (half a byte) — before KV cache and activations.

open as a page

What is the difference between post-training quantization and quantization-aware training?

level: middleimportance: must knowfreq 72%
basics
~20 s

Post-training quantization rounds an already-trained checkpoint, using a short calibration pass to pick scales. Quantization-aware training simulates that rounding inside a training or fine-tuning run so the weights adapt to it: far more expensive, and worth it mainly at 4 bits and below.

open as a page

Why did bf16 displace fp16 as the default 16-bit format for LLMs?

level: middleimportance: must knowfreq 62%
basics
~20 s

Both formats use 16 bits, but bf16 spends more of them on the exponent and fewer on the mantissa. That gives bf16 the same dynamic range as fp32, so large activations and tiny gradients neither overflow nor vanish — at the cost of coarser precision.

open as a page

Why can perplexity stay flat while a 4-bit quantized model fails your task?

level: middleimportance: must knowfreq 55%
basics
~20 s

Perplexity averages next-token loss over a whole corpus, so damage concentrated on a few decisive tokens disappears into the mean. Task accuracy, structured-output validity and per-slice scores expose quantization damage that a flat perplexity curve hides.

open as a page

In GGUF, what does the K in a k-quant such as Q4_K_M mean?

level: juniorimportance: should knowfreq 35%
basics
~20 s

The K marks llama.cpp's k-quant family, where weights sit in super-blocks whose per-block scales and minimums are themselves stored in low precision. The trailing S or M says how many of the most sensitive tensors are promoted to a higher-bit k-quant; the 4 is the base width.

open as a page

What did the Chinchilla result change about splitting a fixed LLM training budget?

level: middleimportance: must knowfreq 68%
basics
~20 s

Chinchilla showed that for a fixed training-compute budget, parameters and training tokens should grow together — roughly 20 tokens per parameter — rather than spending almost the whole budget on more parameters, as earlier Kaplan-era guidance implied.

open as a page

Why deliberately train a 3B LLM far past its compute-optimal token count?

level: seniorimportance: must knowfreq 54%
basics
~20 s

Because inference, not training, is the bill. Over-training a small model costs more once, but every request afterwards runs on fewer parameters — cheaper, faster and denser to batch across hundreds of millions of monthly calls.

open as a page

What is the data wall in LLM pretraining, and what options remain when it binds?

level: middleimportance: should knowfreq 44%
basics
~20 s

The data wall is the point where high-quality text runs out before compute does, so extra FLOPs no longer buy extra fresh tokens. The remaining levers are repeating data across epochs, generating synthetic data, and pulling in other modalities.

open as a page

Why can a distilled 3B student beat a 3B model trained from scratch?

level: middleimportance: should knowfreq 38%
basics
~20 s

The student learns from a strong teacher's full output distribution and worked outputs rather than from raw one-hot next tokens, so each training example carries far more signal. Part of the teacher's compute is effectively transferred into the smaller model.

open as a page

How would you split extra compute across pretraining, RL post-training and test time?

level: principalimportance: should knowfreq 40%
basics
~20 s

Treat them as three curves with different economics. Pretraining raises the base capability, reinforcement learning after it sharpens behaviour on checkable tasks, and test-time compute buys accuracy on the hard request — but you pay that one on every call, forever.

open as a page