Large Language Models
What is actually inside the model you call: transformer layers and attention, tokenization, the pretraining-then-alignment pipeline, decoding at inference, and the context window and knowledge cutoff that bound it. Interviews use this layer to check that your application knowledge sits on real mechanics.
on this pageshowhide
explore
- Transformer Architecture20 questions
- Self-Attention Mechanism5 questions
- Attention and KV Compression5 questions
- Positional Encoding5 questions
- Mixture of Experts5 questions
- Tokenization11 questions
- Subword Algorithms6 questions
- Token Counting and Practical Effects5 questions
- Pretraining and Alignment12 questions
- Pretraining Objectives and Data6 questions
- Preference and RL Post-Training6 questions
- Inference and Sampling15 questions
- Decoding and Sampling Controls6 questions
- KV Cache, Prefill, Decode5 questions
- Latency, Throughput and Batching4 questions
- Context Windows9 questions
- Long-Context Behavior5 questions
- Context Extension Techniques4 questions
- Capabilities and Limits13 questions
- Hallucination and Grounding5 questions
- In-Context Learning and Emergence4 questions
- Thinking and Test-Time Compute4 questions
- Quantization Theory15 questions
- Numeric Formats, Memory Math5 questions
- Quantization Methods5 questions
- Quality and Serving Trade-offs5 questions
- Scaling Laws and Sizing5 questions
questions
100 · 8 sectionsWhy does a transformer need positional encoding, given that attention sees every token?
basics
~20 sSelf-attention treats its input as a set — reorder the tokens and each one gets the same representation back, just moved. Order has to be injected separately, or "the auditor approved the invoice" and "the invoice approved the auditor" are indistinguishable.
In transformer self-attention, what roles do the query, key and value vectors play?
basics
~20 sEach token is projected into three vectors: a query saying what it is looking for, a key advertising what it offers, and a value carrying the content it contributes. Query-key dot products score matches, and those scores weight a sum of values.
Does FlashAttention change attention's output, and where does its speedup come from?
basics
~20 sNo — FlashAttention computes exactly the same attention, bit-for-bit equivalent up to floating-point reordering. It is faster because it tiles the computation in fast on-chip memory and never writes the full score matrix out to the GPU's main memory, so it moves far less data.
In transformer LLMs, how do GQA and MQA differ from multi-head attention?
basics
~20 sMulti-head attention gives every query head its own key and value projections. Multi-query attention makes all query heads share a single key/value head. Grouped-query attention sits between them: query heads are split into groups, and each group shares one key/value head.
An MoE model has 670B total and 37B active parameters — which number sizes GPU memory?
basics
~20 sTotal parameters size memory. Every expert's weights must be resident because any token may route to any expert. Active parameters size compute — the arithmetic and latency per token. A sparse model is therefore memory-hungry like a huge model but computes like a small one.
How do you count the tokens an LLM prompt will use before sending it?
basics
~20 sRun the target model's own tokenizer over the fully assembled request, or call a provider token-counting endpoint when no offline tokenizer exists. Character heuristics such as roughly four characters per token are planning guardrails, not exact counts.
Why do LLMs tokenize text into subwords instead of whole words or characters?
basics
~20 sSubword tokenization keeps a fixed vocabulary that can still encode any input. Word-level vocabularies fail on words they never saw; character-level ones make sequences several times longer. Subword pieces keep common words whole and split rare ones.
Why do LLMs miscount letters and judge 9.11 as larger than 9.9?
basics
~10 sThe model never sees characters. Text arrives as subword token IDs, so spelling is a memorized association rather than something readable, and long numbers split into chunks whose boundaries differ between similar-looking values.
How does byte-pair encoding learn its merge table when a tokenizer is trained?
basics
~20 sBPE starts with every word split into single symbols, counts every adjacent symbol pair across the corpus, merges the single most frequent pair into a new token, and repeats until the vocabulary target is reached. The ordered list of merges is the tokenizer.
Which hidden tokens does an LLM chat request add beyond your message text?
basics
~20 sChat models are fed one rendered sequence, not a list of strings. A template wraps every message in role and turn-delimiter tokens and prepends the system prompt and tool schemas, all billed as input and invisible to a plain character count.
Why does a base LLM checkpoint continue your prompt instead of answering it?
basics
~20 sA base checkpoint is trained only to continue documents, so a question is most plausibly followed by more document text rather than an answer. Answering, and stopping when finished, come from the post-training stages that produce an instruct checkpoint.
What is an LLM's knowledge cutoff, and why is it fuzzy rather than a hard date?
basics
~20 sA knowledge cutoff is the point after which no training text was collected. It is fuzzy because crawls trail real events, coverage of recent months is thin, later training stages can add newer data, and the model has no reliable sense of its own horizon.
What does next-token cross-entropy actually optimise during LLM pretraining?
basics
~20 sNext-token cross-entropy maximises the probability the model assigns to each real next token in the corpus. It optimises fit to the data distribution, including that data's errors and style, never truthfulness, helpfulness or task success.
When would you choose offline DPO over an online RL run like GRPO?
basics
~20 sChoose DPO when good preference pairs already exist and the goal is bounded - tone, format, a known bad habit. It needs no sampling loop, no reward model and no rollout infrastructure. Online RL wins when the reward is verifiable or the model must be pushed well past its current behaviour.
Why does GRPO drop the value critic that PPO-style RLHF requires?
basics
~20 sPPO-style RLHF trains a second network to predict expected return as a baseline. GRPO samples a group of responses per prompt instead and uses the group's own mean reward as that baseline, removing a whole model from memory and from the failure surface.
What does temperature do to an LLM's next-token distribution during sampling?
basics
~20 sTemperature divides the model's raw scores (logits) before they are turned into probabilities. Below 1 it sharpens the distribution toward the top-scoring tokens; above 1 it flattens it, so rarer tokens get picked. At 0 it collapses to always taking the highest-scoring token.
In LLM serving, what do time-to-first-token and inter-token latency each measure?
basics
~20 sTime-to-first-token is the wait from sending a request until the first output token arrives, and it grows with prompt length. Inter-token latency is the gap between successive tokens after that, and it sets how fast the answer streams.
How do top-k, top-p and min-p differ as token truncation strategies?
basics
~20 sTop-k keeps a fixed number of highest-probability tokens. Top-p (nucleus) keeps the smallest set whose probabilities sum past a threshold, so the set size shrinks when the model is confident. Min-p keeps every token above a fraction of the top token's probability.
Why can an LLM cache keys and values across decoding steps but not queries?
basics
~20 sUnder a causal mask, a past token's key and value vectors never change as new tokens arrive, so every later step reuses them. That token's query was consumed once, to produce its own output, and is never read again.
Why is LLM prefill compute-bound while token-by-token decode is bandwidth-bound?
basics
~20 sPrefill pushes every prompt token through the model in one pass, so each weight read is amortized over many tokens and the math units saturate. Decode produces one token per pass, re-reading all weights and the cache for that single token, so memory bandwidth is the ceiling.
How do you extend a shipped 8K-context LLM to 128K without pretraining from scratch?
basics
~20 sRescale the model's rotary position encodings so 128K positions fall inside the range it was trained on — uniform position interpolation, or frequency-aware variants such as NTK-aware scaling, YaRN and LongRoPE — then continue training briefly on long sequences so the model adapts.
Why does a model's effective context length fall short of its advertised token limit?
basics
~20 sAdvertised length is the maximum input the model accepts without erroring; effective length is how far accuracy actually holds. Attention spreads thinner, long-range dependencies are rare in training data, and more text means more plausible-but-wrong material to attend to.
How does lost-in-the-middle position bias affect where you place key facts in a prompt?
basics
~20 sModels attend most reliably to the beginning and end of a long prompt and least reliably to its middle, producing a U-shaped recall curve. Put instructions and the most critical evidence at the edges, not buried mid-context.
In a sliding-window model, why do attention sinks keep an endless stream coherent?
basics
~20 sAttention distributions dump surplus probability mass onto the first few tokens regardless of their meaning. Evict those tokens as the window slides and the mass redistributes onto real content, destabilising the model. Pinning them keeps generation coherent indefinitely.
Why can extending a model's context window regress its short-prompt quality?
basics
~20 sExtension changes the model everywhere, not only past the old limit. Rescaled positions blunt fine-grained local distance signals, and the continued training on long data shifts the model away from short-prompt behaviour, so short-input accuracy typically drops a few points.
What is hallucination in an LLM, and why doesn't telling it "don't make things up" fix it?
basics
~20 sHallucination is a language model stating false or unsupported claims in the same fluent, confident register as correct ones. A prompt instruction cannot fix it because the model has no internal signal separating what it reliably knows from what it is inventing.
What does in-context learning change in an LLM if no weights are updated?
basics
~20 sIn-context learning changes only the next-token probabilities for that one request. Examples and instructions in the prompt condition the output distribution during the forward pass; the parameters stay frozen, so nothing carries over to the next call.
How do factual and faithfulness hallucinations differ, and why does the distinction change your fix?
basics
~20 sA factual hallucination is wrong about the world; a faithfulness hallucination contradicts or exceeds the source text the model was given, even if the claim happens to be true. The first is fixed by supplying sources, the second by constraining and checking generation against them.
How does a reasoning effort level differ from setting a fixed thinking-token budget?
basics
~20 sAn effort level sets a policy — think shallowly, moderately or deeply — and lets the model decide per request how far to go. A fixed token budget makes you guess a number in advance for a difficulty you cannot see yet.
How did training on verifiable rewards produce models that think before answering?
basics
~20 sTraining rewarded only whether the final answer checked out — an exact string match, a passing test suite, a proof checker. With the outcome graded and the path free, models learned to spend tokens exploring, verifying and backtracking, because that reliably produced correct answers.
How much GPU memory do a 70B model's weights need at bf16, int8 and int4?
basics
~20 sMultiply parameters by bytes per parameter. A 70-billion-parameter dense model needs roughly 140 GB of weights at bf16 (2 bytes), 70 GB at int8 (1 byte) and 35 GB at int4 (half a byte) — before KV cache and activations.
What is the difference between post-training quantization and quantization-aware training?
basics
~20 sPost-training quantization rounds an already-trained checkpoint, using a short calibration pass to pick scales. Quantization-aware training simulates that rounding inside a training or fine-tuning run so the weights adapt to it: far more expensive, and worth it mainly at 4 bits and below.
Why did bf16 displace fp16 as the default 16-bit format for LLMs?
basics
~20 sBoth formats use 16 bits, but bf16 spends more of them on the exponent and fewer on the mantissa. That gives bf16 the same dynamic range as fp32, so large activations and tiny gradients neither overflow nor vanish — at the cost of coarser precision.
Why can perplexity stay flat while a 4-bit quantized model fails your task?
basics
~20 sPerplexity averages next-token loss over a whole corpus, so damage concentrated on a few decisive tokens disappears into the mean. Task accuracy, structured-output validity and per-slice scores expose quantization damage that a flat perplexity curve hides.
In GGUF, what does the K in a k-quant such as Q4_K_M mean?
basics
~20 sThe K marks llama.cpp's k-quant family, where weights sit in super-blocks whose per-block scales and minimums are themselves stored in low precision. The trailing S or M says how many of the most sensitive tensors are promoted to a higher-bit k-quant; the 4 is the base width.
What did the Chinchilla result change about splitting a fixed LLM training budget?
basics
~20 sChinchilla showed that for a fixed training-compute budget, parameters and training tokens should grow together — roughly 20 tokens per parameter — rather than spending almost the whole budget on more parameters, as earlier Kaplan-era guidance implied.
Why deliberately train a 3B LLM far past its compute-optimal token count?
basics
~20 sBecause inference, not training, is the bill. Over-training a small model costs more once, but every request afterwards runs on fewer parameters — cheaper, faster and denser to batch across hundreds of millions of monthly calls.
What is the data wall in LLM pretraining, and what options remain when it binds?
basics
~20 sThe data wall is the point where high-quality text runs out before compute does, so extra FLOPs no longer buy extra fresh tokens. The remaining levers are repeating data across epochs, generating synthetic data, and pulling in other modalities.
Why can a distilled 3B student beat a 3B model trained from scratch?
basics
~20 sThe student learns from a strong teacher's full output distribution and worked outputs rather than from raw one-hot next tokens, so each training example carries far more signal. Part of the teacher's compute is effectively transferred into the smaller model.
How would you split extra compute across pretraining, RL post-training and test time?
basics
~20 sTreat them as three curves with different economics. Pretraining raises the base capability, reinforcement learning after it sharpens behaviour on checkable tasks, and test-time compute buys accuracy on the hard request — but you pay that one on every call, forever.