skip to content

AI & LLM Methods

Everything above the model itself: how you prompt it, ground it in your own data, wire it into tools and agents, and fine-tune it when prompting runs out. This is the applied layer interviews reach for the moment you say you have worked with LLMs — the methods, not the maths.

on this pageshow

explore

→ has its own guide

questions

740 · 16 sections

In an AI agent, what are working, episodic, semantic and procedural memory?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Working memory holds what the agent is using right now for the current task. Episodic memory records what happened in past sessions. Semantic memory stores durable facts. Procedural memory holds learned how-to — reusable skills the agent can apply again.

open as a page

In a ReAct agent loop, what do the thought, action and observation steps each do?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Thought is the model's reasoning about what to do next, action is the tool call it emits, and observation is the real tool result appended back into the prompt. The loop repeats until the model answers instead of acting.

open as a page

Why do agents write the plan to an external todo file instead of keeping it in context?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A written plan survives what the conversation does not. Context gets truncated, summarized or crowded out on long runs, so an external plan file keeps the goal and per-step status stable and re-readable, and doubles as an audit trail of what the agent actually did.

open as a page

In LLM function calling, what does the model emit and what turn must you append?

level: juniorimportance: must knowfreq 82%
basics
~20 s

Instead of prose the model returns a structured tool-call block naming the tool, its arguments and a call id. Your code executes the tool, appends a tool-result turn carrying that same id, and calls the model again with the extended history.

open as a page

In a function-calling tool definition, which parts does the model actually see?

level: juniorimportance: must knowfreq 72%
basics
~20 s

The model sees only three things: the tool's name, its natural-language description, and the JSON Schema describing its parameters. Implementation code, docstrings and internal comments never reach it, so everything it needs must live in those three fields.

open as a page

Why must few-shot examples use the same field names and delimiters as the live query?

level: juniorimportance: must knowfreq 65%
basics
~20 s

Demonstrations teach by pattern continuation. If the examples use one separator and field set but the live query arrives in a different shape, the query stops reading as the next item in the series, and the model's output format drifts or it keeps writing examples.

open as a page

What is dynamic exemplar retrieval in few-shot prompting, and how does it differ from a fixed example block?

level: juniorimportance: must knowfreq 55%
basics
~20 s

Dynamic exemplar retrieval picks the few-shot demonstrations per request, pulling the most similar labelled examples from a pool by vector similarity. A fixed block hardcodes the same demonstrations into the prompt for every request, whatever the input looks like.

open as a page

When picking few-shot exemplars, why must every label class appear?

level: juniorimportance: must knowfreq 58%
basics
~20 s

Demonstrations define the label space the model treats as live. A class with no example is predicted far less often than it should be, even when the instruction names it, so coverage of every label comes before adding more examples.

open as a page

What is prompt chaining, and why split a task across several model calls?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Prompt chaining runs one task as an ordered series of model calls, each stage's output feeding the next stage's prompt. Every stage gets a narrow instruction and a checkable result, so a single step can be validated, retried, or replaced.

open as a page

When do you escalate from a direct answer to chain-of-thought or self-consistency?

level: middleimportance: must knowfreq 72%
basics
~20 s

Escalate only when the cheaper rung measurably fails. Direct answers suit lookup and formatting; chain-of-thought pays on multi-step reasoning; sampling several chains and voting pays only when the answer is a short discrete value worth the multiplied cost.

open as a page

What is hallucination in an LLM, and why doesn't telling it "don't make things up" fix it?

level: juniorimportance: must knowfreq 85%
basics
~20 s

Hallucination is a language model stating false or unsupported claims in the same fluent, confident register as correct ones. A prompt instruction cannot fix it because the model has no internal signal separating what it reliably knows from what it is inventing.

open as a page

What does in-context learning change in an LLM if no weights are updated?

level: juniorimportance: must knowfreq 76%
basics
~20 s

In-context learning changes only the next-token probabilities for that one request. Examples and instructions in the prompt condition the output distribution during the forward pass; the parameters stay frozen, so nothing carries over to the next call.

open as a page

What does temperature do to an LLM's next-token distribution during sampling?

level: juniorimportance: must knowfreq 82%
basics
~20 s

Temperature divides the model's raw scores (logits) before they are turned into probabilities. Below 1 it sharpens the distribution toward the top-scoring tokens; above 1 it flattens it, so rarer tokens get picked. At 0 it collapses to always taking the highest-scoring token.

open as a page

In LLM serving, what do time-to-first-token and inter-token latency each measure?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Time-to-first-token is the wait from sending a request until the first output token arrives, and it grows with prompt length. Inter-token latency is the gap between successive tokens after that, and it sets how fast the answer streams.

open as a page

Why does a base LLM checkpoint continue your prompt instead of answering it?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A base checkpoint is trained only to continue documents, so a question is most plausibly followed by more document text rather than an answer. Answering, and stopping when finished, come from the post-training stages that produce an instruct checkpoint.

open as a page

What does a 429 from an LLM provider mean, and why is an immediate retry wrong?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A 429 means you crossed the provider's rate limit, usually a per-minute cap on requests or on tokens. Retrying instantly adds load to an already-throttled account; wait, honour any Retry-After, and back off exponentially with random jitter.

open as a page

Why stream an LLM response token by token instead of returning it all at once?

level: juniorimportance: must knowfreq 72%
basics
~10 s

Streaming does not make generation faster. It makes the wait visible: the user starts reading the first sentence while the rest is still being produced, so an 18-second answer feels responsive instead of frozen.

open as a page

What is the difference between a public LLM benchmark and a task eval?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A public benchmark is a shared, fixed dataset that ranks models against each other on a general capability. A task eval runs your own inputs through your own system and scores your own success criterion. Only the task eval predicts what your users will see.

open as a page

What is LLM-as-judge evaluation, and why is a judge score not ground truth?

level: juniorimportance: must knowfreq 70%
basics
~20 s

LLM-as-judge means prompting a model with a rubric to score another model's output. The score is a noisy estimate, not truth: the judge has its own biases and blind spots, so it must be validated against human labels before anyone trusts the number.

open as a page

What is the difference between offline and online evaluation of an LLM feature?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Offline evaluation scores a fixed, frozen set of saved examples in a harness before release, so it is repeatable and cheap. Online evaluation measures real user traffic after release, where actual behaviour and business outcomes decide whether the change helped.

open as a page

What does prompt chaining between two LLM agents cost in tokens and fidelity?

level: juniorimportance: must knowfreq 60%
basics
~20 s

Prompt chaining pastes one agent's finished output into the next agent's prompt. It is the cheapest wiring to build, but the receiver pays input tokens for every word on every turn and inherits any error, hedge or ambiguity verbatim.

open as a page

What defines an agent's role in a multi-agent system beyond its persona prompt?

level: middleimportance: must knowfreq 62%
basics
~20 s

A role is three things: a persona stating the job, a tool allowlist bounding what the agent can actually do, and behavioural constraints covering when it stops and what it returns. The allowlist does most of the real work.

open as a page

Why send schema-validated payloads between agents instead of free-text messages?

level: middleimportance: must knowfreq 62%
basics
~20 s

A schema turns an implicit agreement into a checkable one. Required fields, types and enums let the receiving agent reject a malformed message at the boundary with a precise error, instead of a model quietly inventing the missing value and shipping a wrong result downstream.

open as a page

Why do multi-agent systems use single-writer discipline instead of locking concurrent writers?

level: middleimportance: must knowfreq 68%
basics
~20 s

Locking assumes writers that wait, know their own blast radius, and roll back cleanly. LLM agents do none of that. Routing every write through one owner, while the others contribute proposals, removes the conflict rather than arbitrating it.

open as a page

In a multi-agent pipeline, why can every agent score high yet end-to-end success be low?

level: middleimportance: must knowfreq 58%
basics
~20 s

Stage accuracies compound: four agents at 90% each leave roughly 66% end-to-end. Errors also propagate, because downstream agents trust upstream output, and per-agent scores measured on clean inputs never see the messy handoffs real predecessors emit.

open as a page

How does recursive text splitting choose cut points, unlike a fixed-size window?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Recursive splitting walks an ordered list of separators — paragraph break, line break, sentence end, space, then bare characters — and cuts on the highest-priority one that fits. A fixed-size window ignores the text and cuts when a counter runs out.

open as a page

Why split Markdown docs on their heading hierarchy instead of a fixed character count?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Headings mark where one topic ends and the next begins, so splitting on them yields chunks that each cover one subject. Fixed-length cuts merge unrelated sections and throw away the heading path that says what the chunk is about.

open as a page

In a RAG prompt, why tag each retrieved chunk with a source ID?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Source IDs give the model a fixed vocabulary of references it can cite and give the reader a way back to the evidence. Without IDs injected alongside the text, any citation the model produces is invented rather than retrieved.

open as a page

Why does RAG retrieval use a bi-encoder rather than scoring every document with a cross-encoder?

level: juniorimportance: must knowfreq 75%
basics
~20 s

A bi-encoder embeds every document once, offline, so a query is answered by a nearest-neighbour lookup over precomputed vectors. A cross-encoder must run the model on each query-document pair, so scoring a whole corpus per query is infeasible.

open as a page

In RAG evaluation, how does faithfulness differ from answer relevance?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Faithfulness asks whether every claim in the answer is supported by the retrieved context. Answer relevance asks whether the answer addresses the question that was asked. They are independent axes, so a fully grounded answer can still be off-topic.

open as a page

How many training examples does a narrow supervised fine-tune actually need?

level: juniorimportance: must knowfreq 60%
basics
~20 s

A few hundred to a few thousand consistent examples usually move a narrow, well-defined task; broad behaviour change needs tens of thousands. Coverage of the real input distribution and consistent labelling matter far more than raw row count.

open as a page

When is fine-tuning the right call instead of a better prompt or retrieval?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Fine-tune only after prompting and retrieval have been tried and still miss. The cases that justify it are behavioural: an output format or house style hard to describe in words, or a stable high-volume task where per-call cost is the binding constraint.

open as a page

How do you decontaminate fine-tuning training data against the evaluation set?

level: middleimportance: must knowfreq 48%
basics
~20 s

Decontamination removes training rows that overlap the evaluation cases. Normalise the text, drop any training row sharing a long n-gram with an eval item, then catch paraphrases with embedding similarity — and split before generating, so synthetic rows never straddle the boundary.

open as a page

What baselines must a fine-tuned model beat before you ship it?

level: middleimportance: must knowfreq 72%
basics
~20 s

At minimum the same base model prompted properly - a strong system prompt and a few-shot variant - plus whatever runs in production today. Score every arm on one held-out set with identical decoding settings, or the improvement is unattributable.

open as a page

Why is held-out loss a poor yardstick for whether a fine-tune helped?

level: middleimportance: must knowfreq 64%
basics
~20 s

Held-out loss scores the token-level likelihood of one reference wording, so it penalises correct answers phrased differently and rewards imitating training style. It tells you the run is healthy, not that the task got better - a task metric or human preference decides that.

open as a page

When a coding agent compacts its conversation, what actually happens to the session?

level: juniorimportance: must knowfreq 62%
basics
~10 s

Compaction replaces the accumulated history with a written summary plus the most recent turns, then continues the same task in a nearly empty window. Work carries on, but the raw earlier messages are gone.

open as a page

Why do agents load context through tools on demand instead of preloading it?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Preloading spends the attention budget on material the model mostly will not use, and quality degrades as windows fill. Just-in-time loading carries lightweight identifiers — paths, IDs, queries — and pulls full content through a tool only when a step actually needs it.

open as a page

Why does an LLM assistant forget facts between sessions, and how do you fix it?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Everything the model can use in a turn is text in its context window, and that window is rebuilt from scratch for each new session. A fact survives only if the application writes it to a store outside the window and puts it back in later.

open as a page

In an LLM prompt containing a long document, where should the instruction go?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Put the instruction after the document, or at both ends. Language models attend most reliably to the beginning and the end of a prompt, so an instruction stated once before a long body is the easiest part to miss.

open as a page

Why reserve output tokens when budgeting an LLM context window?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A context window is shared by the prompt and everything the model generates. Fill it with input and the answer has nowhere to go: responses get truncated mid-sentence or the request is rejected. Reserve output space first.

open as a page

In LLM prompting, what is chain-of-thought and how does it change the output?

level: juniorimportance: must knowfreq 82%
basics
~20 s

Chain-of-thought means having the model write out intermediate reasoning steps before its final answer instead of answering immediately. The output becomes a short worked solution followed by the result, which raises accuracy on problems needing several dependent steps.

open as a page

Does a confident chain-of-thought trace mean its intermediate steps are true?

level: juniorimportance: must knowfreq 60%
basics
~20 s

No. Every step is predicted text, so a chain can contain invented facts — a factor pair that does not multiply out, a contract subsection that does not exist — while reading as careful and rigorous.

open as a page

What are thinking tokens in an extended-thinking LLM, and are you billed for them?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Thinking tokens are reasoning a model generates before its final answer, on a separate channel. They are generated tokens like any others, so you pay for them and wait for them even when the API returns only a summary, or nothing at all.

open as a page

In self-consistency decoding, how is one final answer chosen from many sampled chains?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Self-consistency throws away the reasoning text and votes only on final answers: each sampled chain contributes its extracted answer, identical answers are grouped into one bucket, and the biggest bucket wins. That is a plurality, not a required majority.

open as a page

In prompting, how does zero-shot CoT differ from few-shot CoT?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Zero-shot CoT adds only an instruction to reason step by step, and the model invents its own reasoning shape. Few-shot CoT puts solved examples in the prompt that demonstrate both how to reason and how the answer should look.

open as a page

In a ReAct agent, what happens in one pass of the reason-act-observe loop?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Each pass has three parts: the model writes a short reasoning step about what it needs next, emits one action against a tool, and gets the tool's result back as an observation. That observation joins the running context, and the next pass begins.

open as a page

In ReAct, what is an observation and who is allowed to write it?

level: juniorimportance: must knowfreq 58%
basics
~20 s

An observation is the tool's actual output, appended to the transcript by the runtime after an action. The model must never produce it: generation stops at the action, the real result is inserted, and only then does the model continue.

open as a page

In a ReAct loop, what does the runtime do with an action before the tool runs?

level: juniorimportance: must knowfreq 65%
basics
~20 s

The runtime resolves the action name against its registered tool catalogue, deserializes the arguments and validates them against that tool's schema, then executes. Anything that fails resolution or validation never reaches the tool — an error is returned instead.

open as a page

What does a Reflexion-style agent carry into its next attempt after a failed one?

level: middleimportance: must knowfreq 62%
basics
~20 s

A short self-written note in plain language about why the attempt failed and what to do differently. That note is appended to the next attempt's context, so the retry starts from an explicit lesson instead of repeating the same action.

open as a page

What ends a ReAct loop, and which stop conditions can the model itself decide?

level: middleimportance: must knowfreq 62%
basics
~20 s

The model ends the loop by producing a final answer instead of an action. Because that signal is only the model's own judgement, the surrounding harness adds external stops it cannot override: a cap on passes, a cost or time limit, and a goal check on the result.

open as a page

In Tree of Thought, how do you design the propose and evaluate prompt templates?

level: middleimportance: must knowfreq 62%
basics
~20 s

Write two separate templates with different jobs: a proposer that receives the serialized path and returns candidate next steps as free text, and an evaluator that receives one candidate state and returns a strictly parseable verdict. Keep their shared prefix byte-identical.

open as a page

When would you run a Tree of Thoughts search depth-first with backtracking rather than breadth-first with a beam?

level: middleimportance: must knowfreq 70%
basics
~20 s

Go depth-first with backtracking when a partial solution can be checked for contradiction, most branches die early, and any complete answer is worth having quickly. Use breadth-first with a beam when the depth is short and known and you want several equal-length plans compared side by side.

open as a page

In Tree of Thoughts, how do branching factor, depth and beam width set the LLM call count?

level: middleimportance: must knowfreq 62%
basics
~20 s

With a fixed beam, each level generates width x branching candidates, so generation calls total width x branching x depth — linear in depth. Without a beam the tree grows as branching^depth, which is why unpruned breadth-first search is unaffordable past a few levels.

open as a page

In Tree of Thought, how does value prompting differ from vote prompting?

level: middleimportance: must knowfreq 68%
basics
~20 s

Value prompting scores each partial state on its own and returns an independent number or label. Vote prompting shows several sibling states together and asks which is most promising, returning a relative ranking rather than absolute scores.

open as a page

How do you choose how large a single thought should be in a Tree of Thought?

level: middleimportance: must knowfreq 52%
basics
~20 s

Size a thought so the model can produce several meaningfully different versions of it, and so a partial solution built from it can already be judged promising or hopeless. Too fine and siblings look identical; too coarse and there is almost nothing to branch over.

open as a page

In prompt optimization, what does a meta-prompt that rewrites another prompt contain?

level: juniorimportance: must knowfreq 45%
basics
~20 s

A meta-prompt is a prompt whose subject is another prompt. It carries four things: a specification of the task, the current prompt verbatim, concrete evidence of how that prompt failed, and an instruction to emit a revised prompt in a fixed format.

open as a page

In automatic prompt optimization, where do candidate prompts come from?

level: middleimportance: must knowfreq 48%
basics
~10 s

Three sources dominate: inducing an instruction from labelled input/output pairs, resampling paraphrases of a seed instruction, and mutating slots in a fixed template. They trade diversity against staying on-task, and most systems combine them.

open as a page

Why is automatic prompt optimization run as a discrete, gradient-free search?

level: middleimportance: must knowfreq 62%
basics
~20 s

A prompt is discrete text, not a continuous parameter vector, so no gradient points toward a better prompt. Optimization instead runs as black-box search: propose edited candidates, score each on a dataset, keep the winners, repeat under a fixed budget.

open as a page

In DSPy, what does a signature declare, and why isn't it just a prompt string?

level: middleimportance: must knowfreq 46%
basics
~20 s

A DSPy signature declares one step's named inputs and outputs plus a short description of the task — for example context, question -> answer. The wording, formatting rules and examples that make up the actual prompt are generated by the framework, not written by you.

open as a page

In automatic prompt optimization, how do you pick the scoring metric for a task?

level: middleimportance: must knowfreq 70%
basics
~20 s

Match the scorer to the output shape: exact match or F1 when one label is right, execution-based checks when the output can be run and verified, overlap scores only as a rough proxy, and a judge model only for genuinely open-ended text.

open as a page

What does a prompt-cache hit change about an LLM request's cost and speed?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A hit bills the reused prefix at a fraction of the normal input price — often around a tenth — and the model skips re-processing it, so the first token arrives far sooner. Output tokens are billed and generated exactly as usual.

open as a page

In prompt caching, how many cache reads repay one cache write?

level: middleimportance: must knowfreq 74%
basics
~20 s

Break-even reads equal the write premium above base divided by the discount below base: with a 1.25x write and a 0.1x read, a single reuse already pays; at a 2x write it takes about two. The real constraint is whether that reuse arrives before the entry expires.

open as a page

Why does a timestamp at the top of a system prompt destroy cache hit rate?

level: middleimportance: must knowfreq 74%
basics
~20 s

Prompt caches match an exact prefix from the first token onward. A line like "Current time: 14:32:07" differs on every request, so the match fails at the very first tokens and the entire prompt behind it is reprocessed and re-stored, forever.

open as a page

Why does editing one word early in a long system prompt void the whole cache?

level: middleimportance: must knowfreq 78%
basics
~20 s

Attention is causal: every later token's cached key and value tensors were computed from all the tokens before it. Change a token near the top and the stored state for everything after it is stale, so the server must rebuild the entire prompt.

open as a page

In prompt caching, what is physically stored for a cached prefix?

level: middleimportance: must knowfreq 72%
basics
~20 s

The per-layer attention key and value tensors for every token in the prefix — not the prompt text, not an embedding, and not the model's answer. A hit reloads those tensors and skips recomputing them.

open as a page

What is the difference between a text embedding model and a generative LLM?

level: juniorimportance: must knowfreq 70%
basics
~20 s

An embedding model maps a piece of text to one fixed-length vector of numbers that can be compared with other vectors. A generative LLM produces new text token by token. Embeddings are for comparing meaning, not for writing.

open as a page

Which parts of a search request should never be answered by vector similarity?

level: juniorimportance: must knowfreq 55%
basics
~20 s

Anything with a truth condition: numeric and date constraints, sorting, counting, permission scoping, and exact identifier lookups. Vector similarity ranks by how alike two texts are, and "alike" cannot express "price under 800,000" or "only records this user may see".

open as a page

In vector search, what does cosine similarity measure, and when is it preferred over Euclidean distance?

level: juniorimportance: must knowfreq 80%
basics
~20 s

Cosine similarity measures the angle between two vectors and ignores their lengths, ranging from -1 to 1. Prefer it over Euclidean distance when only direction is meaningful and vector length reflects something you do not want scored, such as text length.

open as a page

What can a UMAP or t-SNE plot of embeddings actually tell you?

level: middleimportance: must knowfreq 66%
basics
~20 s

UMAP and t-SNE preserve which points sit near each other locally, so tight visible groups usually reflect real neighbourhoods. The gaps between blobs, the relative blob sizes and the axes carry no reliable meaning and shift with hyperparameters and seed.

open as a page

How do you choose between k-means and DBSCAN for clustering document embeddings?

level: middleimportance: must knowfreq 62%
basics
~20 s

k-means forces every document into one of k clusters you fix in advance, so it fits a corpus you want fully partitioned. DBSCAN groups by density, discovers how many clusters exist, and leaves sparse points unlabelled as noise.

open as a page

Why is pasting a full patient record into a chat assistant a leak even without model training?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Text placed in a prompt becomes data your own system holds: it sits in conversation history and is resent every turn, and it lands in request logs, traces and support tooling. Training use is a separate, narrower question.

open as a page

Why doesn't a strict JSON schema on an LLM response make its content safe?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A schema constrains shape, not meaning. It guarantees the reply parses and carries the required fields and types. It says nothing about whether a field holds an invented price, a competitor's name, an unsafe instruction, or text that is dangerous where you paste it.

open as a page

Why should an autonomous agent hold its own credentials instead of the operator's?

level: middleimportance: must knowfreq 62%
basics
~20 s

An agent running on a human's session inherits every permission that human has, and its actions are logged as theirs. Giving the agent its own principal lets you scope it narrowly, revoke it alone, and attribute what it actually did.

open as a page

In an LLM feature, why moderate both user input and model output?

level: middleimportance: must knowfreq 68%
basics
~20 s

Input screening rejects harmful requests before you pay for a generation and before they enter the context. Output screening catches harm the model produces anyway, including harm that arrived through retrieved documents or tool results the input check never saw.

open as a page

Why do off-the-shelf PII detectors miss hospital MRNs, and what does over-redaction cost?

level: middleimportance: must knowfreq 62%
basics
~20 s

Shipped recognizers cover identifiers standardised nationally or industry-wide; a facility-assigned medical record number has no fixed format, so nothing matches it. Misses leak silently, while blanket masking strips the dosages, dates and lab values the answer depended on.

open as a page

Why does copying text from a scanned PDF return nothing, but not from a born-digital PDF?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A born-digital PDF stores real text objects — character codes with fonts and positions — so extraction just reads them out. A scanned page stores only a photograph of the paper, so no characters exist until OCR creates them.

open as a page

How does a text-to-image diffusion model turn random noise into an image?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Generation starts from pure random noise. The model repeatedly predicts how much noise is present and removes a little of it, conditioned on the text prompt at every step. After a fixed number of steps the noise has been shaped into an image.

open as a page

In a vision-language model, why can one image cost more tokens than a page of text?

level: juniorimportance: must knowfreq 72%
basics
~20 s

The vision encoder cuts the image into a grid of fixed-size patches and emits roughly one token per patch, so token count grows with image area. A 1024x1024 image at patch size 16 is 64x64 = 4096 patches.

open as a page

In document extraction, why demand a bounding box and page number per extracted field?

level: middleimportance: must knowfreq 56%
basics
~20 s

Grounding turns an unverifiable string into an auditable claim. A clerk jumps straight to the pixels behind "container MSKU4412345", and a box landing on blank space exposes an invented value before it reaches the downstream system.

open as a page

In document extraction, when do you pick classic OCR over a document VLM or a frontier model?

level: middleimportance: must knowfreq 65%
basics
~20 s

Classic OCR wins on clean, fixed layouts at high volume: cheap, fast, deterministic, and it returns per-word boxes and confidence scores. Purpose-built document VLMs handle messy layout and tables; frontier models are for open-ended reasoning over the page.

open as a page