AI & LLM Methods
Everything above the model itself: how you prompt it, ground it in your own data, wire it into tools and agents, and fine-tune it when prompting runs out. This is the applied layer interviews reach for the moment you say you have worked with LLMs — the methods, not the maths.
on this pageshowhide
explore
- AI Agents (has its own guide)81 questions
- Tool Use and Function Calling20 questions
- Planning and Reasoning14 questions
- Agent Memory15 questions
- Agent Loops and Control Flow13 questions
- Multi-Agent Orchestration5 questions
- Agent Evaluation14 questions
- Prompt Engineering (has its own guide)49 questions
- Instruction Design9 questions
- Few-Shot Prompting17 questions
- System and Role Prompts9 questions
- Chain-of-Thought4 questions
- Output Formatting5 questions
- Evaluation and Iteration5 questions
- Large Language Models100 questions
- Transformer Architecture20 questions
- Tokenization11 questions
- Pretraining and Alignment12 questions
- Inference and Sampling15 questions
- Context Windows9 questions
- Capabilities and Limits13 questions
- Quantization Theory15 questions
- Scaling Laws and Sizing5 questions
- AI Engineering97 questions
- LLM App Architecture4 questions
- Vector Databases6 questions
- Evaluation & Testing25 questions
- Observability & Tracing15 questions
- Cost & Latency Optimization14 questions
- Guardrails & Safety19 questions
- Structured Output & Tool-Call Reliability14 questions
- Multi-Agent Systems (has its own guide)41 questions
- Agent Roles and Specialization4 questions
- Orchestration Topologies19 questions
- Inter-Agent Communication5 questions
- Shared State and Context4 questions
- Coordination and Conflict Resolution5 questions
- Multi-Agent Evaluation4 questions
- Retrieval-Augmented Generation81 questions
- Chunking Strategies14 questions
- Embedding Models5 questions
- Vector Retrieval10 questions
- Reranking9 questions
- Context Injection9 questions
- RAG Evaluation10 questions
- Query Transformation15 questions
- RAG Architecture Patterns9 questions
- Fine-tuning30 questions
- Supervised Fine-Tuning5 questions
- PEFT and LoRA6 questions
- Dataset Curation5 questions
- Training Hyperparameters4 questions
- Evaluation and Overfitting5 questions
- Fine-tune vs Prompting5 questions
- Context Engineering26 questions
- Window Budgeting4 questions
- Information Selection5 questions
- Compression & Summarization5 questions
- Ordering & Recency4 questions
- Memory Management4 questions
- Retrieval Integration4 questions
- Chain of Thought33 questions
- Intermediate Reasoning Steps5 questions
- Zero-Shot vs Few-Shot CoT4 questions
- Self-Consistency Decoding10 questions
- Scratchpad and Thinking Tokens4 questions
- Limitations and Failure Modes5 questions
- CoT Evaluation Methods5 questions
- ReAct26 questions
- Reason-Act-Observe Loop5 questions
- Thought Trace Design4 questions
- Tool Action Dispatch4 questions
- Observation Handling4 questions
- Error Recovery4 questions
- Prompt Structure5 questions
- Tree of Thought24 questions
- Thought Generation4 questions
- State Evaluation5 questions
- Search Strategies5 questions
- ToT vs Chain of Thought4 questions
- Implementation Patterns6 questions
- Automatic Prompt Engineering22 questions
- Candidate Prompt Generation4 questions
- Scoring and Evaluation Metrics5 questions
- Optimization Search Loops5 questions
- Meta-Prompting4 questions
- Compiled Prompt Programs4 questions
- Prompt Caching22 questions
- Caching Mechanics5 questions
- TTL and Invalidation4 questions
- Prefix Design4 questions
- Cost and Latency Impact5 questions
- Hit-Rate Optimization4 questions
- Embeddings37 questions
- Embedding Models5 questions
- Similarity Metrics4 questions
- Dimensionality4 questions
- Vector Indexes (ANN)10 questions
- Semantic Search5 questions
- Clustering & Visualization4 questions
- Embedding Space Geometry5 questions
- LLM Safety & Security33 questions
- LLM Threat Modeling5 questions
- Prompt Injection Defense5 questions
- Guardrails & Output Constraining5 questions
- Content Moderation4 questions
- PII & Data-Leakage Handling4 questions
- Safe Tool Use for Agents5 questions
- Safety Evaluation & Red-Team Testing5 questions
- Multimodal AI38 questions
- Vision-Language Models5 questions
- Image Understanding in Practice5 questions
- Document AI5 questions
- Image Generation6 questions
- Speech-to-Text & TTS6 questions
- Video & Audio Understanding6 questions
- Multimodal Embeddings & RAG5 questions
→ has its own guide
questions
740 · 16 sectionsIn an AI agent, what are working, episodic, semantic and procedural memory?
basics
~20 sWorking memory holds what the agent is using right now for the current task. Episodic memory records what happened in past sessions. Semantic memory stores durable facts. Procedural memory holds learned how-to — reusable skills the agent can apply again.
In a ReAct agent loop, what do the thought, action and observation steps each do?
basics
~20 sThought is the model's reasoning about what to do next, action is the tool call it emits, and observation is the real tool result appended back into the prompt. The loop repeats until the model answers instead of acting.
Why do agents write the plan to an external todo file instead of keeping it in context?
basics
~20 sA written plan survives what the conversation does not. Context gets truncated, summarized or crowded out on long runs, so an external plan file keeps the goal and per-step status stable and re-readable, and doubles as an audit trail of what the agent actually did.
In LLM function calling, what does the model emit and what turn must you append?
basics
~20 sInstead of prose the model returns a structured tool-call block naming the tool, its arguments and a call id. Your code executes the tool, appends a tool-result turn carrying that same id, and calls the model again with the extended history.
In a function-calling tool definition, which parts does the model actually see?
basics
~20 sThe model sees only three things: the tool's name, its natural-language description, and the JSON Schema describing its parameters. Implementation code, docstrings and internal comments never reach it, so everything it needs must live in those three fields.
Why must few-shot examples use the same field names and delimiters as the live query?
basics
~20 sDemonstrations teach by pattern continuation. If the examples use one separator and field set but the live query arrives in a different shape, the query stops reading as the next item in the series, and the model's output format drifts or it keeps writing examples.
What is dynamic exemplar retrieval in few-shot prompting, and how does it differ from a fixed example block?
basics
~20 sDynamic exemplar retrieval picks the few-shot demonstrations per request, pulling the most similar labelled examples from a pool by vector similarity. A fixed block hardcodes the same demonstrations into the prompt for every request, whatever the input looks like.
When picking few-shot exemplars, why must every label class appear?
basics
~20 sDemonstrations define the label space the model treats as live. A class with no example is predicted far less often than it should be, even when the instruction names it, so coverage of every label comes before adding more examples.
What is prompt chaining, and why split a task across several model calls?
basics
~20 sPrompt chaining runs one task as an ordered series of model calls, each stage's output feeding the next stage's prompt. Every stage gets a narrow instruction and a checkable result, so a single step can be validated, retried, or replaced.
When do you escalate from a direct answer to chain-of-thought or self-consistency?
basics
~20 sEscalate only when the cheaper rung measurably fails. Direct answers suit lookup and formatting; chain-of-thought pays on multi-step reasoning; sampling several chains and voting pays only when the answer is a short discrete value worth the multiplied cost.
What is hallucination in an LLM, and why doesn't telling it "don't make things up" fix it?
basics
~20 sHallucination is a language model stating false or unsupported claims in the same fluent, confident register as correct ones. A prompt instruction cannot fix it because the model has no internal signal separating what it reliably knows from what it is inventing.
What does in-context learning change in an LLM if no weights are updated?
basics
~20 sIn-context learning changes only the next-token probabilities for that one request. Examples and instructions in the prompt condition the output distribution during the forward pass; the parameters stay frozen, so nothing carries over to the next call.
What does temperature do to an LLM's next-token distribution during sampling?
basics
~20 sTemperature divides the model's raw scores (logits) before they are turned into probabilities. Below 1 it sharpens the distribution toward the top-scoring tokens; above 1 it flattens it, so rarer tokens get picked. At 0 it collapses to always taking the highest-scoring token.
In LLM serving, what do time-to-first-token and inter-token latency each measure?
basics
~20 sTime-to-first-token is the wait from sending a request until the first output token arrives, and it grows with prompt length. Inter-token latency is the gap between successive tokens after that, and it sets how fast the answer streams.
Why does a base LLM checkpoint continue your prompt instead of answering it?
basics
~20 sA base checkpoint is trained only to continue documents, so a question is most plausibly followed by more document text rather than an answer. Answering, and stopping when finished, come from the post-training stages that produce an instruct checkpoint.
What does a 429 from an LLM provider mean, and why is an immediate retry wrong?
basics
~20 sA 429 means you crossed the provider's rate limit, usually a per-minute cap on requests or on tokens. Retrying instantly adds load to an already-throttled account; wait, honour any Retry-After, and back off exponentially with random jitter.
Why stream an LLM response token by token instead of returning it all at once?
basics
~10 sStreaming does not make generation faster. It makes the wait visible: the user starts reading the first sentence while the rest is still being produced, so an 18-second answer feels responsive instead of frozen.
What is the difference between a public LLM benchmark and a task eval?
basics
~20 sA public benchmark is a shared, fixed dataset that ranks models against each other on a general capability. A task eval runs your own inputs through your own system and scores your own success criterion. Only the task eval predicts what your users will see.
What is LLM-as-judge evaluation, and why is a judge score not ground truth?
basics
~20 sLLM-as-judge means prompting a model with a rubric to score another model's output. The score is a noisy estimate, not truth: the judge has its own biases and blind spots, so it must be validated against human labels before anyone trusts the number.
What is the difference between offline and online evaluation of an LLM feature?
basics
~20 sOffline evaluation scores a fixed, frozen set of saved examples in a harness before release, so it is repeatable and cheap. Online evaluation measures real user traffic after release, where actual behaviour and business outcomes decide whether the change helped.
What does prompt chaining between two LLM agents cost in tokens and fidelity?
basics
~20 sPrompt chaining pastes one agent's finished output into the next agent's prompt. It is the cheapest wiring to build, but the receiver pays input tokens for every word on every turn and inherits any error, hedge or ambiguity verbatim.
What defines an agent's role in a multi-agent system beyond its persona prompt?
basics
~20 sA role is three things: a persona stating the job, a tool allowlist bounding what the agent can actually do, and behavioural constraints covering when it stops and what it returns. The allowlist does most of the real work.
Why send schema-validated payloads between agents instead of free-text messages?
basics
~20 sA schema turns an implicit agreement into a checkable one. Required fields, types and enums let the receiving agent reject a malformed message at the boundary with a precise error, instead of a model quietly inventing the missing value and shipping a wrong result downstream.
Why do multi-agent systems use single-writer discipline instead of locking concurrent writers?
basics
~20 sLocking assumes writers that wait, know their own blast radius, and roll back cleanly. LLM agents do none of that. Routing every write through one owner, while the others contribute proposals, removes the conflict rather than arbitrating it.
In a multi-agent pipeline, why can every agent score high yet end-to-end success be low?
basics
~20 sStage accuracies compound: four agents at 90% each leave roughly 66% end-to-end. Errors also propagate, because downstream agents trust upstream output, and per-agent scores measured on clean inputs never see the messy handoffs real predecessors emit.
How does recursive text splitting choose cut points, unlike a fixed-size window?
basics
~20 sRecursive splitting walks an ordered list of separators — paragraph break, line break, sentence end, space, then bare characters — and cuts on the highest-priority one that fits. A fixed-size window ignores the text and cuts when a counter runs out.
Why split Markdown docs on their heading hierarchy instead of a fixed character count?
basics
~20 sHeadings mark where one topic ends and the next begins, so splitting on them yields chunks that each cover one subject. Fixed-length cuts merge unrelated sections and throw away the heading path that says what the chunk is about.
In a RAG prompt, why tag each retrieved chunk with a source ID?
basics
~20 sSource IDs give the model a fixed vocabulary of references it can cite and give the reader a way back to the evidence. Without IDs injected alongside the text, any citation the model produces is invented rather than retrieved.
Why does RAG retrieval use a bi-encoder rather than scoring every document with a cross-encoder?
basics
~20 sA bi-encoder embeds every document once, offline, so a query is answered by a nearest-neighbour lookup over precomputed vectors. A cross-encoder must run the model on each query-document pair, so scoring a whole corpus per query is infeasible.
In RAG evaluation, how does faithfulness differ from answer relevance?
basics
~20 sFaithfulness asks whether every claim in the answer is supported by the retrieved context. Answer relevance asks whether the answer addresses the question that was asked. They are independent axes, so a fully grounded answer can still be off-topic.
How many training examples does a narrow supervised fine-tune actually need?
basics
~20 sA few hundred to a few thousand consistent examples usually move a narrow, well-defined task; broad behaviour change needs tens of thousands. Coverage of the real input distribution and consistent labelling matter far more than raw row count.
When is fine-tuning the right call instead of a better prompt or retrieval?
basics
~20 sFine-tune only after prompting and retrieval have been tried and still miss. The cases that justify it are behavioural: an output format or house style hard to describe in words, or a stable high-volume task where per-call cost is the binding constraint.
How do you decontaminate fine-tuning training data against the evaluation set?
basics
~20 sDecontamination removes training rows that overlap the evaluation cases. Normalise the text, drop any training row sharing a long n-gram with an eval item, then catch paraphrases with embedding similarity — and split before generating, so synthetic rows never straddle the boundary.
What baselines must a fine-tuned model beat before you ship it?
basics
~20 sAt minimum the same base model prompted properly - a strong system prompt and a few-shot variant - plus whatever runs in production today. Score every arm on one held-out set with identical decoding settings, or the improvement is unattributable.
Why is held-out loss a poor yardstick for whether a fine-tune helped?
basics
~20 sHeld-out loss scores the token-level likelihood of one reference wording, so it penalises correct answers phrased differently and rewards imitating training style. It tells you the run is healthy, not that the task got better - a task metric or human preference decides that.
When a coding agent compacts its conversation, what actually happens to the session?
basics
~10 sCompaction replaces the accumulated history with a written summary plus the most recent turns, then continues the same task in a nearly empty window. Work carries on, but the raw earlier messages are gone.
Why do agents load context through tools on demand instead of preloading it?
basics
~20 sPreloading spends the attention budget on material the model mostly will not use, and quality degrades as windows fill. Just-in-time loading carries lightweight identifiers — paths, IDs, queries — and pulls full content through a tool only when a step actually needs it.
Why does an LLM assistant forget facts between sessions, and how do you fix it?
basics
~20 sEverything the model can use in a turn is text in its context window, and that window is rebuilt from scratch for each new session. A fact survives only if the application writes it to a store outside the window and puts it back in later.
In an LLM prompt containing a long document, where should the instruction go?
basics
~20 sPut the instruction after the document, or at both ends. Language models attend most reliably to the beginning and the end of a prompt, so an instruction stated once before a long body is the easiest part to miss.
Why reserve output tokens when budgeting an LLM context window?
basics
~20 sA context window is shared by the prompt and everything the model generates. Fill it with input and the answer has nowhere to go: responses get truncated mid-sentence or the request is rejected. Reserve output space first.
In LLM prompting, what is chain-of-thought and how does it change the output?
basics
~20 sChain-of-thought means having the model write out intermediate reasoning steps before its final answer instead of answering immediately. The output becomes a short worked solution followed by the result, which raises accuracy on problems needing several dependent steps.
Does a confident chain-of-thought trace mean its intermediate steps are true?
basics
~20 sNo. Every step is predicted text, so a chain can contain invented facts — a factor pair that does not multiply out, a contract subsection that does not exist — while reading as careful and rigorous.
What are thinking tokens in an extended-thinking LLM, and are you billed for them?
basics
~20 sThinking tokens are reasoning a model generates before its final answer, on a separate channel. They are generated tokens like any others, so you pay for them and wait for them even when the API returns only a summary, or nothing at all.
In self-consistency decoding, how is one final answer chosen from many sampled chains?
basics
~20 sSelf-consistency throws away the reasoning text and votes only on final answers: each sampled chain contributes its extracted answer, identical answers are grouped into one bucket, and the biggest bucket wins. That is a plurality, not a required majority.
In prompting, how does zero-shot CoT differ from few-shot CoT?
basics
~20 sZero-shot CoT adds only an instruction to reason step by step, and the model invents its own reasoning shape. Few-shot CoT puts solved examples in the prompt that demonstrate both how to reason and how the answer should look.
In a ReAct agent, what happens in one pass of the reason-act-observe loop?
basics
~20 sEach pass has three parts: the model writes a short reasoning step about what it needs next, emits one action against a tool, and gets the tool's result back as an observation. That observation joins the running context, and the next pass begins.
In ReAct, what is an observation and who is allowed to write it?
basics
~20 sAn observation is the tool's actual output, appended to the transcript by the runtime after an action. The model must never produce it: generation stops at the action, the real result is inserted, and only then does the model continue.
In a ReAct loop, what does the runtime do with an action before the tool runs?
basics
~20 sThe runtime resolves the action name against its registered tool catalogue, deserializes the arguments and validates them against that tool's schema, then executes. Anything that fails resolution or validation never reaches the tool — an error is returned instead.
What does a Reflexion-style agent carry into its next attempt after a failed one?
basics
~20 sA short self-written note in plain language about why the attempt failed and what to do differently. That note is appended to the next attempt's context, so the retry starts from an explicit lesson instead of repeating the same action.
What ends a ReAct loop, and which stop conditions can the model itself decide?
basics
~20 sThe model ends the loop by producing a final answer instead of an action. Because that signal is only the model's own judgement, the surrounding harness adds external stops it cannot override: a cap on passes, a cost or time limit, and a goal check on the result.
In Tree of Thought, how do you design the propose and evaluate prompt templates?
basics
~20 sWrite two separate templates with different jobs: a proposer that receives the serialized path and returns candidate next steps as free text, and an evaluator that receives one candidate state and returns a strictly parseable verdict. Keep their shared prefix byte-identical.
When would you run a Tree of Thoughts search depth-first with backtracking rather than breadth-first with a beam?
basics
~20 sGo depth-first with backtracking when a partial solution can be checked for contradiction, most branches die early, and any complete answer is worth having quickly. Use breadth-first with a beam when the depth is short and known and you want several equal-length plans compared side by side.
In Tree of Thoughts, how do branching factor, depth and beam width set the LLM call count?
basics
~20 sWith a fixed beam, each level generates width x branching candidates, so generation calls total width x branching x depth — linear in depth. Without a beam the tree grows as branching^depth, which is why unpruned breadth-first search is unaffordable past a few levels.
In Tree of Thought, how does value prompting differ from vote prompting?
basics
~20 sValue prompting scores each partial state on its own and returns an independent number or label. Vote prompting shows several sibling states together and asks which is most promising, returning a relative ranking rather than absolute scores.
How do you choose how large a single thought should be in a Tree of Thought?
basics
~20 sSize a thought so the model can produce several meaningfully different versions of it, and so a partial solution built from it can already be judged promising or hopeless. Too fine and siblings look identical; too coarse and there is almost nothing to branch over.
In prompt optimization, what does a meta-prompt that rewrites another prompt contain?
basics
~20 sA meta-prompt is a prompt whose subject is another prompt. It carries four things: a specification of the task, the current prompt verbatim, concrete evidence of how that prompt failed, and an instruction to emit a revised prompt in a fixed format.
In automatic prompt optimization, where do candidate prompts come from?
basics
~10 sThree sources dominate: inducing an instruction from labelled input/output pairs, resampling paraphrases of a seed instruction, and mutating slots in a fixed template. They trade diversity against staying on-task, and most systems combine them.
Why is automatic prompt optimization run as a discrete, gradient-free search?
basics
~20 sA prompt is discrete text, not a continuous parameter vector, so no gradient points toward a better prompt. Optimization instead runs as black-box search: propose edited candidates, score each on a dataset, keep the winners, repeat under a fixed budget.
In DSPy, what does a signature declare, and why isn't it just a prompt string?
basics
~20 sA DSPy signature declares one step's named inputs and outputs plus a short description of the task — for example context, question -> answer. The wording, formatting rules and examples that make up the actual prompt are generated by the framework, not written by you.
In automatic prompt optimization, how do you pick the scoring metric for a task?
basics
~20 sMatch the scorer to the output shape: exact match or F1 when one label is right, execution-based checks when the output can be run and verified, overlap scores only as a rough proxy, and a judge model only for genuinely open-ended text.
What does a prompt-cache hit change about an LLM request's cost and speed?
basics
~20 sA hit bills the reused prefix at a fraction of the normal input price — often around a tenth — and the model skips re-processing it, so the first token arrives far sooner. Output tokens are billed and generated exactly as usual.
In prompt caching, how many cache reads repay one cache write?
basics
~20 sBreak-even reads equal the write premium above base divided by the discount below base: with a 1.25x write and a 0.1x read, a single reuse already pays; at a 2x write it takes about two. The real constraint is whether that reuse arrives before the entry expires.
Why does a timestamp at the top of a system prompt destroy cache hit rate?
basics
~20 sPrompt caches match an exact prefix from the first token onward. A line like "Current time: 14:32:07" differs on every request, so the match fails at the very first tokens and the entire prompt behind it is reprocessed and re-stored, forever.
Why does editing one word early in a long system prompt void the whole cache?
basics
~20 sAttention is causal: every later token's cached key and value tensors were computed from all the tokens before it. Change a token near the top and the stored state for everything after it is stale, so the server must rebuild the entire prompt.
In prompt caching, what is physically stored for a cached prefix?
basics
~20 sThe per-layer attention key and value tensors for every token in the prefix — not the prompt text, not an embedding, and not the model's answer. A hit reloads those tensors and skips recomputing them.
What is the difference between a text embedding model and a generative LLM?
basics
~20 sAn embedding model maps a piece of text to one fixed-length vector of numbers that can be compared with other vectors. A generative LLM produces new text token by token. Embeddings are for comparing meaning, not for writing.
Which parts of a search request should never be answered by vector similarity?
basics
~20 sAnything with a truth condition: numeric and date constraints, sorting, counting, permission scoping, and exact identifier lookups. Vector similarity ranks by how alike two texts are, and "alike" cannot express "price under 800,000" or "only records this user may see".
In vector search, what does cosine similarity measure, and when is it preferred over Euclidean distance?
basics
~20 sCosine similarity measures the angle between two vectors and ignores their lengths, ranging from -1 to 1. Prefer it over Euclidean distance when only direction is meaningful and vector length reflects something you do not want scored, such as text length.
What can a UMAP or t-SNE plot of embeddings actually tell you?
basics
~20 sUMAP and t-SNE preserve which points sit near each other locally, so tight visible groups usually reflect real neighbourhoods. The gaps between blobs, the relative blob sizes and the axes carry no reliable meaning and shift with hyperparameters and seed.
How do you choose between k-means and DBSCAN for clustering document embeddings?
basics
~20 sk-means forces every document into one of k clusters you fix in advance, so it fits a corpus you want fully partitioned. DBSCAN groups by density, discovers how many clusters exist, and leaves sparse points unlabelled as noise.
Why is pasting a full patient record into a chat assistant a leak even without model training?
basics
~20 sText placed in a prompt becomes data your own system holds: it sits in conversation history and is resent every turn, and it lands in request logs, traces and support tooling. Training use is a separate, narrower question.
Why doesn't a strict JSON schema on an LLM response make its content safe?
basics
~20 sA schema constrains shape, not meaning. It guarantees the reply parses and carries the required fields and types. It says nothing about whether a field holds an invented price, a competitor's name, an unsafe instruction, or text that is dangerous where you paste it.
Why should an autonomous agent hold its own credentials instead of the operator's?
basics
~20 sAn agent running on a human's session inherits every permission that human has, and its actions are logged as theirs. Giving the agent its own principal lets you scope it narrowly, revoke it alone, and attribute what it actually did.
In an LLM feature, why moderate both user input and model output?
basics
~20 sInput screening rejects harmful requests before you pay for a generation and before they enter the context. Output screening catches harm the model produces anyway, including harm that arrived through retrieved documents or tool results the input check never saw.
Why do off-the-shelf PII detectors miss hospital MRNs, and what does over-redaction cost?
basics
~20 sShipped recognizers cover identifiers standardised nationally or industry-wide; a facility-assigned medical record number has no fixed format, so nothing matches it. Misses leak silently, while blanket masking strips the dosages, dates and lab values the answer depended on.
Why does copying text from a scanned PDF return nothing, but not from a born-digital PDF?
basics
~20 sA born-digital PDF stores real text objects — character codes with fonts and positions — so extraction just reads them out. A scanned page stores only a photograph of the paper, so no characters exist until OCR creates them.
How does a text-to-image diffusion model turn random noise into an image?
basics
~20 sGeneration starts from pure random noise. The model repeatedly predicts how much noise is present and removes a little of it, conditioned on the text prompt at every step. After a fixed number of steps the noise has been shaped into an image.
In a vision-language model, why can one image cost more tokens than a page of text?
basics
~20 sThe vision encoder cuts the image into a grid of fixed-size patches and emits roughly one token per patch, so token count grows with image area. A 1024x1024 image at patch size 16 is 64x64 = 4096 patches.
In document extraction, why demand a bounding box and page number per extracted field?
basics
~20 sGrounding turns an unverifiable string into an auditable claim. A clerk jumps straight to the pixels behind "container MSKU4412345", and a box landing on blank space exposes an invented value before it reaches the downstream system.
In document extraction, when do you pick classic OCR over a document VLM or a frontier model?
basics
~20 sClassic OCR wins on clean, fixed layouts at high volume: cheap, fast, deterministic, and it returns per-word boxes and confidence scores. Purpose-built document VLMs handle messy layout and tables; frontier models are for open-ended reasoning over the page.