AI Engineering
The production side of LLM work: how the app is wired, how quality is measured, how calls are traced, how cost and latency are held down, and what guards sit around the model. Senior interviews spend most of their time here, because calling the model is the easy part.
on this pageshowhide
explore
- LLM App Architecture4 questions
- Vector Databases6 questions
- Evaluation & Testing25 questions
- Offline vs Online Evaluation5 questions
- LLM-as-Judge Validity5 questions
- Benchmarks vs Task Evals5 questions
- Eval Dataset Curation5 questions
- Regression Suites & CI Gating5 questions
- Observability & Tracing15 questions
- Traces & Spans for LLM Calls5 questions
- Token & Cost Telemetry5 questions
- Quality Signals & Feedback Loops5 questions
- Cost & Latency Optimization14 questions
- Streaming & Perceived Latency5 questions
- Batching & Concurrency Control4 questions
- Model Routing & Cascades5 questions
- Guardrails & Safety19 questions
- Moderation & Content Classifiers5 questions
- Prompt-Injection Mitigation Taxonomy5 questions
- Data-Exfiltration Controls5 questions
- PII Handling in Prompts & Traces4 questions
- Structured Output & Tool-Call Reliability14 questions
- Constrained Decoding5 questions
- Schema Validation & Repair Loops5 questions
- Tool-Call Reliability4 questions
questions
97 · 7 sectionsWhat does an LLM agent harness own that the model itself never will?
basics
~20 sThe harness owns everything durable and enforceable: the session event log, the context builder that assembles each call, the tool router that actually executes calls, checkpoints, permissions and observability. The model only reasons over whatever context the harness hands it.
Why do agent answers degrade long before the context window fills, and what fixes it?
basics
~20 sQuality decays as a window fills with stale turns and superseded tool output — context rot — well inside the nominal limit. The fix is context engineering: compact the session into a structured summary and reinitialize the window, keeping only what the remaining work needs.
A tool returns a 40MB log dump — how do you get it into the agent's context?
basics
~20 sYou do not. Store the payload outside the window and pass a reference — a file path, object key or result ID — plus a short digest such as size, shape and the first lines. The agent reopens or searches it on demand if the digest is not enough.
Why rebuild an agent's context each turn from a durable session event log?
basics
~20 sBecause the context window is derived state, not the system of record. An append-only event log lets the harness resume after a crash, rewind to a checkpoint, re-derive a compacted window, change the context policy without losing history, and reconstruct exactly what the model saw.
In vector search, when do cosine similarity, dot product and L2 distance rank results identically?
basics
~20 sOnce every vector is normalized to unit length, all three produce the same ordering: cosine equals the dot product, and squared L2 distance is 2 minus twice the cosine. Without normalization, the dot product favours long vectors and the three diverge.
In an HNSW vector index, what do m, efConstruction and ef trade off?
basics
~20 sm sets how many neighbour links each node keeps, efConstruction how hard the builder searches while inserting, and ef how wide the search is at query time. Raising them raises recall while costing memory, build time and query latency respectively.
Why add BM25 keyword retrieval alongside dense vectors, and how does reciprocal rank fusion combine them?
basics
~20 sDense embeddings match meaning but smear rare exact tokens such as an identifier, a part number or a docket number; BM25 matches those literally. Reciprocal rank fusion merges the two result lists by rank position, so the systems' incomparable raw scores never have to be reconciled.
Why can a metadata filter make an HNSW vector search return almost nothing?
basics
~20 sA selective filter applied after the search leaves almost nothing, because the top-k nearest vectors rarely satisfy a narrow predicate. Applied before the search, it deletes most nodes from the proximity graph, breaking the connectivity that greedy traversal depends on.
What does a cross-encoder reranker buy over the vector index's own ranking?
basics
~20 sA cross-encoder reads the query and the candidate together, so it can judge fine-grained relevance that independently-encoded vectors cannot. It raises precision at the top of the list but cannot recover anything the first stage failed to retrieve, and its cost grows with the number of candidates scored.
What is the difference between a public LLM benchmark and a task eval?
basics
~20 sA public benchmark is a shared, fixed dataset that ranks models against each other on a general capability. A task eval runs your own inputs through your own system and scores your own success criterion. Only the task eval predicts what your users will see.
What is LLM-as-judge evaluation, and why is a judge score not ground truth?
basics
~20 sLLM-as-judge means prompting a model with a rubric to score another model's output. The score is a noisy estimate, not truth: the judge has its own biases and blind spots, so it must be validated against human labels before anyone trusts the number.
What is the difference between offline and online evaluation of an LLM feature?
basics
~20 sOffline evaluation scores a fixed, frozen set of saved examples in a harness before release, so it is repeatable and cheap. Online evaluation measures real user traffic after release, where actual behaviour and business outcomes decide whether the change helped.
How do you sample production traces into a golden eval set without losing rare cases?
basics
~20 sSample by strata, not uniformly. Bucket traces by the dimensions you care about — request type, customer segment, known failure mode — then take a quota from each bucket, so rare but costly cases appear in numbers large enough to score.
Which biases distort LLM-as-judge scores, and how do you control each one?
basics
~20 sThe recurring four are position bias (order of presentation sways pairwise verdicts), verbosity bias (longer wins), self-preference (a judge rates its own family higher), and formatting halo (bullets and confident tone read as quality). Controls: swap orders, normalise or cap length, judge across families, and strip presentation from the rubric.
Why can't thumbs-up/down ratings alone measure quality of a production LLM assistant?
basics
~20 sExplicit ratings cover a tiny, self-selected slice of traffic, often under 1% of sessions, and the people who bother to click skew angry or delighted. Implicit signals such as retries, abandonment and human edits are emitted by every session, so they carry the trend.
Which token counters belong on an LLM cost dashboard beyond input and output?
basics
~20 sReasoning (thinking) tokens, which are billed at generation rates but never shown, and the cache-write versus cache-read split of input tokens, which carry very different unit prices. Without those three extra counters, a dashboard cannot explain the invoice.
How would you model one turn of a tool-using LLM agent as a trace of spans?
basics
~20 sMake the turn one trace: a root run span, agent spans for each reasoning cycle, and child spans for every model call, tool call and retrieval. Each span's parent is whatever caused it and outlives it.
How would you design a drift alert on hourly sampled quality scores for a live assistant?
basics
~20 sScore a fixed small sample of live sessions each hour, compare the rolling mean against a same-hour-of-week baseline rather than a flat threshold, and fire only when the deviation persists across several windows. Pin the scorer version, or scorer drift will masquerade as quality drift.
A trip-planning agent recommended a fully booked hotel — how do you localize the fault in its trace?
basics
~20 sOpen the turn's trace and split the question in two: did the availability fact ever reach the model? The retrieval and tool spans answer that; the model-call span's recorded input answers whether the fact was present and ignored.
What does a 429 from an LLM provider mean, and why is an immediate retry wrong?
basics
~20 sA 429 means you crossed the provider's rate limit, usually a per-minute cap on requests or on tokens. Retrying instantly adds load to an already-throttled account; wait, honour any Retry-After, and back off exponentially with random jitter.
Why stream an LLM response token by token instead of returning it all at once?
basics
~10 sStreaming does not make generation faster. It makes the wait visible: the user starts reading the first sentence while the rest is still being produced, so an 18-second answer feels responsive instead of frozen.
When does a small-first LLM model cascade stop saving money?
basics
~20 sA cascade only pays when the cheap tier answers enough traffic outright. Escalated requests are billed for both tiers plus the escalation check, so savings disappear once the escalation rate approaches the point where the small tier's spend no longer offsets the large tier's.
For a streamed LLM endpoint, which latency numbers do you track besides total time?
basics
~20 sTrack three separately: time to the first token the user can actually see, inter-token latency once the stream is flowing, and total time to the last token. Measure the first one on the client, at p50 and p99.
When does an LLM provider's batch API beat synchronous calls for bulk work?
basics
~20 sUse it when the job has a deadline but no single item needs a fast answer. Batch endpoints trade a completion window of up to a day for roughly half price and much higher throughput off the provider's spare capacity.
Why do content moderation classifiers return per-category scores, not one flag?
basics
~20 sDifferent harms need different responses. Per-category scores let a system hard-block one category, route self-harm signals to a crisis flow, and only warn on mild insults. A single unsafe flag forces one action and one cutoff for every kind of harm.
Which output channels let an LLM app leak data without running code?
basics
~20 sAny path where model output causes a fetch is an egress channel: rendered image and link URLs, outbound tool or webhook calls, redirect targets, even hostname lookups. Controls only work once every such sink is enumerated.
Why bind the tenant filter on RAG retrieval server-side, not in the prompt?
basics
~20 sAnything the model can influence can be widened by text that reaches it. Derive the tenant scope from the authenticated session and bind it at the query layer, so the model contributes only a search string and never the filter that decides which corpus is readable.
Why is the instruction/data split inside an LLM prompt not a real trust boundary?
basics
~20 sA model receives one flat token sequence. System text, user text and retrieved text carry no enforced privilege difference — only a learned tendency to prefer operator wording. Prompt-level separation shifts odds; the enforceable boundary has to live outside the model.
In an LLM app, why screen model output when the user input already passed moderation?
basics
~20 sInput screening only sees what the user typed. A model can still emit harmful text from a benign prompt, from retrieved documents, or from another user's content pulled into context, so the output is a separate surface that needs its own classifier.
Structured Output & Tool-Call Reliability
all 14 Structured Output & Tool-Call Reliability questions →Why validate an LLM's JSON output against a strict schema before your code uses it?
basics
~20 sParsing only proves the text is syntactically JSON. Validation proves the fields your code reads actually exist, with the right types and allowed values — so bad output fails loudly at the boundary instead of corrupting logic downstream.
How does logit masking during decoding make invalid JSON impossible rather than unlikely?
basics
~20 sThe decoder tracks its position in a state machine compiled from the schema and, at every step, drives the logits of all tokens that cannot legally continue to negative infinity. An illegal token has zero probability, so it is never sampled.
What is the difference between JSON mode, a strict schema mode, and grammar-constrained decoding?
basics
~20 sJSON mode guarantees only that the text parses as JSON. A strict schema mode additionally guarantees the object conforms to your schema — required keys, types, enums. A custom grammar constrains any formal language, JSON or otherwise.
How do you shape a JSON schema so a model actually complies with it?
basics
~20 sPrefer flat objects over deep nesting, enums over free text, and a single shape with a discriminator field over unions of alternative shapes. Give fields self-describing names and descriptions that state units and format. Compliance is a property of the schema's design, not only of the model.
Why do LLM tool calls arrive with invented ids or out-of-enum values, and how do you catch them?
basics
~20 sTool arguments are generated as text, so a model fills in a plausible order id, category or amount rather than leaving a field blank. Validate every call at the tool boundary — types, enum membership, units, and whether the referenced record actually exists — before executing anything.