skip to content

AI Engineering

The production side of LLM work: how the app is wired, how quality is measured, how calls are traced, how cost and latency are held down, and what guards sit around the model. Senior interviews spend most of their time here, because calling the model is the easy part.

on this pageshow

explore

questions

97 · 7 sections

What does an LLM agent harness own that the model itself never will?

level: middleimportance: must knowfreq 70%
basics
~20 s

The harness owns everything durable and enforceable: the session event log, the context builder that assembles each call, the tool router that actually executes calls, checkpoints, permissions and observability. The model only reasons over whatever context the harness hands it.

open as a page

Why do agent answers degrade long before the context window fills, and what fixes it?

level: seniorimportance: must knowfreq 62%
basics
~20 s

Quality decays as a window fills with stale turns and superseded tool output — context rot — well inside the nominal limit. The fix is context engineering: compact the session into a structured summary and reinitialize the window, keeping only what the remaining work needs.

open as a page

A tool returns a 40MB log dump — how do you get it into the agent's context?

level: middleimportance: should knowfreq 48%
basics
~20 s

You do not. Store the payload outside the window and pass a reference — a file path, object key or result ID — plus a short digest such as size, shape and the first lines. The agent reopens or searches it on demand if the digest is not enough.

open as a page

Why rebuild an agent's context each turn from a durable session event log?

level: principalimportance: should knowfreq 35%
basics
~20 s

Because the context window is derived state, not the system of record. An append-only event log lets the harness resume after a crash, rewind to a checkpoint, re-derive a compacted window, change the context policy without losing history, and reconstruct exactly what the model saw.

open as a page

In vector search, when do cosine similarity, dot product and L2 distance rank results identically?

level: juniorimportance: must knowfreq 58%
basics
~20 s

Once every vector is normalized to unit length, all three produce the same ordering: cosine equals the dot product, and squared L2 distance is 2 minus twice the cosine. Without normalization, the dot product favours long vectors and the three diverge.

open as a page

In an HNSW vector index, what do m, efConstruction and ef trade off?

level: middleimportance: must knowfreq 66%
basics
~20 s

m sets how many neighbour links each node keeps, efConstruction how hard the builder searches while inserting, and ef how wide the search is at query time. Raising them raises recall while costing memory, build time and query latency respectively.

open as a page

Why add BM25 keyword retrieval alongside dense vectors, and how does reciprocal rank fusion combine them?

level: middleimportance: must knowfreq 56%
basics
~20 s

Dense embeddings match meaning but smear rare exact tokens such as an identifier, a part number or a docket number; BM25 matches those literally. Reciprocal rank fusion merges the two result lists by rank position, so the systems' incomparable raw scores never have to be reconciled.

open as a page

Why can a metadata filter make an HNSW vector search return almost nothing?

level: seniorimportance: should knowfreq 44%
basics
~20 s

A selective filter applied after the search leaves almost nothing, because the top-k nearest vectors rarely satisfy a narrow predicate. Applied before the search, it deletes most nodes from the proximity graph, breaking the connectivity that greedy traversal depends on.

open as a page

What does a cross-encoder reranker buy over the vector index's own ranking?

level: seniorimportance: should knowfreq 48%
basics
~20 s

A cross-encoder reads the query and the candidate together, so it can judge fine-grained relevance that independently-encoded vectors cannot. It raises precision at the top of the list but cannot recover anything the first stage failed to retrieve, and its cost grows with the number of candidates scored.

open as a page

What is the difference between a public LLM benchmark and a task eval?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A public benchmark is a shared, fixed dataset that ranks models against each other on a general capability. A task eval runs your own inputs through your own system and scores your own success criterion. Only the task eval predicts what your users will see.

open as a page

What is LLM-as-judge evaluation, and why is a judge score not ground truth?

level: juniorimportance: must knowfreq 70%
basics
~20 s

LLM-as-judge means prompting a model with a rubric to score another model's output. The score is a noisy estimate, not truth: the judge has its own biases and blind spots, so it must be validated against human labels before anyone trusts the number.

open as a page

What is the difference between offline and online evaluation of an LLM feature?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Offline evaluation scores a fixed, frozen set of saved examples in a harness before release, so it is repeatable and cheap. Online evaluation measures real user traffic after release, where actual behaviour and business outcomes decide whether the change helped.

open as a page

How do you sample production traces into a golden eval set without losing rare cases?

level: middleimportance: must knowfreq 60%
basics
~20 s

Sample by strata, not uniformly. Bucket traces by the dimensions you care about — request type, customer segment, known failure mode — then take a quota from each bucket, so rare but costly cases appear in numbers large enough to score.

open as a page

Which biases distort LLM-as-judge scores, and how do you control each one?

level: middleimportance: must knowfreq 68%
basics
~20 s

The recurring four are position bias (order of presentation sways pairwise verdicts), verbosity bias (longer wins), self-preference (a judge rates its own family higher), and formatting halo (bullets and confident tone read as quality). Controls: swap orders, normalise or cap length, judge across families, and strip presentation from the rubric.

open as a page

Why can't thumbs-up/down ratings alone measure quality of a production LLM assistant?

level: middleimportance: must knowfreq 62%
basics
~20 s

Explicit ratings cover a tiny, self-selected slice of traffic, often under 1% of sessions, and the people who bother to click skew angry or delighted. Implicit signals such as retries, abandonment and human edits are emitted by every session, so they carry the trend.

open as a page

Which token counters belong on an LLM cost dashboard beyond input and output?

level: middleimportance: must knowfreq 66%
basics
~20 s

Reasoning (thinking) tokens, which are billed at generation rates but never shown, and the cache-write versus cache-read split of input tokens, which carry very different unit prices. Without those three extra counters, a dashboard cannot explain the invoice.

open as a page

How would you model one turn of a tool-using LLM agent as a trace of spans?

level: middleimportance: must knowfreq 58%
basics
~20 s

Make the turn one trace: a root run span, agent spans for each reasoning cycle, and child spans for every model call, tool call and retrieval. Each span's parent is whatever caused it and outlives it.

open as a page

How would you design a drift alert on hourly sampled quality scores for a live assistant?

level: seniorimportance: must knowfreq 52%
basics
~20 s

Score a fixed small sample of live sessions each hour, compare the rolling mean against a same-hour-of-week baseline rather than a flat threshold, and fire only when the deviation persists across several windows. Pin the scorer version, or scorer drift will masquerade as quality drift.

open as a page

A trip-planning agent recommended a fully booked hotel — how do you localize the fault in its trace?

level: seniorimportance: must knowfreq 52%
basics
~20 s

Open the turn's trace and split the question in two: did the availability fact ever reach the model? The retrieval and tool spans answer that; the model-call span's recorded input answers whether the fact was present and ignored.

open as a page

What does a 429 from an LLM provider mean, and why is an immediate retry wrong?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A 429 means you crossed the provider's rate limit, usually a per-minute cap on requests or on tokens. Retrying instantly adds load to an already-throttled account; wait, honour any Retry-After, and back off exponentially with random jitter.

open as a page

Why stream an LLM response token by token instead of returning it all at once?

level: juniorimportance: must knowfreq 72%
basics
~10 s

Streaming does not make generation faster. It makes the wait visible: the user starts reading the first sentence while the rest is still being produced, so an 18-second answer feels responsive instead of frozen.

open as a page

When does a small-first LLM model cascade stop saving money?

level: middleimportance: must knowfreq 62%
basics
~20 s

A cascade only pays when the cheap tier answers enough traffic outright. Escalated requests are billed for both tiers plus the escalation check, so savings disappear once the escalation rate approaches the point where the small tier's spend no longer offsets the large tier's.

open as a page

For a streamed LLM endpoint, which latency numbers do you track besides total time?

level: middleimportance: must knowfreq 58%
basics
~20 s

Track three separately: time to the first token the user can actually see, inter-token latency once the stream is flowing, and total time to the last token. Measure the first one on the client, at p50 and p99.

open as a page

When does an LLM provider's batch API beat synchronous calls for bulk work?

level: middleimportance: should knowfreq 50%
basics
~20 s

Use it when the job has a deadline but no single item needs a fast answer. Batch endpoints trade a completion window of up to a day for roughly half price and much higher throughput off the provider's spare capacity.

open as a page

Why do content moderation classifiers return per-category scores, not one flag?

level: juniorimportance: must knowfreq 55%
basics
~20 s

Different harms need different responses. Per-category scores let a system hard-block one category, route self-harm signals to a crisis flow, and only warn on mild insults. A single unsafe flag forces one action and one cutoff for every kind of harm.

open as a page

Which output channels let an LLM app leak data without running code?

level: middleimportance: must knowfreq 62%
basics
~20 s

Any path where model output causes a fetch is an egress channel: rendered image and link URLs, outbound tool or webhook calls, redirect targets, even hostname lookups. Controls only work once every such sink is enumerated.

open as a page

Why bind the tenant filter on RAG retrieval server-side, not in the prompt?

level: middleimportance: must knowfreq 58%
basics
~20 s

Anything the model can influence can be widened by text that reaches it. Derive the tenant scope from the authenticated session and bind it at the query layer, so the model contributes only a search string and never the filter that decides which corpus is readable.

open as a page

Why is the instruction/data split inside an LLM prompt not a real trust boundary?

level: middleimportance: must knowfreq 72%
basics
~20 s

A model receives one flat token sequence. System text, user text and retrieved text carry no enforced privilege difference — only a learned tendency to prefer operator wording. Prompt-level separation shifts odds; the enforceable boundary has to live outside the model.

open as a page

In an LLM app, why screen model output when the user input already passed moderation?

level: middleimportance: must knowfreq 62%
basics
~20 s

Input screening only sees what the user typed. A model can still emit harmful text from a benign prompt, from retrieved documents, or from another user's content pulled into context, so the output is a separate surface that needs its own classifier.

open as a page

Why validate an LLM's JSON output against a strict schema before your code uses it?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Parsing only proves the text is syntactically JSON. Validation proves the fields your code reads actually exist, with the right types and allowed values — so bad output fails loudly at the boundary instead of corrupting logic downstream.

open as a page

How does logit masking during decoding make invalid JSON impossible rather than unlikely?

level: middleimportance: must knowfreq 64%
basics
~20 s

The decoder tracks its position in a state machine compiled from the schema and, at every step, drives the logits of all tokens that cannot legally continue to negative infinity. An illegal token has zero probability, so it is never sampled.

open as a page

What is the difference between JSON mode, a strict schema mode, and grammar-constrained decoding?

level: middleimportance: must knowfreq 74%
basics
~20 s

JSON mode guarantees only that the text parses as JSON. A strict schema mode additionally guarantees the object conforms to your schema — required keys, types, enums. A custom grammar constrains any formal language, JSON or otherwise.

open as a page

How do you shape a JSON schema so a model actually complies with it?

level: middleimportance: must knowfreq 54%
basics
~20 s

Prefer flat objects over deep nesting, enums over free text, and a single shape with a discriminator field over unions of alternative shapes. Give fields self-describing names and descriptions that state units and format. Compliance is a property of the schema's design, not only of the model.

open as a page

Why do LLM tool calls arrive with invented ids or out-of-enum values, and how do you catch them?

level: middleimportance: must knowfreq 68%
basics
~20 s

Tool arguments are generated as text, so a model fills in a plausible order id, category or amount rather than leaving a field blank. Validate every call at the tool boundary — types, enum membership, units, and whether the referenced record actually exists — before executing anything.

open as a page