skip to content

ML System Design

The 'design a recommender' or 'design a fraud detector' round: how training pipelines, feature platforms, serving, monitoring and retraining fit into one system that stays correct in production.

part ofDistributed & scalable systemsoverview, primer and where to startread it →
on this pageshow

explore

questions

245 · 11 sections

In a design round for checkout delivery-date estimates, what counts as the baseline a proposed model must beat?

level: juniorimportance: must knowfreq 62%
basics
~20 s

The baseline is whatever already decides the number in production - here a lane-and-service-level median transit table refreshed weekly from recent actuals - read on live traffic. A comparator invented for the write-up, such as a network-wide mean, is not the bar.

open as a page

In a churn model for a subscription service, what must one row of the training table fix before any feature is picked?

level: juniorimportance: must knowfreq 66%
basics
~20 s

One row must fix three things: the prediction unit (one account in one billing month), the horizon the label covers (cancels within the next 30 days), and the as-of timestamp that separates feature territory from label territory.

open as a page

In a warehouse pick-path service, how would you split a 250 ms p99 budget across its stages before choosing a model?

level: middleimportance: must knowfreq 64%
basics
~20 s

Reserve first what you cannot spend — the caller's network round trip, fixed request handling and explicit headroom — then allocate the remainder across feature lookup, inference and post-processing. The inference slice becomes a ceiling any candidate model must fit.

open as a page

In an inpatient deterioration-risk service, which decisions do the offline sweep metric and the online launch metric each gate?

level: middleimportance: must knowfreq 72%
basics
~20 s

The offline sweep metric ranks candidate models during training - a ranking score over a frozen retrospective cohort. The online launch metric decides whether the service ships: what the deployed worklist changed in care, read on live shifts.

open as a page

Why ship the rule-based arrival window first and instrument it, rather than waiting until the model is ready?

level: seniorimportance: must knowfreq 55%
basics
~20 s

Shipping the rule first delivers the product immediately and starts the clock on the evidence a model needs: every promise logged with the inputs as they stood at decision time, joined later to the actual delivery. Without that log, a model has no honest training set and no measured bar.

open as a page

What must a training run's record hold before two quarterly claim-severity runs can be compared at all?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A run record must pin every input and the yardstick: the dataset snapshot identifier, the code commit, the full parameter set, a content digest of the model artifact, and the evaluation set identifier with its exact metric definition. Two runs compare only when that evaluation set and metric are identical.

open as a page

In a content-moderation labeling operation, why is each reported post judged by three annotators instead of one?

level: juniorimportance: must knowfreq 60%
basics
~20 s

One verdict on a moderated post carries no error signal of its own. Three verdicts make disagreement visible, route the contested post into an adjudication path, and turn agreement into a running health check on the written guideline.

open as a page

Why can a model trained from 'the latest transcript export' not be rebuilt by rerunning the same code?

level: juniorimportance: must knowfreq 62%
basics
~20 s

'Latest' is a moving pointer, not a version. Between the two runs utterances were appended, transcripts corrected and recordings withdrawn, so the second job trains on a different set. Rebuilding needs an immutable snapshot identifier that resolves to fixed bytes.

open as a page

A nightly ad job keeps every click and one non-click in a hundred — what must ship with that dataset to keep scores calibrated?

level: middleimportance: must knowfreq 62%
basics
~20 s

The negative sampling rate, recorded per stratum and per build. Thinning non-clicks a hundredfold multiplies the odds by about a hundred, and without that factor nobody can convert the model's scores back into real click probabilities.

open as a page

A retraining graph's aggregate step is retried after a timeout and one region's daily units double — what went wrong?

level: middleimportance: must knowfreq 68%
basics
~20 s

The step appends. Scheduled steps run at least once, so a retry re-executed work whose rows were already written, and the partition now holds two copies. A step must replace its declared output, not add to it.

open as a page

In an online feature store, why does a value that stopped updating hours ago still serve a prediction without any error?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A feature read is a key lookup that returns whatever was last written, with no age check. The stale value is well-formed, so the model scores it and the wrong answer shows up only in the output.

open as a page

In a ride-hailing dispatch platform, why is the same driver feature stored twice — in a columnar history and in a key-value row?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Training and dispatch read the same feature with opposite access patterns: a training job scans months of one column across millions of rows, while dispatch fetches every feature for one driver in milliseconds. No single store serves both shapes well.

open as a page

A trip driven Tuesday is uploaded Friday: in a telematics feature pipeline, what distinguishes its event time from its ingestion time?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Event time is when the driving happened; ingestion time is when the trip summary reached the feature pipeline and became readable. Devices buffer and retry, so the two differ by hours or days and records arrive out of order.

open as a page

A checkout wait-time model gets a feature computed differently in the serving path than in training - why is there no error?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A model validates nothing about what its inputs mean: any number in the right slot is scored. Training-serving skew therefore lands as a confidently wrong estimate rather than an exception, and shows up only when real outcomes arrive.

open as a page

In a player lifetime-value model that reads a stored history embedding per player, why must each stored vector carry its encoder's version?

level: middleimportance: must knowfreq 62%
basics
~20 s

An embedding's coordinates only mean something relative to the encoder weights that produced them. The version stamp lets the serving path prove the vector came from the encoder the value model was trained against; without it, a mismatch is silent.

open as a page

What does batching audio chunks into one accelerator pass buy a live captioning service, and what does it cost?

level: juniorimportance: must knowfreq 64%
basics
~20 s

Batching amortises the accelerator's large fixed per-pass cost over many chunks, so throughput rises several-fold. It costs latency: the first chunk in a batch waits for the batch to fill or the window to expire before any scoring starts.

open as a page

A bidder's 75 ms internal p99 budget is decode 5, feature fetch 25, scoring 25, post-processing 10 and 10 ms reserve, and the measured p99s are 4, 38, 22 and 9 - which stage is over its line?

level: juniorimportance: must knowfreq 58%
basics
~20 s

Feature fetch, at 38 ms against a 25 ms line. The measured total still fits only because the other three stages came in under theirs, and the 10 ms reserve is down to 2 ms.

open as a page

In a photo library that auto-tags every upload, how do nightly batch tagging, tagging on the upload event, and tagging on first view differ?

level: juniorimportance: must knowfreq 72%
basics
~10 s

They differ in when the scorer runs. A nightly job tags the stored library on a schedule, event tagging scores each photo as it arrives, and first-view tagging scores only photos somebody actually opens.

open as a page

A legal-document classifier caches finished labels; why does keying on upload id nearly never hit while a content hash does?

level: juniorimportance: must knowfreq 62%
basics
~20 s

An upload id is unique to one upload, so every lookup misses. A content hash collapses the identical templates and boilerplate clauses that recur across customers onto one key, and repeats are the only thing a prediction cache can serve.

open as a page

A keyboard's next-word strip must never be empty; through which prediction sources would you fall back, and what does each rung give up?

level: middleimportance: must knowfreq 64%
basics
~20 s

Fall back down a ladder of prediction sources: the personalised scorer, then a cached prediction for this prefix, then a global next-word frequency list, then a constant on-device default. Each rung gives up freshness, then personalisation, then the prefix context.

open as a page

An inbound spam filter is live but no message has a confirmed verdict yet - which signals can you chart today?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Everything the model saw and emitted is available immediately: per-feature summary statistics and null rates, the prediction-score histogram, the blocked rate at the current decision cutoff, and traffic volume - each cut by tenant, inbound path and language.

open as a page

A predictive-maintenance model predicts failure within 30 days - why does this week's precision chart read optimistically high, and what corrects it?

level: middleimportance: must knowfreq 76%
basics
~20 s

Recent cohorts are still right-censored: a failure resolves the moment it happens while survival is only confirmed at day 30, so dropping unresolved rows over-counts failures. Bucket the metric by prediction date, publish only matured cohorts, and restate as labels land.

open as a page

In an inbound spam filter, why is the prediction-score histogram the first label-free signal most designs watch?

level: middleimportance: must knowfreq 60%
basics
~20 s

It compresses every input the model uses into a single curve the system already produces, so a change in any feature - including ones nobody thought to chart - can show up as a changed shape, for the cost of one series and no labels.

open as a page

A weather-feature shift test trips daily while served forecast error is unchanged - should that test page the on-call?

level: middleimportance: must knowfreq 57%
basics
~20 s

No. The pager belongs to the served forecast error, the effect the operator feels. A feature whose distribution moved without moving error is a diagnostic: attach it to the quality page as context, and let it file a ticket at most.

open as a page

Why does a drift test whose reference window rolls forward with recent traffic miss the slow shift it should catch?

level: middleimportance: must knowfreq 66%
basics
~20 s

A rolling reference re-estimates normal from data that already contains the drift, so each week's comparison is against last week's already-moved distribution. Gradual movement never accumulates in the statistic, though a step change still trips it.

open as a page

In a weekly-refreshed query autocomplete suggester, why must each candidate beat the live model before it replaces it?

level: juniorimportance: must knowfreq 60%
basics
~20 s

Newer training data does not guarantee a better suggester, and a refresh can be worse without failing. The model already serving is the baseline every candidate must clear on a stated rule, so promotion is earned rather than automatic.

open as a page

In a music playlist ranking service, why is a completed play not proof the listener wanted that track?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A completed play records what the listener tolerated in a context the service chose: autoplay, background listening and the slot given to the track all produce plays. It is a proxy label, not a verdict.

open as a page

A hotel pricing model still returns a nightly rate for every request, so why does it need refreshing at all?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A live pricing model keeps applying the demand pattern it learned from old bookings, and nothing errors when the market moves. Refreshing refits it on recent bookings and cancellations so its rates track today's demand, not last quarter's.

open as a page

A priority-inbox ranker is refit each month on the mail recipients opened — why does each refresh harden its demotion of a sender class?

level: middleimportance: must knowfreq 62%
basics
~20 s

Demoted mail is rarely seen, so it is rarely opened, so the next training set holds almost no positives for that sender class. Each refresh learns a lower score for it, which suppresses exposure further, and the loop tightens.

open as a page

What must a playlist ranking service log at serve time so a later play can be joined to the ranking that produced it?

level: middleimportance: must knowfreq 62%
basics
~20 s

One impression record per serving request: a request id the client echoes on every later event, each slot's track id, position and rendered flag, the serve-time score and any randomisation propensity, plus the model and feature-spec versions.

open as a page

In an offline gate for a candidate model, why must the shipping score come from a held-out data split, not the training rows?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A model scored on the rows it was fitted on is being asked to recall, not predict, so that score improves with capacity and understates the errors it will make on the next order. Only rows the fit never saw estimate that.

open as a page

A candidate ticket router scores live tickets in shadow mode - why can that window not show whether it routes better than the incumbent?

level: juniorimportance: must knowfreq 64%
basics
~20 s

Shadow mode compares mechanics, not outcomes. The incumbent's queue choice is what agents actually work, so no ticket is ever placed where the candidate said, and resolution time and reroute rate are never produced for the disagreements.

open as a page

In a sponsored-listing strip, how does team-draft interleaving compare two ranking policies within a single user's list?

level: middleimportance: must knowfreq 58%
basics
~20 s

Team-draft interleaving fills the strip by alternating drafts: a coin flip picks who starts, then each policy takes its highest-ranked listing not already placed. Every slot carries a hidden team tag, and a click credits the team that drafted that slot.

open as a page

How do you wire the mirrored call to a shadow ticket-routing model so it cannot add latency or failures to the live request?

level: middleimportance: must knowfreq 55%
basics
~20 s

Keep the candidate off the critical path. Answer the live request from the incumbent first, hand a copy to a bounded background worker with its own pool and timeout, and let shadow failures only increment a counter.

open as a page

Why can a candidate send-time model ramped to 10% of users deliver two notifications to one user in a day?

level: middleimportance: must knowfreq 62%
basics
~20 s

Because the ramp assignment was not sticky. The same user fell on the candidate side in one planning run and on the incumbent side in another, so both sides queued a send for that day. Deterministic per-user assignment with one claimed send key prevents it.

open as a page

In a jobs-for-you rail over four million postings, why retrieve a few hundred candidates before scoring?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Scoring every posting is impossible inside the rail's budget: four million postings at ten microseconds each is forty seconds of CPU per request. Retrieval trades exhaustive coverage for a few hundred plausible candidates found in milliseconds.

open as a page

In a marketplace serving both a browse feed and a typed product search, what does the typed query change about the candidate set?

level: juniorimportance: must knowfreq 64%
basics
~20 s

A typed query is a retrieval constraint applied before any model runs, so the ranker only sees items that already match the stated intent. A browse feed has no such constraint and must manufacture candidates from the viewer's profile and context.

open as a page

For a listener with no listening history, what decides which rung of a podcast feed's fallback ladder serves the request?

level: middleimportance: must knowfreq 68%
basics
~10 s

A signal test on the request, not the account's age. Each rung declares a precondition; the funnel serves the first rung whose precondition holds, re-checked on every request, and logs which rung answered.

open as a page

A news feed's scored candidates hold forty near-identical stories on one event - how does the list-editing layer avoid showing ten?

level: middleimportance: must knowfreq 70%
basics
~20 s

Cluster the near-duplicates first, then select greedily. Content similarity over title and lead text groups the forty filings into one event cluster; the list-editing pass then picks one representative and, using maximal marginal relevance, scores each further slot on relevance minus similarity to what is already placed.

open as a page

On a marketplace search surface, why is applying an in-stock facet inside retrieval different from filtering the ranked list afterwards?

level: middleimportance: must knowfreq 57%
basics
~20 s

Inside retrieval, the shortlist comes back full of eligible items. Filtering afterwards spends the shortlist on ineligible ones, so a selective facet leaves too few results to fill the page and wastes the scoring stage's budget.

open as a page

In an inline login-risk scorer, what is a per-device velocity counter and why is it kept continuously updated?

level: juniorimportance: must knowfreq 58%
basics
~20 s

A per-device velocity counter is a running count of that device's recent login attempts, held in a low-latency key-value store and folded forward by the login event stream, so the inline scorer reads it in one lookup instead of scanning history.

open as a page

In a marketplace's fake-review defence, why send a middle band of risk scores to human reviewers instead of one allow/remove cutoff?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A single cutoff forces every borderline listing into an automatic allow or an automatic removal. The middle band is where the score is genuinely uncertain, so routing it to an analyst buys a cheap decision instead of a costly wrong one.

open as a page

Why does a card-not-present fraud model retrained on the last 30 days of checkouts learn that recent traffic is almost fraud-free?

level: middleimportance: must knowfreq 72%
basics
~10 s

A settled dispute confirms fraud weeks after the checkout, so recent rows carry no dispute yet. Joining absence-of-dispute to a legitimate training label marks immature fraud as good and deflates the recent fraud rate.

open as a page

In a refund-abuse decision layer, what changes if the rules run as a cascade before the model instead of entering it as features?

level: middleimportance: must knowfreq 74%
basics
~20 s

A cascade makes a matching rule an absolute veto and hides the traffic it decided from the model. Rule hits as features make the rule one weighable piece of evidence the model can discount. One is policy, the other is signal.

open as a page

When an allowlist entry and a hard-block rule both match one refund request, what decides the action the decision layer emits?

level: middleimportance: must knowfreq 66%
basics
~20 s

A precedence order declared in advance - not source order, not whichever rule was authored last. The layer collects the matches, resolves them by that documented ranking, emits exactly one action, and records which branch won.

open as a page

A scoring fleet takes 12,000 sensor readings per second at peak, each instance sustains 200, target utilisation 60% - how many instances?

level: juniorimportance: must knowfreq 62%
basics
~10 s

One hundred instances. Dividing the 12,000-per-second peak by 200 gives 60 instances running flat out; dividing again by the 0.6 utilisation target gives 100, whose 20,000-per-second ceiling leaves peak sitting at 60%.

open as a page

A bulk catalogue re-embedding job runs 400 accelerator-hours a month at $3 an hour for 20 million embeddings - what is the cost per embedding?

level: juniorimportance: must knowfreq 62%
basics
~10 s

Six hundredths of a cent. 400 hours at $3 is $1,200 of compute, divided by 20 million embeddings gives $0.00006 each, or $0.06 per thousand. That figure covers marginal compute only.

open as a page

In a video-moderation service, which half of the spend grows with upload volume: the quarterly training run or per-clip scoring?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Per-clip scoring grows with upload volume; the quarterly training run does not. Training is paid once per model version, while scoring is paid again for every clip that arrives, so only the serving half is multiplied by traffic.

open as a page

A hosted invoice extractor charges per document — how do you find the monthly volume where running your own becomes cheaper?

level: middleimportance: must knowfreq 62%
basics
~20 s

Divide the fixed monthly spend of running your own by the per-document saving it earns: crossover volume = fixed monthly cost / (hosted price per document - your own marginal cost per document). Below that volume, buying is the cheaper of the two recurring bills.

open as a page

Which costs does a per-embedding figure computed from the catalogue re-embedding workers' compute bill alone leave out?

level: middleimportance: must knowfreq 58%
basics
~20 s

The amortised training run, the reserved-but-idle share of the accelerator fleet, vector storage and index writes, and orchestration and monitoring. On the worked catalogue refresh they turn $0.00006 of marginal compute into about $0.0006 fully loaded.

open as a page

What must a decision-log row record when a risk model selects a tax return for audit?

level: juniorimportance: must knowfreq 64%
basics
~10 s

A decision-log row records the decision identifier and timestamp, the snapshot of inputs the model was given, the model version and feature-definition version, the score returned, the threshold in force, and the action taken.

open as a page

A mortgage underwriting scorer is retrained from an unchanged code revision a month later and comes out different - why?

level: juniorimportance: must knowfreq 52%
basics
~20 s

A code revision pins the training instructions, not the values they read: the applications the training query returns, the resolved library versions, the run-time configuration and the seed all moved. Rebuilding a model version means pinning those inputs too.

open as a page

Many advertiser tenants share one ad-screening inference tier; one tenant triples its traffic and all tenants slow down - why?

level: juniorimportance: must knowfreq 60%
basics
~20 s

Every tenant draws on the same finite pool of workers, accelerator time and queue slots. The extra requests wait in the same line, so queueing delay rises for everyone as occupancy climbs. Nothing has to fail for this to happen.

open as a page

Which must an audit-selection decision log keep to survive a feature-definition change: the raw return figures or the computed features?

level: middleimportance: must knowfreq 57%
basics
~20 s

The raw submitted figures survive, because they do not depend on your code. Keep them, keep the vector the model was actually fed as proof of what it saw, and stamp the feature-definition version linking the two.

open as a page

Rolling the dropout early-warning model back to its previous version mid-incident: what must that rollback cover beyond the trained artifact?

level: middleimportance: must knowfreq 68%
basics
~20 s

A model version is a pin set, not a file: the artifact plus the preprocessing and feature-definition versions it reads, its decision threshold and its output contract. Roll back all of them together, by flipping one pointer to a retained, still-loadable version.

open as a page