ML System Design
The 'design a recommender' or 'design a fraud detector' round: how training pipelines, feature platforms, serving, monitoring and retraining fit into one system that stays correct in production.
part ofDistributed & scalable systemsoverview, primer and where to startread it →on this pageshowhide
explore
- ML Problem Framing20 questions
- Prediction Target & Label5 questions
- Offline & Online Metrics5 questions
- Latency & Cost Envelope5 questions
- Heuristic Baselines5 questions
- Data & Training Pipelines25 questions
- Example Assembly & Sampling5 questions
- Labeling Operations5 questions
- Snapshot Versioning & Lineage5 questions
- Restartable Job Graphs5 questions
- Experiment Tracking & Promotion5 questions
- Feature Engineering & Feature Stores26 questions
- Offline and Online Split5 questions
- Point-in-Time Joins5 questions
- Training-Serving Skew5 questions
- Freshness SLAs & Backfill6 questions
- Embedding Vectors as Inputs5 questions
- Model Serving & Inference25 questions
- Precompute vs Live Scoring6 questions
- Latency Budget Allocation5 questions
- Prediction & Embedding Caches5 questions
- Fallback Tiers & Staleness5 questions
- Accelerator Batching Tradeoffs4 questions
- Monitoring & Drift Detection21 questions
- Input & Prediction Signals5 questions
- Shift Tests & Baselines6 questions
- Delayed & Proxy Labels5 questions
- Model Quality Alarms5 questions
- Retraining Loops & Feedback20 questions
- Refresh Triggers & Cadence6 questions
- Labels from User Actions5 questions
- Self-Reinforcing Cycles5 questions
- Champion-Challenger Promotion4 questions
- Experimentation & Safe Rollout21 questions
- Offline Gate5 questions
- Shadow Scoring5 questions
- Traffic Ramp & Rollback5 questions
- Interleaving & Long-Term Holdouts6 questions
- Recommender & Ranking Archetype27 questions
- Candidate Retrieval Tier6 questions
- Scoring Cascade5 questions
- Cold Start Bootstrapping5 questions
- Diversity & Business Rules6 questions
- Feed vs Query Intent5 questions
- Fraud & Abuse Detection Archetype20 questions
- Streaming Risk Scoring5 questions
- Rules and Model Blending5 questions
- Thresholds and Review Capacity5 questions
- Delayed Labels and Adversaries5 questions
- Cost & Capacity for ML20 questions
- Fixed vs Per-Request Spend5 questions
- Serving Fleet Sizing5 questions
- Per-Prediction Unit Economics5 questions
- Build-Buy-Baseline Decision5 questions
- ML Platform & Governance20 questions
- Model Lifecycle & Reproducibility5 questions
- Multi-Model & Multi-Tenant Serving5 questions
- Auditability & Decision Logging5 questions
- ML Incidents & Kill Switches5 questions
- AI & Data Scientistrole
- AI Engineerrole
- API Designskill
- Android Developerrole
- Backend Developerrole
- Blockchain Developerrole
- Data Engineerrole
- DevOps / SRE Engineerrole
- Forward Deployed Engineerrole
- Frontend Developerrole
- Full Stack Developerrole
- Game Developerrole
- Java Backend Developerrole
- Kotlin Backend Developerrole
- MLOps Engineerrole
- Machine Learning Engineerrole
- PostgreSQL DBArole
- Server-Side Game Developerrole
- Software Architectrole
- Software Design & Architectureskill
- System Designskill
- iOS Developerrole
questions
245 · 11 sectionsIn a design round for checkout delivery-date estimates, what counts as the baseline a proposed model must beat?
basics
~20 sThe baseline is whatever already decides the number in production - here a lane-and-service-level median transit table refreshed weekly from recent actuals - read on live traffic. A comparator invented for the write-up, such as a network-wide mean, is not the bar.
In a churn model for a subscription service, what must one row of the training table fix before any feature is picked?
basics
~20 sOne row must fix three things: the prediction unit (one account in one billing month), the horizon the label covers (cancels within the next 30 days), and the as-of timestamp that separates feature territory from label territory.
In a warehouse pick-path service, how would you split a 250 ms p99 budget across its stages before choosing a model?
basics
~20 sReserve first what you cannot spend — the caller's network round trip, fixed request handling and explicit headroom — then allocate the remainder across feature lookup, inference and post-processing. The inference slice becomes a ceiling any candidate model must fit.
In an inpatient deterioration-risk service, which decisions do the offline sweep metric and the online launch metric each gate?
basics
~20 sThe offline sweep metric ranks candidate models during training - a ranking score over a frozen retrospective cohort. The online launch metric decides whether the service ships: what the deployed worklist changed in care, read on live shifts.
Why ship the rule-based arrival window first and instrument it, rather than waiting until the model is ready?
basics
~20 sShipping the rule first delivers the product immediately and starts the clock on the evidence a model needs: every promise logged with the inputs as they stood at decision time, joined later to the actual delivery. Without that log, a model has no honest training set and no measured bar.
What must a training run's record hold before two quarterly claim-severity runs can be compared at all?
basics
~20 sA run record must pin every input and the yardstick: the dataset snapshot identifier, the code commit, the full parameter set, a content digest of the model artifact, and the evaluation set identifier with its exact metric definition. Two runs compare only when that evaluation set and metric are identical.
In a content-moderation labeling operation, why is each reported post judged by three annotators instead of one?
basics
~20 sOne verdict on a moderated post carries no error signal of its own. Three verdicts make disagreement visible, route the contested post into an adjudication path, and turn agreement into a running health check on the written guideline.
Why can a model trained from 'the latest transcript export' not be rebuilt by rerunning the same code?
basics
~20 s'Latest' is a moving pointer, not a version. Between the two runs utterances were appended, transcripts corrected and recordings withdrawn, so the second job trains on a different set. Rebuilding needs an immutable snapshot identifier that resolves to fixed bytes.
A nightly ad job keeps every click and one non-click in a hundred — what must ship with that dataset to keep scores calibrated?
basics
~20 sThe negative sampling rate, recorded per stratum and per build. Thinning non-clicks a hundredfold multiplies the odds by about a hundred, and without that factor nobody can convert the model's scores back into real click probabilities.
A retraining graph's aggregate step is retried after a timeout and one region's daily units double — what went wrong?
basics
~20 sThe step appends. Scheduled steps run at least once, so a retry re-executed work whose rows were already written, and the partition now holds two copies. A step must replace its declared output, not add to it.
In an online feature store, why does a value that stopped updating hours ago still serve a prediction without any error?
basics
~20 sA feature read is a key lookup that returns whatever was last written, with no age check. The stale value is well-formed, so the model scores it and the wrong answer shows up only in the output.
In a ride-hailing dispatch platform, why is the same driver feature stored twice — in a columnar history and in a key-value row?
basics
~20 sTraining and dispatch read the same feature with opposite access patterns: a training job scans months of one column across millions of rows, while dispatch fetches every feature for one driver in milliseconds. No single store serves both shapes well.
A trip driven Tuesday is uploaded Friday: in a telematics feature pipeline, what distinguishes its event time from its ingestion time?
basics
~20 sEvent time is when the driving happened; ingestion time is when the trip summary reached the feature pipeline and became readable. Devices buffer and retry, so the two differ by hours or days and records arrive out of order.
A checkout wait-time model gets a feature computed differently in the serving path than in training - why is there no error?
basics
~20 sA model validates nothing about what its inputs mean: any number in the right slot is scored. Training-serving skew therefore lands as a confidently wrong estimate rather than an exception, and shows up only when real outcomes arrive.
In a player lifetime-value model that reads a stored history embedding per player, why must each stored vector carry its encoder's version?
basics
~20 sAn embedding's coordinates only mean something relative to the encoder weights that produced them. The version stamp lets the serving path prove the vector came from the encoder the value model was trained against; without it, a mismatch is silent.
What does batching audio chunks into one accelerator pass buy a live captioning service, and what does it cost?
basics
~20 sBatching amortises the accelerator's large fixed per-pass cost over many chunks, so throughput rises several-fold. It costs latency: the first chunk in a batch waits for the batch to fill or the window to expire before any scoring starts.
A bidder's 75 ms internal p99 budget is decode 5, feature fetch 25, scoring 25, post-processing 10 and 10 ms reserve, and the measured p99s are 4, 38, 22 and 9 - which stage is over its line?
basics
~20 sFeature fetch, at 38 ms against a 25 ms line. The measured total still fits only because the other three stages came in under theirs, and the 10 ms reserve is down to 2 ms.
In a photo library that auto-tags every upload, how do nightly batch tagging, tagging on the upload event, and tagging on first view differ?
basics
~10 sThey differ in when the scorer runs. A nightly job tags the stored library on a schedule, event tagging scores each photo as it arrives, and first-view tagging scores only photos somebody actually opens.
A legal-document classifier caches finished labels; why does keying on upload id nearly never hit while a content hash does?
basics
~20 sAn upload id is unique to one upload, so every lookup misses. A content hash collapses the identical templates and boilerplate clauses that recur across customers onto one key, and repeats are the only thing a prediction cache can serve.
A keyboard's next-word strip must never be empty; through which prediction sources would you fall back, and what does each rung give up?
basics
~20 sFall back down a ladder of prediction sources: the personalised scorer, then a cached prediction for this prefix, then a global next-word frequency list, then a constant on-device default. Each rung gives up freshness, then personalisation, then the prefix context.
An inbound spam filter is live but no message has a confirmed verdict yet - which signals can you chart today?
basics
~20 sEverything the model saw and emitted is available immediately: per-feature summary statistics and null rates, the prediction-score histogram, the blocked rate at the current decision cutoff, and traffic volume - each cut by tenant, inbound path and language.
A predictive-maintenance model predicts failure within 30 days - why does this week's precision chart read optimistically high, and what corrects it?
basics
~20 sRecent cohorts are still right-censored: a failure resolves the moment it happens while survival is only confirmed at day 30, so dropping unresolved rows over-counts failures. Bucket the metric by prediction date, publish only matured cohorts, and restate as labels land.
In an inbound spam filter, why is the prediction-score histogram the first label-free signal most designs watch?
basics
~20 sIt compresses every input the model uses into a single curve the system already produces, so a change in any feature - including ones nobody thought to chart - can show up as a changed shape, for the cost of one series and no labels.
A weather-feature shift test trips daily while served forecast error is unchanged - should that test page the on-call?
basics
~20 sNo. The pager belongs to the served forecast error, the effect the operator feels. A feature whose distribution moved without moving error is a diagnostic: attach it to the quality page as context, and let it file a ticket at most.
Why does a drift test whose reference window rolls forward with recent traffic miss the slow shift it should catch?
basics
~20 sA rolling reference re-estimates normal from data that already contains the drift, so each week's comparison is against last week's already-moved distribution. Gradual movement never accumulates in the statistic, though a step change still trips it.
In a weekly-refreshed query autocomplete suggester, why must each candidate beat the live model before it replaces it?
basics
~20 sNewer training data does not guarantee a better suggester, and a refresh can be worse without failing. The model already serving is the baseline every candidate must clear on a stated rule, so promotion is earned rather than automatic.
In a music playlist ranking service, why is a completed play not proof the listener wanted that track?
basics
~20 sA completed play records what the listener tolerated in a context the service chose: autoplay, background listening and the slot given to the track all produce plays. It is a proxy label, not a verdict.
A hotel pricing model still returns a nightly rate for every request, so why does it need refreshing at all?
basics
~20 sA live pricing model keeps applying the demand pattern it learned from old bookings, and nothing errors when the market moves. Refreshing refits it on recent bookings and cancellations so its rates track today's demand, not last quarter's.
A priority-inbox ranker is refit each month on the mail recipients opened — why does each refresh harden its demotion of a sender class?
basics
~20 sDemoted mail is rarely seen, so it is rarely opened, so the next training set holds almost no positives for that sender class. Each refresh learns a lower score for it, which suppresses exposure further, and the loop tightens.
What must a playlist ranking service log at serve time so a later play can be joined to the ranking that produced it?
basics
~20 sOne impression record per serving request: a request id the client echoes on every later event, each slot's track id, position and rendered flag, the serve-time score and any randomisation propensity, plus the model and feature-spec versions.
In an offline gate for a candidate model, why must the shipping score come from a held-out data split, not the training rows?
basics
~20 sA model scored on the rows it was fitted on is being asked to recall, not predict, so that score improves with capacity and understates the errors it will make on the next order. Only rows the fit never saw estimate that.
A candidate ticket router scores live tickets in shadow mode - why can that window not show whether it routes better than the incumbent?
basics
~20 sShadow mode compares mechanics, not outcomes. The incumbent's queue choice is what agents actually work, so no ticket is ever placed where the candidate said, and resolution time and reroute rate are never produced for the disagreements.
In a sponsored-listing strip, how does team-draft interleaving compare two ranking policies within a single user's list?
basics
~20 sTeam-draft interleaving fills the strip by alternating drafts: a coin flip picks who starts, then each policy takes its highest-ranked listing not already placed. Every slot carries a hidden team tag, and a click credits the team that drafted that slot.
How do you wire the mirrored call to a shadow ticket-routing model so it cannot add latency or failures to the live request?
basics
~20 sKeep the candidate off the critical path. Answer the live request from the incumbent first, hand a copy to a bounded background worker with its own pool and timeout, and let shadow failures only increment a counter.
Why can a candidate send-time model ramped to 10% of users deliver two notifications to one user in a day?
basics
~20 sBecause the ramp assignment was not sticky. The same user fell on the candidate side in one planning run and on the incumbent side in another, so both sides queued a send for that day. Deterministic per-user assignment with one claimed send key prevents it.
In a jobs-for-you rail over four million postings, why retrieve a few hundred candidates before scoring?
basics
~20 sScoring every posting is impossible inside the rail's budget: four million postings at ten microseconds each is forty seconds of CPU per request. Retrieval trades exhaustive coverage for a few hundred plausible candidates found in milliseconds.
In a marketplace serving both a browse feed and a typed product search, what does the typed query change about the candidate set?
basics
~20 sA typed query is a retrieval constraint applied before any model runs, so the ranker only sees items that already match the stated intent. A browse feed has no such constraint and must manufacture candidates from the viewer's profile and context.
For a listener with no listening history, what decides which rung of a podcast feed's fallback ladder serves the request?
basics
~10 sA signal test on the request, not the account's age. Each rung declares a precondition; the funnel serves the first rung whose precondition holds, re-checked on every request, and logs which rung answered.
A news feed's scored candidates hold forty near-identical stories on one event - how does the list-editing layer avoid showing ten?
basics
~20 sCluster the near-duplicates first, then select greedily. Content similarity over title and lead text groups the forty filings into one event cluster; the list-editing pass then picks one representative and, using maximal marginal relevance, scores each further slot on relevance minus similarity to what is already placed.
On a marketplace search surface, why is applying an in-stock facet inside retrieval different from filtering the ranked list afterwards?
basics
~20 sInside retrieval, the shortlist comes back full of eligible items. Filtering afterwards spends the shortlist on ineligible ones, so a selective facet leaves too few results to fill the page and wastes the scoring stage's budget.
In an inline login-risk scorer, what is a per-device velocity counter and why is it kept continuously updated?
basics
~20 sA per-device velocity counter is a running count of that device's recent login attempts, held in a low-latency key-value store and folded forward by the login event stream, so the inline scorer reads it in one lookup instead of scanning history.
In a marketplace's fake-review defence, why send a middle band of risk scores to human reviewers instead of one allow/remove cutoff?
basics
~20 sA single cutoff forces every borderline listing into an automatic allow or an automatic removal. The middle band is where the score is genuinely uncertain, so routing it to an analyst buys a cheap decision instead of a costly wrong one.
Why does a card-not-present fraud model retrained on the last 30 days of checkouts learn that recent traffic is almost fraud-free?
basics
~10 sA settled dispute confirms fraud weeks after the checkout, so recent rows carry no dispute yet. Joining absence-of-dispute to a legitimate training label marks immature fraud as good and deflates the recent fraud rate.
In a refund-abuse decision layer, what changes if the rules run as a cascade before the model instead of entering it as features?
basics
~20 sA cascade makes a matching rule an absolute veto and hides the traffic it decided from the model. Rule hits as features make the rule one weighable piece of evidence the model can discount. One is policy, the other is signal.
When an allowlist entry and a hard-block rule both match one refund request, what decides the action the decision layer emits?
basics
~20 sA precedence order declared in advance - not source order, not whichever rule was authored last. The layer collects the matches, resolves them by that documented ranking, emits exactly one action, and records which branch won.
A scoring fleet takes 12,000 sensor readings per second at peak, each instance sustains 200, target utilisation 60% - how many instances?
basics
~10 sOne hundred instances. Dividing the 12,000-per-second peak by 200 gives 60 instances running flat out; dividing again by the 0.6 utilisation target gives 100, whose 20,000-per-second ceiling leaves peak sitting at 60%.
A bulk catalogue re-embedding job runs 400 accelerator-hours a month at $3 an hour for 20 million embeddings - what is the cost per embedding?
basics
~10 sSix hundredths of a cent. 400 hours at $3 is $1,200 of compute, divided by 20 million embeddings gives $0.00006 each, or $0.06 per thousand. That figure covers marginal compute only.
In a video-moderation service, which half of the spend grows with upload volume: the quarterly training run or per-clip scoring?
basics
~20 sPer-clip scoring grows with upload volume; the quarterly training run does not. Training is paid once per model version, while scoring is paid again for every clip that arrives, so only the serving half is multiplied by traffic.
A hosted invoice extractor charges per document — how do you find the monthly volume where running your own becomes cheaper?
basics
~20 sDivide the fixed monthly spend of running your own by the per-document saving it earns: crossover volume = fixed monthly cost / (hosted price per document - your own marginal cost per document). Below that volume, buying is the cheaper of the two recurring bills.
Which costs does a per-embedding figure computed from the catalogue re-embedding workers' compute bill alone leave out?
basics
~20 sThe amortised training run, the reserved-but-idle share of the accelerator fleet, vector storage and index writes, and orchestration and monitoring. On the worked catalogue refresh they turn $0.00006 of marginal compute into about $0.0006 fully loaded.
What must a decision-log row record when a risk model selects a tax return for audit?
basics
~10 sA decision-log row records the decision identifier and timestamp, the snapshot of inputs the model was given, the model version and feature-definition version, the score returned, the threshold in force, and the action taken.
A mortgage underwriting scorer is retrained from an unchanged code revision a month later and comes out different - why?
basics
~20 sA code revision pins the training instructions, not the values they read: the applications the training query returns, the resolved library versions, the run-time configuration and the seed all moved. Rebuilding a model version means pinning those inputs too.
Many advertiser tenants share one ad-screening inference tier; one tenant triples its traffic and all tenants slow down - why?
basics
~20 sEvery tenant draws on the same finite pool of workers, accelerator time and queue slots. The extra requests wait in the same line, so queueing delay rises for everyone as occupancy climbs. Nothing has to fail for this to happen.
Which must an audit-selection decision log keep to survive a feature-definition change: the raw return figures or the computed features?
basics
~20 sThe raw submitted figures survive, because they do not depend on your code. Keep them, keep the vector the model was actually fed as proof of what it saw, and stamp the feature-definition version linking the two.
Rolling the dropout early-warning model back to its previous version mid-incident: what must that rollback cover beyond the trained artifact?
basics
~20 sA model version is a pin set, not a file: the artifact plus the preprocessing and feature-definition versions it reads, its decision threshold and its output contract. Roll back all of them together, by flipping one pointer to a retained, still-loadable version.