skip to content

questions

5

A keyboard's next-word strip must never be empty; through which prediction sources would you fall back, and what does each rung give up?

level: middleimportance: must knowfreq 64%

answer

  1. never an empty strip
  2. a ladder, not one spare
  3. each rung drops one input class
  4. cache drops freshness, list drops personalisation
  5. bottom rung needs no network

basics

~20 s

Fall back down a ladder of prediction sources: the personalised scorer, then a cached prediction for this prefix, then a global next-word frequency list, then a constant on-device default. Each rung gives up freshness, then personalisation, then the prefix context.

solid answer

~50 s

I would define four rungs and render from whichever one answers inside the strip's deadline. The top rung is the personalised scorer, using this user's typing history plus the current prefix. Below it sits a cached prediction for the same prefix and context: the same personalisation, but produced earlier and possibly by an earlier scorer version, so it gives up freshness. Below that, a global next-word frequency list keyed on the prefix alone keeps the context but loses personalisation entirely; because next-word frequencies are roughly Zipfian, a modest table covers a large share of keystrokes. The last rung is a constant list that needs no network, no feature lookup and no per-user state, which is what lets the strip promise it is never empty. Every response carries a tag naming the rung that produced it.

code

pseudocode · 14 lines
pseudocode
on keystroke(session, prefix, deadline):
    if scorerHealthy and admitted(session):
        result = scorer(session, prefix, deadline)
        if result.arrived:
            return render(result.candidates, rung = "personalised")

    cached = predictionCache.lookup(prefix, contextOf(session))
    if cached.found:
        return render(cached.candidates, rung = "cached")

    if frequencyTable.reachable:
        return render(frequencyTable.lookup(prefix), rung = "frequency_list")

    return render(onDeviceDefaults(prefix), rung = "constant")

go deeper

for a junior

Recall that a prediction path needs more than one source: a live scorer, something cached, something generic, and a constant that always answers.

for a middle

Explain what each rung gives up — freshness at the cache, personalisation at the frequency list, the prefix itself at the constant — and why that is also the quality ordering.

for a senior

Show that you check each rung for independence from the failure above it, and that you tag responses with their rung so degraded traffic stays measurable afterwards.

for a principal

Argue what quality the bottom rung must still deliver to be worth rendering, and what the product gives up by promising a never-empty strip at any quality.

## What a rung actually is A fallback tier is not a spare copy of the scorer. It is a **different way of producing the same output** — three candidate words for the suggestion strip — using different inputs, at a different cost, with a lower expected quality. Designing the ladder is therefore an exercise in ranking prediction *sources* by how much they still know about this user and this sentence, and in checking that each source survives the failure of the one above it. Three properties make a rung usable: - **Strictly cheaper** than the rung above it, in latency and in the dependencies it touches. A rung that costs about the same buys nothing on a saturated path. - **Independently reachable.** A cached-prediction rung that reads from the same online store the scorer reads its features from is not a separate rung; it is the same failure wearing a different name. - **Total at the bottom.** For any prefix at all, the last rung returns candidates. That is what turns "never empty" from an aspiration into an invariant. ## The four rungs and what each one costs | rung | what it uses | what it gives up | what makes it fire | |---|---|---|---| | personalised scorer | this user's typing history, the current prefix, app and locale context | nothing | the default path | | cached prediction for this prefix | a result produced earlier for the same prefix and context | freshness of the inputs, and possibly the current scorer version | scorer unavailable, or the request was shed | | global next-word frequency list | the typed prefix alone, aggregated over a population | personalisation to this user, entirely | the online store or the scorer is unavailable | | constant on-device list | only the letters already typed | the population signal and every piece of server state | network loss, or everything above has failed | Read top to bottom, each step removes one class of input. That is exactly why the ordering is also the quality ordering: the strip's acceptance rate — the share of rendered strips whose suggestion the user taps — falls as inputs drop away. ## The decision order on a single keystroke 1. If the scorer is healthy and admitted the request, score live and render. 2. If it is unhealthy or the request was shed, look for a cached prediction for this prefix and context. 3. If there is none, take the global frequency list for the prefix. 4. If even that is unreachable — no network at all — render the on-device constant list. The order is decided by what is *known to be unavailable*, not discovered by trying each rung in turn: walking the ladder serially spends the deadline more than once, so the slowest responses would also be the worst ones. Either decide up front from health state, or start the cheap rung alongside the scorer and take whichever answer is in hand when the deadline arrives. ## Staleness here means three different things When you report that the strip was served from the cached rung, say **which** staleness you mean. The cached entry may be stale in its inputs (the user's typing history moved on since it was produced), stale as an output (the same candidates are rendered again for a sentence that has since changed direction), or stale in the scorer version that produced it. They have different consequences: stale inputs cost a little relevance, a stale scorer version means some traffic is quietly being served by a model you thought you had replaced. ## The bottom rung is a product decision, not an engineering one A constant list is not free: it consumes the strip's real estate and the user's attention with something that ignores the sentence. Teams sometimes want the strip to go blank instead. Either choice is defensible, but it must be chosen deliberately, because it decides whether "never empty" is an invariant the rest of the design can lean on. ## Tag the rung on the way out Every rendered strip should carry the rung that produced it and the reason that rung was chosen. Without that tag the tiers are indistinguishable after the fact: the quality of the degraded path cannot be measured, an incident leaves no trace in the quality metrics, and the logs that feed later work cannot tell a personalised suggestion from a generic one. The tag costs one field and is the difference between a ladder you can operate and one you merely hope is working.

  • What makes the bottom rung different in kind from the rungs above it?
    It produces candidates with no network call, no feature lookup and no per-user state — a table that ships with the client. Every rung above it shares at least one dependency with the scorer. The bottom rung is the one that still answers when all of those are gone, which is the only reason the strip can promise it is never empty.
  • Should the ladder try each rung in turn on the same request, or pick one up front?
    Trying them in turn spends the deadline more than once: the request pays the scorer's full wait before the cheap rung even starts, so the slowest answers end up being the worst ones. Decide up front when you already know the scorer is unhealthy, or start the cheap rung concurrently and take whichever answer is in hand when the deadline arrives.
  • Where does a smaller, cheaper scorer sit on this ladder, and what decides how small it has to be?
    It sits between the personalised scorer and the non-model rungs: the same inputs, less capacity, so it can answer inside the deadline on whatever resources remain when the primary path is shed or unavailable. Its size is set by the deadline it must clear and the quality floor it must hold. How the smaller scorer is produced is a modelling question; the ladder only cares that it is cheaper and still clears the floor.

A kitchen that can never send an empty plate: cook to order, else plate this morning's prep, else the standing staff dish, else bread. Each step is faster, less personal, and never fails.

saying these in an interview costs you the question

  • Treats the fallback as one spare rather than a graded ladder
  • Puts the cached rung behind the same store the scorer depends on
  • Thinks a global frequency list still reflects this user's habits
  • Serves a fallback without recording which rung produced it
  • Assumes every rung can be tried in sequence within one deadline
open as a page

During a scorer outage the suggestion strip still fills from the frequency list and no errors fire — how would you see the quality loss?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Tag every rendered strip with the rung that produced it and the reason, then split acceptance rate by rung. A blended number absorbs the loss; the share of keystrokes served below the top rung is the signal that actually moves.

open as a page

The per-user typing-history feature is missing for a keystroke the next-word scorer must score now — what are your options, and how do you choose?

level: seniorimportance: should knowfreq 56%

basics

~20 s

Three options: impute with the exact convention the training data used, route to a variant trained without that feature group, or skip the scorer and take the next fallback rung. Never substitute an in-range value the model will read as real evidence.

open as a page

Your keyboard promises a never-empty suggestion strip; how do you decide the quality floor below which a degraded suggestion is worse than the safe default?

level: principalimportance: should knowfreq 36%

basics

~20 s

Price a wrong tap against a missed one. Express the floor as an acceptance rate the lowest rung must hold against the top rung, plus a budget for how much traffic may sit below the top rung; under the floor, render the neutral default.

open as a page

Load doubles and the personalised scorer cannot serve every keystroke inside the strip's deadline — which requests do you shed to the frequency list, and why those?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Shed the keystrokes where the rungs differ least — long, nearly unambiguous prefixes — and shed before the work queues rather than after the deadline passes. Pick the shed set deterministically per session so the strip does not flip rungs mid-sentence.

open as a page