skip to content

A session API, a push fan-out, an image thumbnailer and a nightly cleanup job all run on one hosting tier by habit - which would you move, and on what evidence?

level: seniorimportance: should knowfreq 46%

answer

  1. which ceiling is being fought
  2. one workload at a time
  3. longest run, not the average
  4. idle fraction against startup budget
  5. a move buys new failure modes

basics

~20 s

Move only the workloads that are currently paying for a ceiling they fight: a run-duration limit, a first-request penalty on latency-sensitive traffic, or state the tier will not keep. Decide each workload separately on measured run length, idle fraction, statefulness and packaging unit.

solid answer

~40 s

Treat it as four separate placements, not one platform decision, and start from the ceiling each workload is actually fighting rather than from the tier being unfashionable. For each one ask: how long is its longest run, how much idle time sits between bursts, does it hold anything between requests, what artifact can it be handed over as, and who is patching the host today. The nightly cleanup job is the clearest move if it is fighting a maximum run duration; the push fan-out is short, bursty and stateless and suits an event-driven runtime; the session API is steady and latency-sensitive, so scale-to-zero is a liability rather than a saving. Where nothing is being fought, leave it: a move costs a repackaging and a fresh set of failure modes.

code

pseudocode · 18 lines
pseudocode
function placeWorkload(w, tiers):
    fit = tiers

    # a ceiling that stops a run is disqualifying, not a cost
    fit = fit.where(tier.maxRunDuration >= w.longestObservedRun)

    # the penalty after idle lands on whoever is waiting
    fit = fit.where(tier.firstRequestPenalty <= w.startupBudget)

    if w.keepsStateInProcessBetweenRequests:
        fit = fit.where(tier.instanceLifetime == "controlled by us")

    fit = fit.where(tier.accepts(w.packagingUnit))

    if fit.isEmpty():
        return "change the workload, not the tier"

    return fit.lowestOperationalLoad()

go deeper

for a junior

Know that different workloads suit different hosting tiers, and that a long job and a latency-sensitive API have opposite needs. Being able to name one criterion per workload is enough at this stage.

for a middle

Explain the criteria that separate the tiers - longest run, idle fraction and startup penalty, state held between requests, the artifact handed over - and apply them to one workload end to end.

for a senior

Show the review discipline: measured evidence per workload, a willingness to leave workloads alone, and an explicit account of what the move costs in repackaging and in new failure modes to operate.

for a principal

Frame the standard rather than the migration: how many tiers the organisation should support at once, what a team must demonstrate before adopting another, and who carries the operational cost of the extra variety.

## Start from the ceiling being fought, not the tier being unfashionable A placement review goes wrong in the same way twice: a team standardises on one hosting tier for consistency, then discovers two of its workloads are fighting that tier's limits - and the correction overshoots, moving everything onto whatever the last conference talk recommended. The disciplined version asks one question per workload: **which ceiling is this workload currently paying for?** If the honest answer is "none", leaving it alone is a result, not an omission. ## The five questions that place one workload 1. **How long is its longest run?** Not its average - the longest observed, with headroom. Averages hide exactly the runs that fail against a maximum run duration. 2. **What is its startup budget, and how idle is it?** A workload with long idle gaps can afford a tier that scales to nothing; a workload whose users are waiting on the first request cannot, and the penalty falls precisely on the traffic that arrives after quiet. 3. **Does it keep anything between requests?** If it holds state in the process, the tier's instance lifetime becomes part of the correctness argument, not just the cost argument. 4. **What can it be handed over as?** A machine's worth of setup, a container image, a handler with an entry point, or source plus configuration. Unusual host needs decide more migrations than anything else. 5. **Who patches the host today, and who should?** Real, but the weakest of the five on its own - it explains why a team wants to move, rarely which workload should move first. ## The four services, placed | Workload | Demand shape | Tier that fits | Deciding criterion | |---|---|---|---| | Session API | Steady, latency-sensitive, no state held in the process | Managed container platform or a fully managed application tier | A first-request penalty after idle lands on real users, so scale-to-zero is a liability | | Push fan-out | Bursty, many short independent units, nothing kept between them | Event-driven runtime | Short runs and long idle are exactly what that tier is shaped for | | Image thumbnailer | Bursty, seconds to minutes per item, stateless, needs a large set of native dependencies | Event-driven runtime if one item fits the ceiling and the dependencies fit the unit; otherwise a container platform | The packaging unit decides this one, not the demand shape | | Nightly cleanup | One long run, once a day | Managed container platform on a schedule, or a machine that stays up | Its longest run against the tier's maximum run duration | Two of the four have a genuine reason to move; two probably do not, and saying so is the sign of a real review rather than a migration plan looking for a justification. ## The evidence to bring - **Longest observed run**, per workload, over a period long enough to include a bad day - month-end, a retry storm, a backlog. - **The arrival shape**: how long the idle gaps are, how steep the burst is, and what latency the first request after quiet actually shows. - **Whether anything is held in the process** between requests, including caches and counters that nobody calls state. - **What lives outside the artifact today** - system libraries, native tools, files on the host - because that is where migrations stall. - **The current operational load**: who is patching, how often a host problem reaches a human, and how much of that a tier change would actually remove. ## What the move costs A tier change is a repackaging, and its price is paid twice. Once in the migration itself, and once in the new failure modes the team has to learn: an instance the platform ends on its own schedule, a ceiling that stops a run, a build convention you inherit, a first request that is slower than every other. None of that is a reason never to move. It is a reason to move the workloads that are paying a real price now, to move them one at a time, and to keep one tier's worth of habits rather than four. The other failure mode is the mirror image: deciding the tier first and then selecting the evidence that agrees. Writing the five answers down per workload before naming a target tier is a cheap guard against that, and it is what a reviewer is really testing for. ## What an interviewer is listening for - Four separate decisions rather than one platform preference. - The longest run, not the average, as the measurement that settles a duration question. - A willingness to leave a workload where it is. - Naming the cost of the move, including the operating habits the team takes on.

  • You review a workload and find no ceiling being fought. Should you still move it for consistency?
    Usually not on its own. The move costs a repackaging and a new set of failure modes, so consistency has to be worth more than that - which it sometimes is, when the estate is small enough that one tier's habits are the whole operational saving. Say which of the two you are buying.
  • Which of the five criteria is hardest to get honest evidence for before the move?
    The startup budget. Steady-state latency tells you nothing about the first request after a quiet period, so you need the arrival shape and the latency of requests that follow idle - and on the current tier that penalty may not exist at all, which means estimating rather than measuring it.
  • The thumbnailer's items each finish inside the ceiling, but it needs a large set of native dependencies. What decides its tier?
    The packaging unit. If those dependencies can be expressed inside the artifact the event-driven runtime accepts, the demand shape suits that tier well. If they cannot, a container image is the unit that carries them, and the container platform wins on a packaging constraint rather than on demand.

saying these in an interview costs you the question

  • Moves a workload because a tier is cheaper per request, with no fit check
  • Uses an average run time to prove a job fits a ceiling
  • Puts every service on one tier so the deploy story stays uniform
  • Treats the move as finished once the code builds on the new tier
  • Picks the target tier first, then gathers evidence that agrees