skip to content

Renting Compute

The decisions you make renting someone else's machines: which tier, which size, bought how, and what happens when the pool is empty. Asked because each one is priced differently.

on this pageshow

questions

22

A nightly cleanup job runs about twenty minutes - why does an event-driven runtime rule itself out, and which tier fits?

level: juniorimportance: must knowfreq 70%

answer

  1. duration is a property of the tier
  2. ask what the ceiling is first
  3. minutes, not hours
  4. stopped mid-run, not paused
  5. long jobs need a tier that stays up

basics

~20 s

An event-driven runtime caps how long one invocation may run, so a twenty-minute job is stopped part-way. Put it on a tier that keeps a process alive: a managed container platform running it on a schedule, or a machine that stays up.

solid answer

~40 s

An event-driven runtime bills for execution and packs many tenants onto a shared fleet, so it bounds a single invocation's wall-clock time. That bound is part of the tier's contract rather than an allowance your account can grow, and a twenty-minute run will simply be stopped mid-work, losing anything held only in memory and possibly being retried from the start. The tiers that hold a long run in one piece are plain machines and a managed container platform, either running the job on a schedule. If the job has to live on the event-driven tier anyway, it stops being one job: split it into pieces that finish well inside the ceiling, record progress per piece in a store, and make a piece safe to run twice.

go deeper

for a junior

Recall that an event-driven runtime limits how long one invocation may run, and that long jobs belong on a tier that keeps a process alive. Naming the criterion is most of the answer at this level.

for a middle

Explain why the ceiling exists at all - the tier bills for execution and reclaims instances on a known schedule - and describe what the platform does to a run that reaches it, including what happens to work held only in memory.

for a senior

Show the production consequence: a stopped run that a trigger retries from the beginning, work that must therefore be safe to repeat, and a progress record that makes a restart resume. Say how you would detect this failure in the first place.

for a principal

Frame it as a fit question for the estate: whether reshaping the job into a small batch system is worth the operational saving of staying on one tier, or whether a second tier for long-running work is the cheaper standard for everyone.

## Four hosting tiers, four different promises about time A workload has to run somewhere, and in practice there are four rungs, each of which every large provider sells under a name of its own: - **Plain machines** - you rent a virtual machine, install and supervise the process yourself, and patch the host. The process runs until you stop it or something fails. - **A managed container platform** - you hand over a container image plus a declaration of how to run it. The platform places it, restarts it, and replaces instances during maintenance. - **An event-driven runtime** - you hand over a handler. The platform starts an instance when there is work, runs the handler, and may discard the instance afterwards. **Each invocation has a maximum duration.** - **A fully managed application tier** - you hand over source or a build artifact and a little configuration. The platform builds it, runs it as a long-lived application, routes traffic to it, and patches underneath it. Only the event-driven rung puts a hard clock on a single unit of work, and that one fact decides a surprising number of placements. ## Why that ceiling exists, and why it is not a quota An event-driven runtime multiplexes many tenants' short units of work over a shared fleet and charges for execution rather than for uptime. A bounded invocation is what makes that economical: the platform can reclaim an instance at a known moment instead of holding it for an unknown one. So the maximum duration is a **property of the tier's contract**, not a per-account allowance. That distinction is worth stating precisely, because interviews probe it. A **soft quota** is a number the provider will raise on request, and the cost of hitting one is *time* - a ticket and a wait. A ceiling like maximum invocation duration is the shape of the product, and the cost of hitting it is *an architecture change*: you reshape the workload or you change tier. Providers set different ceilings, and some offer a longer-running variant of the same idea; that variant is another rung with its own limit, which is the same decision again one step over. ## What a twenty-minute run actually does there 1. The platform stops the invocation. Your code does not get a vote; the run simply ends part-way. 2. Anything held only in memory is gone. Anything already written to a store survives - which is why *where* progress is written matters more than how fast the job runs. 3. Depending on how the run was started, it may be started again from the beginning. A job that is not safe to run twice then becomes a data problem, not merely a latency problem. 4. The symptom is a run that stopped, not a message saying "too long". You find the ceiling by knowing it is there. ## Where a twenty-minute job belongs | Tier | Holds one twenty-minute run? | What you take on | |---|---|---| | Plain machine | Yes | patching and supervising the host yourself | | Managed container platform | Yes | packaging the job as an image, and surviving host replacement | | Fully managed application tier | Sometimes | these tiers are shaped around request handling, so a long run needs whatever background worker shape the tier supports | | Event-driven runtime | Not as one invocation | splitting the job and recording progress between the pieces | The straightforward answers are a scheduled run on a managed container platform, or a machine that stays up and runs the job on a timer. Both keep the run in one piece. Note the honest caveat: no tier promises a run is never interrupted, because hosts are replaced everywhere, so a long job should be restartable wherever it lives. The difference is that on three of these rungs interruption is an exception, and on the fourth it is the contract. ## If it has to be event-driven anyway - Split the work by a natural key range or a batch size, so one piece finishes well inside the ceiling. - Record a progress marker per piece in a store, so a restart resumes instead of repeating. - Make a single piece safe to run twice, because retries will happen. - Keep something that knows what is left, and can tell "finished" from "stopped half-way". That is a small batch system you now own and operate. It is a fair trade when the rest of the estate already lives on that tier and the operational saving is real; it is a poor trade when the only argument is that the job ought to be event-driven. ## What an interviewer is listening for - Naming the criterion - maximum run duration - rather than naming a product. - Not offering a bigger size as the cure for a wall-clock ceiling. - Knowing the difference between a ceiling that cannot be raised and a quota that can. - Saying what happens to the partial work, and what a retry then does to it.

  • You split the job into shorter runs to fit the ceiling. What do you now have to design that you did not before?
    A progress record in a store, so a restart resumes rather than repeats; a piece that is safe to run twice, because retries happen; and something that can tell a finished job from one that stopped half-way. You have taken on a small batch system.
  • A colleague says the job just needs more memory or a bigger size to finish inside the ceiling. When is that true?
    Only when the run is bound by a resource that tier scales and your work parallelises inside one invocation. Wall-clock time is often set by a downstream system's throughput or by the number of records, and neither shrinks when you buy a larger size.
  • If the job moves to a managed container platform, is it safe from interruption?
    Safer, not safe. Hosts are replaced for maintenance and failure on every tier, so a twenty-minute run should still be restartable. The difference is that interruption is an occasional event there rather than a guaranteed ceiling on every run.

saying these in an interview costs you the question

  • Says a bigger size or more memory will make a long job fit
  • Treats a maximum run duration as a quota support can raise
  • Assumes a stopped invocation resumes where it left off
  • Splits the job into chained runs with no record of progress
  • Believes any stateless job belongs on the event-driven tier
open as a page

You are moving a measured search service onto rented machines — which numbers set the machine size, and which do you ignore?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Size from the service's own observed processor, memory, disk and network use over a full demand cycle, taken at a high percentile with headroom added. The specification of the server being replaced records a purchase decision, not the workload, and is ignored.

open as a page

A launch request for a new machine is refused in one zone — how do you tell an empty capacity pool from an administrative ceiling?

level: middleimportance: must knowfreq 62%

basics

~20 s

A refused launch means one of two shortages: the provider has no machine of that shape free in that zone, or your account is not permitted more. A ceiling is a number you can look up; pool depth is never published.

open as a page

Machine families are weighted toward processor, memory, local disk or accelerators — which measurements pick one?

level: middleimportance: must knowfreq 58%

basics

~20 s

The ratio in a measured profile picks the family — which resource saturates first while the others idle. Memory-heavy work wants a memory-weighted family, compute-bound work a processor-weighted one, heavy local input and output a disk-weighted one, dense parallel numeric work an accelerator.

open as a page

A reporting service has a flat weekday base and a sharp month-end peak - which purchase posture suits each part of that demand shape?

level: middleimportance: must knowfreq 62%

basics

~20 s

Buy the always-on base with a term commitment, serve the month-end peak with metered on-demand capacity, and put only interruption-tolerant extra work on reclaimable capacity. Price each layer of the shape separately rather than picking one posture for the fleet.

open as a page

How do you choose how often an hours-long training job checkpoints when its machine can be reclaimed at any minute?

level: middleimportance: must knowfreq 50%

basics

~20 s

Balance two costs that pull in opposite directions: redone work, which is about half the interval every time capacity is reclaimed, against the runtime each checkpoint write itself consumes. Measure both, then pick the interval where they are roughly equal.

open as a page

A provider can reclaim a batch job's machine with minutes of warning - what must fit inside that notice, and what does it not promise?

level: middleimportance: must knowfreq 62%

basics

~20 s

The reclaim notice buys just enough time to stop taking new work, flush a checkpoint to durable shared storage and hand the unfinished unit back. It promises no time to finish, no extension, and not even that it will arrive.

open as a page

Your payments ledger fails over into its standby zone and half its replacement machines will not launch — why is capacity scarcest exactly then?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Failover is the moment every tenant wants the same thing at once. The loss of a zone pushes many organisations onto the same surviving zones and the same popular shapes within the same minutes, so the pool the plan intended to launch into is drained by every other plan.

open as a page

Your preferred machine family is unavailable in a zone while a smaller profile launches there — what does that reveal about capacity pools?

level: middleimportance: should knowfreq 46%

basics

~20 s

Capacity is not one fleet-wide reservoir but many narrow pools, roughly one per zone per machine family and size. Scarcity therefore has an address, and a smaller or differently weighted shape often launches where the preferred one will not.

open as a page

Your session API keeps sessions in local process memory - which hosting tiers does that choice rule out, and why?

level: middleimportance: should knowfreq 58%

basics

~20 s

Local memory lasts exactly as long as one instance, so it rules out every tier where the platform, not you, decides when an instance goes away - most sharply an event-driven runtime, and in practice any tier that replaces instances on deploy.

open as a page

Moving one service from plain machines to a container platform to an event-driven runtime - what changes about the unit you hand over?

level: middleimportance: should knowfreq 52%

basics

~20 s

The unit shrinks: from a machine you install a process on, to a container image plus a run declaration, to a handler with a defined entry point. Everything outside that unit becomes the platform's decision rather than yours.

open as a page

A term commitment is billed for every hour of its term - how many run-hours a day does a workload need before it beats metered capacity?

level: middleimportance: should knowfreq 50%

basics

~20 s

Break-even is a utilisation, not a number of hours you can memorise: if the committed rate is a given fraction of the metered rate, the workload must run at least that same fraction of the hours. Compare monthly totals, never the two hourly rates.

open as a page

A team buys a multi-year spend commitment expecting machines to be waiting in its standby zone — what has it actually bought?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A term commitment is a billing instrument: it lowers the price of machines you manage to obtain and sets none aside. Guaranteeing an allocation in a named zone takes a capacity reservation, which is a separate purchase billed from the moment it exists.

open as a page

A session API, a push fan-out, an image thumbnailer and a nightly cleanup job all run on one hosting tier by habit - which would you move, and on what evidence?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Move only the workloads that are currently paying for a ceiling they fight: a run-duration limit, a first-request penalty on latency-sensitive traffic, or state the tier will not keep. Decide each workload separately on measured run length, idle fraction, statefulness and packaging unit.

open as a page

An admin console on a fractional-core tier ran fine for weeks, then crawled during a campaign — why?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A fractional-core machine accrues credit while it stays below its baseline share of a processor and spends that credit to burst above it. Weeks of idling built a balance; sustained campaign traffic drained it, and the machine dropped to its baseline share.

open as a page

Your reporting fleet never drops below forty machines of one profile and triples for three days each month - how much of it would you commit to?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Commit to the floor and meter the spike. The forty machines that are up every hour justify a committed rate; the eighty that exist three days a month run about a tenth of the hours and would cost more committed than metered.

open as a page

Every machine in a reclaimable training fleet disappeared inside one minute - why does that happen, and what spreads the risk?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Reclamation is not an independent per-machine accident: it follows the provider's demand for one machine shape in one place, so identical machines from one capacity pool share a single fate. Spreading means several shapes and several zones, and a workload able to run on all of them.

open as a page

You cannot afford held capacity in the standby zone for every service — how do you decide which workloads get a guaranteed allocation?

level: principalimportance: should knowfreq 33%

basics

~20 s

Rank workloads by what a refused launch costs in the worst hour, not by how important the owning team feels. A small set gets held allocation or warm machines; the rest get shape flexibility, a written degradation plan, and an acknowledged possibility of waiting.

open as a page

A multi-year commitment was signed weeks before a re-platforming halved the fleet it covered - what are your realistic options now?

level: principalimportance: should knowfreq 36%

basics

~20 s

Start by asking what the instrument is expressed in: a commitment tied to a machine shape strands, one expressed as spend per hour can often be consumed by other workloads. Then look for exchange, resale or renegotiation routes, and sequence the migration against the term.

open as a page

How would you set the split between guaranteed and reclaimable machines for a fleet, and which workloads would you bar from the reclaimable half?

level: principalimportance: should knowfreq 34%

basics

~20 s

Size the guaranteed baseline at the demand that must still be served when every reclaimable machine is taken back at once, and put resumable, deadline-tolerant work on the surge. Keep coordinators, sole copies of state and the recovery path itself off the reclaimable half.

open as a page

Rented machine sizes usually step by doubling — what does rounding a service up to the next size buy?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

Rounding up buys roughly double the resources for roughly double the hourly charge — a purchase, not a free safety margin. A service measured just past one rung spends most of the next rung idle, which is where a fleet's bill quietly grows.

open as a page

Some reclaimed machines vanish without ever delivering the warning - what must the work distribution do so a unit is not lost?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Hand work out under a lease that expires unless the worker keeps renewing it, so a silent disappearance returns the unit automatically. That makes delivery at-least-once, which in turn requires restarts and outputs to be idempotent.

open as a page