A nightly cleanup job runs about twenty minutes - why does an event-driven runtime rule itself out, and which tier fits?
answer
- duration is a property of the tier
- ask what the ceiling is first
- minutes, not hours
- stopped mid-run, not paused
- long jobs need a tier that stays up
basics
~20 sAn event-driven runtime caps how long one invocation may run, so a twenty-minute job is stopped part-way. Put it on a tier that keeps a process alive: a managed container platform running it on a schedule, or a machine that stays up.
solid answer
~40 sAn event-driven runtime bills for execution and packs many tenants onto a shared fleet, so it bounds a single invocation's wall-clock time. That bound is part of the tier's contract rather than an allowance your account can grow, and a twenty-minute run will simply be stopped mid-work, losing anything held only in memory and possibly being retried from the start. The tiers that hold a long run in one piece are plain machines and a managed container platform, either running the job on a schedule. If the job has to live on the event-driven tier anyway, it stops being one job: split it into pieces that finish well inside the ceiling, record progress per piece in a store, and make a piece safe to run twice.
go deeper
Recall that an event-driven runtime limits how long one invocation may run, and that long jobs belong on a tier that keeps a process alive. Naming the criterion is most of the answer at this level.
Explain why the ceiling exists at all - the tier bills for execution and reclaims instances on a known schedule - and describe what the platform does to a run that reaches it, including what happens to work held only in memory.
Show the production consequence: a stopped run that a trigger retries from the beginning, work that must therefore be safe to repeat, and a progress record that makes a restart resume. Say how you would detect this failure in the first place.
Frame it as a fit question for the estate: whether reshaping the job into a small batch system is worth the operational saving of staying on one tier, or whether a second tier for long-running work is the cheaper standard for everyone.
## Four hosting tiers, four different promises about time A workload has to run somewhere, and in practice there are four rungs, each of which every large provider sells under a name of its own: - **Plain machines** - you rent a virtual machine, install and supervise the process yourself, and patch the host. The process runs until you stop it or something fails. - **A managed container platform** - you hand over a container image plus a declaration of how to run it. The platform places it, restarts it, and replaces instances during maintenance. - **An event-driven runtime** - you hand over a handler. The platform starts an instance when there is work, runs the handler, and may discard the instance afterwards. **Each invocation has a maximum duration.** - **A fully managed application tier** - you hand over source or a build artifact and a little configuration. The platform builds it, runs it as a long-lived application, routes traffic to it, and patches underneath it. Only the event-driven rung puts a hard clock on a single unit of work, and that one fact decides a surprising number of placements. ## Why that ceiling exists, and why it is not a quota An event-driven runtime multiplexes many tenants' short units of work over a shared fleet and charges for execution rather than for uptime. A bounded invocation is what makes that economical: the platform can reclaim an instance at a known moment instead of holding it for an unknown one. So the maximum duration is a **property of the tier's contract**, not a per-account allowance. That distinction is worth stating precisely, because interviews probe it. A **soft quota** is a number the provider will raise on request, and the cost of hitting one is *time* - a ticket and a wait. A ceiling like maximum invocation duration is the shape of the product, and the cost of hitting it is *an architecture change*: you reshape the workload or you change tier. Providers set different ceilings, and some offer a longer-running variant of the same idea; that variant is another rung with its own limit, which is the same decision again one step over. ## What a twenty-minute run actually does there 1. The platform stops the invocation. Your code does not get a vote; the run simply ends part-way. 2. Anything held only in memory is gone. Anything already written to a store survives - which is why *where* progress is written matters more than how fast the job runs. 3. Depending on how the run was started, it may be started again from the beginning. A job that is not safe to run twice then becomes a data problem, not merely a latency problem. 4. The symptom is a run that stopped, not a message saying "too long". You find the ceiling by knowing it is there. ## Where a twenty-minute job belongs | Tier | Holds one twenty-minute run? | What you take on | |---|---|---| | Plain machine | Yes | patching and supervising the host yourself | | Managed container platform | Yes | packaging the job as an image, and surviving host replacement | | Fully managed application tier | Sometimes | these tiers are shaped around request handling, so a long run needs whatever background worker shape the tier supports | | Event-driven runtime | Not as one invocation | splitting the job and recording progress between the pieces | The straightforward answers are a scheduled run on a managed container platform, or a machine that stays up and runs the job on a timer. Both keep the run in one piece. Note the honest caveat: no tier promises a run is never interrupted, because hosts are replaced everywhere, so a long job should be restartable wherever it lives. The difference is that on three of these rungs interruption is an exception, and on the fourth it is the contract. ## If it has to be event-driven anyway - Split the work by a natural key range or a batch size, so one piece finishes well inside the ceiling. - Record a progress marker per piece in a store, so a restart resumes instead of repeating. - Make a single piece safe to run twice, because retries will happen. - Keep something that knows what is left, and can tell "finished" from "stopped half-way". That is a small batch system you now own and operate. It is a fair trade when the rest of the estate already lives on that tier and the operational saving is real; it is a poor trade when the only argument is that the job ought to be event-driven. ## What an interviewer is listening for - Naming the criterion - maximum run duration - rather than naming a product. - Not offering a bigger size as the cure for a wall-clock ceiling. - Knowing the difference between a ceiling that cannot be raised and a quota that can. - Saying what happens to the partial work, and what a retry then does to it.
- You split the job into shorter runs to fit the ceiling. What do you now have to design that you did not before?A progress record in a store, so a restart resumes rather than repeats; a piece that is safe to run twice, because retries happen; and something that can tell a finished job from one that stopped half-way. You have taken on a small batch system.
- A colleague says the job just needs more memory or a bigger size to finish inside the ceiling. When is that true?Only when the run is bound by a resource that tier scales and your work parallelises inside one invocation. Wall-clock time is often set by a downstream system's throughput or by the number of records, and neither shrinks when you buy a larger size.
- If the job moves to a managed container platform, is it safe from interruption?Safer, not safe. Hosts are replaced for maintenance and failure on every tier, so a twenty-minute run should still be restartable. The difference is that interruption is an occasional event there rather than a guaranteed ceiling on every run.
saying these in an interview costs you the question
- Says a bigger size or more memory will make a long job fit
- Treats a maximum run duration as a quota support can raise
- Assumes a stopped invocation resumes where it left off
- Splits the job into chained runs with no record of progress
- Believes any stateless job belongs on the event-driven tier