skip to content

How would you set preload and worker-recycling policy across a fleet of long-lived Python services?

level: principalimportance: should knowfreq 35%

answer

  1. Three separable decisions, not one
  2. Standardise form, not numbers
  3. Someone must own the growth
  4. Restarts cost capacity somewhere
  5. Make the recycle interval a signal

basics

~20 s

Standardise the mechanism and the contract - what a preload may construct, how a worker drains before exiting, which metrics it emits - and leave the numbers per service, because a job limit follows that service's measured growth per job and its tolerance for a restart-time capacity dip.

solid answer

~50 s

Split the decision into three parts. Whether to preload at all is a per-service call: preloading buys shared pages and fast worker starts but forces every worker to reopen inherited connections and makes a bad import fail the whole pool at once, while re-importing per worker buys isolation at N times the import cost. How to recycle should be standard in form - a job count plus a memory ceiling, jittered, retiring at a job boundary - and per-service in its numbers, since growth per job is a property of the workload. What it costs the fleet is capacity: recycling always trades a little throughput for a memory bound, so the platform needs headroom for the workers that can be restarting at once. Standardise the drain contract and the telemetry, publish defaults rather than mandates, and require an owner for any growth being absorbed.

go deeper

for a junior

Know that these are policies someone sets, not defaults that appear on their own: whether the app is imported before workers are forked, and how long a worker is allowed to live before being replaced.

for a middle

Be able to state the tradeoff both ways. Preloading shares memory and speeds worker startup but hands every worker inherited state; recycling bounds growth but costs a little capacity every time a worker restarts.

for a senior

An interviewer expects you to derive the numbers rather than quote them: growth per job measured, a limit computed from it, jitter added, drain semantics defined, and telemetry that shows whether the underlying growth is getting worse.

for a principal

Own the split between standard and per-service. Fix the contract - what preload may build, what a graceful exit means, what every worker reports - leave the thresholds to the workload, budget fleet headroom for concurrent restarts, and make sure absorbed growth always has an owner.

## The question is not which setting is right; it is which decisions belong where There are three separable choices here, and treating them as one is how platforms end up with a single mandated job limit that is wrong for every service. ## First: preload or not Importing the application once in the parent and forking workers from it buys three things - - pages shared through copy-on-write, - workers that start without re-importing, - and a single place where an import error surfaces before any worker takes traffic. It costs **discipline**. Every connection, thread, pool or seeded generator built during the import is inherited by every worker and must be recreated, and the class of bug that produces is subtle and production-only. It also means the preload is a **single point of failure**: a bad import fails all workers simultaneously rather than one at a time. Not preloading inverts the trade - full isolation, per-worker reload, no inherited-descriptor discipline - at the cost of N import times, N copies of the imported heap, and slower scale-up. A service whose memory is dominated by per-request data rather than by a large static heap gains little from sharing, so the discipline buys nothing. This is a per-service judgement, and the platform's job is to make the tradeoff legible, not to pick for everyone. ## Second: how to recycle The *form* should be standard, because every service that gets it wrong gets it wrong the same way: - retire on a job count, - add a memory ceiling as a second trigger for uneven workloads, - jitter the threshold so the pool never restarts in lockstep, - and always leave at a job boundary after the in-flight work is finished and acknowledged. The *numbers* cannot be standard. A job limit is derived from measured growth per job and the memory the service is allowed, and those differ by an order of magnitude across a fleet. A platform that mandates one limit is really mandating that most services recycle far too often or far too rarely. Publish defaults with the derivation attached - here is how to measure growth per job, here is how to turn it into a limit - and let each service land its own number. ## Third: what it costs the fleet Recycling always trades throughput for a memory bound. If a pool is sized exactly to its load, it is short-handed for the whole restart, so the platform needs an explicit answer to how many workers may be recycling at once and how much headroom that requires. That answer belongs to **capacity planning**, not to individual teams, and it is the reason a per-service knob still needs a fleet-level ceiling on churn. ## Standardise the contract, not the numbers Three things are worth writing down and enforcing. 1. What the preload phase may construct: code and read-mostly data only, never sockets, sessions, threads or seeded generators. 2. What a graceful worker exit means: stop accepting, finish and acknowledge in-flight work, flush, exit with a status the supervisor treats as normal. 3. What every worker must emit: restarts per hour, jobs completed at exit, and resident memory at exit. The first prevents the inherited-state family of bugs entirely. The second makes restarts invisible to callers. The third is the one most fleets skip, and it is what keeps recycling from becoming a way to never fix anything. ## The organisational failure to design against Recycling is genuinely good engineering and it is also the most comfortable place in a system to hide a defect. A team under pressure lowers the job limit, the graph flattens, and the growth is never investigated again. Two quarters later the limit is a tenth of what it was and nobody remembers why. The countermeasure is to treat the **recycle interval as a service-level signal** rather than a configuration value: - alert when it shortens, - require an owner and a ticket for any service absorbing growth this way, - and review the numbers on a schedule. Preloading has its own version of the same trap - a memory saving that decays as workers dirty shared pages, quietly, so a capacity model built on the day-one number is wrong by the end of the first hour. ## When the interpreter's defaults move Finally, mind the platform floor. The interpreter's own defaults move: 3.14 changed the `multiprocessing` default start method to `forkserver` on Unix other than macOS, so the inherit-the-preloaded-app model now has to be requested explicitly. A fleet-wide policy that assumed the old default is a policy that silently stopped applying on upgrade, which is exactly why these decisions want to be written down with their reasons rather than inherited as folklore.

  • When would you deliberately choose not to preload a service?
    When its import graph opens connections, starts threads or seeds generators and cannot be made to stop; when memory is dominated by per-request data rather than a large static heap, so sharing buys little; when per-worker reload or per-worker plugin loading is a requirement; and when the blast radius matters more than the saving, since a bad preload fails every worker at once instead of one.
  • How do you stop worker recycling from quietly hiding a regression?
    Make the mechanism observable and owned. Export restarts per hour, jobs completed at exit and resident memory at exit, then alert on the recycle interval shortening rather than on memory alone. Require a named owner and a tracked item for any service absorbing growth this way, and review those on a schedule. A limit that drops over time is a defect report, not a tuning success.
  • What would you standardise across the fleet, and what would you leave to each team?
    Standardise the contract: what a preload may construct, what a graceful worker exit means, and which metrics every worker emits. Standardise the shape of the policy too - count plus memory ceiling, jittered, retiring at a job boundary. Leave the numbers to the service, because they follow measured growth per job and the memory budget. Publish the derivation alongside the default so teams can land their own figure.

saying these in an interview costs you the question

  • Mandates one job limit for every service in the fleet
  • Preloads everywhere without any inherited-state rule
  • Uses recycling as a permanent substitute for owning the growth
  • Ignores the capacity dip while workers restart
  • Assumes preloading is always the memory win
  • Builds a capacity model on the day-one sharing figure

context