skip to content

Your preloaded metrics-scraper workers grow past their memory budget every day; how do you bound it?

level: seniorimportance: should knowfreq 45%

answer

  1. Bound the lifetime, not the growth
  2. Turn a rising line into a sawtooth
  3. Never all at the same job count
  4. Leave between batches, not during one
  5. The parent's own growth is re-inherited

basics

~20 s

Cap each worker's lifetime. After a set number of scrapes, or above a resident-memory ceiling, the worker stops taking new work, finishes the batch in flight, and exits so the supervisor forks a replacement. Jitter the threshold, and remember recycling bounds a leak rather than fixing it.

solid answer

~50 s

Recycling converts unbounded growth into a sawtooth with a known ceiling. Count completed scrapes in the worker and exit after roughly N of them, with a random jitter of ten percent or so, otherwise every worker crosses the line on the same batch and the scrape stalls all at once. A resident-memory check makes a good second trigger, since growth per job is rarely uniform. Exit only between jobs: finish and acknowledge the 6,800-row batch in hand, flush anything buffered, then exit so the supervisor forks a fresh worker from the preloaded parent. Size the pool so one worker restarting does not drop the scrape interval. Then keep hunting the cause, because recycling only buys time — and note that a leak living in the preloaded parent's own state is re-inherited by every replacement, so recycling never touches it.

code

python · 18 lines
python
import os
import random

MAX_BATCHES = 500

def scrape_one_batch(rows=6800):
    return sum(range(rows))

def worker():
    limit = MAX_BATCHES + random.randrange(MAX_BATCHES // 10)  # jitter
    for _ in range(limit):
        scrape_one_batch()          # finish the batch before retiring
    os._exit(0)                     # supervisor forks a replacement

if os.fork() == 0:
    worker()
os.wait()
print('replacement worker would be forked here')

go deeper

for a junior

Know the idea and the vocabulary: a long-lived worker can be retired after a set amount of work so a fresh one takes its place, which keeps memory use bounded even when something is leaking.

for a middle

Explain the mechanics: a job counter or memory ceiling checked between jobs, a clean exit rather than a kill, and a supervisor that forks a replacement. Be ready to say why a fixed limit shared by every worker causes a synchronised restart.

for a senior

An interviewer expects a lived-through account: how you derived the limit from measured growth per job, why you added jitter, how in-flight work is drained, what headroom the pool keeps during restarts, and what you did afterwards to find the real cause.

for a principal

Own the risk that recycling becomes permanent. Decide what the fleet is allowed to absorb this way, require the growth rate itself to be a tracked signal, and be explicit that a leak in preloaded parent state is invisible to the whole mechanism.

## What recycling actually buys you A scraper worker that grows a little on every batch has a runtime bounded only by the machine. Recycling replaces that open-ended line with a **sawtooth**: growth from a known floor to a known ceiling, then a fresh process back at the floor. The ceiling is a number you choose, which is the entire value of the technique — memory stops being a question about how long the process has been up and becomes a question about how many jobs it has done. This is a **containment mechanism, not a repair**, and stating that distinction is usually what an interviewer is listening for. ## Picking the trigger - A **count of completed jobs** is the predictable trigger: with 6,800-row batches whose per-batch growth you have measured, the arithmetic from growth-per-batch to a batch limit is direct, and it makes worker lifetime proportional to work done rather than to wall-clock time. Its weakness is that growth per job is rarely uniform — one unusually wide scrape target can outgrow twenty normal ones. - A **resident-memory ceiling** covers that: read the process's own usage, and when it crosses the line, retire the worker at the next job boundary. `resource.getrusage(resource.RUSAGE_SELF)` gives a peak in its `ru_maxrss` field, which is a high-water mark rather than a current reading, so a service that wants a live number generally reads the operating system's own per-process accounting. Use both triggers: the count keeps the sawtooth regular, the ceiling catches the outlier. ## Jitter is not optional If every worker is created at the same moment with the same limit and consumes work at roughly the same rate, they all reach the limit within the same few seconds. The scrape then drops to zero throughput while the whole pool restarts, the metrics targets see a synchronised burst of reconnections, and the sawtooth appears on the service's own capacity graph as a periodic cliff. Adding a **random offset of about ten percent** to each worker's limit at startup spreads the restarts across a window and turns the cliff into noise. The same reasoning applies to any memory ceiling, since correlated workloads reach it together too. ## Retire gracefully, at a job boundary The worker must decide to leave, not be killed. 1. It stops pulling new work, 2. finishes and acknowledges the batch in hand, 3. flushes buffered output and any partial write, 4. and only then exits — and the supervisor treats that exit as normal and forks a replacement. Signalling from outside works the same way: a `signal.SIGTERM` handler sets a flag that the job loop reads between batches, rather than raising in the middle of one. Anything that terminates a worker mid-batch either loses that batch or requires the work source to redeliver it, which is a much stronger requirement on the pipeline than simply waiting a few seconds. Capacity has to be planned for it too: if the pool is exactly sized, the pool is short-handed for the whole restart, so either keep headroom for the number of workers that can be recycling at once, or cap concurrent recycles explicitly. ## What recycling does not fix Three things in particular. 1. First, growth that lives in the *preloaded parent* — a cache the parent keeps, an interned registry it appends to — is re-inherited by every replacement, so the sawtooth floor rises even though each worker looks well behaved. 2. Second, resources that are not process-local: file handles, connections to the scrape targets, or kernel-side buffers exhausted at the host level survive the churn or, worse, are recreated faster than they are reclaimed. 3. Third, growth faster than the recycle interval; if a worker can reach the ceiling inside one batch, no lifetime cap saves it. And recycling that is never revisited **institutionalises the leak**: the honest posture is to bound the growth now and own the cause afterwards, with `tracemalloc` snapshots taken at batch boundaries inside one long-lived worker and diffed to see which allocation site keeps growing. ## Instrument the mechanism itself Export: - restarts per hour, - jobs completed per worker, - and resident memory at exit. Those three turn recycling from a hidden crutch into a signal: a recycle interval that shortens week over week is a regression report, delivered before the fleet notices, and it is the number that tells you whether the underlying growth is being fixed or merely absorbed.

  • Why jitter the recycle threshold instead of using one fixed job count?
    Because workers created together and fed the same rate reach a fixed limit together. Throughput drops to zero while the whole pool restarts, the scrape targets see a synchronised reconnection burst, and the service's capacity graph grows a periodic cliff. A random offset of around ten percent per worker, applied at startup, spreads the restarts over a window so the pool always has most of its capacity.
  • What kinds of growth does worker recycling fail to bound?
    Growth held by the preloaded parent, which every replacement re-inherits, so the sawtooth floor climbs. Resources that are not process-local, such as host-level file handles or connections to the scrape targets. And growth fast enough to hit the ceiling inside a single batch, where no lifetime cap helps. Recycling assumes a slow, per-job leak in per-worker state; outside that shape it hides the problem instead of bounding it.
  • How do you retire a worker without losing the batch it is currently processing?
    Let the worker decide. It stops pulling new work, completes and acknowledges the batch in hand, flushes anything buffered, then exits with a status the supervisor treats as normal so a replacement is forked. External signals feed the same path: a `signal.SIGTERM` handler sets a flag the job loop checks between batches. Terminating mid-batch is only acceptable if the work source can redeliver, which is a much stronger requirement on the pipeline.

Recycling a worker is replacing a leaky bucket on a schedule instead of patching it: the floor stays dry, but nobody has fixed the hole, and the schedule only holds while the leak stays the same size.

saying these in an interview costs you the question

  • Kills the worker mid-batch and calls it a graceful restart
  • Gives every worker the same exact job limit
  • Treats recycling as a fix rather than a bound
  • Assumes a fresh fork resets memory the parent is holding
  • Sets a job limit without measuring growth per job
  • Forgets the capacity dip while workers are restarting

context