skip to content

A dashboard your team ships runs unattended on wall displays for weeks, and the current mitigation for its steadily growing memory is a nightly location.reload(). As the lead, how do you decide whether that is an acceptable answer or whether the underlying leak has to be fixed?

level: principalimportance: nice to knowfreq 22%

answer

  1. mitigation is not the same as fix
  2. growth rate against device headroom
  3. who else runs this code path
  4. bounded cache or unbounded retention
  5. measure before anyone argues

basics

~20 s

Decide on evidence: measure the growth rate per hour against the device's headroom, and check who else runs this code. A scheduled reload is a legitimate control for a kiosk with slow, bounded growth; it is a cover-up when ordinary user sessions hit the same ceiling.

solid answer

~60 s

Turn it into three questions with measurable answers. First, what is the growth rate — megabytes per hour of realistic use, and how many hours until the device's budget is exhausted? A reload every twelve hours against a forty-eight-hour runway is an ordinary operational control; against a six-hour runway it is a tripwire waiting to fail. Second, who else runs this code path? Kiosks are the visible victim, but the same leak is usually punishing analysts with all-day tabs on lower-memory laptops, where the browser discards the tab and they lose state silently. Third, is this growth bounded or unbounded — a cache that will plateau is not the same defect as retention that scales with interactions. I would keep the reload as a guardrail either way, because unattended displays need one regardless, but require the fix when normal sessions approach the ceiling, when growth is superlinear, or when crash telemetry already shows discards. And I would fund the measurement first: without a growth rate, both positions are opinions.

go deeper

for a junior

Know that restarting a page frees its memory, and that doing so on a schedule hides a leak rather than removing it. Say that the growth needs to be measured before anyone decides.

for a middle

Explain what makes a reload adequate or not: growth per hour against available memory, and whether the growth has a ceiling. Note that a reload discards in-memory state, so it is not free in a user-facing session.

for a senior

Demonstrate the diagnostic priority — instrument first, then decide — and connect the kiosk symptom to ordinary long sessions where the same leak causes silent tab discards and lost work.

for a principal

Own the criteria and the follow-through: state the thresholds that convert the workaround into a funded fix, keep the reload as an operability control regardless, and make sure the mitigation cannot quietly become the design by attaching an alert and an expiry condition to it.

## Why this is a judgment call and not a rule "Fix the leak" is the reflexively correct answer and is often the wrong prioritisation. Leak hunting is expensive and open-ended — you cannot estimate it honestly before you have found the retainer — while a scheduled reload is cheap, reliable, and something long-running displays need anyway for other reasons (recovering from a wedged socket, picking up a deploy, resetting a stale view). The question is not whether restarting is distasteful; it is whether it is a *sufficient control* for the risk the leak actually poses. ## The three inputs **1. Growth rate versus headroom.** Convert the problem into a number: megabytes per hour under realistic use, divided into the memory the device can spare. That yields a runway in hours, and the mitigation interval must sit well inside it — with margin for the days when the data volume doubles. A nightly reload against a two-day runway is engineering; against an eight-hour runway it is a coin flip, because a busy day quietly moves the deadline. **2. Blast radius.** Wall displays are simply where you *noticed*. The same leaked code runs in every ordinary session, and the ones to worry about are the eight-hour analyst tabs on memory-constrained laptops, where exceeding the budget does not produce a tidy reload but an unloaded or crashed tab and lost unsaved state. Ask crash and reload telemetry whether that is already happening; if it is, the kiosk workaround protects the least important user and the decision makes itself. **3. The shape of the growth.** Distinguish a cache that will plateau — legitimate accumulation with a ceiling you can name — from retention that scales with interactions and has no ceiling at all. Superlinear growth is a third case and the most urgent: it usually means work is being *re-registered* per visit, so the app gets slower as well as heavier, and no reload interval is safe for long. ## What the reload actually costs A reload is not free, and the cost decides how comfortable a permanent control it is. It throws away in-memory state, forces a cold start of the app's data, produces a visible flash on a display someone is looking at, and — if it fires while a user is mid-task rather than at 4am on a kiosk — is a data-loss event. Schedule it around idleness, make it recover the view it was on, and it is unobtrusive. Ship it as a blanket timer in the user-facing app and you have converted a memory bug into a UX bug. ## The decision I would defend Keep the scheduled reload — unattended displays need a restart control regardless of this leak, and removing it after a fix would be a regression in operability. Then set explicit criteria for funding the fix: - **Fix now** if ordinary user sessions approach the ceiling, if telemetry already shows tab discards, if growth is superlinear, or if the growth rate is fast enough that the reload interval has less than roughly a 2x margin. - **Accept and monitor** if the growth is slow, confined to the kiosk profile, and the runway is comfortably longer than the reload interval — with an alert on the growth rate so that "slow" is a measured claim and not a memory of one afternoon. Either way I would spend the first day on instrumentation rather than debate: sample memory from real sessions, record it against session length, and watch discard rates. Almost every argument in this discussion collapses once someone can say "forty megabytes an hour, and analysts hit the ceiling at hour six." ## The organisational failure mode The real risk is not the reload; it is that the reload removes the *symptom* and therefore the pressure, and the leak keeps growing under the next three features. Guard against that: write down what the mitigation is compensating for, put the growth-rate alert somewhere it will be seen, and treat a widening interval — "we moved it to every six hours" — as an incident rather than a tuning change. A workaround with an expiry condition attached is engineering; an undocumented one is technical debt that has learned to hide.

  • What single number would you demand before this discussion goes any further?
    Memory growth per hour of realistic use, measured on a representative device. Divided into the available headroom it gives a runway, and the runway is what tells you whether the reload interval has margin or is a coincidence. Without it, "it grows slowly" and "it will crash" are both unfalsifiable, and the loudest person wins.
  • Why is a scheduled reload more dangerous in the user-facing app than on a kiosk?
    Because it can fire mid-task. On an unattended display nobody loses anything; in a user session a reload discards unsaved input and in-memory state and interrupts work. If you do it at all, gate it on genuine idleness, restore the route and scroll position, and never let it run while a form is dirty.
  • How do you keep the workaround from quietly becoming permanent?
    Attach an expiry condition and a monitored signal: document what the reload compensates for, alert on the measured growth rate, and treat any shortening of the interval as an incident rather than routine tuning. A mitigation with a stated threshold for escalation stays a mitigation; an undocumented one becomes the design.

saying these in an interview costs you the question

  • Says fix every leak regardless of measured impact
  • Treats a scheduled reload as a permanent architectural answer
  • Ignores ordinary user sessions running the same leaky code
  • Assumes kiosk memory limits match a typical laptop
  • Judges severity by heap size instead of growth rate

context