skip to content

Your fleet of metered, sleeping field controllers reports twice an hour and the team wants persistent connections; how do you decide?

level: principalimportance: should knowfreq 42%

answer

  1. price freshness against change rate
  2. an idle connection still costs to hold
  3. who must speak first, and how often
  4. a sleeping client holds nothing open
  5. commands can ride the poll response

basics

~20 s

Price freshness against change rate before picking a transport. A device that sleeps on a metered link and speaks twice an hour cannot hold anything open, and a connection that is idle 99% of its life buys latency nobody asked for.

solid answer

~50 s

Five questions decide it. How fresh must an update be, against how often it actually changes? Must either side speak at an unpredictable moment? Can the client keep a process and a connection alive at all? Will every path between them hold a response open for minutes? And who operates the tier that holds those connections, at what cost? For controllers that wake twice an hour on a metered link, the answers are clear: a held connection spends its life idle, paying keepalives, a radio that cannot sleep and a tier that has to be run, in exchange for latency the workload does not need. Poll on wake, batch everything since the device's last position, and carry pending commands back in the poll response. Revisit the decision when freshness requirements drop to seconds or the device stops sleeping.

go deeper

for a junior

Know that polling is still a legitimate answer, and that the deciding questions are how often the data changes and how fresh it must be.

for a middle

Explain what an idle held connection actually costs — a process and radio kept awake, traffic sent purely to keep it alive, and a tier measured in concurrent connections.

for a senior

Show the operational side: the failure modes differ, a device that cannot stay awake vetoes the option outright, and command latency can ride the poll response instead of justifying a channel.

for a principal

State the decision as explicit conditions with the thresholds that would reverse it, so a change in freshness requirement or power budget triggers a deliberate revisit rather than a surprise on the bill.

The instinct to reach for a persistently held connection whenever a brief says "live" is the thing being tested here. The defensible answer starts from the workload, not the transport. ## Start from the workload For this fleet, the numbers are small and the constraints are hard: - Something to report **twice an hour** per device, plus the occasional alarm. - The device **sleeps** between readings; its radio is the largest item in its power budget. - The link is **metered**, so bytes and radio wakes both appear on a bill. - Commands to a device are rare and are not urgent to the second. - The freshness anyone has actually asked for is measured in minutes. Nothing in that list rewards sub-second delivery, and two items on it actively punish a connection that must be maintained. ## What a held connection costs when it is idle - **The device cannot sleep the way it does now.** Holding a connection means keeping a process, a socket and a radio in a state where they can receive, which is the opposite of what the power budget is built on. - **Idleness is not free on the path.** A connection carrying no bytes gets reaped by intermediaries, so something has to send periodic traffic purely to prove the connection still exists — traffic that is billed and that wakes the radio anyway. - **Someone has to run the tier.** Holding connections means capacity measured in concurrent connections rather than requests per second, plus draining them on deploys. That is an operational commitment, and on a fleet that speaks twice an hour it buys nothing the team can point at. - **The failure modes change.** A connection that looks healthy until the next write fails is harder to reason about than a request that either completed or did not. ## What polling costs here Two exchanges per hour per device, timed to wakes the device was making anyway. The per-request overhead that dominates a tight polling loop barely registers when the poll rate is set by the device's own schedule rather than by a latency target. The bill is the field sets on 48 exchanges a day, and the freshness is bounded by the wake interval, which is the same budget the reporting already lives inside. ## The conditions, stated as a table | condition | poll | hold a connection | |---|---|---| | updates rarer than the cost of holding | yes | no | | client sleeps or cannot keep a process alive | yes | no | | link is metered or power-constrained | yes | no | | something on the path will not hold a response open | yes | no | | no appetite to operate a connection tier | yes | no | | freshness needed within a second, at unpredictable times | no | yes | | either side must speak first, often | no | yes | | a burst clears many events at once, continuously | no | yes | The rows are not weighted equally. A client that cannot stay awake is a veto, not a preference: no amount of latency benefit matters if the device is powered down when the update is published. ## The command path, which is where the argument usually lands "But we need to send commands to the device" is not by itself an argument for holding a connection. A polling device receives commands **in the poll response**: it wakes, sends its position, and gets back both the events and any queued commands, acts on them, and reports on the next wake. The latency of a command is then bounded by the wake interval — the same bound the reporting already accepts. Holding a connection only wins if commands must land faster than the device wakes, and that is a requirement someone has to state and justify. ## What would change the decision 1. **Freshness collapses to seconds.** If an alarm must reach an operator within a few seconds of the sensor seeing it, the wake interval is no longer an acceptable bound and the device has to stay reachable. 2. **The device stops sleeping.** Mains power removes the veto and changes the arithmetic entirely. 3. **The event rate rises until every poll returns a batch.** Then polling is paying an exchange per batch anyway, and the gap between cycles starts costing real delivery complexity. 4. **A second consumer appears with different needs.** An operator console watching one pivot in real time is a different workload from 5,000 controllers reporting twice an hour, and it can be served differently rather than dragging the fleet's transport with it. The principal-level move is stating the decision as those conditions rather than as a preference, so that when one of them changes the team revisits the transport deliberately instead of discovering it through a bill.

  • How does a device that only polls receive a command?
    In the poll response. It wakes, sends its current position, and receives both new events and any queued commands in one answer; it acts, then reports on its next wake. Command latency is then bounded by the wake interval — the same bound the reporting already accepts — and no extra channel is needed.
  • What is the strongest single argument against polling in general?
    An update whose value decays in under a second and whose arrival time is unpredictable. Polling bounds delivery by the interval, and driving the interval to sub-second turns request volume into the dominant cost while still not beating a connection that is already open.

saying these in an interview costs you the question

  • Calls polling legacy and reaches for a held connection by reflex
  • Ignores that a sleeping device cannot keep anything open
  • Counts server CPU only and forgets the metered link and the tier to operate
  • Assumes sub-second freshness is needed when data changes twice an hour
  • Believes commands to a device require their own always-open channel
  • Treats keepalive traffic on an idle connection as costing nothing