A peak-concurrency sweep over a production session log never returns to zero — what went wrong?
answer
- the scan has a built-in checksum
- every interval contributes two opposite deltas
- what must the total be at the end?
- positive residual versus a negative dip
- unmatched starts and windows cutting sessions
basics
~20 sThe algorithm is fine; the event stream is not. A running sum ending above zero means logins without matching logouts — crashed or still-open sessions, duplicate starts, or a window that cut sessions in half.
solid answer
~50 sIn a delta sweep every session contributes `+1` and `-1`, so over a complete, well-formed log the running sum must end at exactly zero and must never dip below it. A positive residual means unmatched starts: sessions still open when the log was cut, crashed clients that never emitted a logout, or duplicated start events. A negative dip means the mirror problem — logouts whose start happened before the window you loaded. Both distort the peak, and both are silent. Treat the end-at-zero and never-negative properties as assertions the pipeline checks on every run, deduplicate by session identity before flattening, reject records where the end coordinate precedes the start, and decide an explicit policy for open sessions: clamp them to the window end, drop them, or seed the running sum with the carried-in count from before the window. Then report the peak alongside how many events were repaired.
go deeper
Know that each interval contributes one positive and one negative delta, so a complete log's running sum has to finish at zero. A non-zero residual is a signal about the input, not a reason to change the loop.
Explain both directions: a positive residual means unmatched starts, a negative dip means end events whose start falls outside the slice. Be able to say how each one skews the reported peak.
Show the diagnosis and the repair policy together: deduplicate, quarantine impossible records, seed the carried-in count at a window boundary, and decide explicitly what an unterminated session means.
Own the metric's definition. Decide what 'concurrent' means for unterminated sessions, where the failure threshold sits before a run is rejected, and how the repair counts are published so downstream consumers can trust the number.
## The symptom is the diagnostic A delta sweep has a built-in checksum that most people never use. Every well-formed interval contributes exactly `+1` and exactly `-1`. Therefore, over a complete log: - the running sum **must** be zero after the last event, and - the running sum **must never** go negative at any point. When either property fails, the scan is not wrong — the input violated the contract the scan assumes. Recognising this instantly, rather than re-reading the accumulate loop, is what separates a senior answer here. ## The two failure directions and what each means **Ends above zero — unmatched starts.** Every one of these is a `+1` with no partner: - *Still-open sessions.* Users logged in when the log was cut. Perfectly legitimate data, and the most common cause. - *Crashed or disconnected clients.* The logout event was never emitted at all. A session that should have lasted eight minutes now inflates concurrency for the rest of the window. - *Duplicate start events.* At-least-once delivery in the ingestion pipeline replays a login. The peak is inflated by the duplicate count, and the residual tells you how many. - *Filtering asymmetry.* A downstream filter (a bot exclusion, a region predicate, a schema-version check) that happens to match starts and ends at different rates. This one is nasty because the log looks intact. **Dips below zero — unmatched ends.** These are `-1` events whose `+1` lies outside the slice: sessions that began before the window you loaded. If you are sweeping a single day out of a long-running service, this is guaranteed, not exceptional. It also means your peak is *under*-counted, because the sessions carried in from before midnight are invisible to the scan. ## The repairs, and the policy question underneath them There is no universally correct fix, only an explicit choice you must be able to defend: 1. **Deduplicate first.** Group by session identity, keep one start and one end. This is the only repair that is unambiguously right, because a duplicate is not data. 2. **Reject impossible records.** End coordinate before start coordinate is either clock skew across machines or a timezone mix. Such a record makes the running sum go negative locally and can push the peak either way; quarantine and count them. 3. **Decide what an open session means.** Clamping an unterminated session to the window end says "it was still open when we stopped looking", which is right for still-live sessions and wrong for crashed ones — a crashed client's session was over long before the window ended. Dropping it says the opposite. Many pipelines split the difference with a maximum-duration cap: any session still open after `T` is closed at `start + T`, on the grounds that a real session that long is a leak. Whatever you choose, it is a *product* decision about what "concurrent" means, and it belongs in writing next to the metric. 4. **Seed the carried-in count.** For a windowed sweep, initialise the running sum with the number of sessions known to be open at the window start, instead of zero. Without that seed the first hours of every window under-report, and the daily peak metric quietly depends on where you drew the window. ## Making it permanent All of this is one extra linear pass at most, and most of it fits inside the existing scan: track `minimum` alongside `maximum`, count repairs, and emit them with the result. Good practice is to fail the job loudly when the residual exceeds a threshold — a handful of open sessions is normal, ten percent of the log is an ingestion incident — and to publish the repair counts next to the peak so a reader can tell a real traffic record from a data-quality artefact. ## What a weak answer looks like The weak answer starts debugging the accumulation: adding branches, clamping the sum at zero, or switching the tie order in the hope that it helps. Clamping is actively harmful — it hides the negative dip that was telling you sessions were carried in, and it converts a loud data problem into a quietly wrong number. The tie order matters only when two boundaries share a coordinate; it cannot make a running sum drift monotonically upward. The strong answer names the invariant first, uses the residual as a measurement of how broken the input is, and only then talks about policy — because the number that ends up on a dashboard is only as trustworthy as the event contract behind it.
- During the scan the running sum goes negative. What does that tell you?You are seeing end events whose start lies before the slice you loaded — sessions carried in from earlier. It also means the peak is under-counted, because those sessions were genuinely open. The fix is to seed the running sum with the carried-in count at the window boundary rather than starting at zero.
- What is the cheapest guard to add so this never ships silently again?Assert the two invariants in the same pass: the sum ends at exactly zero and never goes below the seeded starting count. Both are one comparison per event. Emit the residual and the repair counts alongside the peak so a reader can distinguish a traffic record from an ingestion incident.
- How do you handle a session with no logout — clamp it, or drop it?It depends on why it is unterminated, and you must say so. Clamping to the window end is right for sessions still live; dropping suits records you believe are crashed clients. A common compromise closes any session exceeding a maximum plausible duration at `start + T`. Whichever you pick, document it with the metric.
saying these in an interview costs you the question
- Starts rewriting the accumulate loop instead of checking the data
- Clamps the running sum at zero to hide negative dips
- Blames the tie-break order for a monotonic upward drift
- Assumes every window slice starts with zero open sessions
- Ignores duplicate start events from at-least-once delivery