An inventory sync hits intermittent timeouts partway through a 6,800-row batch — what do you log, re-raise or swallow?
answer
- There are two scopes, not one
- Absorbed is not the same as ignored
- A threshold decides the run's verdict
- Deferred rows need a durable home
- Exit status must not lie to the scheduler
basics
~10 sDecide per scope: a single row's timeout is absorbed, counted and logged once at WARNING; the batch raises a failure once deferred rows cross an agreed threshold. Never exit successfully after swallowing everything.
solid answer
~50 sSplit the decision by scope. **Per row**, an intermittent timeout is a partial failure you can absorb: catch the narrow exception type, record the row identity, append it to a deferred list and continue — one WARNING per row, with a traceback on the first occurrence only, since 6,800 identical tracebacks tell you nothing the first one did not. **Per batch**, the run has a verdict: if deferred rows exceed the threshold you agreed, raise or return a failure so the scheduler marks the run red, and log that once with `logging.exception` at the boundary that turns it into an exit status. The killer anti-pattern is a job that catches everything, logs, and exits zero — the failure exists only in a log nobody reads while the scheduler reports success. Also decide what the caught state means for correctness: absorbed rows must be recorded somewhere durable, not merely logged.
code
python · 25 linesimport logging
logging.basicConfig(level=logging.INFO)
log = logging.getLogger("inventory.sync")
def push(row):
if row % 1700 == 0:
raise TimeoutError(f"row {row} timed out")
def sync(rows):
deferred = []
for row in range(rows):
try:
push(row)
except TimeoutError:
log.warning("row %d deferred", row, exc_info=not deferred)
deferred.append(row)
if len(deferred) > rows // 100:
raise RuntimeError(f"{len(deferred)} of {rows} rows failed")
return deferred
print(len(sync(6800)))go deeper
Recall that catching an error per item lets a loop continue, and that the error still has to be recorded somewhere. Know that a job which swallows everything looks successful to whatever started it.
Explain catching the narrow timeout type per row, collecting the failures, and deciding the batch verdict from a count. Be able to say why the per-row log is a WARNING and the boundary log is the one with a traceback.
Show the two-scope decision with a threshold between them, a durable home for deferred rows, proportional traceback volume, and an exit status that tells the truth. Mention what the caller should assume about partial progress.
Own the policy: who sets the acceptable failure rate, how partial success is represented across jobs, and how absorbed-failure metrics feed error budgets so an intermittent condition is caught as a trend rather than as a first outage.
## The question behind the question "Log, raise or swallow" is not one decision here; it is two decisions at two scopes, and interviewers ask this scenario because candidates who answer at only one scope get it wrong in a characteristic way. **Row scope.** One row of 6,800 timed out. The remaining 6,799 are fine. Aborting the whole run over one intermittent failure wastes the work already done and, if the sync is not idempotent, may make a partial write permanent. So the row-level failure is *deliberately swallowed* — but swallowed means absorbed and accounted for, never ignored. **Batch scope.** The run as a whole either succeeded or did not, and something above — a scheduler, an orchestrator, a human — needs that verdict. That is where an exception or a non-zero exit status belongs. ## Getting the row scope right Three obligations, in order. 1. **Catch narrowly.** Catch the timeout type you expect, not `Exception`. A `MemoryError`, a `KeyboardInterrupt` or a bug in your own transform is not a deferrable row; absorbing it turns a real defect into a silently skipped record. 2. **Record the row, durably.** Append the row's identity to a list you act on: re-queue it, write it to a dead-letter table, return it to the caller. A log line is *not* a record — nothing reads it back, and nothing retries from it. This is the difference between swallowing and handling. 3. **Log proportionally.** One WARNING per deferred row with the row key is useful; 6,800 tracebacks are not. A workable shape is: full `exc_info` on the first failure of each exception type, then message-only lines, then a single summary. If the failure rate is high, the per-row lines are pure noise and the summary is the whole signal. Note the level. These rows failed and were absorbed by design, so `WARNING` is honest; using `logging.exception()` here would label a handled condition as an ERROR and train everyone to ignore ERRORs. ## Getting the batch scope right The batch needs a **threshold decided in advance**: how many deferred rows out of 6,800 still count as a successful run? Four is noise; four hundred means the upstream system is down and pretending otherwise is a lie told to the scheduler. Encode the threshold in code, and when it is crossed, *raise* — with a message that states the counts, chained from a representative failure so the traceback is not empty of cause. Above that, exactly one handler at the top of the job logs the failure with its traceback and translates it into the outcome the environment understands: an exit status for a cron job, a task-failed signal for a worker, a 5xx for a request. That is the single `logging.exception()` in the whole flow. ## The failure modes this scenario is testing - **The green-but-broken job.** `except Exception: log.exception(...)` around the entire `main()` followed by a normal return. The traceback is in the log; the scheduler recorded success; nobody looks. Any catch-all at the top must end in a failing exit status unless it is genuinely recovering. - **Log-as-storage.** "We log the skipped rows, we can grep them later." Grep is not a retry mechanism, and log retention is usually shorter than anyone assumes. If the rows matter, they go in a durable store. - **Traceback flooding.** Per-row `logging.exception()` on a large batch can dominate the log volume and, with a synchronous handler, materially slow the run. Rate-limit or summarize. - **Losing partial progress.** If you absorb 4 rows and re-raise at the end, state what the caller should assume about the 6,796 that succeeded. Silence forces the operator to guess between "nothing happened" and "most of it happened". - **Wrong level inflation.** Marking absorbed, expected, retriable conditions as ERROR is how a team ends up with an alert channel it mutes. ## What a strong answer sounds like Name the two scopes, put a threshold between them, insist the absorbed rows land somewhere durable, keep exactly one traceback-bearing log at the boundary, and make sure the process's exit status tells the truth. Then add the operational detail: an intermittent timeout that trips a handful of rows every run is itself a signal worth a metric even when the run is green, because the day it stops being intermittent you want the trend, not a first alert.
- How many tracebacks should a run that defers 400 of 6,800 rows write?Roughly one per distinct failure mode, not 400. Emit full exc_info on the first occurrence of each exception type, message-only lines with the row identity afterwards, and one summary at the end with the counts. Beyond the noise, per-row traceback formatting on a synchronous handler can dominate the run's wall time, and 400 near-identical blocks bury the one detail that differs.
- Why is a log line not an acceptable record of a skipped row?Because nothing reads it back. Retries, reconciliation and reporting all need a queryable store — a dead-letter table, a re-queue, or a returned list the caller acts on. Logs are also retained for a shorter window than most people assume and are usually not writable by the process that would replay the work. Log for humans; store for machines.
- The batch finishes with 4 deferred rows, under threshold. What should the caller be told?That the run succeeded and that 4 rows are outstanding, with their identities available. Return the deferred list or record it, and emit a metric for the count even on green runs, so the trend is visible before the day it crosses the threshold. Reporting only pass or fail throws away the early warning that the intermittent failure is becoming persistent.
saying these in an interview costs you the question
- Aborting the whole batch on one row's timeout
- Catching Exception around the per-row work
- Logging skipped rows instead of storing them
- Catch-all at the top that still exits zero
- A traceback per failed row on a large batch
- Logging absorbed, expected failures at ERROR