skip to content

Why can datetime.timestamp() shift a stored epoch by hours in a translation-memory updater's sync window?

level: seniorimportance: should knowfreq 42%

answer

  1. The epoch itself carries no zone
  2. Both conversions have a silent default
  3. No tzinfo means the host's local clock
  4. One argument makes the read-back explicit
  5. Relabel with replace, do not shift

basics

~20 s

Called on a datetime whose tzinfo is None, timestamp() assumes the value is local time on the host and applies that host's UTC offset. Change the host or its TZ setting and the same wall-clock value maps to a different epoch.

solid answer

~50 s

`datetime.timestamp()` converts an instant to POSIX seconds. When `tzinfo` is set it uses that offset; when `tzinfo` is `None` it **assumes the value is local time on the host** and applies the platform's offset and DST rules. So an updater that records `datetime(2026, 9, 4, 10, 0).timestamp()` writes a different epoch on a UTC container than on a developer machine offset from UTC - the sync high-water mark moves by the offset, and the next run re-fetches or skips a window of entries. `datetime.fromtimestamp(ts)` has the mirror-image problem: with no `tz` argument it returns *local* time. The fixes are mechanical: attach the zone before converting with `dt.replace(tzinfo=timezone.utc).timestamp()`, and always read back with `datetime.fromtimestamp(ts, tz=timezone.utc)`. And never let a bare `except` swallow the resulting `ValueError` or `OSError` - that is what turns an offset bug into an invisible one.

code

python · 9 lines
python
from datetime import datetime, timezone

naive = datetime(2026, 9, 4, 10, 0, 0)
local_epoch = naive.timestamp()
utc_epoch = naive.replace(tzinfo=timezone.utc).timestamp()
print(local_epoch, utc_epoch)
print("offset applied (seconds):", utc_epoch - local_epoch)
print(datetime.fromtimestamp(utc_epoch, tz=timezone.utc))
print(datetime.fromtimestamp(utc_epoch))

go deeper

for a junior

Remember that an epoch is seconds since 1970 in UTC, and that fromtimestamp takes an optional tz argument. Get into the habit of passing tz=timezone.utc rather than accepting the default.

for a middle

Explain both silent defaults - timestamp() assuming local time when tzinfo is None, and fromtimestamp() returning local time without tz - and show the replace-then-convert fix plus the pure-arithmetic alternative.

for a senior

Diagnose it from the symptom: an incremental job re-processing or skipping a fixed window that matches a host's UTC offset. Show where conversions belong, why exceptions there must not be swallowed, and how boundary comparisons and float rounding produce the same symptom independently.

for a principal

Set the policy that removes the class: aware UTC everywhere internally, one documented storage representation for marks, conversions only at boundaries, and no host TZ configuration as a load-bearing dependency. Decide what an incremental job does when its mark is unusable, before it happens.

## The epoch has no zone; the conversion does A POSIX timestamp is a count of seconds since 1970-01-01T00:00:00Z. It is unambiguous by construction - there is nothing zone-shaped inside a float. All of the risk sits in the two **conversions** at the edges, and both of them have a silent default. **Outbound**, `datetime.timestamp()`: - if the object is aware, it subtracts the known offset and returns the correct epoch; - if `tzinfo` is `None`, the documentation is explicit that the value is *assumed to represent local time*, and the platform's zone rules for that date are applied. **Inbound**, `datetime.fromtimestamp(ts, tz=None)`: - with `tz` given, you get an aware `datetime` in that zone; - with `tz` omitted, you get a **naive local** `datetime` for the host. Neither default raises. Both are the wrong default for a service. ## How that becomes the reported symptom Picture a translation-memory updater that pulls entries changed since the last run. It stores its high-water mark as an epoch float produced by `timestamp()` on a value it built without a zone, and on the next run it re-derives a cutoff with `fromtimestamp()`. On a container with `TZ=UTC` those two conversions cancel and everything looks right - which is exactly why the bug survives the test suite. Deploy the same image to a host whose local zone is two hours ahead, or run it once from a developer laptop, and the mark is written two hours off. If it is written *earlier* than the true instant, every run re-fetches and re-embeds two hours' worth of entries it already has; if it is written *later*, entries changed inside that window are never picked up again and the memory quietly rots. Two things then hide it. First, a broad `except Exception:` around the parse-and-update step - a `ValueError` from a malformed mark, or an `OSError` from `fromtimestamp()` on an out-of-range value, is logged at debug level or not at all, so the run reports success while doing the wrong work. Second, the extra work is not fatal, only expensive: the re-fetch inflates each run until it crosses the 92nd-percentile time budget the job is measured against, and the first real signal anyone sees is a scheduling alert, not a data alert. By then the cause is three deploys back. ## The corrections **Attach the zone before converting out.** For a value you know is UTC: ```python from datetime import datetime, timedelta, timezone dt = datetime(2026, 9, 4, 10, 0) # no tzinfo epoch = dt.replace(tzinfo=timezone.utc).timestamp() # explicit, host-independent ``` `replace()` **relabels** without shifting the clock fields, which is what you want when the value was UTC all along. The pure-arithmetic form documented for the same purpose avoids the platform entirely: ```python epoch = (dt - datetime(1970, 1, 1)) / timedelta(seconds=1) ``` Both operands there are naive, so the subtraction is legal and the result is an exact duration. **Always pass tz on the way back in.** `datetime.fromtimestamp(ts, tz=timezone.utc)` returns an aware UTC value on every host. Reserve the no-`tz` form for the moment you are formatting something for a human on this machine, and nowhere else. **Do not swallow the exception.** `fromtimestamp()` raises `OSError` for values the platform cannot represent and `OverflowError` beyond the type's range; a corrupt mark raises `ValueError` on parse. Each of those should fail the run loudly or fall back to an explicitly logged safe cutoff. A silent `except` here converts a bounded correctness bug into an unbounded one. ## Precision and boundaries `timestamp()` returns a float. For present-day epochs a double still resolves well under a microsecond, but the round trip through `fromtimestamp()` rounds to the nearest microsecond, so `fromtimestamp(dt.timestamp()) == dt` is not something to rely on at the last digit. If the mark is compared with a strict `>`, a one-microsecond wobble at the boundary re-selects or drops the boundary row - the same visible symptom as a zone bug, from a different cause. Two habits remove it: store the mark as an integer count of microseconds, or as an ISO string with an explicit offset; and decide deliberately whether the window is half-open (`>` on the low end, `<=` on the high end) rather than discovering it from a duplicate. ## What a strong answer sounds like Name the default - naive means local, in both directions - then show the two-line fix, then say where the conversion belongs: at the process boundary, once, with everything in between aware and UTC. Finish on the operational half: an exception around a time conversion is a signal about your data, and swallowing it is how a two-hour offset becomes a quarter of a percentile budget.

  • Where in a service should conversions between epochs and datetimes live?
    At the process boundary, and only there: parse inbound epochs once with an explicit `tz`, keep everything internal as aware UTC `datetime` objects, and convert back only when writing out or rendering. Conversions scattered through business logic multiply the number of places a default can silently apply, and make the offset bug impossible to locate from a single stack trace.
  • Why is a swallowed exception around this conversion worse than letting the job crash?
    A crash is bounded and visible: the run fails, the mark is not advanced, and someone reads a stack trace. A swallowed `ValueError` or `OSError` lets the job report success while advancing a wrong high-water mark, so the damage compounds every run and the first symptom is a latency or cost alert rather than a data alert. Fail loudly, or fall back to an explicitly logged cutoff.
  • Does dt.timestamp() round-trip exactly through datetime.fromtimestamp()?
    Not reliably at the last digit. `timestamp()` returns a float and `fromtimestamp()` rounds to the nearest microsecond, so equality at microsecond resolution is not guaranteed. If exactness matters, store integer microseconds - `(dt - epoch_start) // timedelta(microseconds=1)` - or an ISO string with an explicit offset, and choose half-open window boundaries so a boundary row cannot be duplicated or skipped.

It is like writing '10:00' on a shipping label without the city: correct in the warehouse that wrote it, and two hours wrong the moment another warehouse reads it back.

saying these in an interview costs you the question

  • Says a naive datetime is treated as UTC by timestamp()
  • Uses fromtimestamp with no tz and calls the result UTC
  • Thinks the epoch float itself carries a time zone
  • Reaches for replace(tzinfo=...) on a value that was local time
  • Wraps the conversion in a bare except and continues
  • Sets the container TZ variable instead of fixing the code

context