skip to content

The same commit produces byte-different OS packages on two build machines - what non-determinism do you hunt for?

level: middleimportance: should knowfreq 47%

answer

  1. something varies that is not an input
  2. work through it axis by axis
  3. clock, path, order, environment
  4. archives record ownership and modes too
  5. diff the unpacked trees, not the digests

basics

~20 s

Hunt the values that differ between machines but are not build inputs: embedded build timestamps, absolute workspace paths, filesystem file ordering, locale and timezone, hostname and username, and the uid, gid and permission bits recorded in the archive.

solid answer

~60 s

I work through the axes one at a time. **Time**: build dates baked into headers, archive member mtimes, and compressor metadata that stores a timestamp. **Path**: the absolute workspace directory ending up in debug information, in generated config, or in an embedded build banner. **Order**: file lists built from directory iteration or shell globbing, which differ by filesystem and by inode order, plus parallel steps that append to an output in completion order. **Environment**: locale changing sort order and date formatting, timezone shifting a rendered date, hostname and building user embedded in metadata. **Identity metadata**: uid, gid and permission bits captured into the package. The fixes are normalisations, not post-processing: derive one fixed timestamp from the commit and feed it everywhere, sort every file list explicitly, force a canonical locale and UTC, remap build paths to a fixed prefix, and set archive ownership to a constant. To find them, I build twice in deliberately different environments and diff the unpacked trees member by member rather than comparing only the final digest.

go deeper

for a junior

Know that a build can bake in the current time and the directory it ran in, and that those are the first two things to look for when two builds of one commit differ.

for a middle

Be able to walk the axes without prompting - time, path, ordering, locale and timezone, ownership and modes, concurrency, compressor metadata - and name the normalisation for each.

for a senior

Demonstrate the diagnostic loop: amplify the suspected differences deliberately, diff unpacked trees member by member, fix at the point of entry, then guard it with a double-build check in the pipeline.

for a principal

Decide where this investment stops. Byte-level normalisation is worth real effort for artifacts someone must re-derive or audit, and rarely worth it for an internal service image, so set that boundary explicitly rather than mandating it everywhere.

## The shape of the problem When one commit yields two different packages on two machines, nothing mysterious is happening. Something that varies between the machines got captured into the output, and it was not one of the declared inputs. The job is to enumerate the ways ambient state leaks into bytes, and to remove each one at the point where it enters rather than patching the finished artifact. ## The axes, in the order they usually bite **Time.** The largest single source. A build stamps the current date into a package header or changelog, an archive stores an mtime per member, a compressor writes a timestamp into its own header, and generated code embeds a build date for a version banner. The convention that solved this broadly is a single environment-supplied fixed timestamp, `SOURCE_DATE_EPOCH`, carrying a UNIX time that build tooling honours instead of reading the clock. Derive it from the commit itself, so it is a function of the inputs rather than of the run. **Path.** The absolute directory the build ran in ends up in debug information, in generated headers, in embedded assertion messages, and sometimes in a linker search path recorded in the binary. Two runners with different workspace roots then produce different bytes. Fix it either by always building in a fixed path, or by remapping recorded paths to a canonical prefix so the recorded string does not depend on where the checkout landed. **Order.** A file list assembled by walking a directory, or by a glob, comes back in an order the filesystem chose. Different filesystems, different creation histories and different inode allocation give different orders, which changes archive member order, link order, and sometimes the resulting code layout. Sort every list explicitly with a stable, locale-independent comparison. The same applies to any map iterated to produce output: iterate in sorted key order. **Environment.** Locale changes collation, so "sorted" is not the same sort on two machines, and it changes number and date formatting in generated text. Timezone shifts a rendered date across a day boundary. Hostname and the building user get embedded by well-meaning tooling. Pin the locale to a canonical value, pin the timezone to UTC, and strip or fix any host and user metadata. **Identity and mode metadata.** Package archives record uid, gid and permission bits per member. A build running as a different user, or under a different umask, records different values. Normalise ownership to a single constant and set explicit modes rather than inheriting whatever the checkout produced. **Concurrency and randomness.** Parallel steps that append to a shared output produce completion-order results. Anything seeded from a random source, a process id, or a temporary directory name with random characters that leaks into output does the same. Make concatenation order explicit and seed generators from a declared value. **Compression and packaging.** Compressors may store the original filename and a timestamp alongside the data, and different library versions may make different, equally valid encoding choices. Pin the compressor as an input like any other tool, and use the options that suppress stored metadata. ## The method for actually finding it Comparing final digests tells you only that the artifacts differ. Instead, build twice with the differences you suspect deliberately amplified: different workspace path, different hostname, different date, different core count, different umask. Then unpack both artifacts and compare recursively - metadata first, then member by member, then inside each differing member. Nearly always the diff points straight at one of the axes above, and the noise collapses fast: fixing the timestamp axis alone typically removes most of the differing members and leaves the interesting ones visible. Once you have a normalisation in place, keep it honest by building the same commit twice in a job that deliberately varies path and clock, and comparing. Non-determinism regresses the moment somebody adds a generator that prints the date. ## Two things to be careful about Do not solve this by post-processing the finished artifact - stripping metadata after the fact hides the leak instead of removing it, and it will reappear in the next output format you add. And do not fix the timestamp to a hardcoded constant unrelated to the source: derive it from the commit, so the value is still a function of the inputs and two different commits do not claim the same build time.

  • Why derive the fixed build timestamp from the commit rather than hardcoding a constant?
    Because you want the timestamp to be a function of the inputs. A hardcoded constant makes every build of every commit claim the same moment, which destroys useful ordering information and hides mistakes. Deriving it from the commit keeps the value stable for a given source state and different across source states, which is exactly the property the rest of the build already has.
  • Why is stripping metadata from the finished artifact a poor fix?
    It treats the symptom. The leak is still in the build, so it reappears in every new output format, in intermediate artifacts other jobs consume, and in anything the stripping pass does not know about. It also makes the artifact harder to debug. Normalise where the value enters - the generator, the archive writer, the environment - not at the end.
  • How do you keep a build deterministic once you have made it so?
    Make it a checked property rather than a claim. Build the same commit twice in one job with the workspace path, hostname and clock deliberately different, then compare the unpacked outputs and fail on any difference. Without that, the first generator someone adds that prints a date reintroduces the problem quietly.

saying these in an interview costs you the question

  • Blames the compiler before checking timestamps and paths
  • Post-processes the artifact to strip metadata instead of fixing the source
  • Forgets archives record uid, gid and modes
  • Overlooks locale and timezone changing sort and formatting
  • Compares only final digests instead of diffing unpacked trees

context