skip to content

Teams want to rewrite compiled dependencies in production; what standard do you set once running code stops matching reviewed source?

level: principalimportance: should knowfreq 35%

answer

  1. the gap is epistemic, not technical
  2. make the divergence discoverable
  3. silence is the worst failure
  4. fail the build on no match
  5. an owner and an expiry each

basics

~20 s

Make every rewrite reviewable, declared and loud: defined in the build alongside code, failing the build when its target stops matching, recorded in a manifest the artifact carries, and given an owner and an expiry.

solid answer

~40 s

The technical risk is manageable; the standing cost is that review, diagnosis and reproduction all assume the artifact that ran is the source someone read. So the standard restores that assumption by other means. Rewrites live in the build definition, versioned and reviewed like code, not applied by hand to a running process except under an incident procedure that leaves a record. Every rule declares the shape it targets and **fails closed**: if an upstream release stops matching, the build fails rather than quietly shipping an uninstrumented artifact. The artifact carries a manifest of what was rewritten, and the process can report it, so a diagnosis starts from the truth. Finally each rewrite has an owner and an exit condition, because the set only ever grows otherwise.

code

json · 14 lines
json
{
  "transforms": [
    {
      "id": "timing-store-lookup",
      "target": { "unit": "store.Client", "member": "lookup", "shape": "(text) -> record" },
      "moment": "post-build",
      "onNoMatch": "fail",
      "owner": "platform-observability",
      "expiresAfter": "upstream release with a timing hook"
    }
  ],
  "appliedAt": "2026-04-02T10:15:00Z",
  "matched": 1
}

go deeper

for a junior

Take away one habit: if code was changed after compilation, that fact must be written down somewhere the next reader will actually find it.

for a middle

Be able to argue why a rule that matches nothing should fail the build, and what a manifest of applied transforms gives whoever debugs the service later.

for a senior

Design the controls: rules in version control, failure on a miss, a manifest the process can report, and a load check so a malformed artifact never leaves the pipeline.

for a principal

Own the trade explicitly — which cases may buy reach at the cost of reproducibility, who may change a live process, and how every rewrite is retired rather than accumulated.

## What is actually at risk The technique works. What it costs is an assumption that everything else in the organisation rests on: **the code that ran is the code that was reviewed**. Once a rewrite is in the pipeline, a reviewer approving the source has approved something the deployment does not contain, an engineer reading the dependency is reading a body that was replaced, and a reproduction from the published artifact does not reproduce production. None of that is visible from any of the places people look. A standard for rewriting is therefore mostly a standard for **making the gap discoverable**. ## The five rules worth mandating 1. **Rewrites are code.** The rule that selects targets and the transform it applies live in the repository, go through review, are versioned with the service and are applied by the pipeline. An edit applied by hand to a running process is an incident action, not a deployment mechanism. 2. **Fail closed on a miss.** Every rule declares the shape it expects and what to do when nothing matches. The default must be to fail the build. A rule that silently matches nothing is the worst outcome available: the artifact ships, the measurement disappears, dashboards go flat and everyone reads the flatness as good news. 3. **The artifact declares what was done to it.** Ship a manifest — unit, member, transform identity, version, moment of application — inside the artifact, and let the running process report the same list. On-call then discovers the rewrite by reading the process, not by knowing the folklore. 4. **Every rewrite has an owner and an expiry.** Name the team and the condition under which it goes away: upstream ships the hook, the vendor fixes the defect, the measurement moves into first-party code. Without an expiry the set becomes permanent, and permanence is what eventually blocks a dependency upgrade nobody can explain. 5. **Prefer the earliest moment and the narrowest scope.** An edit applied in the build to one artifact is inspectable and reproducible; a hook that rewrites everything on the way in is a platform-wide dependency with platform-wide blast radius, and should be justified as one. ## What the standard trades | Choice | Favours reach and speed | Favours reproducibility and review | |---|---|---| | Where the rule lives | applied ad hoc when needed | versioned in the repository | | On no match | continue silently | fail the build | | When applied | while the process runs | in the pipeline, before release | | Scope | broad pattern over all units | one named artifact | | Lifetime | indefinite | expires on a stated condition | A principal is not choosing a column; they are choosing which cases may cross to the left and what evidence those cases must leave behind. ## The judgment calls that stay open - **Rewrite at all, or hold out for a fix?** If the need is one measurement for one release, a rewrite is cheaper than anything else. If it is a behaviour change several teams will depend on, you are maintaining a private version of somebody else's component and should say so out loud. - **How much drift is tolerable?** Stale position information in a diagnosis is a real cost, paid by whoever is debugging at three in the morning, not by whoever approved the rewrite. Decide deliberately whether that cost is accepted or engineered away. - **Who is allowed to do it in a live process?** The capability that lets you reach a process you cannot restart is also the capability that lets code change under a running system with no record. Restrict who holds it, require a written record afterwards, and require that whatever was learned is baked into an earlier moment. ## How you can tell the standard is working Ask three questions of any production service six months later. Can someone list every rewrite it carries without asking a person? Did the last dependency upgrade fail loudly rather than silently dropping a transform? Has anything on the list been **removed** because its exit condition was met? An organisation that answers yes three times has bought the reach without losing the ability to reason about what it runs.

  • Why is a rewrite that silently matches nothing worse than one that fails the build?
    A failing build is visible, localised and fixed by the team that owns the rule. A silent miss ships an artifact that looks instrumented and is not, so the absence of data is read as the absence of a problem. The failure mode arrives later, during an incident, disguised as good news.
  • When would you refuse the rewrite and carry a maintained fork of the component instead?
    When the change is broad, must alter declared members, or will outlive several upstream releases. A fork is a visible, diffable artifact with an obvious upgrade cost; a rewrite of that size is an invisible fork that breaks quietly whenever upstream shifts shape, and it hides its own size from everyone reviewing the service.
  • What do you require after someone replaces a method body in a running production process?
    A written record of what was replaced and why, an artifact-level change that reproduces it at an earlier moment before the next deploy, and confirmation the process was restarted or redeployed so nothing keeps running code that exists in no artifact. Treat it as a temporary state with a deadline.

saying these in an interview costs you the question

  • Treats a rewrite that matches nothing as a harmless no-op
  • Says reproducibility is fine because the rewriting rule is in the repository
  • Leaves rewrites permanent with no owner and no exit condition
  • Assumes a process-wide rewriting hook has the same blast radius as one edited artifact
  • Believes live replacement counts as a deployment because the change took effect