skip to content

A shared-schema field rename would break forty live detections — how do you ship it?

level: principalimportance: nice to knowfreq 30%

answer

  1. the mapper is one file; consumers are many
  2. history keeps the name it was written with
  3. dual-write, announce a date, then remove
  4. redefinition is worse than rename
  5. raw retention is the only undo

basics

~20 s

Publish both names for a fixed, announced window, migrate the consumers you can enumerate, then remove the old name on a stated date. A silent cut breaks content nobody warned; a permanent alias quietly becomes the schema.

solid answer

~50 s

As the platform owner you have real consumers — detection content, saved searches, dashboards, exports, possibly a co-managing provider — and only some of them are enumerable. One technical fact shapes the whole plan: stored events keep whatever field name they were written with, so for the length of retention any search crossing the cut-over needs both names unless you reindex. A clean cut is never clean for history. So: dual-write both names, publish the change with a date and an owner list, migrate the content repository yourself where you can, then remove the old name and delete the alias on the announced date rather than letting it live forever. Keep the raw events, because re-normalising history from raw is the only genuine undo. The judgment that is not yours alone is what happens when a team cannot migrate in time: slipping the date or breaking their coverage is a risk decision, not a platform one.

go deeper

for a junior

Know that detection content and dashboards reference field names directly, so changing one in the mapping breaks whatever was written against it, and that already-stored events keep the old name.

for a middle

Explain the mechanics of a safe change: dual-writing both names for a window, the difference between rename, re-type and redefinition, and why a search spanning the cut-over needs both names until history ages out.

for a senior

Show the operating plan — enumerate consumers and owners, migrate the content you control, verify by reading the repository rather than waiting for things to stop working, remove on the announced date, and re-normalise history from raw where it matters.

for a principal

Own the call itself: alias forever versus one cut, who is told and by when, what the raw retention budget buys you in reversibility, and what happens when a team cannot migrate in time — a coverage-risk decision that is not the platform team's to make alone.

## Why this is a decision, not a task Renaming a field in a shared schema is trivial in the mapper and expensive everywhere else. The mapper is one file; the consumers are detection content, correlation logic, saved searches, dashboards, scheduled reports, downstream exports, and any partner or provider reading your normalised events. You own the first and have influence over the second. The rest you may not even be able to list. That asymmetry is what makes it a lead's call rather than a change ticket. ## Know which kind of change you are making The three kinds fail differently and deserve different plans: - **Rename.** The old name disappears; consumers reference something that is not there. This breaks **loudly** in the sense that the field is plainly absent — though what a defender actually experiences is content producing nothing, which is why the announcement matters more than the mechanism. - **Re-type.** The name stays, the type changes — a string becomes an integer, a scalar becomes an array. Some consumers keep working, some break in odd ways, comparisons and aggregations change behaviour. - **Redefinition.** The name and type both stay valid and only the meaning changes. This is by far the worst, because nothing anywhere fails and every consumer keeps running while quietly meaning something else. **A redefinition should be a new field, not a new definition of an old one.** If you take nothing else into the room, take that. ## The fact that decides the plan Events already stored carry the field name they were written with. Changing a mapping changes what arrives from now on; it does not rewrite history. So for the whole retention window, any search that spans the cut-over must reference both names, whatever you do — unless you reindex or re-normalise the historical data, which costs compute and time proportional to the store. That single fact kills the idea that a hard cut is simple. Somebody is going to be looking at a ninety-day window that straddles the change on the day after an incident starts, and they need the answer to be right. ## A plan that survives contact 1. **Inventory the consumers you can.** The content repository is enumerable — grep it. Saved searches and dashboards usually are too, imperfectly. Write the list down and name an owner for each entry; unowned content is its own finding. 2. **Dual-write.** Emit both the old and the new field for a defined window. Everything keeps working and new content can be written against the new name immediately. 3. **Announce with a date, not an intention.** Publish the mapping diff, the reason, the removal date and the migration instruction. "Soon" produces nothing; a date on a calendar produces migrations. 4. **Migrate what you own.** The platform team can and should convert the bulk of the content repository itself rather than asking forty owners to do forty small edits. 5. **Verify migration by reading the content, not by waiting.** Confirm each consumer references the new field before you remove the old one, by inspecting the repository and the saved-search inventory. Do not plan to discover stragglers by noticing what has stopped producing results. 6. **Remove on the date, and remove the alias too.** The removal is the point of the exercise. An alias kept "just in case" becomes permanent, and the next engineer inherits two names for one concept and no way to know which is authoritative. 7. **Handle history explicitly.** Either re-normalise the retained raw for the retention window, or publish, alongside the change, the guidance that searches crossing the boundary must include both names until a stated date. Say which you chose. ## Alias forever, or one clean cut? Both are defensible and the answer depends on who the consumers are: - **A permanent compatibility alias is right when the consumer is outside your control** — an external customer, a managed provider co-running the SOC, an appliance whose export format you cannot change. You cannot migrate what you do not own. - **A single coordinated cut is right when every consumer is enumerable and lives in one repository.** The alias is not free: it must be documented, maintained, and eventually removed anyway, and each surviving alias makes the next schema change harder to reason about. Estates that never cut end up with three names per concept and a schema nobody can describe. Say the tradeoff out loud in the interview rather than picking one and defending it as universal. ## Reversibility Retaining the raw events is what makes the whole change safe. Within raw retention you can re-normalise history under either mapping and rebuild the normalised store. Without raw, the mapping you shipped is permanent for everything already ingested, and a mistake found in month two is a mistake you live with. ## The part that is not yours to decide When the date arrives and one team has not migrated, the choice is to slip the date for everyone or to break that team's coverage. That is a risk decision about detection coverage, not a platform decision about a field name, and it belongs to whoever owns that risk. Bringing it to them with the list, the date and the consequences already written down is the difference between a lead and an engineer with a merge button.

  • When is a permanent compatibility alias actually the right answer?
    When the consumer is outside your control — an external customer, a provider co-managing the SOC, an appliance with a fixed export — you cannot migrate what you do not own, so the alias is the cost of the relationship. Inside an estate where every consumer is enumerable and in one repository, an alias is deferred work that makes the next schema change harder and eventually has to be removed anyway.
  • Why is redefining a field's meaning worse than renaming it?
    A rename removes the field, so consumers reference something absent and the problem is at least discoverable. A redefinition keeps the name and type valid, so every rule, dashboard and export keeps running while quietly meaning something else, and nothing prompts anyone to check. A changed meaning should get a new field name and the old one deprecated on a schedule.
  • What does retaining the raw events change about the risk of this change?
    Raw is the undo. Within the retention window you can re-normalise history under a corrected mapping and rebuild the normalised store, so a mapping decision found wrong in month two is recoverable rather than permanent. Without raw, every already-ingested event is frozen in whatever shape the mapper gave it, and the only remaining option is to reason around the defect forever.

saying these in an interview costs you the question

  • Renames the field and announces it afterwards
  • Assumes historical events update themselves when the mapping changes
  • Carries a compatibility alias with no removal date
  • Redefines a field's meaning while keeping its name and type
  • Treats the change as purely technical with no named consumer owners

context