After a merger, two proxy feeds share one normalised action field with different meanings — what breaks?
answer
- same name, two vocabularies
- nothing errors, nothing goes quiet
- confidently wrong, not silent
- check the other fields in the same record
- enforcement mode deserves its own attribute
basics
~10 sEvery rule, dashboard and verdict reading that field silently averages two vocabularies. One feed's deny means the request was blocked; the inherited feed's deny means a monitor-mode policy matched and the traffic still completed.
solid answer
~50 sNothing errors and nothing goes quiet — the rules fire and are confidently wrong. Analysts close cases as prevented on the inherited LEEF proxy feed while the connections actually completed, and hunts filtered to allowed traffic miss half the estate. You find it by enumerating the distinct values of the field per feed and cross-checking against other evidence in the same record: a genuinely blocked request should not carry a success status and a megabyte of response body. The fix is not to pick a better collapse but to stop collapsing — keep enforcement mode as its own attribute, map monitor-mode matches to a value that means observed rather than reusing the blocked value, and retain an observer vendor and product identifier on every event so content can qualify per feed. Then re-review the closed-as-prevented verdicts from the window the wrong mapping was live.
go deeper
Know that a normalised field name does not guarantee a shared meaning, and that two products can use the same word for different things. Be ready to say why you would keep the source product on the event.
Explain why this fails silently: the value is a legal enum member, so nothing errors and no rule stops matching. Describe how enforcement mode and action are different facts that need different fields.
Walk through the diagnosis and the cleanup: per-feed value distributions, contradictions inside the same record, corroboration from another observation surface, a corrected mapping, and a bounded re-review of verdicts closed on the wrong reading.
Own the onboarding standard that prevents it across an estate of many feeds, and the call on how far back to re-open closed cases when the answer costs analyst weeks.
## The shape of the failure After an acquisition, the SOC inherits a second web proxy. Its records arrive as LEEF, with an `action` attribute whose values are allow, deny, observed and alerted. The incumbent proxy has its own action vocabulary. To get one set of dashboards, the mapping author collapses both onto a single normalised action field, matching value names that looked equivalent. They were not equivalent. On the incumbent, deny means the request was refused and never reached the site. On the inherited proxy, most of the estate runs in monitor mode: deny means the policy matched and was recorded, and the request went through anyway. This is a **semantic** defect, not a syntactic one, and that is what makes it hard. There is no parse failure — the value is a legal member of the target enumeration. No rule goes silent — content matches the field and produces results. Every dashboard renders. The system is working exactly as built and telling you something false. ## What it costs - **Verdicts invert.** A responder triaging a hit to a known-bad domain sees action deny, concludes the control prevented it, and closes the case. On the inherited feed the connection completed. - **Hunts miss.** A hunter looking for successful egress filters to allowed traffic and excludes exactly the population that was allowed but recorded as denied. - **Metrics lie upward.** A blocked-versus-allowed ratio reported to leadership now mixes enforcement with observation and reads as better coverage than exists. - **The wrongness is durable.** Because nothing failed, nobody is prompted to look. Silent absence at least eventually raises the question "why is this quiet?"; silent wrongness raises nothing. ## Finding it Three checks, in increasing cost: 1. **Value distributions per feed.** Group the normalised field by source product. Two feeds with wildly different proportions of the same value — one 40% deny, the other 2% — is not proof of anything, but it is the question that starts the investigation. 2. **Internal consistency in the same record.** A blocked request should look blocked in its other fields: no meaningful response size, a status consistent with refusal, no downstream fetch. A record claiming denial while carrying a successful status and a large body is the contradiction that settles it. 3. **External corroboration.** Correlate a sample of supposedly denied destinations against a different observation surface — DNS query records or network flow records showing bytes actually moving to that destination. Flow records carry the five-tuple, counts and timestamps and no payload at all, which is enough here: the question is only whether bytes moved. ## Fixing it properly The instinct is to remap the inherited feed's deny onto the nearest better-fitting value and move on. Do more than that: - **Model enforcement mode explicitly.** Whether a control was enforcing or observing is a fact about the event and deserves its own attribute rather than being smuggled into the action value. OCSF separates disposition from activity for this reason; ECS separates `event.action` (what happened) from `event.outcome`, whose allowed values are success, failure and unknown. - **Write down which reading you chose.** Even `event.outcome` is ambiguous for a block: success can mean the block succeeded or the fetch succeeded. A mapping that does not state its reading will be misread by the next person as surely as this one was. - **Never drop provenance.** Keep the observer vendor and product on every normalised event. Two feeds whose meanings differ can then be qualified in content, and any future divergence is addressable rather than invisible. - **Treat an absent equivalent as a schema request, not a rounding problem.** If a source has a value with no honest home in the target schema, the answer is to extend the schema or add a field, not to map it to the nearest neighbour. Enumerations and booleans are where collapse hurts most, because a wrong member is indistinguishable from a right one. ## Cleaning up behind it A corrected mapping only fixes events ingested from now on. Two things remain: - **Re-normalise history from the retained raw events** for the window the wrong mapping was live, if raw retention covers it. This is the only reason the raw copy exists. - **Re-triage the affected verdicts.** List the cases from the inherited feed closed on the reading "prevented, no impact" during that window, and work them again. Some will be genuinely fine; some are unremediated successful connections that the SOC has already told itself it handled. Scope both by the window, and say so — an open-ended re-review of everything is how this work gets abandoned. ## Onboarding discipline that prevents it When a new feed arrives, require the mapping author to enumerate every distinct value of every enumerated field, write one sentence about what each means in that product, and name the target schema value. Any row where the sentence and the target disagree is the defect, caught before content is built on it. It is an hour of work per feed, and it is the cheapest hour in the pipeline.
- How would you catch this class of defect at onboarding, before content is built on the feed?Require the mapping author to enumerate every distinct value of every enumerated field in the source, write one sentence on what each means in that product, and name the schema value it maps to. Any value whose sentence does not match its target is the defect, found before it becomes a verdict. Cross-check a sample against another field in the same record.
- The inherited feed only ever emits three of the four values. Is it safe to collapse them?No. A value that has not appeared yet is not a value that cannot appear — a configuration change on the source can start emitting it tomorrow, into content built on the assumption it never would. Collapse also destroys the ability to tell which feed produced which meaning, so keep the source vendor and product on every event regardless.
- What do you do about the cases already closed on the wrong reading?Bound it by the window the wrong mapping was live, re-normalise that window from retained raw, then list the cases from the affected feed closed as prevented and re-triage them. Some will be fine; some are completed connections the SOC believes it already handled. An unbounded re-review of everything gets started and never finished.
saying these in an interview costs you the question
- Assumes two fields with the same name carry the same meaning
- Collapses an enum to the nearest existing value to avoid a schema change
- Drops the source vendor and product once events are normalised
- Believes mapping errors always surface as parse failures
- Fixes the mapping forward and ignores verdicts already closed