skip to content

Shared Field Schemas

Every source spells the same idea differently, so detections are written against a normalised field set, and a lossy or renamed mapping quietly changes what that field set means.

on this pageshow

explore

questions

4

Why normalise vendor logs onto a shared schema like OCSF or ECS?

level: juniorimportance: must knowfreq 62%

answer

  1. one name per concept, across products
  2. content written once, not once per vendor
  3. a mapping is somebody's assertion
  4. wire format versus event model
  5. the raw copy is the only appeal

basics

~10 s

A shared schema gives every product one field name for the same idea, so a single detection, search or pivot works across all feeds instead of being rewritten once per vendor dialect.

solid answer

~50 s

A firewall calls the client address `src`, a proxy calls it `c-ip`, an endpoint agent calls it `SourceIp`, a cloud audit trail calls it `sourceIPAddress`. Without normalisation, every detection, dashboard and pivot has to be written once per product, and correlating across products means hand-written joins between dialects. A shared schema fixes one name and one type per concept — the Elastic Common Schema's `source.ip` and `user.name`, OCSF's `src_endpoint.ip` — so content is written once and runs everywhere. The cost is that a mapping is an assertion: a human decided that this vendor field means that schema field, and if the decision was wrong or lossy, everything downstream inherits it with no error anywhere. That is why normalisation is only safe alongside keeping the original event — OCSF carries a `raw_data` attribute and an `unmapped` object for exactly this reason. The raw copy is the only thing that can be re-normalised later.

go deeper

for a junior

Be ready to name a couple of real schemas (ECS, OCSF) and one wire format (CEF, LEEF, RFC 5424 syslog) and say plainly what normalisation buys: one field name per concept so a search or rule works across every feed.

for a middle

Explain the mechanics: what a mapper does at ingest, why an event model and a wire format fail differently, and why a mapping is an untested assertion until someone compares raw events against normalised output.

for a senior

Show the operating judgment — which fields you validate first on a new feed, why raw retention is non-negotiable, and how you would catch a mapping error that produces legal-looking values and no errors at all.

for a principal

Own the tradeoff between a rich event model and a cheap flat format across an estate: what a small vocabulary costs in future investigations, and who pays when the schema is chosen for ingest convenience rather than for what detections and responders will need.

## The problem A security operations centre reads from products that were never designed to agree with each other. A network firewall writes the client address as `src`. A web proxy writes `c-ip`. An endpoint agent writes `SourceIp`. A cloud provider's control-plane audit trail writes `sourceIPAddress`. A Windows Security event does not have an obvious single field for it at all — the meaning is spread across event-specific fields. All four are recording the same fact. Without a shared vocabulary, three things get expensive fast: - **Detection content multiplies.** "Alert when this user authenticates from a new country" has to be written once per identity source, and each copy has to be maintained. - **Correlation becomes bespoke.** Joining a proxy record to an endpoint record on the same user means knowing both dialects and any case or format differences between them. - **Analysts pay a tax per source.** A responder at 03:00 should be pivoting on an address, not remembering which of four spellings this particular index uses. A shared field schema is the answer: pick one name, one type and one meaning per concept, and translate every source into it at ingest. ## What the common schemas actually are It is worth separating two different kinds of thing that get lumped together. **Event models** describe the shape and meaning of an event. The **Elastic Common Schema (ECS)** defines dotted, nested fields — `source.ip`, `destination.port`, `user.name`, `process.command_line`, `event.category`, `event.outcome` — with types and, for some fields, allowed value sets. **OCSF (Open Cybersecurity Schema Framework)** goes further: events belong to classes and categories, activities and severities are integer enumerations paired with a name, and the schema explicitly reserves an `unmapped` object for vendor attributes the mapping could not place, plus a `raw_data` attribute for the original record. **Wire formats** describe how bytes are laid out on the way in. **CEF** is a pipe-delimited header (version, vendor, product, version, event class, name, severity) followed by flat `key=value` extensions drawn from a fixed dictionary of short keys. **LEEF** is similar in spirit, with a header and delimited attributes. **RFC 5424** is the modern syslog framing: priority, version, timestamp, hostname, app-name, procid, msgid, an optional structured-data section of bracketed elements, then a free-text message. The distinction matters because their failure modes differ. A wire format's vocabulary is finite and flat: if the source has a nested object or a field with no dictionary key, it must be flattened, crammed into a generic custom slot, or dropped. An event model can hold richer structure but still only holds what the mapper chose to put there. ## A mapping is an assertion, not a fact The most important thing to say in an interview is that normalisation is a **translation performed by code someone wrote**. Every mapped field is a claim: "this vendor attribute means this schema concept." Nothing validates that claim. If the mapper puts the destination address into `source.ip`, no parser errors, no rule fails to compile, and every dashboard built on top is confidently wrong. So mappings deserve the same scrutiny as detection logic: - Validate a sample of raw events field by field against the normalised output rather than trusting the vendor's shipped mapping document. - Pay closest attention to identity, host, address and timestamp fields, because everything correlates on those, and to enumerations such as action or outcome, because a wrong enum member is silent. - Record which fields were **not** mapped. OCSF's `unmapped` makes that explicit; in other pipelines it is often invisible, and "invisible" means the field is gone. ## Timestamps deserve their own paragraph A mapper choosing the wrong time is one of the most common real defects. ECS distinguishes `@timestamp` (when the event happened), `event.created` (when the agent noticed it) and `event.ingested` (when the pipeline stored it). If a source's local, un-zoned device time is mapped onto `@timestamp` without a zone, a timeline built later can put effect before cause. A mapper can only carry what the device wrote — it cannot fix a wrong clock — but it can and should record which clock it used. ## Keep the raw record Normalisation is not reversible. Anything the mapper truncated, flattened, dropped or coerced is gone from the normalised copy forever. Retaining the original event — as `raw_data`, or in a separate raw store — is what makes the mapping a decision you can revisit rather than a decision you are stuck with. Within the raw retention window you can re-normalise history under a corrected mapping; outside it, the mapper's mistakes are permanent. ## What good sounds like "One name per concept so content is written once; a mapping is an assertion I should test, not trust; and keep the raw event because normalisation only goes one way."

  • Is CEF the same kind of thing as OCSF?
    No. CEF is a wire format — a pipe-delimited header plus flat `key=value` extensions from a fixed dictionary of short keys. OCSF is an event model: typed, nested objects with classes, categories and enumerations. You can carry an OCSF-shaped event as JSON over almost any transport; you cannot express a nested object in CEF without flattening it or losing it. Choosing CEF is choosing a small, flat vocabulary, and the vocabulary is where the loss happens.
  • Which fields would you check first when validating a new source's mapping?
    The keys everything correlates on: user, source and destination address, host name, and the timestamp with its zone. Then every enumerated field — action, outcome, disposition — because a wrong enum member produces a legal value and never errors. Compare a sample of raw events against the normalised output field by field, rather than accepting the vendor's own mapping document as correct.
  • What goes wrong if the mapper picks the wrong timestamp field?
    Cross-source timelines stop being trustworthy. ECS separates `@timestamp` (when it happened) from `event.ingested` (when the pipeline stored it); if a delayed feed is stamped with ingest time, its events land minutes or hours after events they actually preceded. A reconstruction that puts effect before cause is worse than no timeline, because it reads as evidence.

A shared schema is the customs form every shipment must be rewritten onto. It makes everything comparable at a glance — and whatever does not fit a box on the form simply is not in the country's records, no matter what was in the crate.

saying these in an interview costs you the question

  • Says normalisation is lossless if the schema is big enough
  • Treats the vendor's shipped mapping as authoritative and untested
  • Keeps only normalised events and discards the raw record
  • Confuses CEF, a wire format, with an event model like OCSF
  • Claims one schema removes any need to know the source product

context

open as a page

A CEF mapping caps process command lines at 1023 characters — what does that destroy?

level: middleimportance: should knowfreq 48%

basics

~20 s

Everything past the cap is gone with no error, so executions that differ only in their tail arrive as byte-identical strings. They dedupe into one repetitive-looking event, and the bytes that would have distinguished them never left the mapper.

open as a page

After a merger, two proxy feeds share one normalised action field with different meanings — what breaks?

level: seniorimportance: should knowfreq 38%

basics

~10 s

Every rule, dashboard and verdict reading that field silently averages two vocabularies. One feed's deny means the request was blocked; the inherited feed's deny means a monitor-mode policy matched and the traffic still completed.

open as a page

A shared-schema field rename would break forty live detections — how do you ship it?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Publish both names for a fixed, announced window, migrate the consumers you can enumerate, then remove the old name on a stated date. A silent cut breaks content nobody warned; a permanent alias quietly becomes the schema.

open as a page