skip to content

In a STIX 2.1 bundle, what does an indicator object claim that an observed-data object does not?

level: middleimportance: should knowfreq 52%

answer

  1. one looks forward, one looks back
  2. pattern versus first_observed
  3. a third object records the match
  4. claim, observation, sighting
  5. the meaning lives in the edges

basics

~20 s

An indicator asserts a detection pattern its author believes signals malicious activity, so it is a forward-looking claim. Observed-data asserts only that particular artefacts were seen, when, and how many times, with no claim about malice.

solid answer

~40 s

An `indicator` carries a `pattern` plus `pattern_type` and `valid_from`: its author is saying "if you see this, treat it as suspicious". It is a judgement, and it can be wrong. An `observed-data` object carries `first_observed`, `last_observed`, `number_observed` and references to the observable objects themselves: it is a factual record that these artefacts appeared, and asserts nothing about whether they were bad. The third piece people forget is the `sighting` relationship object, which says a given indicator matched somewhere, with a `count` and `where_sighted_refs`. Keep the directions straight: an indicator is a claim, observed-data is an observation, a sighting is a match. None of the three is a compromise. Why the producer thinks the domain matters lives in the `relationship` objects around it, such as an `indicates` edge to a `malware` object.

code

json · 20 lines
json
{
  "type": "indicator",
  "spec_version": "2.1",
  "id": "indicator--8e2e2d2b-...",
  "created": "2026-03-04T11:20:00.000Z",
  "name": "Consent page for look-alike mail add-in",
  "indicator_types": ["malicious-activity"],
  "pattern_type": "stix",
  "pattern": "[domain-name:value = 'secure-mail-addin.example']",
  "valid_from": "2026-03-04T11:20:00.000Z",
  "object_marking_refs": ["marking-definition--... (TLP:AMBER)"]
},
{
  "type": "relationship",
  "spec_version": "2.1",
  "id": "relationship--1c4f...",
  "relationship_type": "indicates",
  "source_ref": "indicator--8e2e2d2b-...",
  "target_ref": "malware--3f1a..."
}

go deeper

for a junior

Recall that an indicator carries a pattern the author believes is suspicious, while observed-data records that artefacts were actually seen at a time. Do not use the word IOC for both.

for a middle

Explain the fields that carry the difference: pattern, pattern_type and valid_from against first_observed, last_observed and number_observed, plus the sighting relationship with its count. Say what a producer asserts in each case.

for a senior

Demonstrate the downstream consequence: expiry driven by valid_until, refusing to push bare observables to blocking controls, and preserving the relationship edges so an analyst can later answer why a value is on a list.

for a principal

Own how the community models its content in the first place, including whether members are asked to submit observations or judgements, and what a shared sighting count is allowed to be used for.

## Three different assertions that look alike STIX 2.1 separates things most tooling smears together, and interviews probe exactly this because getting it backwards is the classic threat-intelligence error. **`indicator` — a claim about the future.** Required content is a `pattern`, a `pattern_type` (usually `stix`, the built-in patterning language) and a `valid_from` timestamp. Optional `indicator_types` says what kind of claim it is (`malicious-activity`, `anomalous-activity`, `benign`), and `valid_until` gives it an expiry. The semantics are: *the author believes that matching this pattern means something*. A producer can publish an indicator having never seen the pattern fire anywhere; it can be derived from a sample, from a report, or from a guess. It carries the producer's confidence, not truth. **`observed-data` — a claim about the past.** It carries `first_observed`, `last_observed`, `number_observed` and, in 2.1, `object_refs` pointing at the STIX Cyber-observable Objects (SCOs) that were seen: a `domain-name`, an `email-addr`, a `file`. The semantics are: *these artefacts existed here in this window, this many times*. It makes no malice claim whatsoever. A `domain-name` observable inside observed-data is not an accusation; it is a note that the name appeared. **`sighting` — a relationship, not an object about a thing.** It is an SRO with `sighting_of_ref` (the indicator or other SDO that was matched), `count`, `first_seen`/`last_seen` and `where_sighted_refs` (an identity, i.e. who saw it). Semantics: *someone's tooling matched this*. A sighting count of forty across nine member organisations raises the community's belief that the indicator is worth acting on. It does not establish that any of the nine was compromised, and it does not establish the pattern is precise; forty matches on a shared marketing domain is forty false positives. ## Why the distinction has consequences An ingest pipeline that flattens all three into "an IOC row" loses the ability to answer the questions that matter: - **Can I block this?** An indicator with a tight pattern and a relationship to a malware family, marked with a confidence, is blockable. A bare observable lifted out of observed-data is not — it may be a legitimate host the intruder abused. - **Should I still be looking for it?** `valid_from` and `valid_until` on the indicator say when the claim was believed to hold. Observed-data timestamps do not carry that meaning, so a pipeline that ages content by observation time expires the wrong things. - **Why is this here?** The reason lives in the edges. An `indicator` with an `indicates` relationship to a `malware` object, which in turn has a `uses` relationship from an `intrusion-set`, is intelligence. The same indicator with no edges is a string. ## How the two platforms model this **MISP** organises around an **event** — a container with a date, a distribution level, tags, and a list of **attributes** (typed value plus category, e.g. `domain|ip-src|md5`). An attribute has an `to_ids` flag that decides whether it is meant to be pushed to detection tooling, which is MISP's version of the indicator-versus-observation distinction. MISP **sightings** are a separate count of who reported seeing the attribute. There is no direct STIX equivalent of the event container itself; exporters typically render it as a `report` or `grouping` object, and some of the event's context becomes free text. **OpenCTI** is STIX-native: what it stores *is* the object graph above, with connectors ingesting MISP, TAXII collections and vendor feeds into it. That makes relationship-heavy work natural and makes round-tripping a MISP event through it slightly lossy in the other direction, because a graph does not want to be re-flattened into a dated container. The practical takeaway for an interview: know that a MISP attribute with `to_ids` set is close in intent to a STIX indicator, that a MISP attribute without it is closer to an observation, and that neither platform's model maps to the other without decisions being made. ## The direction errors to avoid saying out loud - "An indicator means it was observed." No — that is observed-data. - "A sighting means the org was compromised." No — it means a pattern matched. - "An IOC and a TTP are the same thing at different sizes." No — an indicator is an artefact-level pattern, a technique is a behaviour, and only the behaviour survives the adversary changing infrastructure. - "The bundle told us it was malicious." A bundle told you what a producer asserted, under a marking, at a time.

  • A shared STIX sighting object reports a count of 40 across nine member organisations. What does that prove?
    That nine members' tooling matched the pattern forty times. It raises confidence that the indicator is worth acting on, and it says nothing about whether any of them was compromised. If the pattern is loose, forty matches are forty false positives and the count is actively misleading. Sightings are a popularity signal for the indicator, not an incident count.
  • Where do MISP's event-and-attribute model and the STIX object graph disagree?
    MISP's event is a dated container of typed attributes with a distribution level; STIX has no container object with those semantics, so exports usually become a `report` or `grouping` and some context degrades to free text. The `to_ids` flag on an attribute is roughly the indicator-versus-observation distinction. OpenCTI stores the STIX graph natively, so the lossy step is normally the flattening back into events.
  • You ingest an indicator whose pattern is a single IP address. What should you check before blocking?
    Whether the address is dedicated adversary infrastructure or shared hosting the adversary merely rented, and when the claim was made. `valid_from` and `valid_until` bound the producer's belief, and IP-level indicators age in days. Blocking a shared cloud egress address on a partner's claim is how a feed causes your outage, so check the surrounding relationships and your own telemetry first.

saying these in an interview costs you the question

  • Says an indicator means the artefact was observed somewhere
  • Reads a sighting count as a count of compromised organisations
  • Treats a bare observable as a blockable indicator
  • Ages ingested indicators by ingest time, ignoring valid_until
  • Claims MISP events map one-to-one onto STIX objects

context