skip to content

In a causal DAG, what does an arrow assert, and what does a missing arrow assert?

level: juniorimportance: should knowfreq 55%

answer

  1. the drawing is an assumption, not a result
  2. arrows are permissive claims
  3. absences do the real work
  4. no loops; unroll feedback over time
  5. missing node means no unmeasured common cause

basics

~10 s

An arrow from X to Y asserts X may directly cause Y. A missing arrow is the stronger claim: no direct effect at all. Acyclicity forbids a variable causing itself through a loop.

solid answer

~50 s

A causal directed acyclic graph has one node per variable and a directed edge `X -> Y` whenever X is a possible direct cause of Y relative to the variables drawn. The edge asserts direction and possibility only: nothing about size, sign, or functional form. The real assumptions live in the edges you did **not** draw. A missing arrow is a hard claim that there is no direct causal effect between two nodes, and a missing node claims no other common cause matters. Acyclic means you cannot follow arrows and return to where you started, so genuine feedback is drawn as time-indexed nodes such as `Sales(t)` and `Spend(t+1)`. Because the graph is a set of assumptions rather than a result, expect to be asked where each arrow came from: domain knowledge and temporal order, not the data.

go deeper

for a junior

Be ready to read a small graph out loud: name the nodes, state what each arrow claims, and point out that a missing arrow claims no direct effect. Know that acyclic means no loops.

for a middle

Explain why absences are stronger assumptions than presences, and how feedback is handled by unrolling variables over time. Be able to say what an arrow does not claim about magnitude or sign.

for a senior

Show where you actually get arrows in a real project: assignment mechanism, temporal order, expert elicitation. Demonstrate that you draw unmeasured common causes explicitly rather than hoping they are absent.

for a principal

Own the process: who reviews the graph, how disagreements about a single edge are recorded, and how you communicate that the conclusion rests on stated assumptions rather than on the estimator.

## What the picture is A causal directed acyclic graph (DAG) is a drawing with two ingredients. **Nodes** are variables: things that vary across the units you study, such as ad spend, seasonality, a customer's tenure, or whether a purchase happened. **Directed edges** (arrows) connect nodes. The graph is *acyclic*: starting at any node and following arrows in the direction they point, you can never come back to where you started. A DAG used for causal inference is not a flowchart, a data pipeline, or a picture of correlations. It is a compact statement of causal assumptions that you can reason over with rules. ## What an arrow says An arrow `X -> Y` says: X is a possible **direct** cause of Y, where "direct" is relative to the other variables drawn in the graph. If you intervened to change X while holding the other nodes at fixed values, Y might change through no intermediary shown in the picture. Equally important is what the arrow does **not** say: - It does not say how large the effect is. - It does not say the effect is positive, negative, linear, or monotone. - It does not say the effect is present for every unit; it says it is not ruled out. So an arrow is a weak, permissive claim. Drawing one costs you almost nothing in assumptions. ## What a missing arrow says The assumptions that do real work are the arrows you leave out. A missing arrow between X and Y is the exact-zero claim: X has **no** direct causal effect on Y at all, once you account for the other variables in the graph. That is much stronger than any arrow you draw. The same holds for missing **nodes**. Every DAG implicitly claims that any variable not drawn is either irrelevant or not a common cause of two drawn variables. Analysts often draw an explicit unobserved node (commonly labelled U) precisely to make that claim visible instead of silent: writing `U -> Spend` and `U -> Sales` for an unmeasured seasonality driver is an honest admission, whereas simply omitting the node would quietly assert that no such driver exists. This asymmetry is why reviewers interrogate the sparse parts of a graph. "Why is there no edge here?" is a sharper question than "why did you draw this edge?". ## Paths and direction A **path** between two nodes is any sequence of edges connecting them, followed without regard to arrow direction. A **directed path** follows arrows all pointing the same way; a directed path from T to Y is what carries the causal effect of T on Y. Direction matters enormously. `Spend -> Sales` and `Sales -> Spend` produce the same association in observational data but demand completely different analyses. A DAG forces you to commit to a direction before you touch the data, which is exactly the discipline it is there to impose. A node reachable from X by following arrows forward is a **descendant** of X; X is its **ancestor**. Descendants of the treatment matter later, when choosing which variables you are allowed to adjust for. ## Why acyclic Acyclicity rules out `A -> B -> A`. Real systems obviously do contain feedback: sales this quarter change next quarter's budget, which changes sales again. The DAG handles this by indexing variables by time, so `Spend_1 -> Sales_1 -> Spend_2 -> Sales_2` is a perfectly ordinary acyclic graph. The restriction is on the drawing, not on the world: an instantaneous loop makes the graph's probability semantics ill-defined, whereas a time-unrolled chain is well behaved. ## Where the arrows come from A DAG is not estimated from the data in the usual workflow. It comes from: - **Temporal order** — a variable measured before another cannot be caused by it. - **Domain knowledge** — how the mechanism works, elicited from people who run the system. - **Design facts** — what actually determined who got treated (targeting rules, eligibility thresholds, a randomiser). The data can *refute* a graph in a limited way: a DAG implies certain conditional independencies, and if the data plainly violate them, the graph is wrong. It can rarely confirm one, because several different graphs typically imply the same independencies. ## Why interviewers start here Every later tool — reading which paths carry bias, deciding which variables to control for, judging whether an effect is identified at all — is machinery applied to a graph. If a candidate is loose about what an arrow means, the machinery downstream is decorative. Being crisp that arrows are permissive and absences are strong is the single most useful thing to carry out of this topic.

  • Your system has genuine feedback between ad spend and sales. How do you draw that in an acyclic graph?
    Index the variables by time and unroll the loop: `Spend_1 -> Sales_1`, `Sales_1 -> Spend_2`, `Spend_2 -> Sales_2`. Each node is a variable at a specific period, so no arrow ever returns to its own node and the graph stays acyclic. The feedback is fully represented; only the instantaneous self-loop is forbidden, because it has no well-defined probabilistic meaning.
  • Can you learn the arrows from the data instead of assuming them?
    Only partially. Data reveal conditional independencies, and several distinct graphs usually imply the same set of them, so the data narrow the candidates rather than pick one. Direction in particular is frequently unidentifiable from association alone. In practice the graph comes from temporal order, mechanism knowledge and how treatment was assigned, and the data are used to check the independencies it implies.
  • Why do people draw an unobserved node U instead of leaving it out?
    Because omitting it silently asserts it does not exist. Drawing `U -> T` and `U -> Y` for an unmeasured common cause makes the assumption visible and lets the path rules show its consequence: a route between treatment and outcome that no measured variable can close. Naming the threat explicitly is what turns an untested hope into a stated, arguable assumption.

An arrow is like saying "this wire might carry current"; a missing arrow is like saying "these two components are definitely not connected". The second claim is the one an electrician would challenge you on.

saying these in an interview costs you the question

  • Treats arrows as correlations rather than causal claims
  • Thinks a drawn arrow claims a large or positive effect
  • Believes the graph is learned from the data
  • Says feedback cannot be represented in a DAG at all
  • Ignores that leaving a node out is itself an assumption

context