How do you log the trigger condition in the control arm when the feature exists only in treatment?
answer
- control needs a would-have-qualified record
- same check, same point in the flow
- log without changing anything visible
- logging must not ship only with the feature
- verify trigger rates and timing match
basics
~20 sEvaluate the same trigger condition inside the control code path and record that the user would have qualified, while changing nothing they see. Without that counterfactual record there is no comparable control group for a triggered readout.
solid answer
~50 sCounterfactual logging means the control arm runs the same eligibility check the treatment arm runs, at the same point in the flow, and emits the same exposure event with the same timestamp, then does nothing visible. The check has to sit before any treatment-specific behaviour, so it cannot be influenced by the change itself, and it has to ship to both arms rather than living inside the new feature's code. The classic failure is instrumenting only the new path: you then end up comparing triggered treatment users with the whole control arm, which is a comparison between an engaged subpopulation and a general one, not an experiment. Once it is in place, validate it before trusting any triggered number: trigger rates and trigger timing should look alike across arms, and the two triggered populations should match on attributes fixed before assignment.
go deeper
Be ready to say that a triggered comparison needs a record in the control arm of who would have hit the condition, and that control must look unchanged to the user.
Explain the mechanics: one shared eligibility check placed before any treatment-specific behaviour, emitting the same event with the same timing in both arms.
Show the operational habits, naming the failure modes you have hit and the checks you run on trigger rate, timing and pre-assignment balance before you trust the number.
Argue for exposure logging as platform infrastructure rather than per-experiment work, so no team can produce a triggered readout without a validated control-side record.
## The problem A triggered analysis needs two comparable sets: users in treatment who met the exposure condition, and users in control who *would have* met it. The second set is counterfactual. Control users never saw the feature, so nothing in the product naturally marks them, and if the instrumentation lives inside the new code path, the mark does not exist for them at all. Counterfactual logging is the engineering answer: make the control arm evaluate the same condition and record the same event, while remaining behaviourally identical to what it was before the experiment. ## What correct instrumentation looks like - **One shared check, both arms.** The eligibility evaluation is written once, outside the feature code, and executed on every request or session that reaches the relevant point. It reads the assignment only to decide what to *render*, never to decide whether to *log*. - **Placed before treatment can act.** The check sits at the point in the flow that is identical in both arms. If the feature has already changed the layout, the sequencing, or what loads, an exposure check downstream of that point is measuring something the treatment has touched. - **A no-op in control.** The control branch emits the exposure event and returns. Nothing renders, nothing is prefetched, no extra request goes out. Any side effect turns the logging itself into a second treatment. - **Same fields, same timestamp semantics.** Both arms record the exposure at the same moment relative to the user's journey, so post-trigger outcome windows are comparable. - **The same delivery path.** If the treatment logs from a client bundle that only ships with the new feature, and control logs from the server, the two arms have different loss rates and different latency, and the resulting sets differ for reasons that have nothing to do with the change. ## Validating it before you trust the readout Instrumentation is a claim, and it should be checked rather than assumed. 1. **Trigger-rate parity.** The share of assigned users who trigger should be close between arms. A material gap means either the logging differs or the treatment is influencing who qualifies. Either way, the triggered readout is not yet trustworthy. 2. **Trigger-timing parity.** The distribution of time from assignment to exposure should look alike. A shifted distribution usually means the check sits at different points in the two flows. 3. **Pre-assignment balance within the triggered sets.** Compare the triggered users in each arm on attributes that were fixed before assignment, such as tenure or platform. They should look alike; if they do not, the two exposed populations are not the same kind of user. 4. **A quiet period or a no-op variant.** Running the exposure check with no behavioural difference in either arm for a short window shows whether the logging alone produces a difference. It should produce none. ## Common failure modes - **Treatment-only telemetry.** The most common one. It quietly forces the analysis into comparing exposed treatment users to the whole control arm, which is not a like-for-like contrast and will typically flatter whichever arm has the more engaged users. - **Late evaluation.** The exposure check placed after the feature has already rearranged the page, so control and treatment are being asked a slightly different question about a slightly different state. - **Lazy loading differences.** The new feature pulls in a module that also carries the logging, so exposure events in treatment arrive for a different, narrower set of clients than in control. - **Retrofitting.** Reconstructing the control-side exposed set afterwards from loosely related events, such as anyone who visited a nearby page. That reconstruction is a modelling assumption wearing the costume of a measurement, and it usually cannot be validated. ## When the trigger cannot be evaluated without the feature Sometimes the condition genuinely depends on code that only exists in the new build. Three ways out, in order of preference: ship the eligibility check to both arms while gating only the visible behaviour behind the flag; fall back to a coarser trigger that is determined before treatment can act, accepting a lower trigger rate precision in exchange for a valid comparison; or accept that only the intent-to-treat readout is available for this experiment and size the traffic accordingly. What you should not do is run the triggered comparison anyway and hope the mismatch is small. ## The one-line summary The triggered analysis is not a filter you apply at read time. It is a measurement you have to build into both arms before launch, and then verify, before any of its numbers mean anything.
- How do you sanity-check counterfactual logging before trusting a triggered readout?Compare trigger rates and the distribution of time from assignment to exposure across arms; both should look alike. Then check that the two triggered sets balance on attributes fixed before assignment, such as tenure or platform. A short no-op window, where the check runs but neither arm behaves differently, confirms the logging alone produces no difference.
- What if the trigger condition can only be evaluated inside code that ships with the feature?Split the check from the behaviour: ship the eligibility evaluation to both arms and gate only what the user sees behind the flag. If that is impossible, fall back to a coarser condition settled before treatment can act, or accept that this experiment supports only the intent-to-treat readout and size the traffic for the diluted effect.
- Is it acceptable to reconstruct the control exposed set from historical events after the fact?Only as a rough sensitivity check, never as the primary readout. A reconstruction from adjacent events is an assumption about who would have qualified, not a measurement, and its errors are usually correlated with the very engagement the experiment is measuring. If the reconstruction is all you have, report intent-to-treat as primary and label the triggered number as exploratory.
It is like a placebo pill in a drug trial: the control group has to go through the identical motions at the identical moment, so that you know exactly who in that group corresponds to each treated participant.
saying these in an interview costs you the question
- Compares triggered treatment users to the entire control arm
- Adds exposure logging only inside the new code path
- Evaluates the trigger after treatment has changed the flow
- Lets the control-side logging cause a visible side effect
- Reconstructs the control exposed set retroactively and reports it as primary
- Never compares trigger rates between arms before reading the result