skip to content

Why can a Heap Connect event table's historical counts change after a definition edit?

level: seniorimportance: nice to knowfreq 30%

answer

  1. the event table is a view, not a log
  2. someone can edit the rule at any time
  3. backfills arrive on the next sync
  4. late identify rewrites past user counts
  5. snapshot anything that must not move

basics

~20 s

Because the synced event is derived from Heap's retained raw interactions via an editable definition. Change the rule and past interactions are reclassified, so the warehouse copy is restated for dates you already loaded and reported on.

solid answer

~50 s

Heap Connect delivers autocaptured data and your defined events into a cloud warehouse. The critical property for a data engineer is that a defined event is **not** an immutable fact stream — it is a view over retained raw interactions produced by a rule someone can edit in Heap's UI. Narrow the selector, add a page filter or rename the event and the historical population changes; the next sync brings restated numbers for periods your dbt models already published. Identity behaves the same way: when `heap.identify` runs late, earlier anonymous activity is merged onto the identified user, so past-period user counts and first-touch attribution shift after the fact. Treat these tables as mutable upstream, not as append-only logs. If a number must be stable — a board metric, a billing input — snapshot it in your warehouse with the date it was computed, and restrict who may edit the definitions that feed it.

code

sql · 11 lines
sql
-- Freeze a reporting number instead of re-reading a mutable event table
-- Table naming follows your own Heap Connect schema and event names
insert into reporting.checkout_started_daily_snapshot
select
    date_trunc('day', e.event_time) as event_date,
    count(*)                        as events,
    count(distinct e.user_id)       as users,
    current_timestamp               as computed_at
from analytics_heap.checkout_started e
where e.event_time >= dateadd('day', -1, current_date)
group by 1;

go deeper

for a junior

Know that Heap can sync its data into a warehouse, and that the events you see there come from definitions someone wrote in Heap rather than from fixed tracking code.

for a middle

Explain why a defined event is derived data: the rule is applied to retained raw interactions, so editing or adding a definition changes what lands in the warehouse for past dates.

for a senior

Show the modelling response — rebuildable models, snapshots for numbers that must not move, awareness that identity merges also restate history, and deliberate reconciliation with the vendor UI.

for a principal

Own where the semantic layer lives: vendor-UI definitions for speed versus version-controlled warehouse models for reproducibility, plus edit permissions and ownership for the events that feed reporting.

## What the warehouse gets Heap Connect syncs Heap's data into a cloud warehouse — Snowflake, Redshift, BigQuery and similar targets — on a schedule, so analysts can join product behaviour to orders, subscriptions and support tickets in SQL rather than living inside the vendor UI. Two kinds of thing arrive: the underlying autocaptured activity (interactions, pageviews, sessions, users and their properties) and the events you have **defined**, materialised as their own tables or views so a query can select "Checkout Started" directly. That second category is where the engineering surprise lives. ## Why the numbers are not frozen A defined event is metadata over retained raw data. It exists because someone wrote a rule — a selector, some text or attribute filters, a page filter — and Heap applies that rule to every matching interaction it holds, past and future. Three consequences follow: 1. **Editing the rule restates history.** Narrow the definition to exclude clicks on a disabled button and last quarter's conversion denominator changes. The edit happens in a UI, in seconds, with no code review and no migration. 2. **Defining a new event backfills.** A table that did not exist yesterday can arrive fully populated with a year of history on the next sync. 3. **Deleting or renaming a definition removes or moves a table.** Downstream models referencing it break in a way that looks like a pipeline failure but is actually a governance failure. So a dbt model reading a Heap Connect event table is reading a **mutable upstream**. If it computes monthly conversion and someone edits the definition, a full refresh produces different numbers for months already reported — with no row-level change in your own repository to explain it. ## Identity adds a second source of restatement Heap assigns an anonymous identity as soon as a visitor arrives. When `heap.identify` later runs — at signup, at login — the prior anonymous activity is merged onto the identified user. That is desirable behaviour (you want the pre-signup browsing attributed to the person who signed up), but it means user-level aggregates for *past* periods can change after the fact: a user counted as anonymous last week becomes a known user this week, and first-touch attribution moves with them. Heap surfaces the merge history so downstream joins can be reconciled rather than left dangling on stale ids. Practically, any warehouse model keyed on Heap's user identity needs to be rebuildable rather than incrementally appended on the assumption that identity is stable. ## How to model around it - **Do not treat the sync as append-only CDC.** Build the models so a full rebuild is normal and cheap, or key incremental logic on something Heap does not retroactively change. - **Snapshot what must not move.** For board metrics, billing inputs or anything that appears in an external report, write the computed number into your own table with a `computed_at` timestamp. A restated dashboard is fine; a restated invoice is not. - **Version and govern the definitions.** Give each reporting-critical event a named owner, document what it matches, and restrict edit permissions. Export or record the definition set so "what did this event mean in March?" has an answer. - **Rebuild the semantics you care about.** Some teams sync the raw interaction data and define the important events in dbt instead, so the semantic layer sits in reviewed, version-controlled SQL. You lose the point-and-click speed and gain reproducibility — a genuine trade-off, not an obvious win. - **Reconcile deliberately.** Expect the vendor UI and the warehouse to disagree at the margins — timezone handling, session boundaries, sync lag, definition edits between runs. Pick one as the number of record for each metric and say so out loud. ## What an interviewer is checking This is a data-engineering question wearing a product-analytics costume. They want to know whether you recognise a mutable upstream when you see one and whether your instinct is to snapshot, govern and rebuild rather than to assume immutability and be embarrassed when last month's figure moves. Mentioning identity merges as a second restatement channel is the detail that separates someone who has actually run this sync from someone who has read the marketing page.

  • How does late identification change historical user counts in Heap?
    A visitor is anonymous until `heap.identify` runs. When it does, the earlier anonymous activity is merged onto the identified user, so someone counted as anonymous last week becomes a known user afterwards and first-touch attribution moves. Warehouse models keyed on Heap's user identity therefore need to be rebuildable rather than assuming identity is stable.
  • Would you define events in Heap's UI or rebuild them in dbt from the raw data?
    It depends on who needs to move fast. The UI gives analysts point-and-click definitions with instant backfill; dbt gives you version control, code review and reproducible history. A common split is UI definitions for exploratory work and dbt-defined events for the small set of metrics that feed reporting or billing.
  • Your dashboard and the Heap UI disagree by a few percent. What do you check?
    Sync lag between the last export and the live UI, timezone and session-boundary handling, whether a definition was edited between runs, and any filtering your model applies — bots, internal users, test accounts. Then pick one as the number of record for that metric and document the choice rather than reconciling forever.

saying these in an interview costs you the question

  • Treats synced Heap event tables as append-only immutable facts
  • Builds incremental models assuming history never changes
  • Ignores that anyone with UI access can restate a metric
  • Forgets identity merges reattribute past anonymous activity
  • Assumes the warehouse copy and the Heap UI must always agree

context