When would you publish a fact table only at a summarized grain rather than the atomic grain?
answer
- one direction only: detail to summary
- start from what the default is
- three exceptions, one of them weak
- what question can you never ask again?
- the loss is discovered later by someone else
basics
~10 sAlmost never by choice. Publish only summaries when atomic rows are legally restricted, unavailable from the source, or genuinely unaffordable. Accept that any question below the summary grain becomes unanswerable until history is reloaded.
solid answer
~50 sThe default is atomic grain, because dimensionality is richest at the finest level and every coarser summary can be derived from it while the reverse never works. Publishing only a summary is a deliberate exception with three honest justifications: the atomic rows are **restricted** (privacy regulation, a data-sharing contract, personally identifying detail you may not retain), they are **unavailable** (the source only ever emits a daily roll-up), or they are **unaffordable** at the retention the business needs. Cost alone is the weakest of the three and usually dissolves under measurement. The price you pay is permanent: any question below the summary grain — this customer's sequence of events, a re-cut by a dimension you dropped, a restatement after a definition change — becomes unanswerable until history is reloaded, if it can be reloaded at all. If you take the exception, document the grain as a contract, keep atomic detail somewhere cheaper for as long as you legally can, and record what was dropped.
code
text · 9 linesATOMIC one row per product per ticket
ticket 88421 | 09:14 | store 12 | product A | qty 2 | promo SUMMER
SUMMARY one row per product per store per day
2026-03-04 | store 12 | product A | qty 940
Derivable from atomic : the summary row above
Lost with summary only: basket contents, time of day, promotion split,
any restatement after a metric definition changego deeper
Know that a summary can always be built from detail but detail can never be rebuilt from a summary, so the finest available level is the safe default.
Explain why dimensionality is richest at the atomic grain and name concrete questions a daily roll-up can no longer answer, such as basket-level analysis.
Weigh the real exceptions — regulatory retention, a source that only emits roll-ups, measured cost — and handle the fallout: published grain contract, documented dropped dimensions, a retention window for detail.
Own it as an irreversible platform decision: it destroys optionality for teams you will never meet, so demand measurement behind cost claims and revisit the choice as volumes and prices change.
## Why atomic is the default A fact table at the lowest level the business process actually captures has two properties nothing else has. **Dimensionality is maximal.** The finer the row, the more descriptors take exactly one value for it. At one row per order line you can slice by product, promotion, cashier, payment method and channel. At one row per store per day, most of those stop applying, because a day's total spans many products and many tickets. Every descriptor lost at design time is a class of question the mart can never answer. **Summaries are derivable; detail is not.** You can always compute a daily total from line rows. You can never recover line rows from a daily total. That asymmetry is why atomic grain is the resilient choice: it absorbs questions nobody had thought of when the model was built, which is the usual reason a warehouse survives its first reorganisation. ## The three defensible reasons to publish only a summary **1. You are not allowed to keep the detail.** Privacy regulation, a retention policy, a data-sharing contract with a partner, or an internal rule about personally identifying data can all make atomic rows something you may not retain past a short window. Aggregation is then a control, not a shortcut — summarised rows that no longer identify an individual may be retained when the underlying events may not. This is the strongest reason and it is not negotiable by engineering. **2. The detail does not exist.** Some sources only ever emit a roll-up: a partner sends daily totals, a legacy system was decommissioned and left monthly extracts, a third-party advertising platform reports per-campaign-per-day and nothing finer. You cannot publish a grain the source never produced, and inventing one by allocation manufactures precision the data does not have. **3. The detail is genuinely unaffordable.** Extremely high-volume telemetry at the retention the business demands can cost more than the questions it answers are worth. This is the reason most often claimed and least often true — it deserves an actual measurement of storage and query cost against the value of the questions, not an assumption. Modern columnar storage is cheap enough that "too big" is frequently a guess. ## What you give up, precisely - **Unanticipated questions.** Any analysis below the summary grain is impossible, permanently. "Did customers who bought A also buy B in the same basket?" cannot be answered from daily product totals, ever. - **Re-cuts by dropped dimensions.** If the summary drops promotion, no future report can be split by promotion, even though the events had one. - **Restatement.** When a metric definition changes — revenue now excludes returns, active user now means something else — atomic detail lets you recompute history. A summary computed under the old definition cannot be re-derived, so history becomes a mixture of two definitions. - **Debuggability.** When a number is disputed, atomic rows let you drill to the individual events and settle it. A summary offers nothing to drill to, and disputes end in "the pipeline says so". - **Future summaries.** Every additional roll-up someone asks for later has to come from the atomic layer. Without it, each new summary needs a new load from the source, if the source still has the data. ## Making the exception responsibly If you take it, take it explicitly. - **Declare the grain as a published contract**, prominently: "one row per campaign per day; no user-level detail exists". Consumers must not discover the limit by writing a query that silently answers the wrong question. - **Keep atomic detail for as long as you legally and economically can**, even in a cheaper, colder, less-modelled store, so that recomputation is possible in the window that matters most. - **Record what was dropped and why.** A dimension omitted from a summary is a decision with an owner and a date, not an accident. - **Choose the summary grain one level finer than you think you need.** Adding a rarely-used key to a summary is cheap; discovering a year later that you needed it is not. - **Revisit the cost argument on a schedule.** "Unaffordable" was true at one volume and one price. It may not be true next year, and the atomic backfill window closes as retention expires. ## The version of this decision that is not the exception Note that publishing summaries *in addition to* atomic detail is not this decision at all — that is a normal performance practice and carries none of these costs, because the detail remains and every summary can be rebuilt from it. The question here is specifically whether the atomic layer exists at all. Keep those two apart in an interview, because conflating them is the fastest way to look as if you would delete the detail to speed up a dashboard. ## What an interviewer listens for They want the default named first, the three legitimate exceptions distinguished by kind — regulatory, source, cost — and honest scepticism about the cost argument. Above all they want the irreversibility stated: this decision destroys optionality, and the loss is discovered later by someone who cannot fix it.
- A team says atomic rows cost too much to keep. How do you test that claim?Measure it rather than accept it: actual storage cost at the required retention, what the atomic layer is scanned for, and the value of the questions only it can answer. Compare against the cost of losing restatement and drill-down. Cost is the weakest of the three justifications and usually dissolves once someone prices it.
- If regulation forces you to drop atomic detail after 90 days, what do you keep?Keep the atomic rows for the full permitted window so recomputation is possible where it matters most, and design the retained summaries before the window closes — including any dimension a future report might need, since it cannot be added afterwards. Document the cut-off date so consumers know exactly where history changes character.
- How is publishing only a summary different from publishing summaries alongside atomic detail?Entirely different decisions. Summaries built on top of an atomic layer are a performance practice: the detail remains, every summary can be rebuilt, and nothing is lost. Publishing only a summary destroys optionality permanently — no restatement, no drill-down, no new dimension. Only the second one needs the justifications discussed here.
saying these in an interview costs you the question
- Drops atomic detail to make a dashboard faster
- Treats storage cost as self-evident without measuring it
- Assumes detail can be recovered from a summary later
- Never documents which dimensions the summary dropped
- Confuses adding summaries with removing the atomic layer