You own the S3 landing zone for a data lake that receives a few million small events per day from many producer teams. How do you lay out buckets, prefixes and object sizes, and what does getting it wrong cost you later?
answer
- zones are boundaries, not folders
- the prefix is your only index
- partition on something that never changes
- millions of tiny objects is the trap
- rewriting a lake means copying it
basics
~20 sKeep a raw immutable landing zone separate from curated output, partition prefixes by a stable dimension such as source and ingestion date, and aggregate events into files of tens to hundreds of megabytes instead of millions of tiny objects. Reorganising a lake later means copying everything.
solid answer
~50 sI separate zones by bucket — raw, curated, and anything published outward — because zone is the boundary for access control, lifecycle and blast radius, and a bucket is the cheapest place to draw that line. Inside raw I partition by source and by ingestion date, `source=<team>/dt=YYYY-MM-DD/`, so readers can prune whole prefixes instead of listing the world and every consumer sees the same convention. The critical decision is object size: a few million events per day written one-per-object means millions of PUTs and GETs, listings that take minutes, and query engines spending all their time opening files. I buffer upstream so events land as files of tens to hundreds of megabytes, or run a compaction pass that rewrites the small ones. Raw stays append-only and never edited, because everything downstream can be rebuilt from it — but only if it is still intact.
go deeper
Know that S3 has no real folders — the prefix is just part of the key — and that a consistent key convention is what lets readers find data without scanning the whole bucket.
Explain why millions of tiny objects hurt: per-request charges, slow listings and per-file overhead in query engines, and describe buffering or compaction as the remedy.
Show the operating view — compaction jobs, an immutable raw zone enforced by IAM and versioning, and a reconciliation sweep that proves nothing arrived without being processed.
Own the one-way door: zone boundaries drawn at bucket or account level, a single enforced convention across producer teams, and an explicit stance on latency versus file size — because changing any of it later means copying the lake.
## Why layout is a one-way door Everything else in a lake is rebuildable. Derived tables can be recomputed, jobs can be rewritten, catalogs can be regenerated. The physical layout of the landing zone cannot be changed cheaply, because changing it means copying every object — at petabyte scale that is a project with a budget, not a refactor. That asymmetry is why the layout decision belongs to whoever owns the platform, and why it is worth being conservative and boring. ## Zones as buckets Three zones cover most designs: - **Raw / landing** — exactly what producers sent, unmodified, append-only. - **Curated** — cleaned, conformed, partitioned for reading. - **Published** — the subset shared outside the platform. Giving each zone its own bucket (often its own account, for the same reasons) gets you a natural boundary for policy, encryption keys, lifecycle rules and audit. It also gives you a blunt but effective blast-radius control: a mistake in a curated-zone job cannot touch raw, because the job's role has no write access to that bucket. Inside raw, resist the urge to let each producer invent its own structure. One convention, enforced at onboarding, is worth more than any amount of later tooling. ## Prefix design The prefix is the only index the object store gives you, so put the dimensions readers filter on into it — typically source and time: ``` s3://lake-raw/source=checkout/dt=2026-08-21/hour=14/part-0001.parquet s3://lake-raw/source=search/dt=2026-08-21/hour=14/part-0001.parquet ``` Two properties matter. First, **prunability**: a reader interested in one day lists one prefix rather than scanning the bucket, and the cost difference grows with the lake. Second, **stability**: partition by *ingestion* date rather than by an event's business date, because ingestion date is known at write time and never changes, whereas late-arriving data forces you to rewrite business-date partitions. A few things do not belong in a key: anything personally identifying (keys appear in logs, metrics and access reports and are not encrypted the way object bodies are), and anything mutable, such as a status that will change later. ## The small-file problem This is the failure people actually hit. A few million events a day written as one object each produces: - **Request cost that dominates storage cost.** You pay per PUT and per GET, and at millions per day the request line on the bill exceeds the bytes line. - **Slow, expensive listing.** Listings return objects in pages; enumerating millions of keys to plan a query is minutes of wall clock before any data is read. - **Poor query throughput.** Analytical engines pay a fixed cost to open each file. Ten thousand 10 KB files take far longer to scan than one 100 MB file with the same content. - **Worse economics in colder storage classes.** Infrequent-access classes bill a per-object minimum size and a minimum duration, so tiny objects there can cost more than leaving them in Standard. Two cures, and mature platforms use both. **Buffer before writing** — an ingest layer that accumulates events and flushes on a size or time threshold means files arrive already the right size. **Compact after writing** — a scheduled job that rewrites yesterday's many small files into few large ones, deleting the originals once the rewrite is durable. Compaction is unavoidable in practice because some producers will always write small. ## Wiring the pipeline to the layout S3 event notifications are the natural trigger for "a new file landed", and they scale fine — but only if the unit of work matches the file layout. Per-object notifications on a raw zone receiving millions of tiny files means millions of invocations doing almost nothing each. The pattern that holds up is: notifications drive lightweight bookkeeping (record arrival, enqueue work), and the heavy processing runs per partition on a schedule or when a partition is marked complete. Notifications also do not backfill and are at-least-once, so a periodic listing sweep that reconciles what arrived against what was processed is worth building — it is the only mechanism that positively proves nothing was missed. ## The tradeoffs to state out loud - **Partition granularity.** Hourly partitions give tight pruning and many small files; daily gives larger files and coarser pruning. Choose from the read pattern and the daily volume, and expect to differ per source. - **Latency versus file size.** Buffering to 128 MB adds minutes of delay. If a consumer needs seconds, it should read the stream, not the lake — do not degrade the lake's layout to serve a real-time need. - **Producer autonomy versus one convention.** Letting teams choose their own layout is faster on day one and a migration project on day four hundred. ## What a strong answer sounds like Name the zones and why the boundary is a bucket, name the partition scheme and why ingestion time, then spend most of the answer on file size — because that is the decision that costs real money and the one an interviewer is checking you have lived through. Close with the fact that all of it is expensive to change later, which is what makes it a platform decision rather than a team one.
- Why partition by ingestion date rather than by the event's own timestamp?Ingestion date is known when you write and never changes, so a partition is finished once its window closes. Business date is not: an event that arrives three days late belongs in an already-written partition, forcing a rewrite and breaking anything that treated that partition as complete. Curated layers can repartition by business time, where rewriting is expected and cheap.
- Producers refuse to buffer and keep writing one object per event. What do you do?Accept the raw writes and compact on a schedule — a job per source that rewrites each closed partition into a handful of large files and deletes the originals once the rewrite is durable. Point consumers at the compacted output rather than the raw prefix. Then make the request cost visible per producer through cost allocation, because a bill they can see moves teams faster than a standard they cannot.
- How do you keep the raw zone genuinely immutable rather than just conventionally so?Make it structural. No role outside the ingest path gets write or delete on that bucket, versioning is on so an overwrite is recoverable, and Object Lock is available where a regulator or an audit requires provable immutability. Convention alone fails the first time an engineer with broad permissions cleans up something that looked wrong.
- Where do S3 event notifications fit in a lake this size?As the arrival signal, not the processing unit. Notifications record that a file landed and enqueue work; the heavy job runs per partition, so cost tracks partitions rather than objects. Because notifications are at-least-once and never backfill, pair them with a periodic listing sweep that reconciles arrived against processed — that sweep is what actually proves completeness.
saying these in an interview costs you the question
- Treats prefixes as folders with no cost consequence
- Writes one object per event and calls it scalable
- Partitions raw data by business timestamp
- Plans to reorganise the layout later if needed
- Puts user identifiers directly into object keys