What would you pin down in a partner's file-drop convention before accepting daily CSV drops?
answer
- the engineering is easy; the ambiguity is not
- how does the loader know it may start
- what does a second file for the same day mean
- silence must be distinguishable from failure
- every unenforced rule will be broken
basics
~20 sPin the delivery mechanics: a scoped prefix and deterministic file names, staged upload then copy to the final key, a completeness marker or manifest, explicit re-drop semantics, a mandatory zero-row drop on empty days, retention, and a named owner on each side.
solid answer
~50 sTreat it as an operating agreement, not a wiki page. **Location and naming**: a prefix per partner and dataset, partitioned by arrival date, with names carrying dataset, window and a run id so re-drops are recognisable. **Atomicity**: upload to a staging prefix and copy to the final key, so a half-uploaded object is never visible where the loader looks. **Completeness**: a marker or a manifest with file list and row counts, and the rule that nothing is read without it. **Re-drop semantics**: decide explicitly whether a second file for a date replaces or supplements, and make the load a partition overwrite either way. **Empty days must still produce a drop**, so silence and failure are distinguishable. Then add credentials scoped to their prefix, raw-zone retention, a deadline you alert on, an error channel back to them, and a test drop before go-live. Schema and semantic guarantees are a separate agreement layered on top.
code
yaml · 26 linespartner: acme
dataset: orders
delivery:
final_prefix: s3://landing/acme/orders/ingest_date=YYYY-MM-DD/
staging_prefix: s3://landing/acme/_incoming/
upload: write to staging_prefix, copy to final_prefix on completion
filename: orders_YYYYMMDD_run<NN>_part-<NNNN>.csv.gz
completeness: _MANIFEST.json listing object keys and row counts
empty_day: still required, manifest with zero files
redrop: a later run id replaces the whole date, never supplements
schedule:
expected_by: "06:00 UTC"
alert_if_missing_after: "08:00 UTC"
csv_dialect:
delimiter: ","
quote: '"'
escape: doubled quote
newline_in_quoted_fields: allowed
encoding: UTF-8
header: required
null_token: ""
timestamp: "YYYY-MM-DDTHH:MM:SSZ (UTC)"
operations:
credentials: write-only, scoped to acme prefix
raw_retention_days: 90
rejection_channel: [email protected]go deeper
Know that a partner drop needs an agreed location, an agreed file naming pattern, and a signal that the drop is finished before anything is read.
Explain why the upload must be staged and copied, why a completeness marker or manifest gates the load, and why the CSV dialect has to be written down rather than sniffed.
Cover the operational clauses: alert on a missing drop by deadline, define re-drop semantics and enforce them with partition overwrite, retain the raw zone, and provide an error channel for rejected drops.
Own the negotiation and the enforcement. Rank which clauses are non-negotiable for correctness against which are tradeable for adoption, keep delivery separate from schema guarantees, name accountable owners on both sides, and encode the rules in the loader rather than a document.
## Why this is a leadership question The engineering is not hard. What makes partner file drops fail is that nobody wrote down what a drop *is*, so every ambiguity is discovered in production by someone who cannot fix it. The value you add here is deciding the ambiguities in advance and making them enforceable. Keep the scope clear: this agreement governs **delivery** — where files go, how you know they are complete, what a repeat means, when they are late. The promises about columns, types and business meaning are a separate agreement layered on top; conflating them produces a document nobody owns. ## Location, naming and identity One prefix per partner per dataset, with credentials scoped to exactly that prefix — write-only if the platform supports it. Isolation is cheap now and expensive to retrofit after two partners have shared a bucket. Partition by **arrival** date so a prefix, once closed, never changes and is therefore a stable replay unit. File names should carry the dataset, the window they cover, a run or sequence id, and a part number: `orders_20260820_run03_part-0000.csv.gz`. That lets an operator recognise a re-drop, spot a gap in sequence, and attribute a bad file to a run without opening it. Opaque UUID names are collision-safe and diagnostically useless. ## Atomicity of the upload A 4 GB CSV uploaded directly to its final key is visible-in-progress under some upload paths and, at minimum, invites a loader to start reading a prefix whose contents are still changing. Require the partner to write to a staging prefix and copy to the final key on completion, or to upload under a temporary name and copy. The loader watches only the final prefix. ## Completeness Nothing is read without a completeness signal. A zero-byte marker written last is the minimum. A manifest is better: it names every object, carries per-file row counts and a batch id, and lets you verify what you loaded against what was declared and fail loudly on a short batch. The manifest is also what makes orphan files from a failed partner run inert — you load what it names and ignore the rest. ## Re-drops and corrections This is the clause most often missing and most often costly. Two files exist for 20 August: is the second a correction that replaces the first, or additional rows that supplement it? Pick one, write it down, and make the load a partition overwrite so the semantics are enforced by the code rather than by hope. If corrections are real, require them to restate the full window rather than sending deltas, because a delta with no ordering guarantee is unmergeable. ## Absence must be an event An empty prefix is indistinguishable from a broken partner. Require an explicit zero-row drop with its marker on days with no data, so "nothing happened" is a positive statement. Pair that with an expected-by deadline on your side and an alert when the marker has not appeared — a missing drop must page as loudly as a failed load, and by default it does not. ## The file dialect Since the format is CSV, the delivery agreement has to pin the dialect, because it is a property of the bytes rather than of the schema: delimiter, quote and escape characters, whether quoted fields may contain newlines, character encoding, header row presence, the null token, the decimal separator, line terminator, and the timestamp format with timezone. Configure your reader explicitly against those rather than letting it sniff. Compression and whether it is applied per file belongs here too. ## Operations, security and money Retention on the raw zone, and who may read it. Encryption expectations and whether the drop may contain personal data — the answer determines the bucket's access model, its lifecycle policy and possibly its region. An error channel: when you reject a drop, where does that land for them, and what is the resend procedure? Rate and size expectations, so a partner's one-off 400 GB backfill does not arrive unannounced at 09:00 on a Monday. And a named owner on both sides, with an escalation path. An agreement with no accountable human degrades into folklore within two quarters. ## Make it testable and make it enforced Require a test drop against a staging prefix before go-live, exercising the empty-day case, a re-drop, and a deliberately malformed file. Version the agreement itself, so "which version were you sending in July" is answerable. Then encode as much as you can in the loader: reject files whose names do not match the pattern, refuse prefixes without a marker, compare against manifest counts, quarantine rather than crash on a malformed row, and alert on missing drops. Every rule you leave to human diligence will be violated, usually during someone else's on-call. ## The judgment call How much you can demand depends on leverage. A partner who gains nothing from your pipeline will not build a manifest emitter because you asked nicely. Rank your requirements: atomic upload, a completeness signal and unambiguous re-drop semantics are non-negotiable because without them correctness is impossible; manifests, row counts and format upgrades are worth trading away for adoption. Know which line you are defending before the meeting.
- Why require a zero-row drop on days with no data?Because an empty prefix cannot be told apart from a dead producer, a revoked credential or a broken cron. An explicit empty drop with its marker turns "nothing happened" into a positive statement you can verify, and it lets you alert on a missing drop by deadline rather than discovering the gap weeks later in a reconciliation.
- Which clauses would you concede to a partner with no incentive to invest?Concede manifests, row counts and format upgrades — you can compensate with your own validation. Do not concede atomic upload, a completeness signal, or unambiguous re-drop semantics: without those three you cannot be correct at any effort, only lucky. Knowing that ranking before the negotiation is most of the job.
- How do you keep the agreement from decaying into folklore?Encode it in the loader — reject non-conforming names, refuse prefixes without a marker, verify counts, quarantine malformed rows, alert on missing drops. Version the document, name an owner and an escalation path on each side, and require a test drop that exercises the empty day, a re-drop and a malformed file before go-live.
- Where does this agreement stop and a data contract begin?This one governs delivery mechanics: location, naming, atomicity, completeness, re-drop semantics, timeliness, retention. Guarantees about which columns exist, their types and their business meaning, and how changes to them are announced, belong to a separate layered agreement. Merging the two produces a document that no single team owns and therefore nobody maintains.
saying these in an interview costs you the question
- Leaves re-drop semantics undefined and hopes it never happens
- Treats an empty prefix as an acceptable way to say no data
- Lets partners upload directly to the final key
- Writes the rules down but enforces none of them in code
- Demands manifests and Parquet from a partner with no leverage over