In a photo-sharing service, how should the ingest pipeline keep a new upload unpublished until malware and moderation checks pass?
answer
- no unscanned byte is reachable
- private landing zone
- status gates every read path
- at-least-once trigger, idempotent worker
- magic bytes over declared type
basics
~20 sUploads land in a private quarantine location; the completion event queues scan jobs, and only a pass marks the record published and exposes the object for serving. Failures are rejected, and nothing is served before a verdict.
solid answer
~40 sThe signed upload URL writes into a **quarantine** prefix or bucket that has no public or CDN read access. The store's completion event, backed by the client's complete call and a reconciliation sweep, enqueues a job keyed by upload ID. Workers determine the real file type from its leading bytes rather than the declared `Content-Type`, enforce size and dimension limits, run malware scanning and content moderation, and only on a pass copy the object to the serving location or grant read access, then set the metadata status to `published`. Failures become `rejected` and are deleted or held for review, and the uploader is told. Because events are delivered at least once, every step is idempotent. Serving paths check the status, so a slow scan shows a processing placeholder and the pipeline fails closed.
code
json · 9 lines{
"jobType": "scan-upload",
"uploadId": "up_81f3",
"objectKey": "quarantine/up_81f3",
"objectSha256": "3b1f...",
"attempt": 1,
"onPass": {"copyTo": "serving/p/7c2e91", "setStatus": "published"},
"onFail": {"setStatus": "rejected", "retainForReview": true}
}go deeper
Remember the idea: new uploads first land somewhere private, get checked, and only then become visible to other people.
Describe the status machine from pending to published or rejected, and explain why the file type must come from the bytes rather than the declared header.
Show how you would run it: at-least-once triggers with idempotent workers, a reconciliation sweep, poison-file handling, fail-closed behaviour, and metrics on time to publish.
Weigh scan depth against time to publish, decide which checks block publication versus run afterwards, and plan how rescans handle content that newer rules would reject.
## Why the scan happens after the upload With **direct-to-storage uploads**, application servers never see the bytes in flight, so they cannot inspect a file while it streams in. The declared file name and `Content-Type` come from the client and prove nothing. Any check that looks at the content, such as file type, malware or policy violations, must therefore run **after** the bytes land and **before** anyone else can fetch them. The design goal is simple to state: *no unscanned byte is ever reachable by another user.* ## A status machine for each upload | Status | Where the bytes are | Who can read them | |---|---|---| | `pending` | being uploaded to quarantine | nobody | | `uploaded` | complete, in quarantine | nobody (internal workers only) | | `scanning` | in quarantine | scan workers | | `published` | in the serving location | viewers allowed by sharing rules | | `rejected` | deleted, or retained for review | trust and safety reviewers | The status lives in the metadata database, and every read path checks it. A CDN or public URL is only ever pointed at the serving location, never at quarantine. ## Triggering processing 1. The client finishes the upload into the quarantine key. 2. The object store emits a **completion event**, and the client also calls `complete`; either one enqueues a processing job keyed by `uploadId` and the object's version or checksum. 3. A worker claims the job, verifies size and checksum, and moves the record to `scanning`. 4. The worker runs the checks listed below. 5. On pass, it copies the object to the serving location (or grants read access to it) and sets `published` in one idempotent step. 6. On fail, it sets `rejected`, deletes or retains the object according to policy, and notifies the uploader. Completion events are commonly **at-least-once**: they can be duplicated, delayed, or lost through a misconfiguration. So: - Every handler is **idempotent**, using conditional status transitions keyed by upload ID. - A **reconciliation sweep** periodically finds records stuck in `uploaded` or `scanning` beyond a threshold and re-enqueues them. - Two triggers (the event and the client call) reduce the chance that an upload is never processed. ## What the scan stage checks - **Real type** from the file's leading bytes (its magic number), not from the extension or declared header; reject types the product does not accept. - **Limits**: byte size, image dimensions, video duration, and decoding cost, so a tiny file cannot expand into an enormous image. - **Archive safety**: nesting depth and expansion ratio, to stop decompression bombs. - **Malware scanning** with a regularly updated engine. - **Content moderation**: automated classifiers, with borderline cases routed to human review. - **Privacy clean-up**: stripping embedded metadata such as location from photos before publishing, where the product requires it. Heavy media work such as transcoding typically runs after the security checks, as a separate stage. ## Operating the pipeline - **Fail closed.** If the scanner is down, uploads can still be accepted into quarantine, but nothing is published; queue age grows and alerts fire. Publishing unscanned content to keep users happy defeats the purpose. - **User experience.** The uploader sees a processing placeholder in their own view; everyone else sees nothing until publication. - **Timeouts and poison files.** A file that crashes or hangs the scanner is retried a bounded number of times, then moved to a dead-letter queue and treated as rejected pending review. - **Metrics.** Time from upload to publish, rejection rate by reason, and backlog age show whether the stage keeps up with ingest. - **Rescans.** When detection rules improve, already-published content may be re-scanned in the background and taken down if it now fails. ## What interviewers listen for A private landing zone, an explicit status that gates every read path, processing triggered by at-least-once events with idempotent handlers and a reconciliation sweep, content-based type detection, and a fail-closed stance when the scanner is unavailable.
- Should a photo-sharing ingest pipeline fail open or fail closed when the malware scanner is unavailable?Closed for publishing: records stay in `scanning`, the backlog grows, and alerts fire on queue age, but nothing unscanned becomes reachable. Accepting new uploads into quarantine can continue because quarantined bytes are harmless, which keeps the upload experience working while the scanner recovers and the backlog drains afterwards.
- Why not rely only on the object store's completion event to start processing?Such events are typically at-least-once and can be duplicated, delayed or silently lost through a misconfiguration. Handlers must be idempotent, keyed by upload ID and object version, and a reconciliation sweep should re-enqueue records stuck in `uploaded` or `scanning`. The client's own complete call is a useful second trigger.
saying these in an interview costs you the question
- Trust the declared Content-Type and file extension to identify the file.
- Publish immediately and take the file down if a later scan flags it.
- Scan the bytes inside the upload request even though clients upload directly.
- Completion events arrive exactly once, so no deduplication is needed.
- If the scanner is down, publish anyway so users are not blocked.