skip to content

Platform Storage Services

Object, block and file storage as distinct rentals, plus durability, tiers, versioning and access policy. Asked because a badly chosen shape or an open policy is found in production, not in review.

on this pageshow

questions

23

An external scanner reports that your receipts store is readable by anyone on the internet — which grant produces that, and what was not leaked?

level: juniorimportance: must knowfreq 70%

answer

  1. stores are closed until someone opens them
  2. the finding is a grant, not a breach
  3. nobody authenticated, so nothing was stolen
  4. knowing a key differs from discovering keys
  5. a per-object grant hides from a store-level review

basics

~20 s

A grant written on the store side — at store or object level — naming any caller as allowed to read produces public access. Nothing was leaked: the store is doing exactly what its policy says, so no credential was stolen and no defect was exploited.

solid answer

~40 s

Object stores start closed, so a store is only reachable by an anonymous request if someone wrote a grant on the store side naming an unrestricted caller — either on the store's policy or on individual objects. The scanner needed no credential, which is the whole finding: this is not a stolen key, not an exploited bug, and not a platform failure, so rotating secrets fixes nothing. Two separate capabilities are worth telling apart: reading an object whose exact key you already know, and listing the store to discover keys. A store that allows listing turns a guessing problem into a download. Treat the data as copied, remove the grant, then find who wrote it and why.

go deeper

for a junior

Know that object stores deny by default, so public access is always something someone granted on the store side. Be able to say that an anonymous reader means no credential was stolen.

for a middle

Explain the difference between being able to read a known key and being able to list the store, and where a grant can hide when the store's own policy looks clean.

for a senior

Handle it as a data-exposure incident: remove the grant, bound the exposure from access records, trace the change in the management audit trail, then separate public assets from private data.

for a principal

Decide the standing rule. One store per exposure class, a deliberate exception path for genuinely public assets, and a preventive control so the question stops arriving from strangers.

The mobile backend keeps user-uploaded receipts in an object store and serves them to the app. Someone running an internet-wide scan reports that the receipts are readable without any credential. Understanding what that finding is — and is not — is a first-screen question for anyone who has deployed a service that writes files. ## What "public" actually means here Object stores are closed by default: an unmatched request is denied. So a public store is not a gap, it is a **grant**. Somewhere on the store side there is a rule naming an unrestricted caller — *anyone*, or the anonymous caller — as allowed to read. It can sit in two places: - on the **store's policy**, which applies to everything in it (or to a key prefix inside it); - on an **individual object**, granted one at a time, often by whatever code uploads it. The second is the nastier variant, because a store whose policy reads perfectly tight can still contain thousands of individually exposed objects, and the store-level view shows nothing. ## Readable is not the same as listable | Capability granted | What an anonymous caller can do | Practical effect | |---|---|---| | Read an object | fetch an object whose full key it already knows | exposure depends on whether keys are guessable | | List the store | enumerate every key in the store | turns the store into a downloadable archive | Teams often lean on unguessable keys — a long random component in the object name — and call that acceptable. It is a real mitigation and a weak one: keys escape through links, logs, referrer headers and shared screenshots, and a single accidental listing grant deletes the mitigation entirely. If the scanner produced a **list** of your objects, assume complete exposure. If it produced one object it already knew about, exposure is narrower but the grant is the same defect. ## Why nothing was leaked This is the part candidates get wrong. The scanner authenticated as nobody. That means: - **no credential was stolen**, so rotating keys, passwords or tokens changes nothing about the finding; - **no software defect was exploited**, so there is no patch to apply; - **the platform behaved correctly** — it enforced the policy it was given, and the policy said yes. The defect is entirely in the grant, and the fix is entirely in the grant. Rotating secrets in response is a common and expensive misdiagnosis that leaves the store open while everyone is busy. ## How stores end up like this The recurring routes are mundane: 1. Someone loosened the policy to unblock a demo, a browser upload or a mobile client that could not sign requests, and it was never tightened. 2. A rule meant to widen the **resource** — every key under a prefix — was written to widen the **caller** instead. 3. A pattern for serving a public website's assets was copied onto a store that also holds private data, because one store was cheaper to operate than two. 4. The upload path grants each new object out individually, so the exposure grows with traffic and no one edits anything. ## What to do, in order 1. **Remove the grant** and confirm an unauthenticated request now fails. Stop the bleeding before investigating. 2. **Assume the data is copied.** Public means public for as long as the grant existed; the response is a data-exposure response, not a configuration cleanup. 3. **Read the audit trail** of reads against the store to bound what was fetched and from where — accepting that access logging may not have been on, which is itself a finding. 4. **Find the change** in the management API's audit record: who wrote the grant, when, and with what stated purpose. That tells you whether this was one mistake or a pattern. 5. **Separate the stores.** If genuinely public assets and private user data share one store, the long-term fix is two stores with different policies, not a cleverer single policy. The last step is the one that stops a repeat. A store that holds only public assets can be public without anyone losing sleep; a store that holds user documents should be one whose policy nobody has a reason to loosen.

  • The store's policy looks tight, yet objects are still fetchable anonymously. Where else can the grant be?
    On the objects themselves. Many stores allow a grant per object, and an upload path can set one on every write, so exposure accumulates without anyone editing the store's policy. Audit at object level, or use an account-level setting that refuses public grants outright, since that applies whatever individual objects say.
  • Is relying on long random object keys instead of a policy an acceptable control?
    Only as a second layer. Unguessable keys narrow discovery, but keys travel in links, logs, referrer headers and screenshots, and one listing grant makes them all enumerable. The access decision should be a policy the store enforces; treat key entropy as defence in depth, never as the control itself.

saying these in an interview costs you the question

  • Calls it a breach and starts rotating credentials first
  • Says the platform or a bug exposed the store
  • Treats unguessable object keys as equivalent to a policy
  • Assumes a clean store policy proves no object is public
  • Fixes the grant and skips assessing what was downloaded
open as a page

An object store advertises durability with many nines, but ingest requests fail for an hour - which promise did that figure never make?

level: juniorimportance: must knowfreq 76%

basics

~20 s

Durability is about bytes surviving; availability is about bytes being reachable now. A many-nines durability figure estimates how unlikely it is that the store loses an object, and says nothing about an hour of failed requests.

open as a page

Which storage shape fits a transcoder's scratch space while rendering, and which fits the finished videos many clients fetch?

level: juniorimportance: must knowfreq 84%

basics

~20 s

Scratch belongs on a block volume: it attaches to one machine and behaves like a local disk, so seeks and in-place writes are cheap. Finished videos belong in an object store, fetched by key over HTTP by any number of readers.

open as a page

Weekly-read application logs are moved to the coldest archive storage tier to save money — what goes wrong?

level: juniorimportance: must knowfreq 66%

basics

~20 s

A colder tier discounts rent by charging separately for reads and by making them slow. Data read every week pays a retrieval charge every week and waits for a restore each time, so the bill usually goes up, not down.

open as a page

A job overwrote every document in a versioned object store with a corrupt render — how do you get the originals back?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Versioning kept each pre-overwrite copy as a previous version under the same key, so recovery is promoting that version back to current, key by key. Nothing is restored from a backup, and every retained version keeps being stored and charged.

open as a page

A policy attached to your receipts store grants read to a caller with no credential — what can an identity-attached permission never do?

level: middleimportance: must knowfreq 62%

basics

~20 s

A policy attached to the store itself can name callers its owner does not administer, including an anonymous one, so it grants access to a caller carrying no permission of its own. An identity-attached permission only widens what an identity you already manage may do.

open as a page

A nightly job overwrote a day of raw events in a store with many-nines durability - why did that durability figure not help?

level: middleimportance: must knowfreq 58%

basics

~20 s

Durability defends the bytes against hardware, not against you. The store did exactly what it was asked and then replicated the overwrite faithfully to every copy, so a high durability figure is not a backup and never rewinds a bad write.

open as a page

Renaming the prefix of 200,000 finished renders in an object store runs for hours and bills per operation — what is the store doing?

level: middleimportance: must knowfreq 62%

basics

~20 s

There is no rename. The key namespace is flat, so the store lists every key under the old prefix and performs a copy to the new key followed by a delete of the old one — per object. Cost and duration scale with object count, not bytes.

open as a page

A lifecycle rule moves objects to a colder tier on day 30 and deletes them on day 45 — what is billed?

level: middleimportance: must knowfreq 54%

basics

~20 s

Both actions bill in full. A colder tier carries a minimum storage duration, so an object deleted after fifteen days there is still charged for the whole minimum, plus a per-object fee for the move — often more than the rent the rule saved.

open as a page

A delete call on a versioned object store returned success, yet the stored bytes are still billed — what actually happened?

level: middleimportance: must knowfreq 62%

basics

~20 s

The delete wrote a marker that became the key's current version and hid it from ordinary listings and reads. The earlier versions still exist and still occupy paid storage until something removes them by version identifier.

open as a page

Every object is replicated to a second region, yet a mistaken delete removed the documents there too — why?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Replication copies operations, so a deletion is reproduced faithfully at the destination. It defends against losing a failure domain, never against a wrong call. Versioning, a retention lock or a genuinely separate backup are what survive a mistake.

open as a page

What extra failure does an object store survive as its copies move from one zone, to several zones, to another region?

level: middleimportance: should knowfreq 60%

basics

~20 s

Copies inside one zone survive a failed disk or machine. Copies spread across zones in a region survive losing a whole zone. A copy in another region survives losing the region. Each step costs more and is usually a setting you choose.

open as a page

In an object store, why is reading part of a large object cheap while changing part of it usually rewrites the whole object?

level: middleimportance: should knowfreq 48%

basics

~20 s

Reads can be served for any byte span because the store already holds the bytes and can return a slice. Writes create a new immutable object under the key, so there is no place to patch — the store replaces what the key points at.

open as a page

Your team must make a public grant on any store impossible rather than merely reviewed — what does an account-level override give you?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A preventive control. An account-level override refusing public grants sits above every store in the account, so a rule naming an unrestricted caller is rejected when written or ignored when evaluated, whatever a project team puts in a store's policy.

open as a page

A reporting job whose identity allows reading every store is refused on the receipts store alone — how did two permissions combine to produce that?

level: seniorimportance: should knowfreq 52%

basics

~20 s

The request is evaluated against both sides at once. A matching explicit deny in the store's own policy ends the request no matter how broad the caller's identity permission is, and a condition on a store-side allow that the request fails to satisfy has the same effect.

open as a page

Why can an object store's asynchronous copy in a second region be missing objects that ingest already wrote successfully?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Because the write is acknowledged once the near copies are durable, and the far copy is applied afterwards by a background worker. That copy trails by a variable lag, so the newest objects have not reached it yet.

open as a page

A second transcoding worker cannot attach the block volume the first still holds — what is the limit, and how do you get around it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

A block volume attaches to one machine at a time. The filesystem on it — its cache, its free-space map, its journal — is a single-machine structure, so two machines writing would corrupt it. Share through an object store or a shared mount instead.

open as a page

Tiering four hundred million small log objects into a colder class made the storage bill rise — why?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Colder tiers bill a fixed overhead per object on top of its real size, and charge a fee per object moved. Across hundreds of millions of tiny objects that overhead dwarfs the payload, so the cheaper per-gigabyte rate applies to far more billed gigabytes.

open as a page

An audit needs a year of logs by tomorrow and they sit in the coldest tier — what do you check first?

level: seniorimportance: should knowfreq 42%

basics

~10 s

Two independent numbers decide it: the tier's restore latency, which says whether tomorrow is achievable, and the retrieval charge plus per-object restore requests, which says what it costs. Check both before promising the deadline.

open as a page

A retention lock on stored records cannot be shortened, and an erasure request arrives for one of them — what now?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A lock that cannot be shortened keeps the bytes until its retained-until date passes, so the erasure obligation has to be met another way — usually by destroying the key material that makes the record readable, and by never placing erasable personal data under such a lock in the first place.

open as a page

A transition rule with a thirty-day age has not moved an object on day 31 — why might that be?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

Lifecycle rules are evaluated asynchronously on the provider's own schedule, so an object becomes eligible at the age and moves some time afterwards. The other common cause is a filter mismatch: the object's key never matched the rule.

open as a page

Ten workers share one mounted tree holding millions of tiny frame files, and throughput collapses — what is the dominant cost?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Round trips, not bytes. On a shared mount every open, stat, create and directory lookup crosses the network, and with tiny files those metadata operations vastly outnumber the data transferred. The fix is to make the unit of work bigger.

open as a page