Should the team run the payments ledger itself on the container platform or buy a managed data service, and what decides it?
answer
- not a capability question
- obligations, not servers
- who is woken at 3am for this data
- has anyone here restored it?
- price an hour down and an hour lost
basics
~20 sCapability is not the question - the platform can run it. The decision turns on who owns durability, restores, version upgrades and the pager, weighed against the control and the money a managed service costs, and against what an hour down and an hour of lost writes are worth.
solid answer
~60 sThe platform can schedule a data-owning workload with identity, per-copy storage and ordered start-up, so "can we?" is not the decision. What you would be taking on is everything that has nothing to do with running a process: durability of the stored data, a retention policy, a restore that someone has rehearsed and timed, version upgrades of the data engine, recovery when a failure domain disappears, and the on-call rotation that carries all of it. Buying converts that into money plus a fixed boundary - less control over tuning, extension availability, upgrade timing and the network path, and a dependency whose exit cost grows with the data. I would decide on four things: whether anyone here has actually restored this data, what an hour down and an hour of lost writes cost the business, whether the workload needs something the managed offering will not do, and whether there is a plausible exit. Running it yourself is the right answer when the store is small and rebuildable, or when a platform team already does this work for several workloads.
go deeper
Recall that running a data store yourself means owning backups and recovery, while a managed service does much of that for you at a price.
Explain what the platform does and does not provide for a data-owning workload, and name the obligations - backups, restores, upgrades - that stay with the team either way.
Argue the case with numbers: what downtime and lost writes cost, what the restore actually takes, and which constraint would rule an offering out.
Separate capability from obligation, price the obligation in people rather than servers, set the standard for which data workloads the organisation hosts itself, and keep the exit cost visible.
## The question is not whether the platform can run it Every modern container platform can run a data-owning workload properly: durable per-copy identity, storage requested per identity and reattached on replacement, ordered start-up and shutdown. Anyone arguing from capability - in either direction - is answering a question nobody asked. The decision is about **which obligations your team takes on**, and those obligations are almost entirely outside the part the platform automates. ## What you take on by running it yourself - **Durability.** Not "the storage is replicated" - that covers a device failing. Durability means the data survives a bad write, a bad release, an accidental deletion and a lost failure domain, which means backups, retention that outlasts discovery, and copies stored independently of the original. - **The restore.** A written procedure, rehearsed, timed, and kept current as the workload changes. Until someone has restored it, the recoverability of the data is an assumption. - **Version upgrades of the data engine.** These are the changes with no safe rollback, because the data moves forward with them. Scheduling them, testing them and having a plan when one goes wrong is recurring senior work, not a ticket. - **Failure-domain recovery.** Deciding in advance where a second usable copy of the data lives, how current it is, and who is allowed to declare that you are switching to it. - **The pager.** Someone is woken up for this workload, and for a data-owning workload the 3am decisions are the expensive kind - the ones where the wrong move loses writes. ## What you buy, and what you give up | | Run it yourself | Buy a managed service | |---|---|---| | Durability and retention | your design, your policy, your drill | a stated commitment you verify, not build | | Restore | your runbook and your rehearsal | a supported operation, still your decision to invoke | | Engine upgrades | your scheduling and your risk | mostly theirs, on their timing | | Tuning, extensions, version choice | full control | whatever the offering exposes | | Network path and locality | you place it | constrained to their footprint | | Cost shape | infrastructure plus staff time | a visible bill that grows with usage | | Exit | move the data yourself | move the data yourself, plus contractual terms | Two lines there are routinely misread. "Durability" does not become someone else's problem: a managed service protects you from losing the bytes, not from your own application writing the wrong ones, so you still care about point-in-time recovery and you still decide when to use it. And "cost" is not the monthly bill against the storage price; it is the bill against the engineering time the obligations above consume, which is invisible on a spreadsheet and very visible in a quarter. ## How I would actually decide 1. **Has anyone on this team restored this data from a copy and timed it?** If not, the team does not yet own the workload it is proposing to run, and that is the cheapest possible thing to fix before deciding anything. 2. **What do an hour of downtime and an hour of lost writes cost?** For a payments ledger both numbers are large and one of them may be regulatory. That pushes hard toward buying, because it turns a durability commitment into something you can hold someone to. 3. **Does the workload need something the offering will not do?** A specific engine version, an extension, a tuning knob, a data-residency constraint, a latency the managed footprint cannot meet. A real blocker here settles it; an imagined one is how teams talk themselves into running things. 4. **Is there a platform team that already does this?** Running one data-owning workload well is expensive; running the fifth costs much less. The answer differs for an organisation that already has the capability and one that would be building it for this workload alone. 5. **What is the exit, in both directions?** Data has weight either way. Knowing roughly what it takes to move the data out - and what it takes to move it in - keeps this from being a one-way door taken by accident. ## When running it yourself is right - The store is small and **rebuildable from an upstream source** - a cache, a derived index, a search index. It looks stateful and is not; treat it as replaceable and stop paying for durability you do not need. - No managed offering exists for what you need, or the ones that do impose a constraint the product cannot accept. - The organisation already operates data workloads on the platform with a rehearsed restore, and this one joins an existing rotation and an existing runbook. - Scale has made the premium the dominant line in the bill, and the team has both the capability and the appetite to take the obligations back deliberately. The answer an interviewer is listening for is not "buy" or "run". It is whether you separated capability from obligation, priced the obligation in people rather than in servers, and named the one fact you would check first.
- A team argues that buying makes data loss someone else's problem. Where is that wrong?A managed service commits to not losing the bytes it stores. It does not protect you from your own application writing wrong data, from an accidental deletion, or from a bad release, and it cannot decide for you when to recover to an earlier point. You still own the recovery decision, and you still have to know how to invoke it.
- Which workloads look stateful but are not, and should simply be run on the platform?Anything whose entire contents can be regenerated from an upstream source of truth: a cache, a derived index, a search index, a rendered artifact set. Losing one costs a rebuild and some latency, not data. Give it identity and storage only if the rebuild is slow enough to matter, and never pay for durability it does not need.
- What would make you reverse this decision two years later?Cost shape at a much larger scale, a constraint the offering will not lift, or the organisation acquiring a platform team that already runs data workloads with a rehearsed restore. The counter-pressure is the exit cost, which grows with the data, so the reversal is worth pricing while the data is still small.
saying these in an interview costs you the question
- Argues from capability: the platform can run it, so it should
- Says managed services are always more expensive, full stop
- Assumes buying makes data loss entirely someone else's problem
- Ignores who carries the pager for this data at 3am
- Compares the bill against storage price, not engineering time
- Treats it as a one-way door without pricing the exit