What does the 'you build it, you run it' operating model mean, and what conditions have to be in place for it to actually improve reliability rather than just burning out engineers?
answer
- Werner Vogels / Amazon origin
- closes feedback loop: bug writer = bug responder
- sustainable rotation + tooling + authority required
- burnout if adopted without support
- on-call theater failure mode
basics
~20 sThe team that writes a service's code is also the team that gets paged when it breaks in production, instead of handing it off to a separate ops team. The idea is that owning the pain of your own bugs makes you write more reliable code. It only works well if that team also has good tools, reasonable on-call load, and real authority to fix root causes.
solid answer
~60 s'You build it, you run it' (a phrase attributed to Amazon CTO Werner Vogels, describing Amazon's shift away from separate dev and ops teams) means the engineers who write a service are also directly responsible for operating it in production — carrying its pager, doing its incident response, and owning its reliability — rather than throwing a build 'over the wall' to a separate ops or SRE team. The rationale is a feedback-loop argument: when the same people who wrote a bug also get woken up by it, they feel the operational cost of poor code quality directly and are motivated (and empowered) to fix root causes rather than patch symptoms. For this to actually improve reliability rather than just spread burnout, several conditions matter: the team needs a sustainable on-call rotation with enough people to avoid chronic sleep loss; good observability and tooling (so debugging at 2am is tractable); real authority to prioritize reliability work over new features when needed; and organizational buy-in that on-call load is a legitimate reason to say no to feature deadlines. Without those, you-build-it-you-run-it just relocates operational pain onto feature engineers without giving them the means or authority to reduce it.
go deeper
Should state the core idea correctly (the people who write the code also get paged for it) and intuit that this creates an incentive to write more reliable code.
Should list at least two concrete preconditions (sustainable rotation size, good observability tooling) needed for the model to work well, not just describe the slogan.
Should discuss the burnout/attrition failure mode explicitly, the tension between feature roadmap pressure and reliability work inside the same team, and distinguish the model from simply cutting a central ops team's headcount.
Should be able to design a sustainable on-call and escalation structure for a real org (including where SRE/platform teams fit alongside the model), and reason about when the model is a poor fit (e.g., a team too junior or too small for a given service's criticality).
## What the model is 'You build it, you run it' describes an operating model where the same engineering team that designs and writes a service's code also carries production responsibility for it — being on the on-call rotation, responding to its incidents, and owning its reliability metrics — rather than handing a finished build off to a separate, centralized operations team. The phrase is widely attributed to Amazon CTO Werner Vogels, describing how Amazon moved away from a traditional split where developers wrote software and a separate ops organization ran it in production. It's tightly linked to the DevOps movement broadly, but it's a specific, sharper claim within it: **not just** 'developers and ops should collaborate,' **but** 'there should be no separate ops team for this service at all — the builders are the operators.' ## The feedback loop The mechanism is a **feedback-loop argument grounded in incentive alignment**. In a split model, the team that writes fragile code doesn't experience the 3am pages that fragility causes — someone else does — so there's a structural disconnect between the decision that creates operational risk (shipping code without enough tests, skipping proper error handling, ignoring a flaky dependency) and the person who pays the cost of that decision. You-build-it-you-run-it closes that loop: the engineer who decides to skip a retry-with-backoff on a flaky downstream call is the same engineer who will personally be paged when that flakiness cascades into an outage. Two consequences: - In theory this produces better-engineered software over time, because reliability becomes a first-person concern rather than an externality pushed onto another team. - It also tends to shorten mean-time-to-resolution during incidents, because the responder has full context on the code (they wrote it, or work daily alongside whoever did) rather than a generalist ops engineer unfamiliar with the service's internals. ## Why it fits microservices and single-team ownership Why this fits naturally with microservices and single-team ownership: it's the operational half of the same organizational boundary discussed under service ownership. A team that owns a service's code but not its on-call has **incomplete ownership** — someone else's pain (the ops team's pages) doesn't show up in that team's feedback loop at all. Full ownership, in the Team Topologies / Conway's Law sense, means the team that has decision rights over the code also bears the operational consequences of those decisions. ## The trade-offs The trade-offs are real and often underweighted by companies that adopt the slogan without the supporting conditions. - **The most direct cost is on-call burden landing on feature engineers** who may not have previously carried a pager, and who are also expected to keep shipping features — so reliability work competes directly with roadmap pressure inside the same small team, rather than being absorbed by a separate group. If the team is understaffed for its on-call load (say, four engineers rotating a 24/7 pager for a chatty, failure-prone service), you get chronic sleep disruption, burnout, and attrition — the operational pain gets 'felt' as intended, but with no accompanying authority or slack to actually fix root causes, it just produces exhausted engineers rather than better-engineered services. - **A second cost is fit.** Not every engineer wants to (or is good at) incident response; forcing the model uniformly onto every team regardless of service criticality or team seniority can be a poor fit, especially for junior-heavy teams thrown straight into high-stakes on-call. ## The preconditions that make it work The preconditions that make the model actually work, rather than just relocate pain, are fairly specific. 1. **A genuinely sustainable rotation.** First, enough people that any individual isn't on-call more than roughly one week in four to six, with real backup/escalation paths, so on-call doesn't become a standing tax on personal life. 2. **Second, strong observability tooling** (dashboards, tracing, alerting tuned to reduce false-positive pages) so a 2am incident is tractable rather than a blind archaeology dig through unfamiliar logs. 3. **Third, organizational authority.** The team needs to be able to say 'we're pausing feature work this sprint to pay down the reliability debt causing these pages' and have that respected by product management, not overridden by a deadline every time. 4. **Fourth, blameless postmortems** and a genuine root-cause-fixing culture — if incidents just produce quick patches under pressure to get back to feature work, the feedback loop that's supposed to improve code quality never actually closes. ## Failure modes Failure modes show up concretely. - **A company adopts the slogan**, disbands its central ops team, but doesn't invest in tooling or headcount, and within a year the best engineers on the most failure-prone services start quitting or requesting transfers away from on-call-heavy teams — a visible, expensive form of adverse selection against reliability work. - **Another failure mode is 'on-call theater'.** A rotation exists on paper, but escalation always ends up routing to the two most senior engineers regardless of whose week it officially is, because junior engineers weren't given enough context or authority to actually resolve incidents, defeating the model's feedback-loop purpose while still inflicting the sleep cost broadly. - **A well-cited counter-pattern to watch for.** Companies that talk about 'you build it, you run it' as a cost-cutting move (eliminate the ops headcount, don't replace the capacity elsewhere) rather than a reliability-improvement move tend to see reliability get worse, not better, because the model's benefit specifically depends on the preconditions above, not on the reassignment of pages alone.
- How does 'you build it, you run it' differ from having a dedicated SRE team, and are they mutually exclusive?A dedicated SRE team is a specialized group that often owns cross-cutting reliability practices, tooling, and sometimes shared on-call for less-critical services, whereas you-build-it-you-run-it puts primary on-call directly on the feature team itself. They're not mutually exclusive — many companies (including Google, which popularized SRE) run a hybrid where feature teams carry primary on-call for their own services, and an SRE team acts more like a platform/enabling team providing tooling, standards, and consultation rather than being the 24/7 first responder for every service.
- What's a reasonable staffing rule of thumb to keep an on-call rotation sustainable under this model?A commonly cited guideline is that no individual should be primary on-call more often than roughly one week in four to six, which generally means a team needs at least four to six people able to carry the rotation, plus a secondary/escalation path so a single person isn't the sole line of defense. Below that ratio, sleep disruption and burnout risk rise sharply, and the model starts producing attrition rather than better reliability.
- If a team is drowning in low-value pages (noisy alerts, non-actionable incidents), is more on-call staffing the right first fix?Not usually — adding people to a rotation full of noisy, non-actionable alerts just spreads the pain across more engineers rather than reducing it, so alert quality (better thresholds, deduplication, actionable runbooks, fixing root causes of recurring pages) should be fixed first. Staffing is a legitimate second lever once the alerting itself is signal, not noise.
It's like a chef who also has to eat every dish sent back by an unhappy customer: tasting your own mistakes firsthand is a strong incentive to cook carefully, but only if the kitchen also gives that chef enough staff, decent equipment, and the authority to change the recipe — otherwise you've just made the chef sick without fixing the food.
saying these in an interview costs you the question
- Thinks the model is just 'developers should feel bad about bugs,' with no mention of tooling, staffing, or authority as preconditions
- Assumes it's free to adopt (no headcount or observability investment implied)
- Doesn't recognize burnout/attrition as a realistic failure mode
- Conflates this model with simply eliminating a central ops team without replacing its function
- Can't name any precondition for the rotation itself being sustainable