skip to content

A CTO proposes reorganizing every team around microservices boundaries company-wide within one quarter to 'fix' Conway's Law problems. As a principal engineer, what would make you push back on the timing or scope of that plan?

level: principalimportance: should knowfreq 30%

answer

  1. disruption concentration = velocity dip everywhere at once
  2. boundary uncertainty in young domains
  3. measure pain before reorganizing (MTTR, deploy cadence)
  4. stage the rollout, define checkpoints
  5. reorg cost vs coupling debt cost-benefit

basics

~20 s

Reorganizing everyone at once, fast, is risky: people lose their teammates and context all at the same time, work slows down company-wide, and if the target architecture guess turns out wrong, you've paid the cost twice. It's usually better to reorganize in stages, starting where the pain is worst.

solid answer

~60 s

I'd push back on both speed and blanket scope. A full company reorg in one quarter maximizes disruption at a single point in time: every team simultaneously loses tacit knowledge, working relationships, and ramp-up efficiency, which tanks delivery velocity company-wide right when you need stability to execute the migration itself. It also assumes the target architecture is already correctly known everywhere, which is rarely true — domain boundaries for a rapidly-evolving product area are exactly the ones most likely to be wrong in six months, so locking org structure to them early is expensive to undo. I'd propose a staged approach: apply the inverse Conway maneuver first to the one or two domains causing the most acute pain (e.g., the service with the worst shared-ownership incident history), let that prove out the model and reveal unknowns, and only extend it elsewhere once the target boundaries in those areas are validated by real usage, not just a whiteboard diagram. Stable, low-change domains may not need dedicated reorg at all — the cost of a team boundary change there can exceed the coupling cost it would fix.

go deeper

for a junior

Not typically expected to drive this conversation; a reasonable answer notes that big reorgs are disruptive and slow things down temporarily.

for a middle

Should recognize that not every part of the org needs reorganizing at once, and that a big-bang reorg carries real short-term velocity cost.

for a senior

Should articulate the boundary-uncertainty risk for young/volatile domains and propose starting with the highest-pain area rather than accepting or flatly rejecting the CTO's plan.

for a principal

Should produce a structured counter-proposal: concrete metrics to prioritize which domain reorganizes first, a staged rollout with checkpoints, and an explicit acknowledgment of the trade-off between staging (temporary inconsistency) versus big-bang (concentrated disruption and higher wrong-guess cost).

## Why a principal pushes back There are several independent reasons to push back on a fast, company-wide, one-shot reorganization, and a principal engineer should be able to articulate them as concrete costs and risks, not just gut discomfort with change. ## The four objections 1. **The first is disruption concentration.** Any team reorganization has a real, measurable cost: people lose daily proximity to former teammates, tacit knowledge that lived in informal relationships (who knows the weird edge case in this billing code, who to ask about that flaky test) gets scattered or lost outright, and new teams need weeks to months to rebuild the trust and shared context that let them move fast. Doing this to every team simultaneously means the entire company's delivery velocity dips at the same time, which is exactly the wrong moment to also be executing a technically risky microservices migration that needs stable, high-context teams to get right. A staged rollout — reorganizing one or two domains first — contains that velocity dip to a smaller blast radius and lets the organization learn from the first attempt before repeating it elsewhere. 2. **The second is boundary uncertainty.** The inverse Conway maneuver assumes you know the target architecture well enough to organize teams around it durably. For a mature, slow-changing domain (say, a well-understood billing ledger), that assumption often holds. For a young or fast-evolving domain — a new product line still finding its shape, or a domain where the business itself is actively deciding how to segment its offering — the 'correct' service boundary is genuinely unknown and likely to be wrong in six to twelve months as the product evolves. Locking a team's boundary to a guessed architecture in a volatile domain means paying the reorg cost, discovering the guess was wrong, and then paying it again to fix it — worse than not reorganizing that domain yet and accepting some Conway's-Law-driven messiness there until the domain stabilizes enough to commit to boundaries. 3. **The third is uneven pain distribution** — not every service or team actually has a Conway's Law problem worth solving right now. A rarely-changed, single-team-owned internal service that works fine has no coupling debt to pay down; reorganizing around it anyway is pure cost with no corresponding benefit. A principal engineer should push the CTO to identify where the actual pain is concretely measurable — services with contested on-call history, unusually slow deploy cadence relative to change volume, or repeated 'whose bug is this' incidents — and prioritize the reorg where evidence, not architectural idealism, shows the current org structure is actively causing damage. 4. **The fourth is capacity and change-management realism:** a quarter is an aggressive timeline for a full reorg plus the actual technical migration work it's meant to enable, and squeezing both into the same window risks doing neither well — teams spend the quarter in reorg limbo (new manager, new codebase, unclear ownership) rather than either stabilizing under the old structure or productively executing the new one. Sequencing — reorg first, let teams stabilize for a defined period, then execute the heaviest migration work — tends to produce better outcomes than parallelizing both under time pressure, even though it's slower calendar-wise. ## The other side of the trade-off The trade-off a principal has to weigh explicitly, though, is that staged reorgs also have a cost: - prolonged periods where some domains are reorganized and others aren't create inconsistency (two different operating models coexisting, cross-team friction between 'new model' and 'old model' teams); - and moving too cautiously can let a genuinely painful coupling problem fester for another year while waiting for 'enough evidence.' The right answer isn't 'never reorganize broadly' — it's **sequencing by measured pain and boundary confidence**, with an explicit plan for extending the model once the pilot area validates it, rather than either extreme (reorganize everything at once on a guess, or never commit to any organizational change and let Conway's Law's problems compound indefinitely). ## How to operationalize the pushback A concrete way to operationalize the pushback: ask the CTO to name the two or three services with the worst current symptoms, using data — - incident MTTR - deploy frequency versus commit volume - on-call page volume - contested ownership incidents — and propose the inverse Conway maneuver there first, with a defined checkpoint (e.g., one or two quarters) to evaluate before deciding whether and how to extend it company-wide. This reframes the conversation from 'reorg everyone now because Conway's Law says so' to 'apply the lever surgically where the data shows it's needed, learn from it, then scale the successful pattern' — which is both lower-risk and more likely to produce organizational buy-in, since teams see a working proof point rather than a company-wide mandate justified by theory alone.

  • What data would you actually want in hand before deciding which domain to reorganize first?
    Incident data showing contested or ambiguous on-call response (multiple teams paged, or delayed response due to ownership confusion), deploy frequency relative to commit volume as a proxy for coordination drag, and PR review latency for cross-team versus single-team changes on the same service. These turn 'this feels painful' into a ranked, evidence-based list of where the inverse Conway maneuver will pay off fastest.
  • How would you handle the inconsistency of running two different team operating models (reorganized vs. not-yet-reorganized) at the same time?
    Set clear, communicated expectations that the staged rollout is deliberate and temporary, with a named checkpoint for extending it, so teams don't read the inconsistency as indecision or favoritism. It also helps to keep the interaction contracts (how a not-yet-reorganized team's services are consumed) stable during the transition, so the inconsistency is felt as an internal org detail rather than something that leaks into every team's day-to-day integration work.
  • Is there a domain where you'd actively recommend against ever applying the inverse Conway maneuver?
    Yes — a stable, low-change, single-team-owned internal service with no history of ownership conflict or coordination pain has no coupling debt to justify the reorg cost, so applying the maneuver there is pure overhead with no corresponding benefit. The maneuver should be reserved for domains where organizational friction is measurably slowing delivery or causing incidents, not applied uniformly as a matter of architectural principle.

It's like renovating every room in an occupied hotel on the same weekend instead of one wing at a time: you guarantee total chaos and no functioning rooms simultaneously, versus proving the new room layout works in one wing, learning from complaints, and rolling it out gradually while the rest of the hotel keeps operating.

saying these in an interview costs you the question

  • Agrees to the full-scope, one-quarter reorg without asking what pain it's meant to solve or how it's measured
  • Treats reorganizing as free or low-risk regardless of scale
  • Has no counter-proposal (staged rollout, pilot domain) — just objects without an alternative
  • Doesn't distinguish between stable/mature domains and volatile/young ones when discussing reorg timing
  • Can't name any concrete metric that would indicate a domain actually has a Conway's-Law-driven problem worth fixing

context