You join a company whose 'microservices' cannot be deployed independently and whose service boundaries look arbitrary. How would you diagnose whether the organization's structure is the root cause, and what would you propose?
answer
- Symptom: distributed monolith — lockstep deploys, arbitrary seams
- Evidence: deploy coupling, feature spread, change coupling, data write access
- Ask 'why does this boundary exist?' — acquisitions, vendors, old teams
- Ownership first, data second, headcount last
- Validate per context with lead time / blocked work; be willing to stop
basics
~20 sLook for team boundaries hiding inside the architecture: which teams must talk to ship one feature, who owns each database, and which services always release together. If the seams match old team or company boundaries rather than business capabilities, the org is the cause — fix ownership and data boundaries, not just the code.
solid answer
~50 sI'd gather evidence before proposing a reorg. Three sources: (1) delivery data — how many teams and services a typical feature touches, lead time, deploy coupling (which services always ship together), and where work sits blocked; (2) code and history — change-coupling analysis from version control (files/services that change together), ownership concentration, and shared write access to data stores; (3) socio-technical mapping — a value-stream map of one real feature plus a diagram of actual communication paths, compared to the org chart and to the org's history (acquisitions, outsourcing, reorgs). Org-shaped architecture shows up as boundaries that align with old teams or vendors and cut across business capabilities, plus a shared schema. If confirmed, I'd propose an incremental Inverse Conway Maneuver: choose boundaries from domain analysis and observed change coupling, move ownership before moving people, split the data behind each boundary, give each boundary its own pipeline, and validate with lead time, change-failure rate, and cross-team blocked work before proceeding to the next context.
go deeper
Notice that teams and databases are shared and that every feature needs several teams; suggest clearer ownership.
Collect concrete evidence — which services deploy together, who writes to which database, how many teams a feature touches — and propose cross-functional ownership per service.
Separate org-caused coupling from data and process coupling, propose an incremental Inverse Conway Maneuver with data split and ownership-first steps, and name validation metrics.
Run it as a socio-technical change program: evidence gathering, hypothesis testing against alternative causes, reversible sequencing, explicit cost and risk of reorg, funding/political dynamics, stop criteria, and the option of consolidating instead of splitting.
## Framing The symptom — services that must be released together, with boundaries nobody can justify technically — is the classic **distributed monolith**. Conway's Law offers a hypothesis: the boundaries were produced by the *communication structure* (past or present), not by domain analysis. But this must be tested, because the alternative causes (a rushed decomposition, a framework's default layout, an over-eager architect) call for different fixes. ## Step 1 — Establish the symptom precisely Before theorizing, quantify: - **Deploy coupling**: from the release history, which services are deployed within the same window, repeatedly? A correlation matrix of deployments exposes lockstep groups. - **Feature spread**: pick 5-10 recently delivered features. How many services, repositories, and *teams* did each require? A healthy stream-aligned setup shows most features inside one team. - **Flow metrics**: lead time for changes, deployment frequency, change failure rate, time to restore (the four DORA measures), broken down per team and per service. Wide variance across services owned by one team suggests overload; uniformly bad numbers with high cross-team spread suggest structural coupling. - **Blocked work**: how often is work waiting on another team, and for how long? Cross-team wait time is the direct cost of a misplaced boundary. ## Step 2 — Look for the org fingerprint - **History interview**: ask long-tenured staff *why* each boundary exists. Answers like "that was the Bangalore team", "that came with the acquisition", "the vendor built that", "the DBAs owned that schema" are direct evidence of Conway effects. Boundaries usually outlive the teams that made them. - **Compare three maps**: (a) the service dependency graph, (b) the current org chart, (c) the *actual* communication graph (who is in whose chats, reviews, incident channels). If (a) matches an *older* version of (b), you have a fossilized org. - **Skill silos**: separate frontend/backend/DBA/QA/ops teams reliably produce layered boundaries; every feature then crosses every team. That's a Conway signature. - **Data ownership**: who can write to each store? Shared write access is the single strongest predictor of a distributed monolith, and usually traces to a central DBA or data team. - **Change coupling from version control**: mine the history for files/modules that change in the same commit or same pull request across service boundaries. Boundaries that are constantly co-changed are wrong boundaries, regardless of how they were drawn. - **Ownership churn / contributor spread**: components edited by many teams with no clear owner correlate strongly with defects (this is the empirically supported part of the Conway literature). ## Step 3 — Distinguish causes | Evidence | Likely cause | Fix | |---|---|---| | Boundaries match old teams/vendors/acquisitions; features cross all teams | Org-shaped architecture | Reorg + rebound boundaries (Inverse Conway) | | Boundaries match technical layers (api/service/data) | Skill-siloed teams or layered-thinking default | Cross-functional stream-aligned teams | | Shared schema, everything else fine | Data coupling, not org | Data decomposition, ownership rules | | One team owns everything and is drowning | Cognitive overload | Reduce scope, platform, split domain | | Boundaries fine, but a manual release train forces lockstep | Process/tooling | Independent pipelines, contract testing | The honest answer is often *several at once*; sequencing matters more than picking one. ## Step 4 — Proposal I would propose a staged, reversible program rather than a big-bang reorg: 1. **Agree the target boundaries on evidence.** Use domain analysis (bounded contexts, event storming) *cross-checked* against observed change coupling. Where the two disagree, investigate — the history usually knows something the whiteboard doesn't. Prefer fracture planes: business subdomain, change cadence, compliance, risk, performance isolation, user persona. 2. **Fix the cheapest structural blockers first.** Independent pipelines, contract/consumer-driven tests, removing the shared release train. Sometimes this alone restores independent deployability and buys time. 3. **Move ownership before moving people.** Reassign codeowners, review gates, and on-call for a candidate context. This tests the boundary at low cost: if the split immediately generates a stream of cross-team tickets, the boundary is wrong and can be reverted in a week. 4. **Split the data.** No team split is real until each boundary owns its store. Plan dual-write/backfill/cutover explicitly; this is usually the longest pole and where these programs fail. 5. **Then restructure teams** into long-lived cross-functional stream-aligned teams, one per validated context, adding a thin platform capability to strip extraneous load and an enabling capability to close skill gaps created by dissolving the old silos. 6. **Make interaction modes explicit**: time-boxed collaboration while an interface is discovered, then X-as-a-Service as the steady state. 7. **Instrument and stop early.** Success metrics: percentage of features delivered inside one team, deploy-coupling groups shrinking, lead time and change-failure rate, cross-team blocked hours. If the numbers don't move after the first context, re-examine the hypothesis instead of continuing. ## Risks to name explicitly - **Reorg cost is real**: knowledge loss, attrition, a productivity dip lasting a quarter or more. Don't start unless the org shape is demonstrably the bottleneck. - **Boundary mistakes are now people-shaped** and expensive to reverse — hence ownership-first, incremental. - **Political capture**: reorgs get hijacked by headcount and title concerns. Tie every move to a delivery metric and publish the rationale. - **Doing nothing is an option**: if the business is not constrained by delivery speed, the cheaper answer may be to consolidate the distributed monolith back into fewer deployable units rather than to reorganize. ## The senior signal A strong answer resists jumping to "reorg into stream-aligned teams". It gathers evidence, distinguishes org-caused from process- or data-caused coupling, sequences data and ownership changes ahead of headcount changes, keeps changes reversible, and states the conditions under which it would abandon the plan.
- What single piece of evidence most strongly indicates a distributed monolith caused by organizational structure?Multiple teams holding write access to the same data store, combined with services that are always deployed together. Shared writable data means the coordination the team split was supposed to remove is still mandatory; if the sharing traces to a former central DBA/data team, the organizational origin is confirmed.
- How can version-control history help you choose better boundaries?Change-coupling analysis: mine commits/PRs for files or modules that repeatedly change together. Co-changing artifacts on opposite sides of a boundary indicate the boundary is misplaced; clusters of co-change indicate a natural cohesive unit. Ownership churn and number of contributing teams per component additionally predict defect-proneness.
- When would you deliberately NOT reorganize, despite finding an org-shaped architecture?When delivery speed isn't the binding business constraint, when the domain is still poorly understood (you'd hard-code a guess into people), when the org is small enough that communication is already cheap, or when the cheaper fix is to recombine the distributed monolith into fewer deployable units and fix process/tooling coupling first.