skip to content

How should service granularity evolve over the lifetime of a system, and what concrete signals tell an engineering organization it's time to split a service apart or merge services back together?

level: principalimportance: should knowfreq 55%

answer

  1. granularity is living, not fixed
  2. change-coupling mining from VCS history
  3. strangler fig for safe incremental splits
  4. facade for safe incremental merges
  5. Segment 2018 merge-back example

basics

~20 s

Granularity isn't a one-time decision - as a system and its teams grow, the 'right' size for a service changes. You split a service when a real, measurable reason shows up (different scaling needs, a new team owning part of it, slow deploys from unrelated changes). You merge services back when they always change together and the split is just adding overhead without any benefit.

solid answer

~60 s

Granularity should be treated as a living architectural decision, not a one-time design choice: the right boundary for a given capability shifts as traffic patterns, team structure, and business requirements change, so an organization needs both a way to detect when a boundary has gone stale and a safe mechanism to change it. Split signals include divergent scaling needs backed by real metrics, a new or reorganized team taking ownership of part of a service, repeated incidents where an unrelated change inside a shared service caused an outage or blocked a release, and change-coupling analysis showing one part of the codebase changes far more often than the rest. Merge signals include services that are almost always deployed together, a high ratio of cross-service calls relative to internal logic, and services whose combined operational cost clearly outweighs any independence benefit they're delivering. The safe mechanism for both directions is incremental: use the strangler-fig pattern to peel capability out of a monolith-like service gradually behind a stable interface, or, for merging, route traffic through a facade while consolidating the underlying implementation, so the system stays deployable and testable throughout the transition rather than requiring a risky big-bang rewrite.

go deeper

for a junior

Should understand that a service's 'right' size can be a wrong answer later on, and that changing size is possible, without needing to name specific migration patterns.

for a middle

Should name at least one concrete split signal (e.g., a new team taking ownership, or divergent scaling) and one concrete merge signal (e.g., services always deployed together), and understand a gradual migration is safer than a rewrite.

for a senior

Should describe the strangler fig pattern concretely and explain how it keeps the system safely releasable during a split, and should be able to reason about the interim costs of a granularity change in progress.

for a principal

Should treat granularity review as an ongoing organizational practice, cite concrete data sources for detecting drift (change-coupling mining, incident postmortems), and discuss real-world precedent for both splitting and merging directions, including the org-level trade-offs each implies.

## Granularity is a living decision Service granularity is not a property you get right once at design time and then leave alone — it's a decision that should be revisited as three things change over a system's life: 1. **Traffic and scaling patterns** — a feature that was low-traffic at launch can become the system's dominant cost center. 2. **Team structure** — teams split, merge, and reorganize, and service boundaries that don't track team boundaries create friction. 3. **Business requirements** — a capability that used to be a minor feature can become a standalone product needing its own release cadence, compliance boundary, or even separate vendor/SLA. Treating granularity as fixed leads to architecture that's well-matched to the org and business of two years ago and increasingly mismatched to the org and business of today — this mismatch is exactly what produces the macroservice and nanoservice smells covered elsewhere on this topic. ## The signals that a boundary has gone stale Detecting that a boundary has gone stale requires concrete signals, not intuition. **On the 'should split' side:** - a **change-coupling analysis** (mining version-control history for files/modules that are almost always modified together across many commits) can reveal that a chunk of a service's codebase has effectively stopped changing in lockstep with the rest, meaning it's already behaving like a separate capability even though it's still packaged together - **incident postmortems** are another rich signal — repeatedly seeing 'an unrelated change in module X caused an outage in module Y' or 'we couldn't hotfix Y because it required a coordinated release with X's owning team' points directly at a coupling boundary that should become a service boundary **On the 'should merge' side**, the signals mirror the chattiness and macroservice discussions: - a consistently high ratio of cross-service network calls to actual business logic performed - deploy logs showing two services are released together in lockstep on nearly every change - on-call data showing incidents in one service are almost always resolved by also touching the other ## The mechanism matters as much as the decision Once a signal justifies a change, the mechanism matters as much as the decision. A big-bang rewrite — freezing feature work to carve a new service out of an existing one in one large migration — carries high risk: it takes the team out of shipping value for an extended period, and any mistake in the new boundary only surfaces after the full cutover, when it's expensive to reverse. The industry-standard alternative, the **strangler fig pattern** (named by Martin Fowler after the strangler fig vine that grows around a host tree and gradually replaces it), routes traffic through a facade or proxy and incrementally moves individual capabilities behind that facade into the new service one at a time, while the old code path keeps serving whatever hasn't been migrated yet. This keeps the system releasable and rollback-able throughout the migration — if the new piece misbehaves, traffic for that specific capability can be routed back to the old implementation without touching anything else. ## The cost of a change in progress Splitting-in-progress carries real interim cost; merging-in-progress carries a mirrored cost. | Direction | What you pay while it is under way | |---|---| | Splitting-in-progress | for a while you have to maintain both the old and new code paths, keep data consistent across the boundary being carved (often via dual-writes or change-data-capture events during the transition), and carry the cognitive overhead of a boundary that's half-drawn | | Merging-in-progress | consolidating two services' data models usually means one has to be treated as canonical and the other's data migrated or reconciled, and any external consumers of the service being absorbed need a compatibility period (often via an adapter or the same facade approach) before they can be pointed at the new consolidated API | Neither direction is free, which is exactly why granularity changes shouldn't be undertaken on a hunch — the concrete detection signals exist precisely to make sure the migration cost is justified by a real, evidenced problem rather than an aesthetic preference for 'cleaner' boundaries. ## Granularity review as an ongoing practice A well-documented real pattern is large platform companies' publicized experience of extracting high-traffic subsystems (e.g., checkout, inventory) out of a larger core service specifically once those subsystems' independent scaling needs and team ownership became clear from production data — and, conversely, some companies have published accounts (Segment's widely discussed 2018 write-up about consolidating a large microservice fleet is a commonly cited example) of deliberately merging services back together after finding the operational overhead of an over-decomposed service fleet exceeded the independence benefit for their actual team size and traffic. The organizational takeaway is that a mature engineering org should treat granularity review as an ongoing practice — periodically revisiting service boundaries using change-coupling and incident data, similar to how it revisits capacity or dependency-upgrade needs — rather than as a decision made once at a system's inception and never revisited.

  • How does change-coupling analysis from version-control history concretely identify a splitting or merging opportunity?
    By mining commit history for files or modules that are frequently modified together in the same commit or pull request, you can build a coupling graph independent of the current package/service structure - a cluster of files that always change together but currently spans a service boundary suggests a merge, while a cluster that has stopped changing with the rest of its current service suggests it's already a de facto separate capability ready to split out.
  • Why is the strangler fig pattern generally safer than a big-bang extraction, and what does it cost you?
    It's safer because traffic is routed through a facade and moved to the new service incrementally, capability by capability, so a mistake in the new boundary only affects the piece just migrated and can be rolled back by routing that traffic back to the old path. The cost is a longer transition period during which both the old and new implementations, and often duplicated or synchronized data, must be maintained simultaneously, which is real ongoing complexity until the migration completes.
  • What's a concrete reason a company might deliberately merge microservices back together rather than keep splitting?
    If the operational cost of running many small services - pipelines, on-call, monitoring, cross-service debugging - has grown faster than the actual independence benefit the team is realizing (e.g., services still deploy together in practice, or the team is too small to truly own each one separately), consolidating reduces that overhead. Segment's publicly discussed 2018 move away from a large microservice fleet toward more consolidated services is a concrete real-world example of this trade-off being made deliberately.

It's like renovating a house you live in, not building a new one from scratch: you don't tear down the whole structure to move one wall - you shore up a section, cut the new opening, verify it's sound, and only then remove the old wall, so the house stays livable the entire time.

saying these in an interview costs you the question

  • Treats the initial service boundary as permanent rather than something to revisit
  • Recommends splitting or merging without naming any concrete detection signal or data source
  • Assumes only 'split' is ever the right direction and doesn't consider merging back
  • Proposes a big-bang rewrite as the default way to change a service boundary
  • No mention of how to keep the system releasable/rollback-able during the transition

context