skip to content

A team notices that a single request to their order-checkout API triggers 15 synchronous calls across 6 microservices before it can respond. What is this smell usually called, what causes it at the granularity level, and how would you address it?

level: middleimportance: must knowfreq 75%

answer

  1. N+1-style call chains
  2. latency multiplies across hops
  3. granularity symptom, not root cause itself
  4. event-driven / cache to cut sync calls
  5. BFF/aggregator parallelizes fan-out

basics

~20 s

This is called 'chatty' service communication - too many small back-and-forth network calls needed to do one piece of work. It usually means services were split too finely along lines that don't match how work actually flows, so simple operations require calling many of them in sequence. Fixes include combining related services, batching calls, or having one service own more of the workflow.

solid answer

~50 s

Chatty communication happens when a business operation is scattered across too many fine-grained services, forcing a caller (or orchestrator) to make many sequential or fan-out network calls to complete one logical unit of work. At the granularity level it's a direct symptom of over-decomposition: services split along technical or entity lines rather than around cohesive business capabilities, so data and logic needed together live in different services. It matters because each hop adds latency (multiplying tail latency across the chain), a new failure point, and coupling (the caller must know the call sequence and handle partial failures). Fixes include merging services that are always called together, introducing an aggregating/composite service or BFF for read-heavy fan-out, using asynchronous events instead of synchronous request/response where real-time consistency isn't required, and caching read-only reference data locally instead of calling out for it on every request.

go deeper

for a junior

Should recognize that many sequential network calls for one user action is slower and more fragile than a few, and name at least one basic fix like combining calls or caching.

for a middle

Should explain latency multiplication across hops and name at least two distinct mitigations (e.g., caching/events and an aggregating layer).

for a senior

Should connect chattiness explicitly to a granularity root cause, discuss the consistency trade-off of event-driven fixes, and describe cascading-timeout/retry-storm failure modes.

for a principal

Should discuss how to prevent chattiness organizationally (API/domain design review, tracing-based SLO ownership) and weigh when merging services is right versus when parallel aggregation or caching is sufficient without re-coupling ownership.

## What chattiness is Chatty communication is what happens when completing one logical business operation requires a caller to make many small, sequential (or fanned-out) network calls to other services, rather than a single call or a small handful of calls doing meaningful work each. In the checkout example, a single 'place order' request might need to call: - an **inventory service** to check stock - a **pricing service** to compute the total - a **tax service** - a **customer service** to validate the account - a **fraud-check service** - a **discount service** - a **shipping-rate service**, and so on If any of these calls themselves fan out to further calls, the total round trips balloon well past what the logical operation actually warrants. Compare this to a well-sized 'Checkout' service that owns enough of this logic internally, or that makes a small number of coarse-grained calls, to complete the same work in two or three hops. ## Why it is a granularity problem Chattiness is fundamentally a granularity problem: it shows up when services are decomposed along boundaries that don't match how work actually flows through the system. If a business capability's data and logic are spread thin across many small services (see: entity services, nanoservices), any nontrivial workflow has no choice but to touch all of them — the chattiness is the visible symptom of a granularity decision made upstream. It also shows up when a team tries to keep every service 'pure' by refusing to duplicate even small amounts of read-only data, trading a network call for strict normalization. ## What it costs - **The cost of chattiness is compounding latency** — if each of 15 calls takes even 20ms at p50, sequential chaining alone adds 300ms, and at higher percentiles (p99) the tail latencies of each hop combine, so the overall p99 can be dramatically worse than any single service's own p99 (sometimes called the 'fan-out tax'). - **Each hop is also an independent failure point** requiring its own timeout, retry, and circuit-breaker policy — get any of that wrong and one flaky downstream service can take down the whole checkout flow. - **On the other side**, merging services to reduce chattiness reduces this network cost but re-couples logic that might genuinely need independent scaling or ownership, so the fix isn't 'always merge' — it's 'merge or restructure calls specifically where the chatty pattern isn't buying real isolation.' ## How it shows up in production In production, chatty architectures show up as: - **cascading timeouts** — a single slow service in a long synchronous chain causes the whole chain to breach its SLA, even if the slow service itself isn't fully down - **thundering-herd retries** — many callers all retrying the same overloaded downstream service, making the overload worse - **hard-to-debug latency**, since a slow checkout response could be caused by any one of 15 calls and requires distributed tracing (a trace ID propagated through headers, visualized in a tool like Jaeger or Zipkin) to even localize - **brittle deploy coordination**, which it also tends to produce: a schema change in one of the 15 called services can silently break checkout in ways that are hard to catch without contract testing across all consumers ## The complementary fixes There are several complementary fixes, chosen based on which forces matter most for the specific chain. 1. **Where calls are read-only and don't need strict real-time consistency**, replace synchronous request/response with an event-driven approach: services subscribe to domain events and maintain their own local read-optimized copy of the data they need, so checkout reads its own cache instead of calling the pricing service live — this is the CQRS/materialized-view style fix. 2. **Where several services are always called together as one unit of work and rarely evolve independently**, merge them into one coarser service. 3. **Where the calls are genuinely independent but happen to be needed together**, introduce an aggregating layer — a Backend-for-Frontend or a dedicated 'checkout orchestrator' — that fans out calls in parallel rather than sequentially, cutting latency even without reducing the call count, and centralizes the timeout/retry/circuit-breaker policy in one place instead of duplicating it in every caller. The widely documented API Gateway/BFF pattern used by companies with large device-specific mobile fleets is a common real example of solving exactly this class of fan-out problem.

  • Why can switching some of those synchronous calls to asynchronous events reduce chattiness without just moving the problem?
    Because for read-only data that doesn't need to be perfectly real-time, the consuming service can maintain its own local materialized copy updated by events, turning what used to be a live network call into a local read. The trade-off is eventual consistency, so this only works where the business can tolerate a staleness window, e.g., showing a slightly-stale product name is fine, but checking real-time fraud status usually isn't.
  • Does adding a Backend-for-Frontend (BFF) to aggregate the 15 calls actually fix the underlying problem?
    It fixes the caller-facing symptom - a single client request becomes one call to the BFF, which can parallelize the fan-out to cut latency - but the 15 internal calls still happen, so it doesn't reduce total backend load or fully solve tail-latency amplification. It's most useful when the granularity itself is actually correct and the problem is really about who orchestrates the calls, not about over-decomposition.
  • How would distributed tracing help diagnose a chatty-communication problem in production?
    By propagating a shared trace ID through all 15 calls and visualizing the resulting span tree, you can see exactly which hops are sequential versus parallel, which one is the slowest, and where retries are amplifying load - without tracing you're stuck guessing from aggregate latency numbers which of many services is the actual bottleneck.

It's like needing to walk to five different offices in five different buildings, one after another, just to get one form stamped - even if each office is fast, the total trip time is dominated by walking between buildings, not by the stamping itself.

saying these in an interview costs you the question

  • Suggests only 'add caching everywhere' without noting the consistency trade-off
  • Doesn't connect chattiness back to a granularity/decomposition root cause
  • Treats parallelizing calls as equivalent to eliminating them
  • No mention of cascading timeouts or retry storms as production failure modes
  • Can't name any concrete mitigation beyond 'make the services faster'

context