Without reading any code, what deployment and operational signals would tell you a 'microservices' system is actually a distributed monolith?
answer
- deploy-timestamp correlation
- postmortem cascade language
- trace waterfall shows temporal coupling
- on-call/reviewer overlap as Conway's Law proxy
- runtime data beats reading code
basics
~20 sLook at the release calendar and incident history: if services almost always deploy together, or one service's outage always drags others down with it, that's the tell - you don't need to read the code to see the coupling.
solid answer
~30 sMine deployment metadata and incident data rather than code. Correlate deploy timestamps across services - a high correlation between two services' release times, or a shared release train/change-freeze calendar, signals forced coordination. Look at incident postmortems for cascading-failure language like 'Service X outage caused Service Y errors,' and check whether on-call rotations or code reviewers overlap heavily across 'separate' services, an org-level echo of code coupling. Distributed tracing data showing every request fanning out synchronously through several services with no fallback is another strong signal, as is a change-failure-rate spike whenever any one service in the cluster deploys.
go deeper
Should suggest at least one plausible ops signal, like noticing services always release together.
Should name two or more concrete signals, such as deploy correlation and cascading incidents, and roughly how to check them.
Should describe a systematic, multi-signal detection approach combining deploy correlation, postmortems, and tracing waterfalls, and be able to distinguish forced coupling from coincidental correlation.
Should connect technical signals to organizational ones, such as on-call/reviewer overlap as a Conway's Law proxy, and use the combined evidence to prioritize which coupling to remediate first based on business and incident impact.
## Why runtime data beats reading code Detecting a distributed monolith from operational data alone is a legitimate and often faster method than reading code, because code review requires access and expertise across every service and can be misleading: source can look decoupled, with separate repos and separate services, while runtime behavior proves otherwise. Deployment and observability data, centrally available from the CI/CD system, tracing backend, and incident tooling, reflects what the system actually does under real conditions rather than what its structure merely suggests. ## Deploy correlation The first concrete signal is **deploy correlation**. - Pull deploy timestamps per service from the CI/CD system over a meaningful window, such as ninety days, and compute how often one service's deploy falls within a short window, say plus or minus one hour, of another service's deploy. - A high correlation across many separate deploys, not just one shared release, indicates a forced dependency rather than coincidence. - Also look for explicit **release-train calendars** or change-freeze processes that name multiple services together — an organizational artifact that mirrors technical lockstep. ## Incident and postmortem language The second signal is incident and postmortem language. Search postmortems for recurring phrases like 'root cause in Service X caused errors in Service Y or Z.' A distributed monolith shows this pattern chronically, across many unrelated incidents over time, not just once as an unlucky coincidence. It's also worth pulling uptime and error-rate dashboards to check whether a service's historical downtime correlates with SLO breaches in nominally unrelated services during the same windows. ## Trace waterfalls The third signal comes from **distributed tracing**. Trace waterfalls reveal whether an incoming request fans out into a long synchronous chain with no fallback or circuit-breaker spans, versus a shallow call graph with async hops. A distributed monolith's traces typically show deep, wide synchronous fan-outs where a slow leaf span visibly inflates the root span's total duration — **temporal coupling** made visible directly in the trace. ## The organizational echo The fourth signal is organizational, an echo of **Conway's Law** used in reverse to diagnose rather than predict coupling. - Check code-review approver overlap and on-call rotation overlap across services that are nominally separate: if the same handful of engineers approve pull requests and get paged for both Service A and Service B, that's a proxy for coupling even before checking deploy timestamps, because tightly coupled services tend to require the same people's context to change safely. - Combine this with **database access logs** where available, checking which services' credentials query which tables — a service whose credentials touch tables 'owned' by three other services is a strong, concrete signal. ## A worked scenario A worked scenario ties these together. An SRE investigating a persistent Friday deploy-freeze spanning twelve nominally independent services: 1. pulls CI deploy history and finds nine of them deploy within the same two-hour window on every release; 2. cross-references postmortems and finds six incidents in the last quarter where one service's database migration caused timeouts in two unrelated services; 3. and finds via tracing that the checkout service's p99 latency is dominated by a synchronous call four hops deep into a legacy inventory service with no timeout configured. That combination of deploy correlation, cascading-incident history, and a trace waterfall is enough to diagnose a distributed monolith without reading a single line of application code, and it also gives concrete, prioritized remediation targets: the inventory call chain first, since it's both the largest latency and the largest incident contributor.
- Why is deploy-time correlation a better signal than just asking teams if their services are coupled?Teams often underestimate or normalize coupling they've worked around for years, treating something like 'we always deploy Tuesday together, that's just our process' as unremarkable, so self-report is unreliable. Deploy timestamps from the CI/CD system are objective, cover long time windows, and reveal patterns people have stopped consciously noticing.
- What's a limitation of using distributed tracing alone to detect this?Tracing shows the call graph and latency for requests that actually happened, so low-traffic or degraded-path interactions, such as a rarely-hit fallback or a nightly batch job, may not show up in typical trace sampling. You need to pair tracing with deploy history and incident data, and check batch or cron job dependencies separately, to get the full coupling picture.
- If deploy correlation is high but no incidents show cascading failures, is it still a distributed monolith?It's a warning sign worth investigating rather than a confirmed case - high deploy correlation could also come from an unrelated process constraint, such as a shared CI runner queue or a policy requiring all deploys in one weekly window, rather than technical coupling. Check whether the correlation is forced by dependency, meaning deploying B without A actually breaks something, versus merely habitual, before concluding it's the anti-pattern.
Like diagnosing a shared-custody dispute by watching who actually shows up together at every event, rather than reading the custody agreement on paper - the calendar tells the truth the document might not.
saying these in an interview costs you the question
- Says you must read the code to know if something is a distributed monolith
- Doesn't mention using deploy/CI data or incident postmortems as evidence
- Treats one shared release as proof, ignoring the need for a pattern across many releases
- Doesn't connect tracing waterfalls to the temporal-coupling concept
- Ignores the organizational/on-call-overlap signal entirely