A platform team is under pressure to 'go event-driven' across the board after a successful migration of one workflow. As the architect, when would you push back and keep a workflow request-driven instead?
answer
- synchronous needed when caller acts on the result immediately
- saga/compensation for cross-step invariants, not bare events
- operational maturity: correlation IDs, tracing, DLQ monitoring, idempotency
- one pilot success doesn't prove readiness
basics
~20 sNot every part of the system benefits from going async. If people need an instant yes/no answer, if two things must happen together perfectly, or if the team can't yet trace bugs across services, staying with direct calls is often safer.
solid answer
~40 sI'd push back whenever the workflow genuinely needs a synchronous answer in the same interaction (e.g., 'was my payment authorized,' a login check) — deferring that to an event just adds a polling or callback layer that reproduces synchronous semantics with more moving parts. I'd also push back when the operation needs strong, immediate consistency across steps (e.g., transferring money between two accounts) where eventual consistency risks a visibly broken invariant, unless the team is ready to build a proper saga with compensations. And I'd push back organizationally if the team doesn't yet have the operational maturity — correlation IDs, distributed tracing, dead-letter-queue monitoring, idempotent consumers — because adopting events without that maturity trades a well-understood failure mode (cascading synchronous failure) for a poorly-observed one (silent staleness/duplicate processing) that's harder to diagnose.
go deeper
Should be able to give at least one plain-language example of an operation that needs an immediate answer and shouldn't go async.
Should articulate the synchronous-need and consistency-invariant criteria and apply them to a concrete example.
Should distinguish bare event chains from sagas with compensation, and propose concrete criteria (not just intuition) for when to keep something synchronous.
Should treat this as an organizational and operational-maturity judgment as much as a technical one — evaluating whether the team has the tracing/idempotency/monitoring foundation before broad rollout, and pushing back on 'one success proves readiness everywhere' reasoning.
"We successfully went event-driven for one workflow, so let's do it everywhere" is a common but risky generalization, because the case for events is scenario-specific, not universal — it depends on what a given workflow actually needs (an immediate answer vs. eventual propagation) and what the team is actually equipped to operate. As the architect, there are several concrete situations where I'd push back on blanket adoption. ## When the workflow needs a synchronous answer When the interaction genuinely needs a synchronous answer: some operations are, by their nature, request/response — the caller cannot proceed, and often cannot even render a UI, without knowing the outcome right now. - Logging in. - Authorizing a payment. - Checking whether a username is available at signup. These all need an immediate yes/no in the same interaction the user is waiting on. You can technically build an async version of any of these (publish `LoginRequested`, poll or subscribe for `LoginResult`), but doing so just reimplements synchronous request/response with extra latency and extra moving parts (a poll loop, a correlation table to match requests to eventual results) — you've paid the complexity cost of asynchrony without gaining anything, since the caller still can't proceed until it gets the answer. **The tell for this case is simple:** if the very next thing the calling code or the user does depends on the specific result of this operation, keep it synchronous. ## When an invariant must hold at every moment When strong, immediate consistency is a correctness requirement, not just a nicety: some operations have an invariant that must never be visibly violated, even for a moment — a bank transfer where the total of both accounts must stay constant, or a seat-booking system where the same seat must never be sold twice. A bare, uncoordinated event chain (`DebitRequested`, then separately `CreditRequested`) can leave the system in a state where the debit succeeded but the credit hasn't happened yet, or worse, failed silently — an eventually-consistent "we'll fix it later" answer isn't acceptable when a customer can observe money missing from both places at once. This doesn't necessarily mean "go fully synchronous" — a **saga pattern** with explicit compensating transactions can coordinate this correctly while still being asynchronous — but it does mean you can't just naively fire two independent events and call it done; the team needs the orchestration/compensation machinery in place first, and if they don't have the time or expertise to build that correctly, plain synchronous transactions are the safer default until they do. ## When the team is unequipped to operate it When the team lacks the operational maturity to run event-driven systems safely: this is the argument I'd weigh most heavily as an architect, because it's organizational, not just technical. Running an event-driven workflow safely in production requires a set of practices most synchronous-only teams haven't built yet: - Correlation IDs propagated through every event and log line. - Distributed tracing that understands async spans. - Idempotent consumers to survive at-least-once redelivery. - Dead-letter-queue monitoring. - Consumer-lag alerting. A single successful pilot migration doesn't prove the team has institutionalized these practices — it might just mean the pilot workflow was simple enough, or lucky enough, not to expose the gaps yet. Rolling out events broadly before that operational foundation exists trades a well-understood failure mode (a synchronous chain fails loudly and immediately, and everyone already knows how to read a stack trace) for a poorly-observed one (a message silently stuck in a dead-letter queue for two weeks before a customer notices), and the second kind of incident is typically far more expensive to diagnose the first few times it happens. ## The three questions I would ask instead Rather than a blanket "no," I'd ask the team to evaluate each candidate workflow against three questions: 1. Does the calling code need the result before it can proceed? If yes, stay synchronous. 2. Does the operation require an invariant that must never be visibly violated even momentarily? If yes, either stay synchronous/transactional or build a proper saga, not a bare event. 3. Does the team have the tracing/idempotency/DLQ-monitoring foundation in place to safely operate an async version? If no, invest in that foundation before or alongside the migration, don't skip it because the first pilot went fine. ## The case I would push back on A team that successfully moved "send order confirmation email" to events might be tempted to also move "authorize customer payment" to events for consistency's sake. I'd push back there specifically: the checkout page cannot render a success or failure screen to the customer without knowing the payment outcome in that same interaction, so payment authorization stays a synchronous call to the payment provider (with its own timeout/circuit-breaker protection), while post-authorization side effects — the email, inventory update, analytics — remain the event-driven part, exactly the same split most mature e-commerce systems converge on.
- Isn't a saga just 'events done right' for the consistency problem — why call it out separately from 'bare events'?A saga adds explicit orchestration and compensating actions (e.g., 'if credit fails, issue a reversing debit') on top of the same event/message infrastructure, so it's a deliberate design for maintaining an invariant across steps, whereas a 'bare' event chain just fires independent events with no coordinated rollback if a later step fails — the distinction is the presence of compensation logic, not the transport.
- How would you actually measure whether a team has the operational maturity you mentioned, rather than just guessing?Concrete signals: do their services already propagate a correlation ID end to end and can an engineer query one ID across all logs; do they have DLQ depth and consumer-lag as monitored, alerted metrics; have their consumers been audited for idempotency; and has the team run at least one incident retro where they successfully root-caused an async issue using these tools rather than manual timestamp cross-referencing.
- What's a middle-ground option between 'fully synchronous' and 'bare fire-and-forget events' for a workflow that needs an eventual result but not necessarily an instant one?A request/reply-over-events pattern (or a synchronous kickoff plus a webhook/callback), where the caller gets an immediate acknowledgment and a correlation token, then either polls or receives a callback with the final result — useful for operations like a long-running video encode where the caller doesn't need the answer in-line but does need a specific correlated outcome eventually, unlike pure fire-and-forget which offers no result at all back to the originator.
It's like a restaurant deciding to switch every order to a written note left on a board just because it worked well for restocking napkins — but a customer asking 'do you have this dish tonight' still needs someone to answer them at the table, not check a board an hour later.
saying these in an interview costs you the question
- Treats 'event-driven' as strictly superior and can't name a case where synchronous is correct
- Proposes converting a login/payment-authorization check to pure fire-and-forget events
- Assumes eventual consistency is always acceptable regardless of the invariant involved
- Judges organizational readiness for async by whether one pilot succeeded, not by concrete tooling/practice checks
- Confuses a saga (with compensation) with a bare, uncoordinated event chain