Across a service, where should data-access failures be translated into domain outcomes, and what does leaking them cost?
answer
- sort the failures before placing the code
- kind must survive the journey
- map only what the domain models
- one handler for the rest
- coupling to kinds, not to codes
basics
~20 sMap only the few failures the domain genuinely models, close to the code that knows the rule, and let the rest travel outward with their kind intact to a single handler. Blanket wrapping destroys the retry decision; blanket leaking spreads persistence vocabulary through every layer.
solid answer
~50 sThere are three candidate homes: the component that owns statements, the use case, and one outer handler. Each failure kind wants a different one. A constraint violation that encodes a business rule should become a domain outcome near the code that knows the rule, because that is the only place with the context to name it. Transient conflicts should keep their kind and travel outward, because the retry decision is made outside the boundary and a wrapper that erases the kind destroys it. Everything else - conversion errors, unexpected constraints - should reach one outer handler that logs the cause and returns a generic failure. The two failure modes are symmetrical: wrapping everything into one opaque error makes retry impossible and diagnosis expensive, while leaking untranslated failures puts persistence vocabulary into signatures across the service.
go deeper
Know that failures from the data layer are not shown to users as they arrive, and that somewhere in the service they are turned into either a message about a rule or a generic error.
Explain why the failure kind has to survive as far as whoever decides on a retry, and why keeping the original as the cause matters for logs.
Show a concrete division: portable kinds at the data component, an enumerated map for domain outcomes, one outer handler for the rest, with unrecognised cases logged loudly.
Own the tradeoff. Argue why a small failure-kind vocabulary crossing layer boundaries beats perfect isolation, and define what the team measures to know the placement still holds.
## Three candidate homes | Where | Good for | Bad at | |---|---|---| | The component owning statements | making failures portable at all; reading engine detail once | it lacks the context to say which business rule was broken | | The use case | mapping a named constraint to a domain outcome; deciding retry | repeating the same mapping if several use cases share a rule | | One outer handler | logging, correlation, the generic response | it is far from the cause, and can only answer generically | The mistake is picking one and applying it to everything. The three kinds of failure a service actually sees want different treatment. ## Sort the failures first, then place the translation 1. **Failures the domain models.** A named constraint that encodes a rule someone can be told about. These should become domain outcomes, and the mapping belongs near the code that knows what that constraint means. Push it too far out and the handler is guessing; push it too far in and the statement-owning component has to know about product rules. 2. **Failures the caller can recover from mechanically.** Transient conflicts. These must keep their kind while they travel, because the decision to re-run belongs outside the boundary and can only be made by something that can still see what kind of failure occurred. 3. **Failures that are bugs.** Conversion errors, foreign-key violations nobody expected, an unrecognised constraint name. These want one place that logs the cause with full detail and returns something generic, and they want an alert rather than a friendly message. ## The cost of wrapping everything A blanket policy of converting every failure into a single application error is popular because it is simple to state in a review. What it costs: - **The retry decision disappears.** Once the kind is erased, nothing upstream can distinguish a conflict worth re-running from a violation that will fail forever, so the service either retries everything or nothing. - **Diagnosis gets expensive.** If the original is not kept as the cause, the constraint name and the engine's own detail are gone from the logs at exactly the moment they are needed. - **Alerting flattens.** Real defects and ordinary business rejections arrive as the same type, so a threshold that catches the first also fires on the second. ## The cost of leaking The opposite policy - let the layer's exceptions travel everywhere untranslated - has its own bill: - Persistence vocabulary appears in the signatures and tests of layers that should not know a database exists, and it outlives the technology choice that introduced it. - Every caller invents its own interpretation of the same failure, so the same constraint is rendered three different ways in three endpoints. - Failures arrive late - at flush or at commit - so the code that catches them is often nowhere near the code that caused them, and a leaked failure gives the catcher no help in bridging that distance. ## A workable division 1. Translate raw engine failures into portable kinds **once**, where statements are owned. 2. Map the **small, enumerated** set of failures the domain models into domain outcomes, next to the rule they belong to, using constraint names rather than message text. 3. Let everything else propagate **with its kind intact**, so the retry decision remains possible outside the boundary. 4. Terminate at **one outer handler** that logs with the cause chain and a correlation identifier and returns a generic response. 5. Make the enumerated set **reviewable**: an unrecognised constraint is treated as a bug and logged loudly rather than mapped to something plausible. ## What to measure afterwards The design is only as good as what it shows you. Count failures by kind, not just by endpoint; a service that cannot answer *how many conflicts did we retry yesterday* has already lost the kind somewhere. Watch for the same constraint being mapped in more than one place, and for the outer handler's generic bucket growing - the second usually means a rule was added to the schema and never added to the map. ## The tradeoff to state out loud This division is not free. It puts a small amount of persistence knowledge - the notion of a retryable kind - into layers above the data-access component, and it accepts that a few failure types cross more than one layer boundary before anyone acts on them. The alternative, perfect isolation, buys purity by destroying the information that operational behaviour depends on. Naming that tradeoff explicitly is the point: the goal is not zero coupling, it is coupling to a small, stable vocabulary of failure kinds rather than to an engine's error codes.
- How do you keep the retry decision possible without exposing persistence types everywhere?Expose the kind, not the type. A small application-owned vocabulary - transient conflict, permanent rejection, ambiguous outcome - can travel outward in whatever error shape the service already uses, while the engine-specific type stops at the data-access component. Upstream code branches on that vocabulary and never names anything a database defined.
- Two use cases enforce the same uniqueness rule. Where does the mapping live?In one place both call, owned by whatever component owns that rule, rather than copied into each use case. Duplicated mappings drift, and drift shows up as the same violation rendered differently in two endpoints. If no such component exists, that absence is usually the real finding.
- What signals that the translation boundary is in the wrong place?Generic failures rising in the outer handler while user-visible rejections stay flat, the same constraint mapped in several files, or a retry policy that cannot tell a conflict from a violation. Each points at information destroyed too early or a decision made too far from its context.
saying these in an interview costs you the question
- Wrapping every data-access failure into one opaque application error
- Letting untranslated persistence exceptions travel through every layer
- Dropping the original failure so the cause never reaches the logs
- Mapping an unrecognised constraint to a plausible business message
- Deciding placement per layer instead of per failure kind