In a production system, an Anti-Corruption Layer sits in front of a flaky third-party API and has been running fine for months. What are the most likely ways this kind of layer fails or degrades in production, and how would you design against each one?
answer
- translation drift = silent bad mapping
- ACL as single point of failure/bottleneck
- N+1 chatty translation under load
- scope creep without ownership
- fail loudly on unmapped values
basics
~20 sThe translator can quietly mis-map data when the outside system changes without telling you, the extra layer can slow things down or become a single point of failure, and it can turn into a dumping ground for random logic if nobody owns it. Guard against these with tests, monitoring, and clear ownership.
solid answer
~40 sCommon failure modes: (1) translation drift, where the external system changes shape and the mapper silently produces wrong-but-plausible output instead of erroring - guard with contract tests and alerting on unmapped values; (2) the ACL becoming a single point of failure or latency bottleneck for every consumer - guard with its own timeouts, circuit breakers, caching, and independent scaling; (3) N+1 or chatty translation, where a batch operation gets translated into many individual foreign calls - guard by designing batch-aware adapters; (4) scope creep, where unrelated logic gets bolted onto the ACL because it's a convenient shared component - guard with clear ownership and a narrow, documented contract for what belongs inside it.
go deeper
Should recognize that the translator can get things wrong if the outside system changes, even without naming formal terms like 'translation drift'.
Should name at least two distinct failure modes (e.g., silent mis-mapping and added latency) with a plausible cause for each.
Should propose concrete defenses for each failure mode - contract tests, circuit breakers/timeouts, batch-aware adapters - and connect them to how they'd actually be monitored in production.
Should discuss the organizational failure mode (scope creep, ownership) alongside the technical ones, and describe how they'd design incident response and ownership so drift is caught within a bounded time, not discovered by a customer.
## Translation drift An Anti-Corruption Layer that has been quietly working for months is not evidence it's safe - it usually means the failure modes haven't been triggered yet, because most of them are triggered by change on the far side of the boundary rather than by steady-state traffic. The most consequential failure mode is **translation drift**: the external or legacy system changes its response shape (a renamed field, a new status code, a changed date format) without the ACL team being notified, and if the translator wasn't written defensively, it doesn't throw an error - it maps the unfamiliar value to whatever the code happens to fall through to (often a default or null), producing domain objects that are subtly wrong rather than obviously broken. This is dangerous precisely because it doesn't look like an outage: requests succeed, response codes are 200, and the bad data only surfaces downstream as a business-logic bug, a wrong report, or a customer complaint, often days or weeks after the actual upstream change. - **The defense** is to make translators fail loudly on anything they don't explicitly recognize - raise a typed error or emit a distinct 'unmapped value' metric rather than silently defaulting - paired with contract tests that run against captured real payloads from the external system and are re-validated whenever that system's team announces a change, if such notice even exists. ## An availability and latency bottleneck A second failure mode is the ACL becoming an availability and latency bottleneck for everything behind it. Because every consumer is funneled through the same facade, a slowdown or outage in the external system, or in the ACL's own process, now affects every consumer simultaneously, whereas before the ACL existed, at least some consumers might have had independent (if uglier) integrations. This is compounded if the ACL doesn't apply its own timeouts and circuit breakers: a slow legacy system can tie up the ACL's connection pool or thread pool, and that exhaustion then cascades to every consumer waiting on the ACL, turning one flaky dependency into a much wider outage. - **The defense** is to treat the ACL like any other dependency a resilient system talks to - explicit timeouts shorter than the caller's own budget, a circuit breaker that fails fast once the external system is clearly unhealthy, and where the data allows it, a cache or last-known-good fallback so transient external outages don't propagate as user-facing failures. ## Chatty or N+1 translation A third, easily overlooked failure mode is chatty or N+1 translation: a well-intentioned facade method like `getOrders(orderIds: List)` gets implemented by looping over the IDs and calling the legacy adapter once per ID, because the legacy system's own API is single-item, and nobody optimizes the translator to batch. Under low load this is invisible; under real production volume it turns a single logical operation into hundreds of external calls, multiplying latency and load on the fragile downstream system exactly when it's least able to absorb it. - **The defense** is to design the adapter layer to batch aggressively where the legacy protocol allows it, and where it doesn't, to be explicit (in code review and in monitoring) about the N+1 shape so it's a deliberate, bounded trade-off rather than an accidental one that only gets discovered under a load spike. ## Scope creep A fourth failure mode is organizational rather than technical: **scope creep**. Because the ACL sits at a convenient shared chokepoint, teams under deadline pressure tend to bolt unrelated logic onto it - caching for an unrelated feature, authorization checks that belong elsewhere, business rules that have nothing to do with translating the legacy model - especially when the ACL has no single clear owner. Over months or years this turns the ACL into exactly the kind of tangled, hard-to-change component it was built to prevent, and by the time anyone notices, untangling it is its own migration project. - **The defense** is mostly organizational discipline: assign explicit ownership, document what belongs inside the boundary (translation and only translation) versus what doesn't, and treat proposals to add unrelated logic to the ACL with the same scrutiny a team would apply to adding unrelated logic to any other well-scoped service. ## One story that ties several together A concrete example that ties several of these together: a payments team's ACL in front of a third-party processor worked fine for a year, until the processor rolled out a new fraud-status enum value during a minor API update. The translator, written to default unrecognized statuses to 'approved' rather than erroring, silently approved a batch of transactions that should have been flagged, and the bug was only caught when reconciliation numbers didn't match days later - a textbook case of translation drift compounded by a fail-silent default, and exactly the class of bug that contract tests and loud failure on unmapped values are designed to catch before it reaches production traffic.
- How would you specifically detect translation drift before it causes a customer-facing incident?Run contract tests against real captured payloads from the external system on a schedule (not just at build time), and instrument the translator to emit a metric or log event every time it hits a value it doesn't recognize, with alerting on any nonzero rate. Some teams also snapshot-diff the external system's schema/OpenAPI spec periodically to catch upstream changes proactively rather than reactively.
- Should the ACL apply its own retry logic on top of whatever the legacy system already does?Yes, but carefully - the ACL should own timeouts and a circuit breaker so a slow or down legacy system doesn't exhaust its own resources or cascade to callers, but retries need to respect idempotency (don't blindly retry a write that might not be idempotent) and should back off rather than hammering an already-struggling legacy system.
- If the ACL becomes a bottleneck under load, is scaling it horizontally always the right fix?Not necessarily - if the bottleneck is actually the legacy system behind it (e.g., a database with limited connections), adding more ACL instances just shifts more concurrent load onto an already-strained dependency. In that case the real fix is often caching, request coalescing, or rate-limiting at the ACL rather than horizontal scaling of the ACL itself.
Like a customs checkpoint that's worked fine for years because nothing unusual has crossed it - the real test comes when a new kind of cargo shows up that the inspectors' rulebook was never updated for, and if they wave it through by default instead of flagging it, contraband slips in unnoticed.
saying these in an interview costs you the question
- Assumes months of uptime means the ACL has no failure modes
- Doesn't mention what happens when the translator meets an unrecognized value
- Proposes scaling the ACL without considering the legacy system's own capacity
- No mention of timeouts/circuit breakers for the legacy dependency
- Treats scope creep as a purely technical (not ownership/process) problem