skip to content

Beyond class hierarchies, how does the Liskov Substitution Principle apply to service APIs, message schemas, and plugin implementations — and how would you enforce substitutability across teams?

level: principalimportance: nice to knowfreq 30%

answer

  1. Contract, not `extends` — same three rules
  2. Additive-only: optional in, fields never removed
  3. New error code / lost ordering / lost idempotency = weakened
  4. Conformance suite + consumer-driven contracts in CI
  5. Hyrum's Law; shadow-run before cutover

basics

~20 s

Any time a consumer written against one contract receives a different implementation or version, LSP applies: a new version must not demand more of callers or promise less. Enforce it with shared contract tests run in every provider's pipeline.

solid answer

~60 s

LSP is a statement about *contracts*, not about the `extends` keyword — so it governs any substitution a consumer can't see: a v2 of an HTTP endpoint replacing v1 behind the same route, a second implementation of a plugin SPI, a new producer of an event schema, a different vendor behind an abstraction layer. The rules transfer directly: don't add required request fields or narrow accepted values (precondition strengthening); don't remove response fields, drop guarantees like ordering/idempotency/at-least-once delivery, or introduce new error codes (postcondition weakening); don't break invariants consumers rely on such as ID stability or monotonic sequence numbers. Schema evolution rules are literally this — adding optional fields is safe, adding required ones is not. Enforcement is organisational rather than compiler-based: publish an explicit contract, ship a **provider-agnostic conformance/contract test suite** every implementation must pass in CI, use consumer-driven contract testing, run schema-compatibility checks in the registry as a merge gate, and shadow/dual-run new implementations against real traffic before cutover. Guard against Hyrum's Law by measuring what consumers actually depend on, not just what you documented.

go deeper

for a junior

Recognise that API versioning rules — add optional fields, never remove or require new ones — are the same substitutability idea applied to services.

for a middle

Map the three contract rules onto concrete API/schema changes and name the safe vs unsafe evolution list, with schema-registry compatibility checks as the tool.

for a senior

Include non-shape guarantees (ordering, idempotency, delivery semantics, error taxonomy, latency), plugin SPIs, test fakes, and enforcement via shared conformance suites and consumer-driven contract tests in CI.

for a principal

Treat contract breadth as a deliberate, up-front governance decision with costs on both sides; add Hyrum's Law, shadow/diff-testing before cutover, deprecation scheduling with usage measurement, vendor abstractions designed to the intersection of capabilities with explicit capability negotiation, and organisation-wide gates that make incompatible changes unmergeable.

## Restating LSP without inheritance The Liskov Substitution Principle is usually taught as "a subclass must be usable wherever its base class is expected". The general form drops classes entirely: > If clients are written against contract **C**, then any implementation or version substituted behind **C** must not break those clients. Definitions: - **Contract**: the promises a consumer may rely on — request/response shape, accepted value ranges, error taxonomy, ordering, idempotency, delivery semantics, latency envelope, stability of identifiers. - **Substitution**: swapping what sits behind that contract — a new deployment, a different vendor, a second implementation of a plugin interface, a v2 of a schema. - **Consumer**: any caller that cannot be changed in lockstep — another team's service, a mobile app already on users' phones, a downstream stream processor, a third-party integrator. The kicker at system scale: **you usually cannot recompile the consumers.** In a monolith, an LSP violation is a failing test. Across a service boundary, it is an incident. ## The three rules, translated **Preconditions may not be strengthened — don't demand more of callers.** - Adding a *required* request field or header. - Narrowing accepted values (a status enum that used to accept five values now rejects two; a string field gains a stricter length or format rule). - Requiring authentication/scopes that weren't required before. - Requiring a new call ordering ("you must call `prepare` first"). - Tightening rate limits below what consumers were told. **Postconditions may not be weakened — don't deliver less.** - Removing or renaming a response field, or making a previously-always-present field optional. - Returning a new error code or HTTP status the consumer's error handling never anticipated. - Dropping a guarantee: strict ordering → best-effort ordering; exactly-once → at-least-once; synchronous write-through → accepted-and-queued (`200` becoming an effective `202`). - Losing idempotency, so a retry now double-charges. - Materially widening the latency envelope so callers' timeouts start firing — technically outside a naive contract, practically a breaking change. **Invariants must be preserved — don't break the always-true properties.** - Identifier stability (an ID that used to be permanent now gets reissued). - Monotonic sequence numbers or version counters. - Referential guarantees ("every order references an existing customer"). - Units, timezone, or currency semantics of a field silently changing — the most dangerous kind, since types stay identical. And the **history rule**: don't allow state transitions consumers believe impossible — e.g. a resource that was documented as append-only starting to permit deletion, invalidating every cache built on that assumption. ## Where this shows up concretely - **HTTP/gRPC API versioning.** "Backward compatible change" is LSP by another name. Additive-only evolution (new optional request fields, new response fields consumers must tolerate) keeps substitutability; anything else needs a new version and a migration window. - **Message/event schemas.** Avro/Protobuf/JSON-Schema compatibility modes formalise the rules: BACKWARD compatibility (new consumers read old data) and FORWARD compatibility (old consumers read new data) are precisely postcondition/precondition constraints. Protobuf's "never reuse a field number, never change a field's type, new fields must be optional" is the LSP checklist as a wire-format rule. - **Plugin / SPI ecosystems.** Every third-party implementation of an interface you publish is a subtype you did not write. If your SPI has optional operations or vague semantics, implementations will diverge and consumers will start type-sniffing. This is where a published **conformance suite** (like a TCK — technology compatibility kit) earns its keep. - **Vendor abstraction layers.** Wrapping two payment providers behind one interface only works if the *weakest* provider can honour the interface's promises. Designing the interface around the strongest one guarantees the other's adapter will throw, no-op, or lie. The disciplined move is to design the abstraction to the intersection of capabilities, and expose the extras through explicit capability negotiation rather than pretending they're universal. - **Test doubles and fakes.** An in-memory fake substituted for a real repository is an LSP question: if the fake is case-sensitive where the database is not, or doesn't enforce a unique constraint, green tests hide production failures. Fakes should pass the same contract suite as the real implementation. - **Database/storage swaps.** Replacing a strongly consistent store with an eventually consistent one behind the same repository interface weakens a postcondition callers depend on ("read your own write"). ## Enforcement across teams No compiler spans organisational boundaries, so enforcement is process plus tooling: 1. **Write the contract down explicitly** — machine-readable where possible (OpenAPI, Protobuf, AsyncAPI, JSON Schema) plus prose for the parts schemas can't express: ordering, idempotency, retry semantics, error taxonomy, SLOs. 2. **Shared conformance / contract test suite.** One suite expressed purely in the contract's terms, executed against *every* implementation (real, fake, new version, third-party plugin) in that implementation's CI. This is the single highest-leverage practice; it is the abstract-test-suite technique scaled up. 3. **Consumer-driven contract testing.** Consumers publish the expectations they actually rely on; providers verify against them before deploying. Converts "we think this is compatible" into a pipeline gate. 4. **Schema registry compatibility gate.** Enforce BACKWARD/FULL compatibility mode as a required check on merge, so an incompatible schema literally cannot be registered. 5. **Shadow traffic / dual-run / diff-testing.** Run the new implementation alongside the old on real requests, compare responses, and only cut over when divergence is understood. Catches the undocumented behaviours no test suite encoded. 6. **Progressive rollout with fast rollback**, plus contract-level monitoring (error-code distribution, field-null rates, latency percentiles) rather than only infrastructure metrics. 7. **Deprecation policy, not silent change**: version, announce, measure remaining callers, then remove. Substitutability at scale is a scheduling problem as much as a design one. ## Hard nuances - **Hyrum's Law**: with enough consumers, *every observable behaviour* becomes someone's dependency, regardless of what you documented — response field ordering, incidental sort stability, exact error strings. Strict documentation is necessary but not sufficient; measure real usage before changing anything. Some teams deliberately inject controlled jitter or randomised ordering so consumers cannot depend on unpromised behaviour. - **Contract breadth is the lever you actually control.** A deliberately weak contract ("errors of any type may occur; ordering is not guaranteed") preserves your freedom to substitute, at the cost of pushing complexity into every consumer. A strong contract makes consumers simple and locks your implementation options. Choosing this deliberately, up front, is the principal-level decision — retrofitting a weaker contract onto existing consumers is itself the breaking change. - **Bug-compatible substitution**: if the old implementation had a bug consumers worked around, fixing it can break them. Sometimes the right move is to fix behind a version or a flag. - **Non-functional substitution**: security posture, data residency, and cost profile can all differ between implementations behind the same interface without any functional difference — worth treating as part of the contract in regulated contexts.

  • Which schema changes are safe for an event or message consumed by services you don't control?
    Additive and optional: new fields with defaults, new enum values only if consumers are documented to tolerate unknown values, widened accepted ranges. Unsafe: removing or renaming fields, making an optional field required, changing a field's type or units, reusing a field identifier, or narrowing accepted values. Enforce it with a registry compatibility mode as a merge gate rather than by review.
  • How do you handle a case where the new implementation genuinely cannot honour an old guarantee?
    Don't hide it behind the same contract. Introduce an explicitly versioned contract with the weaker guarantee, migrate consumers on a published deprecation schedule, and keep the old contract serving until measured usage reaches zero. Silently weakening a live contract converts a design problem into a production incident with no rollback story for consumers.
  • What is Hyrum's Law and why does it complicate LSP at system scale?
    It observes that with enough consumers, every observable behaviour of your system becomes someone's dependency, whether documented or not — field ordering, error message text, incidental latency. So a technically contract-compliant substitution can still break consumers, which is why shadow traffic, diff-testing, and usage measurement complement the written contract.

Swapping the airline operating a code-shared flight number: passengers booked on the number expect the same route, baggage allowance, and connection guarantees. A carrier that flies it with a tighter baggage rule (demands more) or drops the connecting guarantee (delivers less) breaks every itinerary built on the original promise, even though the flight number on the ticket is identical.

saying these in an interview costs you the question

  • "LSP only applies to inheritance" — it applies to any substitution behind a contract consumers rely on
  • "Adding a required request field is backward compatible" — it strengthens the precondition and breaks existing callers
  • "We only removed a field nobody should have used" — Hyrum's Law says someone did; measure before removing
  • "Same JSON shape, so it's compatible" — ignores ordering, idempotency, delivery, error-taxonomy and unit semantics
  • "Code review will catch schema breakage" — compatibility must be a mechanical merge gate, not a human check
  • "Our in-memory fake behaves close enough to the database" — a fake that skips a constraint is a substitutability violation that hides production bugs

context