skip to content

The resolver, directory and time services every segment must reach now sit behind an inspecting chokepoint, added after one segment was compromised — what does that cost you?

level: seniorimportance: should knowfreq 41%

answer

  1. on the critical path of everything
  2. sized for the aggregate, not one segment
  3. decide what it does when it dies
  4. do not let it depend on what it fronts
  5. every host legitimately produces this traffic

basics

~20 s

It puts one device on every flow in the estate: latency on services that precede every other call, a failure domain wider than any segment, capacity sized for the aggregate, and a fail-open or fail-closed choice where both answers hurt.

solid answer

~50 s

The universal tier is the worst place to add a hop, because its services are prerequisites: resolution precedes every connection, time precedes authentication, the directory precedes authorisation. Latency there multiplies downstream, and the hop must be sized for the whole estate's aggregate rather than one segment's peak — when it fills, the symptom is everything failing at once. The fail posture must then be decided per service and written down: fail-closed on the directory is defensible, fail-closed on resolution is an estate-wide outage an adversary can trigger cheaply, and fail-open means the control is absent exactly when stressed. Watch for the chokepoint authenticating its own administrators against the directory it fronts. And be honest about yield: every segment legitimately produces this traffic, so the real value is constraining which segment reaches which service and watching the tier's outbound.

go deeper

for a junior

Understand that a device placed in front of services everyone depends on affects everyone, and that what it does when it fails is a decision somebody has to make deliberately.

for a middle

Explain why latency in front of resolution, time and directory compounds through downstream calls, and what a stateful device does when its capacity is exceeded.

for a senior

Show the per-service fail posture with reasoning, the aggregate capacity argument, the circular-dependency check, and an honest account of how little content inspection yields on this stream.

for a principal

Be ready to say when the hop is not worth building at all, and what enforcement remains when the shared services are operated by a provider you cannot place anything in front of.

## Why this is the most expensive place in the estate to add a hop A chokepoint in front of a business application costs that application some latency. A chokepoint in front of the universal set costs *everything*, because the services behind it are preconditions rather than destinations: - A host cannot open any connection before it resolves a name. - It cannot authenticate before its clock is close enough to agree. - It cannot authorise before the directory answers. So added milliseconds are not paid once per flow; they are paid before every flow, often several times, and again on each dependent call. This is the compounding cost people underestimate when they draw the diagram. ## Capacity is an aggregate problem A segment boundary is sized for a segment. This hop is sized for the sum of every segment, at their coincident peaks, including the correlated bursts that happen when something else fails and thousands of clients retry at once. When it fills, it queues, and then it drops — and because everything depends on it, the observable symptom is not "the inspection device is degraded", it is "the entire estate is broken and nobody can say why". Design the capacity line from the aggregate and defend it with that number, because it is the number that gets cut in a budget review. ## The fail posture, per service, in writing The worst outcome is a single fail posture applied to all four services because a device only offers one setting. Reason about them separately: | Service | Fail-closed means | Reasonable default | |---|---|---| | Resolution | No host in the estate can start anything | Fail open; the outage is worse than the exposure | | Time | Clocks drift, then authentication fails later | Fail open, with drift alarming | | Directory | New authentications are refused | Fail closed for new sessions is defensible | | Log transport | Records are lost silently | Fail open with client-side buffering | The security point behind the table: an adversary who knows you fail closed on resolution has been handed a cheap estate-wide denial capability — they attack the chokepoint, not the estate. An adversary who knows you fail open knows the control is absent under exactly the conditions they can create. Neither is wrong in the abstract; what is wrong is not having chosen, and not having told the people who own availability which one you chose. ## The circular dependency that ruins recovery Check what the chokepoint itself depends on. If its administrators authenticate against the directory it fronts, or it resolves names through the resolver behind it, then the failure you are designing for is unrecoverable: you cannot log in to fix the thing that is preventing logins. The universal tier and anything guarding it need an out-of-band administration path and credentials that do not traverse the dependency. ## Inspection, specifically If "inspecting" means terminating TLS, you have taken on more than latency: the hop now holds a trust anchor every client accepts, which makes it a high-value target sitting on every flow. Clients that pin certificates or use mutual authentication will break, and the breakage will be attributed to anything but the new hop for the first day. Selecting on the visible handshake instead is cheaper but weaker, and encrypted client hello removes even the name you were selecting on. Decide which you are doing before you promise anyone visibility. ## What the hop actually buys Be honest about yield, because this is what separates a senior answer from a diagram. Every host in the estate is *supposed* to be talking to these services. An intruder in a compromised segment using resolution, the directory and time is producing traffic that is, by design, identical to everyone else's. Content inspection of that stream is a low-yield exercise dressed up as a control. What the position genuinely gives you: 1. **Which segment reaches which service.** Not every segment needs a directory bind. Splitting the universal permit by service is enforcement the hop can actually do. 2. **A place to see the tier's own outbound**, which is low volume and high signal, because the shared services legitimately originate almost nothing. 3. **A single place to change the policy** for a permit that otherwise lives in every segment's rule set — which is a real operational win and also a single place to make an estate-wide mistake. ## And the case where you cannot do it at all In a remote-first estate the shared services are subscriptions. Identity, log ingest and resolution terminate at a provider, at anycast addresses shared with other tenants, over a name set the provider changes without telling you. There is no place to insert a hop, an address-based allow-list cannot distinguish your tenant's endpoint from someone else's on the same infrastructure, and the total-dependency price is paid to a third party rather than to your own device. The chokepoint then has to move to the endpoint or to an identity-aware position, and the honest conclusion is that your enforcement point became authentication rather than reachability. Say that plainly instead of pretending the network diagram still describes the control.

  • Which of these services would you let fail open, and why?
    Resolution almost certainly, because failing closed there stops every connection in the estate and hands an adversary a cheap denial lever. Time and log transport can fail open too, with drift alarming and client-side buffering absorbing the gap. The directory is the one where failing closed for new authentications is defensible, since the whole point of the answer is a verdict you should not fabricate under load.
  • Your chokepoint's administrators authenticate against the directory it fronts. Why is that fatal?
    It makes the outage unrecoverable. The condition you built the device to survive is exactly the condition that locks you out of it, so the first real incident becomes a physical or provider-level recovery. The universal tier and everything guarding it need an administration path and credentials that do not traverse the dependency they protect.
  • How do you size it, and what number do you defend?
    From the aggregate of every segment at coincident peak, plus the retry storm that follows any partial failure, not from an average or from the busiest single segment. Defend that number explicitly in the budget conversation, because a device that is one hop from everything degrades into an estate-wide outage rather than a local slowdown when it is undersized.

saying these in an interview costs you the question

  • Assumes inspection can distinguish an intruder's use of shared services from normal use
  • Sizes the chokepoint from one segment's traffic or from an average
  • Applies a single fail posture to all four shared services
  • Terminates TLS without checking which clients pin or mutually authenticate
  • Lets the chokepoint authenticate its admins against the directory behind it

context