You must set one external-entry standard for an estate of forty containerised services — how do you decide, and what do you exempt?
answer
- a default plus a written exemption
- decidable without a lead
- someone owns the shared edge
- exemption count is the health metric
- self-service, or the standard erodes
basics
~20 sSet a default most workloads never have to argue with — internal-only inside, one shared edge outside — then write the exemption test: traffic the edge cannot route, a blast radius that cannot be shared, volume that would distort the edge. Then fund the edge's owner.
solid answer
~50 sA standard is two things: **a default** and **a written exemption test**. The default I would set is internal-only unless a workload has external callers, and a shared edge routing by hostname and path when it does — because that is the shape whose cost does not multiply by the number of services and whose certificate story has one owner. The exemption test has to be answerable without me: traffic the edge cannot parse into routable requests; an outage that must not be shared, for a reason someone can state; volume large enough to distort capacity planning for everyone else. "We would prefer our own" is not on the list. Then the organisational half, which is what actually decides whether the standard holds: someone owns the edge, publishing a new hostname is self-service with a published lead time, and the exemption count is reviewed — because a rising count means the default no longer fits.
go deeper
Know that estates usually have a default way of exposing a service, and that asking for something different is expected to come with a reason.
Be able to argue why one shared entry point is the cheaper default for many services, and name the two constraints that stop a workload using it.
Design the exemption test so teams can apply it themselves, and say what the shared edge needs in capacity, ownership and alerting before it can be the default.
Own the whole trade, including the organisational half: who runs the edge, how fast a team can publish, and what a rising exemption count is telling you about the standard.
## What a standard is actually deciding It is not deciding which entry shape is best — all four have a case, and the estate will use more than one. It is deciding **which shape a team gets without an argument**, what evidence buys a different one, and who is accountable for the shared thing when it breaks. A standard that only names a preferred shape is a document; a standard that names a default, a test and an owner is a control. ## The default For an estate of this size the defensible default is two rules: - **Internal-only unless the workload has callers outside the platform.** Most services in a forty-service estate are called only by other services. Making external exposure the thing you have to ask for, rather than the thing you have to remember to turn off, is the single highest-value line in the standard. - **A shared edge when it does have outside callers.** Its cost does not multiply by the number of services, its certificate ownership lands in one place, and a new public hostname is a table entry rather than a procurement. What that default deliberately gives up is isolation. Every workload behind the edge shares a failure domain and a change surface, and the standard should say so plainly rather than pretending the shared edge is free. ## The exemption test An exemption must be decidable by the team asking, not negotiated with a lead each time: 1. **Can the edge parse this traffic into requests carrying a hostname?** If not, the shared edge cannot route it at all, and a balancer provisioned for that workload is the correct shape rather than a concession. 2. **Must this workload's availability be independent of the rest of the estate, for a reason that can be written down?** A regulatory boundary, a contractual commitment, a separate tenancy. "It is important to us" is not that reason; every service is important to its team. 3. **Would this workload's traffic distort the edge's capacity planning?** A workload an order of magnitude larger than the rest is a capacity problem for everyone behind the edge, and isolating it is cheaper than sizing the edge around its peaks. Everything else is a request, not an exemption. The fleet-wide port belongs in the standard too, but as a named escape hatch with an owner and a review date, because its coordination cost is paid by humans and grows with every use. ## What the standard must not try to settle - **Per-workload authorisation.** The edge is one place to get a rule wrong for forty services; workloads keep establishing who is calling them. - **Everything about encryption on the internal hop.** Whether the leg from edge to workload is encrypted again is a separate decision with its own cost, and belongs to a different subject than the choice of entry shape. - **Hostname naming, ownership and retirement.** Related, and worth a standard of its own, but folding it in makes this one unreadable. ## The organisational half This is what actually determines whether the standard survives its first year. - **Someone owns the edge** — capacity, upgrades, the renewal alert, the on-call rotation. An unowned shared thing decays into everyone's incident and nobody's roadmap. - **Publishing a hostname is self-service**, with review where it is genuinely needed and a published lead time. If adding a route takes a ticket and a week, teams route around the standard, and the exempt path becomes the fast path. - **Exemptions are recorded with their reason and re-read.** They are the standard's health metric, not its failure. ## How you know it is working - The **exemption count is flat or falling** as the estate grows. A climbing count means the default no longer fits what teams are building, and the standard needs changing rather than enforcing harder. - **Time to publish a new public hostname** is measured in minutes of work, not days of waiting. - **The number of held external addresses** tracks the exemption list and nothing else — stray addresses mean workloads left the standard without anyone noticing, or that removed exposures never released what they allocated. - **Edge incidents have a named owner and a post-incident action**, rather than being absorbed as background noise by forty teams. ## The honest caveat If the exemption test, applied truthfully, exempts half the estate, consolidation was the wrong bet here and the standard should say so. The problem would then be the protocol mix, not the entry shape, and writing a standard that most workloads must escape is worse than having none: it teaches teams that the rules are for people who did not push back.
- A team asks for its own external address "for isolation". How do you test that claim?Ask which outage they are isolating from, and whether the edge's actual failure modes would reach them. Ask whether their traffic is even routable by the edge. Grant it when there is a stateable boundary — regulatory, contractual, tenancy — or a real capacity distortion. Refuse it when "isolation" turns out to mean preferring to run their own, because granting that dissolves the standard for everyone.
- What makes a shared edge erode back into per-service entry points?Change latency, almost always. If publishing a hostname needs a ticket and a wait, the exempt path becomes the fast path and teams take it for reasons that have nothing to do with the exemption test. Self-service publishing, a published lead time for the cases that need review, and an owner with capacity to answer are what keep the default cheaper than the escape.
saying these in an interview costs you the question
- Names a preferred shape without stating the default.
- Writes a standard with no exemption path at all.
- Ignores who operates the shared edge day to day.
- Assumes teams migrate without a self-service publish path.
- Treats cost as the only input to the decision.