skip to content

Workloads in a peered network are refused by the managed store while the endpoint's own network succeeds — what is the usual cause?

level: seniorimportance: should knowfreq 38%

answer

  1. classify the failure before fixing it
  2. refused means it arrived somewhere
  3. routes travel, answers are attached
  4. the peer got the public answer
  5. timeout would have meant routing instead

basics

~20 s

The private answer was attached only to the network holding the endpoint, so the peered network resolves the store's public address, takes the public path, and is refused by the rule that accepts only requests arriving through the endpoint.

solid answer

~50 s

Two things have to travel across a peering link and only one of them does so by default. The **route** to the endpoint's address is a normal route and propagates; the **private answer** for the service's hostname is scoped to the networks it is attached to, and on most platforms the peer is not one of them unless you say so. So the peer resolves the published hostname the ordinary way, gets the provider's public address, leaves through its own internet path and arrives at the public front door — where the store's resource-attached rule refuses anything not arriving through the named endpoint. The tell is the *shape* of the failure: an explicit refusal means the request arrived somewhere, which means it took the public path. A missing route would have produced a timeout instead.

go deeper

for a junior

Note that a call can fail in two very different ways: never arriving at all, or arriving and being turned away — and they point at different layers.

for a middle

Explain that routes propagate across a peering link while a private answer is attached to specific networks, so the peer resolves publicly unless told otherwise.

for a senior

Use the failure signature to locate the layer, fix resolution for the peer, and verify with a lookup from inside the peer rather than from the endpoint's own network.

for a principal

Set the standard that the resolution association ships in the same change as the endpoint, so enforcement later does not take out every network nobody checked.

## Reading the failure before changing anything The first useful move is not a fix, it is a classification. Private service access has three independent layers, and each one fails with a different signature: | Layer | Misconfigured | What the caller sees | |---|---|---| | Resolution | peer gets the public answer | request reaches the public front door and is **refused** | | Routing | private answer, no route to the endpoint address | connection **hangs, then times out** | | Acceptance | arrives privately, credential grants nothing | refused, but with an **authorization** shape naming the principal | A refusal tells you the bytes arrived somewhere. If the store is configured to accept only the private path, a refusal from a peer is near-diagnostic of a public answer: the peer looked the name up, got the provider's public address, and went out through its own internet path. ## Why the answer does not cross the link and the route does A peering link joins two private address ranges at the routing layer. It carries reachability, and a route to the endpoint's address can propagate across it like any other prefix. Name resolution is a different service: the private answer for the service's hostname is an association between that answer and a **set of networks**. Platforms differ in how that association is made — some attach the answer to each network explicitly, some let a designated network share its resolution with peers, some expect an endpoint of its own in each network — but they agree that the network holding the endpoint is not automatically speaking for its peers. The consequence is a class of bug that looks like a permission problem and is a resolution problem, which is why teams reach for the credential first and lose an afternoon. ## The fix, and its two variants 1. **Extend resolution.** Attach the private answer for that hostname to the peered network, or point that network's lookups at a resolver that holds it. This is the cheap fix and it keeps one endpoint serving both networks. 2. **Give the peer its own endpoint.** Where the platform does not share an endpoint across the link, or where the two networks are owned by different teams and should not depend on each other's resolution, place a second endpoint in the peer. This costs another standing charge and more addresses, and it removes a cross-team coupling. Either way, verify in this order: - resolve the service's published hostname **from a workload in the peer** and confirm the address is inside a prefix you own; - confirm the peer has a route to that address and that the endpoint's network permits the traffic at its boundary — a correct answer with no route is the timeout case; - only then look at the store's own rules for the calling identity. ## What this is not - **Not a peering mechanics question.** Whether the link is point-to-point or through a hub, and whether routing is transitive, is a different subject; this failure occurs on a link that is working correctly for everything else. - **Not a pricing question.** The peer's traffic reaching the public front door does change what the bytes cost, but the failure here is a refusal, not a bill. - **Not a credential problem**, even though the error text often reads like one. The identity was never the variable that changed. ## The wider lesson to state in an interview Private service access is the one construct in this category where the **control plane of names** and the **data plane of routes** must agree, and they are configured in different places by different people. An estate that adds private endpoints steadily, without a standard for where the private answer is attached, accumulates networks that are silently on the public path — quietly working until a store starts refusing the public path, at which point every one of them fails on the same afternoon. That is why the resolution association belongs in the same change as the endpoint, and why the verification step is a lookup from inside each network that must use it, not from the network that owns it.

  • The peer resolves the endpoint's address correctly but the call hangs instead of being refused. What changed in the diagnosis?
    Resolution is now fine and reachability is not. A hang points at the routing layer or the boundary filter: the peer may have no propagated route to the endpoint's prefix, or the endpoint's network may be dropping traffic from the peer's range. Nothing about the store's rules is involved yet.
  • When would you place a second endpoint in the peer rather than extend resolution to it?
    When the platform does not let one endpoint serve both networks, or when the two networks belong to different teams and you do not want one owning the other's name resolution. It costs another standing charge and more addresses, and buys independence in exchange.
  • Why does this class of defect usually surface all at once rather than gradually?
    Because networks silently on the public path keep working until the store starts refusing that path. The enforcement change is what converts a latent misconfiguration into an outage, which is why enforcement should follow a verified lookup from every network expected to use the private path.

saying these in an interview costs you the question

  • Starts by rotating the calling identity's credential
  • Assumes the private answer follows any peering link automatically
  • Reads an explicit refusal as evidence of a missing route
  • Blames transitive routing for a resolution failure
  • Concludes the endpoint itself is broken because one network works