skip to content

Networking & Proxy Concepts

The vendor-neutral half of this tree: what a proxy actually does, how traffic gets routed, balanced, secured and kept alive, and what a service mesh buys and costs. Interviewers start here because your answer has to hold whether you run nginx, Envoy, HAProxy or a managed load balancer — every product below is just one implementation of these ideas.

on this pageshow

questions

page 1 of 2

A reverse proxy load-balances across several backend instances and actively health-checks each one. What does the proxy do with a backend that starts failing those checks, and what has to happen before that backend receives traffic again?

level: juniorimportance: must knowfreq 72%

answer

  1. not one failed probe — a count
  2. consecutive failures out, successes back
  3. new requests stop, in-flight ones finish
  4. the proxy never restarts anything
  5. survivors absorb the ejected instance's load

basics

~20 s

After a configured number of consecutive failed checks the proxy marks the backend unhealthy and stops sending it new requests, while continuing to probe it. It rejoins the pool only after a configured number of consecutive successful checks.

solid answer

~40 s

A health check is the proxy's answer to one narrow question: should I send the next request here? A backend that fails a probe is not ejected immediately — proxies require N consecutive failures (an unhealthy threshold), because a single lost packet or a garbage-collection pause is not evidence of a broken instance. Once the threshold is crossed, the instance leaves the selection pool for new requests; requests already in flight are usually allowed to finish. The proxy keeps probing the ejected instance, and after a configured number of consecutive successes it goes back into rotation, often behind a ramp rather than at full share. Note what the proxy does *not* do: it never restarts the instance and never declares it globally broken — health here is per proxy, and another proxy may disagree.

go deeper

for a junior

Be ready to describe the loop out loud: repeated probes, several consecutive failures before ejection, several consecutive successes before return. Say clearly that the proxy stops routing but does not restart anything.

for a middle

Explain why the threshold exists at all and what it costs — every extra required failure adds one interval of detection delay while filtering out noise like a GC pause or a dropped probe packet.

for a senior

Show that you plan capacity for ejection: state how much headroom the pool needs so removing an instance does not push the survivors into the overload that ejects them too.

for a principal

Own the policy question of who defines check semantics across a fleet — what the platform mandates, what a service team may tune, and how you keep a hundred teams from each inventing a different meaning of healthy.

## The question the proxy is actually asking A reverse proxy or load balancer holds a list of backend instances it is allowed to send requests to. An active health check is how the proxy answers one narrow question on its own: *should I send the next request to this instance?* It is not a diagnosis of what is wrong, it is not a restart decision, and it is not a statement that the instance is broken for everyone. Almost every mistake in this area comes from forgetting that framing. ## The three numbers that define the check loop An active check runs on a timer, inside each proxy instance, against each backend. Three settings shape it: - **Interval** — how often a probe is sent (say, every 10 seconds). - **Timeout** — how long the proxy waits for the probe's response before calling that probe a failure. A probe that answers slowly is a failed probe, not a slow one. - **Thresholds** — how many *consecutive* failed probes flip a healthy backend to unhealthy, and how many *consecutive* successful probes flip it back. Every mainstream proxy exposes both; the names differ per product. The threshold is the part people forget. A single failed probe is a poor signal: probes are lost to dropped packets, a stop-the-world pause, a momentary scheduler delay, or a listener that was mid-reload. Requiring several consecutive failures buys a cheap, effective noise filter — at the cost of detection delay, which is the central trade of this whole topic. ``` probe ok -> failures = 0 probe fail -> failures += 1 ; if failures == unhealthyThreshold: mark DOWN (while DOWN) probe ok -> successes += 1 ; if successes == healthyThreshold: mark UP ``` ## What the transition actually changes Marking a backend unhealthy removes it from the pool the load-balancing algorithm chooses from **for new requests**. Requests already in flight to that instance normally run to completion; the proxy does not usually abort them, though it may stop reusing that instance's idle keep-alive connections. Probing continues — an ejected backend that stopped being probed could never come back. What does not happen matters just as much. The proxy has no authority over the process: it cannot restart it, redeploy it, or take it out of an orchestrator's rotation. Deciding to *restart* something unhealthy is a different system's job entirely. The proxy's power ends at "I will not route to you." ## Coming back The return path is deliberately asymmetric in most production configurations: quick to eject, slow to readmit. A common shape is two consecutive failures to eject and three to five consecutive successes to return, so a genuinely flapping instance cannot oscillate in and out of the pool every few seconds. Some proxies additionally lengthen the exclusion period each time the same instance is ejected again, and many ramp its traffic share up gradually rather than handing it a full share the instant it passes. ## The two consequences juniors miss **Ejecting a backend does not reduce the traffic.** The requests that instance was serving are now spread across the survivors. A pool of four instances running at 70% CPU cannot absorb one ejection — losing one puts the other three at roughly 93%, and the pool is one bad minute from ejecting itself into an outage. Health checking is only safe if the pool has headroom. **Health is per proxy, not global.** Each proxy instance runs its own check loop and keeps its own verdict. If you run six proxies, the backend receives six times the probe traffic, and during a partial network problem some proxies may consider a backend healthy while others do not. That is by design — each proxy routes based on reachability *from where it stands* — but it means "is instance 3 healthy?" is not a question with one answer, and a dashboard that shows one is aggregating. ## What a good answer sounds like Name the state machine (consecutive failures out, consecutive successes back), say plainly that in-flight requests usually finish while new ones stop, note that the proxy does not restart anything, and mention that the survivors absorb the load. That is the whole junior-level shape, and it sets up every harder question in this area.

  • Why is the number of successful checks needed to return often higher than the number of failures needed to eject?
    Because the two errors have different costs. Ejecting a healthy instance costs a little capacity; readmitting a broken one costs real user requests. Asymmetric thresholds also damp flapping: an instance that recovers for one probe and fails the next never accumulates enough consecutive successes to rejoin, so it stays out until it is genuinely stable.
  • If every backend in the pool fails its health check at once, what should the proxy do?
    Serving errors from an empty pool is strictly worse than trying a possibly-degraded backend, so many proxies fail open: when the healthy fraction falls below a configured floor, they ignore health status and balance across every host. It is a deliberate admission that a pool-wide failure is far more likely to mean the check is wrong than that every instance died simultaneously.
  • Does the proxy's health status have any effect on the backend process itself?
    None. The proxy only changes its own routing table. The process keeps running, keeps holding its connections, and keeps logging. Restarting or replacing an unhealthy instance is the job of whatever supervises it, and confusing the two leads people to expect a load balancer to fix a hung process, which it never will.

saying these in an interview costs you the question

  • Thinks one failed probe removes a backend immediately
  • Believes the proxy restarts the unhealthy backend
  • Assumes in-flight requests are killed on ejection
  • Treats health status as global rather than per proxy
  • Ignores that survivors inherit the ejected instance's load

context

open as a page

A reverse proxy can run at layer 4 (connection level) or layer 7 (request level). What can each one inspect and route on, and what does each one hand to the backend?

level: juniorimportance: must knowfreq 80%

basics

~20 s

A layer 4 proxy sees only connection data — source and destination IP and port, and for TLS the cleartext SNI hostname — and copies bytes through untouched. A layer 7 proxy parses each request, so it can route on path, method, headers or cookies.

open as a page

A web application runs on two identical instances behind a load balancer that spreads requests round-robin. Users report being logged out at random. What is causing that, and what can the load-balancing layer do about it?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Each instance keeps session state in its own memory, so round-robin sends the next request to an instance that never saw the login and treats the user as anonymous. At the proxy layer the fix is session affinity: pin a client to one backend.

open as a page

Both a forward proxy and a reverse proxy sit between a client and a server and relay traffic. What actually distinguishes them — which side deploys and configures each, and what does each one hide from the other side?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A forward proxy acts for the clients and is configured on the client side, so the origin server sees only the proxy. A reverse proxy acts for the service owner, is deployed in front of the backends, and the client never knows it exists.

open as a page

A load balancer in front of your web servers is configured to "terminate TLS". Which parts of the path from the browser to the application are encrypted after that, and which are not?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Terminating TLS means the load balancer holds the certificate's private key and completes the HTTPS handshake itself, so the browser-to-load-balancer hop is encrypted and the load-balancer-to-application hop is plain HTTP unless you separately encrypt it.

open as a page

A load balancer can decide a backend is bad in two ways: by sending it probe requests of its own, or by watching how real client requests to that backend turn out. Compare the two — what does each detect that the other misses, and why do serious production setups run both?

level: middleimportance: must knowfreq 66%

basics

~20 s

Active probes are synthetic and cheap to run, so they catch a dead or unreachable instance before users do, but they only ever test the probe path. Passive detection judges an instance from real request outcomes, catching failures probes never touch — at the cost of real failed requests.

open as a page

A round-robin load balancer sends each backend the same number of requests, yet one backend sits at 90% CPU while another is nearly idle. Explain how that happens, and what least-connection and weighted balancing each change.

level: middleimportance: must knowfreq 72%

basics

~20 s

Round-robin equalises request counts, not work. When request costs or backend capacities differ, equal counts produce unequal load. Least-connection routes by in-flight requests so busy backends are skipped; weighting fixes only known, static capacity differences.

open as a page

A service reaches its dependency only through a DNS name that resolves to several instance addresses, with no registry and no proxy in between. Which parts of service discovery does that give you, and which does it not?

level: middleimportance: must knowfreq 60%

basics

~20 s

DNS-only discovery supplies naming and resolution: a stable name mapping to a set of addresses. It supplies no health state, no push updates and no real per-request balancing, and its TTL is a hint that clients routinely ignore.

open as a page

Your services already get retries, timeouts and circuit breaking from an in-process resilience library, and public TLS terminates at a shared edge gateway. What does adding a service mesh give you that neither of those does, and what does the library still do better?

level: middleimportance: must knowfreq 60%

basics

~20 s

A service mesh applies retries, timeouts, per-hop identity and uniform telemetry to service-to-service traffic through proxies beside each workload, in any language and without code changes. An in-process library still sees call intent — fallbacks, semantics — that a proxy cannot.

open as a page

You are placing a proxy in front of an HTTPS service and must choose between terminating TLS at the proxy, passing the TLS connection through untouched, and terminating then re-encrypting to the backend. What does each choice buy you, and where does the private key live in each?

level: middleimportance: must knowfreq 64%

basics

~20 s

Termination puts the key on the proxy and lets it read and route HTTP; passthrough leaves the key on every backend and reduces the proxy to forwarding bytes; re-encryption terminates, routes, then opens a second TLS connection to the backend, so both hops are encrypted at the cost of two handshakes and two certificate sets to manage.

open as a page

Users get a 504 after about 30 seconds on a request that crosses a CDN, an edge load balancer and an ingress proxy, but the application log shows that same request completing successfully after 45 seconds. Explain what happened, and how you would set the timeouts across those hops.

level: seniorimportance: must knowfreq 62%

basics

~20 s

One hop's timeout fired at 30 seconds and synthesised the 504 while the application kept working to 45. Proxy hops run independent fixed timers with no shared deadline, so the effective timeout is the shortest hop, not the sum.

open as a page

A reverse proxy terminates the client's connection and opens its own connection to the backend. Why does the client keeping its connection alive not give you connection reuse on the backend leg, and what has to be true at the proxy for that reuse to happen?

level: middleimportance: should knowfreq 55%

basics

~20 s

A reverse proxy runs two independent connections per request, client-side and upstream-side, so client keep-alive says nothing about the backend leg. Upstream reuse needs an explicit pool, HTTP/1.1 on that leg, and no Connection: close forwarded through.

open as a page

A reverse proxy probes each backend every 10 seconds with a 2-second check timeout, and marks a backend unhealthy after 3 consecutive failed probes. Roughly how long can a dead backend keep receiving requests, and what goes wrong if you shrink those numbers to make detection nearly instant?

level: middleimportance: should knowfreq 54%

basics

~20 s

Roughly 22 to 32 seconds: up to one interval before the next probe, then three probes ten seconds apart, the last costing its two-second timeout. Tightening the numbers shortens that window but multiplies probe load and makes brief pauses eject healthy backends.

open as a page

A team replaces a connection-level (layer 4) proxy in front of an HTTP service with a request-level (layer 7) proxy. What extra work does the proxy now do, and what does that cost in CPU, memory and failure surface?

level: middleimportance: should knowfreq 58%

basics

~20 s

A layer 7 proxy terminates TLS, parses and validates every request, holds per-request state and buffers, and maintains its own upstream connection pool. That costs handshake and parsing CPU, memory per in-flight request, added latency, and a new component that can reject or time out requests itself.

open as a page

A cache tier of N nodes sits behind a proxy that picks a node with hash(key) mod N. What happens to the hit rate when one node is removed, and how does consistent hashing change the outcome?

level: middleimportance: should knowfreq 45%

basics

~20 s

Modulo hashing remaps most keys when N changes, so nearly the whole cache misses at once and the origin takes the load. Consistent hashing maps keys and nodes onto a ring, so removing a node moves only that node's share — roughly 1/N of keys.

open as a page

In a design discussion one engineer says "put a load balancer in front of it", another says "no, an API gateway", and a third says "it is just a reverse proxy". What actually distinguishes those three roles, and can a single process play all of them?

level: middleimportance: should knowfreq 55%

basics

~20 s

All three are reverse proxies; the words differ in emphasis. Load balancer stresses spreading traffic over a pool, API gateway stresses per-caller concerns such as authentication and quotas, and reverse proxy is the plain role underneath both. One process can be all three.

open as a page

In a sidecar-based service mesh, how many extra proxies does one service-to-service request cross, and where do the added latency and the extra memory actually come from?

level: middleimportance: should knowfreq 45%

basics

~20 s

Two: the caller's sidecar and the callee's sidecar, so a chain of N calls crosses 2N proxies. Latency is paid per hop and shows up worst in the tail; sidecar memory tracks how much mesh configuration the proxy holds, not how much traffic it carries.

open as a page

You are taking a backend out of a load balancer's pool for maintenance. Why is "stop sending it new requests" not the same thing as draining it, and what does the proxy have to do to remove it without failing work that is already in flight?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Marking a backend down only stops new requests. Draining also waits out requests already dispatched and connections already pooled to it, bounded by a deadline after which the proxy force-closes. Long-lived connections never end on their own and must be cut.

open as a page

A reverse proxy is configured to try the next upstream when a request fails, allowing up to two extra attempts with a 10-second read timeout each, while the calling client gives up after 15 seconds. Walk through what the client and the backends actually experience.

level: seniorimportance: should knowfreq 45%

basics

~20 s

Worst case is 30 seconds of proxy attempts against a 15-second client, so later attempts run after the client has already left. The backends do up to three times the work per user request, and slow backends turn retries into a load multiplier.

open as a page

During a traffic spike every instance in a backend pool slows down, health checks begin timing out, and the load balancer ejects instances one after another until almost nothing is left in rotation — turning a slowdown into a total outage. Explain the feedback loop, and what safeguards stop health checking from causing this.

level: seniorimportance: should knowfreq 47%

basics

~20 s

Overload makes checks time out, ejection moves that instance's traffic onto the survivors, the survivors slow further and are ejected in turn. The loop is closed because ejection increases the load that caused it. Safeguards: fail open below a healthy floor, cap the ejectable fraction, and never eject on latency alone.

open as a page

Ten replicas of a gRPC service sit behind a connection-level (layer 4) load balancer, but two replicas serve almost all the traffic and replicas added by autoscaling stay idle. Why does connection-level balancing produce this, and what fixes it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

A connection-level balancer picks a backend once per connection, and a gRPC client keeps one long-lived multiplexed connection carrying every call. Load follows connections, not requests, so a handful of pinned connections concentrate traffic and new replicas never receive one.

open as a page

A team enabled cookie-based session affinity at their load balancer to fix a stateful service. A year later, instances added by autoscaling during peak receive almost no traffic, and every rolling deploy produces a burst of errors. Explain both symptoms.

level: seniorimportance: should knowfreq 58%

basics

~20 s

Affinity pins existing sessions, so a new instance can only receive brand-new sessions — during a spike most traffic is already pinned and scale-out adds unusable capacity. A deploy destroys every session pinned to each replaced instance at once, producing error bursts.

open as a page

A web application sits behind a reverse proxy that terminates TLS and forwards to the app over plain HTTP. Every HTTPS request now ends in an endless redirect loop: the app believes the scheme is http, redirects to https, and the next request reaches it as http again. What did the proxy hide from the app, and what has to change on each hop?

level: seniorimportance: should knowfreq 44%

basics

~20 s

TLS ended at the proxy, so the app's own connection is plain HTTP and its force-https rule fires forever. Fix both hops: the proxy must send X-Forwarded-Proto, and the app must honour that header only on connections coming from the proxy.

open as a page

An instance deregisters itself from the service registry before it exits, yet callers keep sending it requests for tens of seconds afterwards. What determines the length of that window, and how do you shut the instance down without dropping those requests?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Deregistration is a write, not a stop: the removal must replicate, every caller must refresh its cached member list, and open connections outlive the entry. Keep serving after deregistering, for longer than that whole window, then exit.

open as a page

A JVM service keeps opening connections to an instance address that was replaced twenty minutes ago, even though `dig` run on the same host already returns only the new address and the record's TTL is 60 seconds. What is still holding the old address, and how would you fix it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The runtime and the connection pool are holding the old address, not DNS. A JVM caches lookups under its own networkaddress.cache.ttl setting, ignoring the record's TTL, and an already-open pooled connection never re-resolves at all.

open as a page

After enrolling a service in a sidecar-based service mesh, the application's first outbound calls at pod startup fail with connection refused or an immediate 503, and then everything works normally. What is happening, and how do you fix it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Traffic redirection is in place before the sidecar proxy is serving, so the application's earliest calls are steered into a proxy that has no listener or no configuration yet and are refused. Fix it by making the proxy start and become ready before the application container runs.

open as a page

Compare the two service-mesh data-plane models — a proxy injected into every pod versus a shared proxy running on each node — in terms of enrolment, resource cost, blast radius and L7 features.

level: seniorimportance: should knowfreq 32%

basics

~20 s

A per-pod sidecar isolates each workload and handles L7 itself, but enrolling or upgrading means restarting the pod and paying memory for every pod. A node-shared proxy avoids the restart and scales with nodes, at the cost of a shared blast radius.

open as a page

An edge proxy authenticates callers with mutual TLS, and the application behind it needs to know which client certificate was presented. How should that identity reach the application, and what is the classic way this arrangement is defeated?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The proxy verifies the client certificate and then asserts the resulting identity to the application in a request header, because the certificate itself cannot survive termination. That assertion is defeated whenever the header is not stripped from inbound requests or the application is reachable without going through the proxy.

open as a page

You are deciding whether to front a service with connection-level (layer 4) or request-level (layer 7) proxying. How do you make that call, and what situations make layer 4 the right answer even though layer 7 is more capable?

level: principalimportance: should knowfreq 42%

basics

~20 s

Choose by the decisions you need to make per request. If routing, rewriting, per-request retries or quotas depend on reading the request, you need layer 7. If the payload is not HTTP, the proxy tier may not hold the key, or throughput and latency dominate, layer 4 is correct.

open as a page

A compliance requirement states that traffic must be encrypted in transit. Your platform terminates TLS at the edge and speaks plain HTTP internally. How would you decide between leaving it as is, re-encrypting to every backend, and passing TLS through to the origin?

level: principalimportance: should knowfreq 33%

basics

~20 s

Start from what the control actually says and which network segments it covers, then pick the cheapest topology that satisfies it: plaintext inside a genuinely bounded segment, re-encryption where wires leave a trust boundary, and passthrough only when the edge must never hold the key or see plaintext.

open as a page

showing 1–30 of 38