skip to content

Your platform runs about forty services on AWS. Would you put them all behind one shared Application Load Balancer using host- and path-based rules, or give each service its own ALB? What drives the decision?

level: principalimportance: should knowfreq 30%

answer

  1. not one, not forty
  2. what is shared when they share a load balancer
  3. hourly charge versus quotas and blast radius
  4. cut on boundaries that already exist
  5. some settings are load-balancer-wide

basics

~20 s

Neither extreme. Group services into a handful of ALBs along boundaries that already exist - public versus internal, team or account ownership, and traffic profile - because a shared load balancer saves money but shares quotas, load-balancer-wide settings, blast radius and change control.

solid answer

~60 s

I would not run forty load balancers, and I would not run one. A shared ALB is cheaper - you pay one hourly charge instead of forty and capacity units only track real traffic - and it gives one place for TLS, WAF and access logs. But everything on it is shared: the rules quota, the certificate quota, load-balancer-wide attributes like the idle timeout and desync mode, one WAF web ACL, one failure domain, and one rule table that forty teams have to change safely. A dedicated ALB buys isolation, per-service tuning and clean ownership, at forty hourly charges and forty things to keep patched and monitored. So I segment on boundaries that already exist: public edge versus internal, one per team or per AWS account so IAM can delegate rule changes cleanly, and a separate load balancer for anything with an unusual profile - long-lived streaming connections needing a large idle timeout, or a workload whose traffic spikes would disturb its neighbours. That usually lands at a handful of ALBs, not one and not forty.

go deeper

for a junior

Understand that one ALB can front many services through host- and path-based rules, and that running a separate load balancer per service costs more because each one carries its own hourly charge.

for a middle

Explain concretely what is shared: rules and certificate quotas, load-balancer-wide attributes such as the idle timeout, one WAF web ACL, one set of access logs and one failure domain.

for a senior

Argue the operational consequences - per-service alarming from target-group metrics, a rule table that shadows itself when priorities are careless, and which single-service requirements justify pulling a workload onto its own load balancer.

for a principal

Own the seam definition and the governance: where the boundaries fall (exposure, account ownership, traffic profile, compliance), who may change routing and through what pipeline, and how the convention survives another year of growth.

## Why this is a real question and not a preference An Application Load Balancer is both a routing device and a unit of blast radius, cost, quota and ownership. Consolidating trades isolation for economy; fragmenting does the reverse. At forty services the two extremes are both visibly wrong, so the interesting work is choosing the seams. ## What a shared ALB gives you **Cost.** ALB pricing is an hourly charge per load balancer plus capacity units. The hourly charge is the part that multiplies with instance count regardless of traffic, so forty mostly-idle internal load balancers are forty standing charges for the privilege of being idle. Capacity-unit consumption tracks new connections, active connections, processed bytes and rule evaluations - it follows real traffic, so consolidating does not make it disappear, but it also does not penalise you for sharing. **Fewer moving parts.** One DNS name (or a wildcard), one certificate story, one web ACL, one access-log destination, one set of alarms. Onboarding a new service is a rule, not a new piece of infrastructure. ## What a shared ALB costs you **Shared quotas.** Rules per load balancer, certificates per load balancer and target groups per load balancer are all service quotas. They are adjustable, but they are real ceilings and they are consumed collectively. A convention of fifteen path rules per service hits the rules ceiling long before forty services are onboarded; one host-header rule per service scales much further. Design the routing convention with the quota in mind rather than discovering it during an onboarding. **Shared load-balancer-wide settings.** Some attributes belong to the load balancer, not the rule or the target group: the idle timeout, the desync mitigation mode, HTTP header handling, access-log configuration. One service that needs a five-minute idle timeout for streaming imposes that timeout on every other service sharing the load balancer, which in turn means idle connections are held longer everywhere. When one tenant's requirement would degrade the rest, that tenant needs its own load balancer. **Shared blast radius and change control.** A rule table is a shared mutable resource. A mis-prioritised catch-all rule added by one team can shadow every rule beneath it, and a listener misconfiguration takes down all forty services at once. IAM does not slice this finely - permission to modify rules on a listener is not really permission to modify only *your* rules - so a shared load balancer implies either a central owner who makes all routing changes or a pipeline that owns the rule table and accepts declarations from teams. **Shared observability.** Metrics are per load balancer and per target group. Load-balancer-level signals - overall 5xx counts, rejected connections, capacity-unit consumption - become aggregates across forty services, so per-service alarms must be built from target-group dimensions instead, and log analysis needs to split by target group or host. That is workable, but it is work. ## The seams I actually cut on 1. **Exposure.** Internet-facing and internal are different load balancers by construction - the scheme is fixed at creation and cannot be changed - so that split is free. 2. **Account and ownership.** A load balancer lives in one account and VPC. If teams own accounts, the load balancer boundary follows the account boundary and IAM delegation becomes trivial instead of contorted. 3. **Traffic profile.** Anything with a materially different shape gets its own: long-lived streaming or WebSocket traffic needing a large idle timeout; a public endpoint requiring an aggressive WAF rule set that would false-positive on internal tools; a service whose spikes are large enough to matter to neighbours. 4. **Compliance.** A workload with a stricter TLS policy or its own logging destination is cleaner on its own load balancer than as an exception on a shared one. What is left after those cuts - typically a public ALB, an internal ALB, and one or two specials per environment - is where consolidation is genuinely safe. ## What I would say about the migration The move from forty to a handful, or from one to a handful, is done service by service using DNS and weighted routing, not in a single cutover. Each service moves, gets validated on the new path, and only then loses its old entry point. And I would write down the convention that keeps the design honest afterwards: one rule per service keyed on host header, priorities allocated in blocks per team with gaps, a default action that fails closed with a fixed response, and a review before anyone adds a wildcard rule at a low priority.

  • Which ALB settings are load-balancer-wide and therefore inherited by every service sharing it?
    The idle timeout, desync mitigation mode, access-logging configuration, deletion protection and the scheme are attributes of the load balancer itself. TLS policy and certificates sit on the listener, and health checks and stickiness on the target group. Anything in the first group is a shared decision, which is often what forces a split.
  • How would you let forty teams change their own routing on a shared ALB without letting them break each other?
    Not with raw IAM - rule-level permissions do not carve up a listener cleanly. Put the rule table in a pipeline: teams declare their host and target group in their own repository, the pipeline renders the full rule set, allocates priorities from a per-team block, and validates that no wildcard shadows another team before applying.
  • Does consolidating forty load balancers into a few reduce the capacity-unit portion of the bill?
    Barely. Capacity units track new connections, active connections, bytes processed and rule evaluations - all of which follow real traffic and move with it. What consolidation removes is the per-load-balancer hourly charge, which is why the saving is largest for many low-traffic internal services and smallest for a few busy ones.

saying these in an interview costs you the question

  • Always give every microservice its own load balancer for isolation.
  • One ALB can hold unlimited rules and certificates.
  • The idle timeout can be set per listener rule.
  • IAM can restrict a team to editing only its own rules on a shared listener.
  • Consolidating load balancers eliminates the traffic-based part of the cost.

context