skip to content

What does a CloudFront origin group with origin failover actually do, and which requests and failure modes does it not cover?

level: seniorimportance: nice to knowfreq 28%

answer

  1. primary plus secondary, in order
  2. status codes you nominate as failure
  3. reads only, never writes
  4. no background probing of the primary
  5. two origins, two copies of the content

basics

~20 s

An origin group pairs a primary and a secondary origin: when the primary returns one of the configured failover status codes, or the connection fails or times out, CloudFront retries that same request against the secondary. It covers only GET, HEAD and OPTIONS, and it is a per-request retry, not a health check.

solid answer

~50 s

You define an origin group containing exactly two origins — primary and secondary — plus failover criteria: a chosen set of status codes (from 4xx and 5xx values such as 500, 502, 503 and 504) that count as failure. A cache behavior then targets the *group* instead of an origin. When CloudFront's request to the primary hits a connection error, a timeout, or a response matching the criteria, it retries the identical request against the secondary and serves that. The limits matter as much as the feature: failover applies only to **GET, HEAD and OPTIONS** requests, so writes are never retried; there is **no background health checking**, so every request pays the primary's failure latency before falling back; and both origins must be able to serve the same content, which for two buckets means you keep them in sync yourself. It is a read-path resilience feature for one distribution, not a regional failover strategy — routing users away from a broken region is Route 53's job.

code

json · 20 lines
json
{
  "OriginGroups": {
    "Quantity": 1,
    "Items": [
      {
        "Id": "assets-og",
        "FailoverCriteria": {
          "StatusCodes": { "Quantity": 3, "Items": [500, 502, 503] }
        },
        "Members": {
          "Quantity": 2,
          "Items": [
            { "OriginId": "primary-bucket" },
            { "OriginId": "secondary-bucket" }
          ]
        }
      }
    ]
  }
}

go deeper

for a junior

Know that a distribution can hold a primary and a secondary origin, and that CloudFront can retry a failed read against the secondary automatically.

for a middle

Explain the configuration — failover status codes, two ordered members, a behavior targeting the group — and that failover triggers on connection errors and timeouts as well as on listed statuses.

for a senior

Show the operating reality: reads only, no health checking so every miss pays the primary's timeout, content parity is your problem, and a distribution silently running on its secondary looks healthy from the edge.

for a principal

Position it correctly in a resilience strategy — what the edge retry covers versus DNS-level health-checked routing versus load-balancer health — and set the expectation that an untested failover path is not a failover path.

## The mechanism An origin group is a small construct in the distribution config: an id, a `FailoverCriteria` list of HTTP status codes, and exactly two members in priority order. ```json { "Id": "assets-og", "FailoverCriteria": { "StatusCodes": { "Quantity": 3, "Items": [500, 502, 503] } }, "Members": { "Quantity": 2, "Items": [{ "OriginId": "primary-bucket" }, { "OriginId": "secondary-bucket" }] } } ``` A cache behavior's target may be an origin **or** an origin group. On a cache miss CloudFront goes to the primary. If the connection cannot be established, the origin read times out, or the response status is in the failover criteria, CloudFront issues the same request to the secondary and serves whatever comes back — including an error, since there is no third attempt. ## What it is good for - Static assets replicated into two buckets in two regions, so a regional S3 disruption does not blank the site. - A custom origin backed by an alternate stack, where the secondary can serve a degraded but valid response. - Absorbing the brief unavailability of a primary during a deployment, for read traffic. ## The four limits worth stating in an interview **1. Read methods only.** Failover happens for GET, HEAD and OPTIONS. A POST that fails against the primary is not retried against the secondary — and that is deliberate: a blind retry of a non-idempotent write is a correctness bug, not resilience. So an origin group in front of an API protects almost nothing. **2. No health checks.** Nothing probes the primary in the background and marks it down. Every failing request tries the primary first, so during a full primary outage every single miss pays the primary's timeout — potentially seconds — before the secondary is even contacted. Latency degrades badly even though the site technically stays up. **3. Only the statuses you list, and only from the origin.** If the primary answers 200 with wrong or empty content, that is a success as far as CloudFront is concerned. Likewise a 403 is only a failure if you put 403 in the criteria — and adding it can be dangerous with an S3 origin under origin access control, where 403 is also what a genuinely missing object returns. **4. Content parity is on you.** Two S3 origins are two independent buckets; CloudFront does not copy anything between them. Keeping the secondary current is a separate replication concern, and a stale secondary silently serves old assets exactly when you are least able to notice. ## Where it sits against the alternatives | Need | Right tool | |---|---| | Retry a failed read at the edge, same distribution | origin group | | Route users to a healthy region, DNS level | Route 53 with health checks | | Spread load across many backends | a load balancer as the origin | | Survive a single unhealthy instance | load balancer target health checks | Origin groups compose with these rather than replace them: a common shape is Route 53 health-checked records handling the regional decision, and origin groups handling the narrower case of one origin service degrading while the region is otherwise fine. ## Operational notes CloudFront's access logs record the result the viewer received, so a distribution quietly running on its secondary can look perfectly healthy from the outside. If you rely on failover, alarm on the primary's own health independently rather than inferring it from edge metrics, and rehearse the failover — an untested secondary tends to be an empty or stale one.

  • Why does origin failover deliberately exclude POST and PUT?
    Because retrying a non-idempotent write against a different backend can duplicate the effect — two orders, two charges — and CloudFront cannot know whether the primary applied the request before failing. Excluding writes keeps the feature safe by default; write-path resilience belongs in the application's idempotency design, not in a blind edge retry.
  • During a full outage of the primary origin, why does the site feel slow even though failover works?
    There is no background health check marking the primary down. Every cache miss attempts the primary first and only falls back after a connection error or the origin read timeout, so each request carries that full delay. Failover preserves availability, not latency, which is why DNS-level health-checked routing is the better answer for a sustained regional failure.
  • What goes wrong if you add 403 to the failover criteria for an S3 origin behind origin access control?
    Under origin access control, a request for a key that does not exist returns 403 rather than 404. Treating 403 as failure makes every missing object trigger a second full round trip to the secondary, which then returns 403 as well — doubled latency and request cost for what is simply a typo or a stale link.

saying these in an interview costs you the question

  • Thinks CloudFront health-checks the primary continuously
  • Expects POST requests to fail over to the secondary
  • Assumes CloudFront replicates content to the secondary origin
  • Treats an origin group as load balancing across origins
  • Uses origin groups as the whole multi-region strategy

context