In Amazon S3, what does Replication Time Control (RTC) add to a replication rule, and how would you measure replication lag on a rule that does not use it?
answer
- default is best-effort, not bounded
- buying a tail guarantee, not raw speed
- fifteen-minute commitment, priced per GB
- metrics can be enabled without the SLA
- alarm on a growing backlog, not one object
basics
~20 sDefault S3 replication is best-effort with no time commitment. Replication Time Control adds a service level agreement covering replication of 99.99% of objects within 15 minutes, for an extra per-GB charge, and surfaces replication metrics and events so breaches are visible.
solid answer
~50 sWithout RTC, S3 replication is asynchronous and best-effort: most objects land at the destination within minutes, but nothing contractual bounds the tail, and a large object or a burst can sit behind others. **Replication Time Control** turns that behaviour into a commitment — an SLA covering replication of 99.99% of objects within 15 minutes — and it comes with **S3 Replication metrics** in CloudWatch plus replication event notifications, so a missed threshold is something you can alarm on rather than discover during a failover. You pay a per-GB replication-time charge on top of the normal replication costs. For measurement on a rule without RTC, you can enable S3 Replication metrics on the rule itself and watch `ReplicationLatency`, `BytesPendingReplication` and `OperationsPendingReplication` in CloudWatch; for a single object, `HeadObject` returns a replication status of PENDING, COMPLETED or FAILED. Reach for RTC when a recovery point objective is written down and someone will be held to it.
code
json · 11 lines{
"Bucket": "arn:aws:s3:::example-dr-bucket",
"ReplicationTime": {
"Status": "Enabled",
"Time": { "Minutes": 15 }
},
"Metrics": {
"Status": "Enabled",
"EventThreshold": { "Minutes": 15 }
}
}go deeper
Know that S3 replication is asynchronous and best-effort by default, and that Replication Time Control is the paid option that attaches a time commitment to it.
Explain what RTC actually delivers — a 15-minute commitment for 99.99% of objects, replication metrics, and threshold events — and name the per-GB charge that comes with it.
Tie it to a recovery point objective and build the operational picture: alarm on a rising pending backlog, use an object's replication status to separate authorisation failures from volume, and know that metrics do not require RTC.
Decide where the commitment is worth paying for: which datasets have an RPO someone will be held to, what evidence an auditor needs, and how the per-GB charge scales against the volume the bucket actually ingests.
## The default: best-effort S3 replication is asynchronous. A `PutObject` on the source returns as soon as the source write is durable, and replication happens afterwards. In practice the great majority of objects appear at the destination within minutes, but that is an observed behaviour rather than a guarantee. Large objects take longer. A sudden burst of writes queues. Nothing in the default configuration tells you how far behind you are, and nothing entitles you to anything if the tail stretches. That is fine for "we would like a second copy" and not fine for "our recovery point objective is fifteen minutes and the auditor has it in writing". ## What RTC adds **S3 Replication Time Control** is a per-rule option in the destination configuration. It provides: - **A service level agreement**: AWS commits to replicating 99.99% of objects within 15 minutes of upload, with service credits if the commitment is missed. Verify the current SLA text before quoting numbers in a design document — SLA terms are versioned documents that change. - **S3 Replication metrics** in CloudWatch, which give you the operational picture: `ReplicationLatency` (how far behind the rule is running), `BytesPendingReplication` and `OperationsPendingReplication` (the size and count of the backlog). - **Replication event notifications**, including an event raised when an object misses the 15-minute threshold and another when it replicates after having missed it. Those are the hooks you alarm on. It costs a per-GB **Replication Time Control data transfer** charge on top of the ordinary replication request charges, cross-Region transfer and destination storage. On a high-volume bucket that is not a rounding error, which is why RTC is a decision rather than a default. ## Measuring without RTC You do not have to buy the SLA to get visibility. **S3 Replication metrics can be enabled on a replication rule independently**; RTC simply includes them. Once on, the CloudWatch metrics above let you build the alarm that actually matters operationally: not "did one object miss", but "is the pending backlog growing", which is the shape of every real replication incident — a permissions change, a KMS key policy edit, a destination bucket policy tightened by another team. A good alarm pair: - `OperationsPendingReplication` above a threshold **and rising** over several periods — the rule has stalled rather than momentarily queued. - `ReplicationLatency` above your RPO — the tail has stretched past what the design assumes. For a single object under investigation, the replication status is exposed directly: ```bash aws s3api head-object --bucket example-source --key data/latest.json \ --query ReplicationStatus ``` `PENDING` means in flight, `COMPLETED` means replicated, `FAILED` means S3 tried and could not — almost always a permissions or key-policy problem rather than a transient one. On the destination object the same field reads `REPLICA`. ## When to turn RTC on Ask what changes if replication runs an hour late. If the answer is "nothing, the copy is for a compliance checkbox and a Regional disaster is the only scenario we care about", default replication with metrics and an alarm is proportionate. If the answer is "we fail over to the second Region and serve reads from it, so an hour of lag is an hour of missing data for customers", then the RPO is real, and RTC is how you make it a commitment with an alarm attached rather than an assumption. There is a third case worth naming: a **regulated workload** where you must be able to demonstrate the replication objective to an auditor. RTC's SLA and its metrics are evidence; "it usually takes about five minutes" is not. ## The trap in the question Candidates often answer that RTC "makes replication faster". It is better to say it makes replication **bounded and observable**. Throughput is broadly the same; what you are buying is the tail guarantee, the metrics, and the events that tell you when the tail was missed.
- Does enabling RTC make replication throughput higher?Not meaningfully — the value is a bounded tail rather than raw speed. What you gain is a commitment covering 99.99% of objects within 15 minutes, plus the metrics and events that make a breach visible. Describing RTC as "faster replication" is the common miss; describing it as "bounded and observable replication" is the accurate framing.
- Which CloudWatch alarm would you actually build on a replication rule?Alarm on `OperationsPendingReplication` or `BytesPendingReplication` staying elevated across several periods, because a stalled rule shows up as a backlog that grows rather than drains. Pair it with a `ReplicationLatency` alarm set at your recovery point objective. A single object missing a threshold is noise; a rising backlog is an incident.
- Replication metrics show a growing backlog with no configuration change on your side. Where do you look?Almost always at permissions the rule depends on but does not own: the destination bucket policy, the replication role's policy, and the KMS key policy on either side if objects are encrypted with SSE-KMS. Check `ReplicationStatus` on a stuck object — FAILED points at authorisation, while a long-running PENDING points at volume or object size.
saying these in an interview costs you the question
- Saying RTC makes replication faster rather than bounded
- Assuming default replication carries a time guarantee
- Believing replication metrics require paying for RTC
- Alarming on one late object instead of a growing backlog
- Quoting SLA numbers without checking the current terms