An archival lifecycle rule moved a few hundred million small log objects into S3 Glacier Deep Archive, and the monthly S3 bill went up rather than down. Which AWS-specific charges explain that, and what would you do differently?
answer
- the bill follows object count here
- an index entry per archived object
- the move itself is billable
- six months of rent, paid up front
- archive bundles, not files
basics
~20 sArchiving is billed per object, not just per gigabyte. Glacier Flexible Retrieval and Deep Archive add fixed per-object metadata overhead, each transition is a paid request, and Deep Archive carries a 180-day minimum duration — so hundreds of millions of tiny objects lose on every axis.
solid answer
~50 sThree per-object charges swamp the per-GB saving. First, for `GLACIER` and `DEEP_ARCHIVE` S3 stores extra metadata for every archived object — a small amount billed at S3 Standard rates for the index entry plus a larger allowance billed at the archive rate — which is tens of kilobytes per object regardless of how small the object is. At a few hundred million objects that is terabytes of pure overhead. Second, every lifecycle transition is a billable request, and transitions into the Glacier classes are among the pricier request types, so the migration itself costs real money once. Third, Deep Archive has a 180-day minimum storage duration, so deleting the data sooner does not recover anything. The fix is to stop archiving objects and start archiving *archives*: roll the logs into large compressed bundles before transition, or expire them outright if nobody will ever read them.
go deeper
Know that S3 Glacier classes are billed per object as well as per gigabyte, so archiving a very large number of very small files does not automatically save money.
Explain the three per-object charges — archive metadata overhead, the billable transition request, and the minimum storage duration — and show how they scale with object count rather than with stored bytes.
Do the arithmetic before shipping the rule: pull object count and size distribution from S3 Inventory, compare overhead and request cost against the payload, and reach for compaction into large bundles rather than a wider transition rule.
Own the retention and data-layout standard that prevents this class of bill — where compaction happens in the pipeline, which datasets are expired rather than archived, and a review gate that costs an archival lifecycle change before it reaches production.
## The mental model: archive classes charge per object Every S3 cost intuition people carry is per-GB. The Glacier classes break it, because S3 has to keep an online index of objects whose data is offline. That index is billed to you, per object, and it does not shrink because your objects are small. Once object count rather than object size drives the bill, a rule that looks like a 95% discount can be a net increase. ## Per-object metadata overhead For objects in `GLACIER` (Glacier Flexible Retrieval) and `DEEP_ARCHIVE`, S3 stores additional metadata per object: a small block held at S3 Standard rates so the object name and metadata stay listable, plus a larger block billed at the relevant archive rate for the archive index. Together this is a few tens of kilobytes per object. Do the arithmetic out loud in the interview — it is the whole point of the question. For 300 million objects at roughly 40 KB of overhead each, that is about 12 TB of billed storage that contains none of your data. If the objects average 8 KB, your actual payload is around 2.4 TB. You are paying for five times more overhead than content, and part of it at S3 Standard rates. Note that Glacier Instant Retrieval (`GLACIER_IR`) behaves like a normal online class here and does not carry that archive-index overhead — but it has its own 128 KB minimum billable object size, so tiny objects lose there too, just differently. ## Transition requests A lifecycle transition is an API operation and it is billed per object moved, per 1,000 requests, with transitions into the Glacier classes priced well above ordinary requests. Moving 300 million objects is therefore a one-off charge in the thousands of requests-units, incurred before you save anything. It also lands as a spike in the month the rule is enabled, which is often how the problem gets noticed. ## Minimum storage duration Deep Archive bills a 180-day minimum per object; Glacier Flexible Retrieval and Glacier Instant Retrieval bill 90 days. So the escape hatch — "we will just delete it all" — does not work for six months. If a later expiration rule fires at day 30, you still pay the remaining 150 days at the archive rate, plus you have already paid the transition. ## And then retrieval If anyone ever wants the data back, restores are billed per request as well as per GB, so a per-object restore of hundreds of millions of tiny objects is expensive again, and slow — even under S3 Batch Operations, which is the only sane way to drive it. ## What to do instead **Aggregate before you archive.** The single highest-leverage change is reducing object count: roll a day or an hour of logs into one compressed bundle — gzip, or a columnar format such as Parquet if the data will ever be queried — and archive the bundle. Ten million objects become ten thousand. Every per-object charge collapses by the same factor, restores become tractable, and the data stays queryable by tools that read from S3 directly. **Ask whether it should exist at all.** Log data with no retention requirement and no reader should be expired, not archived. Archiving is a cost decision people reach for to avoid making a retention decision. **Check the size distribution first.** An S3 Inventory report gives object count and size per prefix, which is exactly the input this decision needs and is almost never consulted before a rule ships. Wire the projected cost of the rule — objects times overhead, plus transition requests, plus minimum duration — into the review of any archival lifecycle change. **Filter the rule.** Lifecycle filters accept `ObjectSizeGreaterThan`, so a rule can archive only objects above a threshold where the economics work and leave the rest to a compaction job or an expiration rule. ## The takeaway Glacier is cheap for large, cold objects and hostile to small ones. Anyone proposing an archival lifecycle rule should be able to state the object count and the median object size before the per-GB rate is mentioned at all.
- Does S3 Glacier Instant Retrieval suffer the same per-object overhead?Not the archive-index overhead — it behaves like an online class for listing and reads. But it applies a 128 KB minimum billable object size and a 90-day minimum duration, so a few-kilobyte object is still billed as 128 KB. Small objects are uneconomic across every cold class; only the mechanism differs.
- How would you size the saving before enabling an archival rule?Pull an S3 Inventory report for the prefix to get object count and size distribution, then compute three terms: payload GB at the archive rate, object count times per-object overhead, and object count times the transition request price. If overhead and requests are the same order as the payload, compact first and archive afterwards.
- What is the practical way to compact objects that are already archived?Restore them in bulk with S3 Batch Operations against an S3 Inventory manifest, run a job that concatenates them into large bundles, write the bundles back, then delete the originals — accepting that you still owe the remaining minimum duration on each. It is cheaper to compact before archiving than after, which is the argument for reviewing the rule up front.
saying these in an interview costs you the question
- Assumes archive cost is purely the per-GB rate
- Thinks lifecycle transitions are free because S3 performs them
- Believes deleting archived data early stops the charges
- Plans to archive billions of tiny objects without compaction
- Confuses Glacier Instant Retrieval's size minimum with archive overhead