On AWS, the name "Glacier" shows up both as a standalone archival service with vaults and archives and as a set of Amazon S3 storage-class names. Where does archival data actually live in each case, and which of the two should a new design target?
answer
- two products share one name
- vaults and archive IDs versus buckets and keys
- opaque ID means you own the index
- storage classes are a property on a normal object
- GLACIER_IR, GLACIER, DEEP_ARCHIVE
basics
~20 sGlacier today is a set of Amazon S3 storage classes — GLACIER_IR, GLACIER and DEEP_ARCHIVE — applied to ordinary objects in ordinary S3 buckets. The older standalone S3 Glacier service stores opaque archives in vaults under its own API and is legacy; new designs should use the S3 storage classes.
solid answer
~50 sThere are two different things wearing the same name. The original service, launched as Amazon Glacier and now called S3 Glacier, has its own API surface: you create a *vault*, upload an *archive*, and get back an opaque archive ID — there are no keys, no prefixes, and no bucket. The modern form is a set of **storage classes on normal S3 objects**: `GLACIER_IR` (Glacier Instant Retrieval), `GLACIER` (Glacier Flexible Retrieval, which was simply called "S3 Glacier" before the 2021 rename) and `DEEP_ARCHIVE` (Glacier Deep Archive). Those objects stay in the same bucket under the same key, keep their versions, policies, tags and event notifications, and are listed by the same `ListObjectsV2` call — only the `StorageClass` field changes. For anything new, use the storage classes: you keep the entire S3 feature set and the vault API's opaque archive IDs are a data-management burden with no upside.
code
bash · 3 linesaws s3api put-object --bucket my-archive-bucket --key reports/2024-q4.tar --body 2024-q4.tar --storage-class GLACIER_IR
aws s3api list-objects-v2 --bucket my-archive-bucket --prefix reports/ --query 'Contents[].[Key,StorageClass]' --output table
aws s3api head-object --bucket my-archive-bucket --key reports/2024-q4.tar --query 'StorageClass'go deeper
Know that Glacier means cheap archival storage on AWS, and that today you normally get it by setting a storage class on a normal S3 object rather than by using a separate service.
Be ready to name the three storage-class values and to explain that an archived object keeps its bucket, key and version — the class is a property of the object, not a different place it was moved to.
Show that you resolve the naming ambiguity before designing: ask whether the data is objects or vault archives, and point out that the vault side loses bucket policies, replication, inventory and event notifications.
Own the migration argument. Explain why an archive addressed by an opaque ID makes your own index a single point of failure, and frame moving a legacy vault estate onto S3 storage classes as consolidating two control planes into one.
## Two products, one name "Glacier" is one of AWS's most confusing names because it refers to two distinct things, ten years apart in design. **Amazon S3 Glacier, the standalone service** (launched 2012 as "Amazon Glacier") is a separate API with its own vocabulary. You create a **vault** — a container roughly analogous to a bucket — and upload a **archive**, which is an opaque blob. The service hands back an **archive ID**: a long, machine-generated string. There is no key, no prefix, no folder-like listing. If you want to know what an archive contains, you have to keep your own index somewhere else, because AWS will only give you an inventory of IDs and sizes, refreshed periodically. Reading an archive back is not a synchronous `GET`; you start a retrieval job and collect the output later. **The Glacier storage classes** are the modern shape. Here nothing new is created — an ordinary S3 object in an ordinary bucket simply carries a different `StorageClass` value: - `GLACIER_IR` — S3 Glacier Instant Retrieval - `GLACIER` — S3 Glacier Flexible Retrieval (before the 2021 rename this class was just called "S3 Glacier", which is the root of most of the confusion) - `DEEP_ARCHIVE` — S3 Glacier Deep Archive The object keeps its key, its version ID, its metadata and tags, its object-level ACL or bucket-policy coverage, and its participation in event notifications. `ListObjectsV2` still lists it and reports the class: ```bash aws s3api list-objects-v2 --bucket my-archive-bucket \ --query 'Contents[].[Key,StorageClass]' --output table ``` ## Why the distinction matters in practice The practical consequence is **which API surface applies**. With the storage classes, everything you already built against S3 keeps working: the same SDK client, the same bucket policy, the same `s3:GetObject` and `s3:PutObject` permissions, the same replication and inventory tooling. The one behavioural difference is that objects in the asynchronous classes (`GLACIER` and `DEEP_ARCHIVE`) are not directly readable — a plain `GetObject` on one fails with the `InvalidObjectState` error until you have called `RestoreObject` on it. `GLACIER_IR` objects are read with an ordinary `GetObject`. With the vault service, none of that transfers. Bucket policies do not apply — vaults have their own vault access policies. S3 lifecycle rules do not target vaults. S3 Inventory, S3 Replication, S3 Event Notifications and presigned URLs are all bucket features and simply do not exist on the vault side. You are running a second, parallel storage system with its own identity model and its own operational habits. ## What to choose For any new workload the answer is the storage classes, and interviewers expect you to say so without hedging. The reasons are concrete rather than stylistic: 1. **Addressability.** Objects have meaningful keys you chose. Archives have IDs the service chose, so you must operate and back up your own ID-to-meaning index — a database whose loss makes the archive unreadable in practice. 2. **One control plane.** Access control, encryption configuration, logging and auditing all go through the S3 mechanisms you already run, rather than a second set. 3. **Transitions are declarative.** Moving cold data into an archival class is a property change on an object you already have, not an export into a foreign system. The vault API still exists and continues to serve deployments built on it years ago, so "we have a Glacier vault" is a legitimate legacy answer — but it is a migration candidate, not a design choice you make today. ## The trap in the naming Because the class now called Flexible Retrieval used to be called plain "S3 Glacier", older runbooks, blog posts and even internal wikis say "Glacier" when they mean one specific storage class, and other documents say "Glacier" when they mean the vault service. In a design discussion, resolve the ambiguity explicitly: ask whether the data is objects in a bucket or archives in a vault. That single clarifying question is most of what this topic is worth in an interview.
- Someone tells you their backups are "in Glacier". What do you ask next before you can reason about the design?Whether those are S3 objects in a bucket carrying a Glacier storage class, or archives in an S3 Glacier vault. The answer decides which control plane applies — bucket policies, replication, inventory and event notifications exist only on the bucket side, and vault archives are addressed by opaque IDs, so the team must be maintaining its own ID index somewhere.
- If an object is in the GLACIER storage class, what happens when your application calls GetObject on it?The request fails rather than blocking: S3 returns an `InvalidObjectState` error, because the object's data is not in a directly readable state. You call `RestoreObject` on the key first, then read it once the restore completes. Code that assumes every key in a bucket is immediately gettable is the usual source of this surprise.
- Does putting an object in a Glacier storage class change anything about its identity in the bucket?No. Same bucket, same key, same version ID, same tags and metadata, same coverage by the bucket policy. `ListObjectsV2` and `HeadObject` still return it and report the class in the `StorageClass` field. Only the readability of the payload and the billing rate change — the object's identity is untouched.
saying these in an interview costs you the question
- Claiming Glacier is a separate bucket type in S3
- Thinking archived objects move out of the bucket entirely
- Assuming a Glacier vault supports S3 bucket policies
- Saying GetObject just blocks until an archived object is ready
- Treating the vault API as the current recommended path