You must back up the Amazon EBS volumes of a busy database instance without stopping it. What consistency does a snapshot of a running instance give you by default, and what would you do to get more?
answer
- a snapshot is a power cut
- the engine's log does the rest
- freeze only for the API call
- two volumes, two moments
- the plural API coordinates them
basics
~20 sA snapshot of a running instance is crash-consistent: it captures the volume as if power were cut, so recovery relies on the database's own journal replay. For application consistency you flush and quiesce writes briefly before the API call, and use CreateSnapshots to capture multi-volume instances at one point in time.
solid answer
~50 sBy default a snapshot taken while the instance runs is **crash-consistent** — equivalent to pulling the plug. For a journaling filesystem and a database with a write-ahead log that is usually recoverable, because the engine replays or rolls back on start, but it is a recovery, not a clean backup, and anything still buffered in the application is lost. To do better I get the writers quiet for the instant of the call: flush and freeze the filesystem, or quiesce the database, issue `CreateSnapshot`, then release — the freeze only has to cover the API call, not the copy, because the snapshot's point in time is when the call is made and the data copies in the background. The other half of the answer is multi-volume: if data and log sit on separate volumes, two independent `CreateSnapshot` calls are not the same instant. `CreateSnapshots` takes an instance and produces coordinated, point-in-time-consistent snapshots of all its attached volumes.
code
bash · 5 lines# coordinated point-in-time snapshots of every volume on one instance
aws ec2 create-snapshots \
--instance-specification InstanceId=i-0abc,ExcludeBootVolume=false \
--description "coordinated db backup" \
--tag-specifications 'ResourceType=snapshot,Tags=[{Key=Backup,Value=db}]'go deeper
Know that you can snapshot a volume while the instance is running, and that the result is like a sudden power loss rather than a clean shutdown copy.
Explain crash consistency versus application consistency, and that the snapshot's point in time is the API call while the data copies in the background.
Demonstrate the production sequence: quiesce and freeze briefly around the call, use CreateSnapshots for an instance whose dataset spans volumes, and plan for lazy-load slowness on restore.
Own the recovery objective end to end — which datasets justify hand-rolled snapshot backups at all, where copies live, who may delete them, and how often a restore is actually rehearsed against a stated RTO.
## What you get for free, and what it is worth Calling `CreateSnapshot` against a volume attached to a running instance is completely safe for the volume — nothing is blocked, the instance keeps serving I/O, and the snapshot's point in time is the moment the call is accepted, with the data copied out in the background. What you get is a **crash-consistent** image: exactly the state a power cut would have left on disk. Whatever had reached the volume is there; whatever was still sitting in the OS page cache or an application buffer is not. Partially written application state is possible, though EBS preserves the ordering of writes it has acknowledged. For most modern software that is recoverable. A journaling filesystem replays its journal at mount; a relational engine with a write-ahead log replays committed transactions and rolls back the rest at startup. So the honest framing is: crash consistency usually works, at the price of a recovery cycle whose duration you have not measured and whose success you have not tested. ## Getting application consistency Application consistency means the data on disk is a clean, self-describing state, not one that needs recovery. You reach it by making the writers still for the instant of the snapshot call: 1. Tell the application to flush its buffers and hold new writes — for a database, that is the engine's own quiesce or read-lock mechanism after a checkpoint. 2. Flush and freeze the filesystem so nothing is left in the page cache. 3. Call `CreateSnapshot` (or `CreateSnapshots`). 4. Unfreeze and release the lock as soon as the call returns. The critical property that makes this practical: **you only hold the freeze for the API call**, not for the duration of the copy. The snapshot is a point-in-time reference established immediately; the block copying happens afterwards against that reference while the volume is fully in use again. A well-written hook holds writes for a moment, not for the minutes the snapshot takes to reach `completed`. Amazon Data Lifecycle Manager can run pre- and post-scripts around scheduled snapshots via Systems Manager, which is how you get this quiesce sequence to run on a schedule instead of living in a hand-rolled cron job. On Windows, the equivalent is a VSS-coordinated snapshot driven through Systems Manager, so VSS-aware applications flush themselves. ## The multi-volume trap This is the part that separates a senior answer from a textbook one. Databases commonly separate data files and the transaction log onto different EBS volumes for performance. If you snapshot them with two separate `CreateSnapshot` calls, you have two *different* points in time — seconds apart under load — and restoring them together can produce a log that does not correspond to the data files. It may still recover; it may not. The fix is the plural API: ```bash aws ec2 create-snapshots \ --instance-specification InstanceId=i-0abc,ExcludeBootVolume=false \ --description "coordinated db backup" ``` `CreateSnapshots` takes an **instance** and produces snapshots of all its attached volumes at a single coordinated point in time. Any design where one logical dataset spans several volumes should use it. ## Restore is the other half of the exercise A backup you have never restored is a hypothesis. Two AWS-specific realities to plan for: - **Lazy loading.** A volume created from a snapshot serves reads immediately but fetches untouched blocks from snapshot storage on first access, so a restored database is slow until warmed. If your recovery time objective is tight, either pre-read the device after restore or enable **Fast Snapshot Restore** on the snapshot in the AZ you would restore into — it makes volumes fully initialised at creation, and is billed per snapshot, per AZ, per hour, so you enable it on the snapshots you would actually use. - **The recovery itself takes time.** Journal replay on a large, busy database is not instant. Measure it, because it is part of your RTO whether you measured it or not. ## When to stop doing this by hand If the workload is a supported engine, a managed database gives you coordinated, application-aware backups and point-in-time recovery without any of this choreography. Hand-rolled EBS snapshot backups of a self-managed database on EC2 are a legitimate choice — but the tradeoff you are accepting is that consistency, retention and restore testing are now your job.
- Why does the write freeze only need to last for the API call rather than the whole snapshot?Because the snapshot's point in time is fixed when the call is accepted; the block copying happens afterwards against that reference while the volume continues serving I/O. So the quiesce window is milliseconds to a second or two, not the minutes the snapshot spends reaching completed state.
- Your restored database volume comes up but runs at a fraction of normal speed for an hour. What is happening?Blocks are being lazily fetched from snapshot storage on first touch, so every cold read pays a round trip. Warm it by reading the whole device after restore, or enable Fast Snapshot Restore on that snapshot for the target AZ so the volume is fully initialised at creation — the latter is billed per snapshot per AZ per hour.
- Would you still hand-roll snapshot backups for a database you could run as a managed service?Usually not. A managed engine gives coordinated backups, point-in-time recovery and tested restores without any quiesce choreography. Rolling your own on EC2 is defensible when you need an engine, version or extension the managed service does not offer — but then consistency, retention and restore rehearsals are explicitly your responsibility.
saying these in an interview costs you the question
- Says a snapshot of a running instance is always application-consistent
- Thinks the instance must be stopped to snapshot safely
- Holds the filesystem freeze until the snapshot reaches completed
- Snapshots data and log volumes with separate independent calls
- Never rehearses the restore or measures recovery time