Raspberry Pis deployed as always-on devices frequently end up with a corrupted or worn-out microSD card. Why does that happen, and how would you design a deployment to avoid it?
answer
- finite erase cycles, cheap controller
- two failure modes, not one
- the translation layer, not your file
- every write you skip is lifetime
- or stop writing to the card at all
basics
~20 sConsumer microSD cards have limited write endurance and simple controllers, so constant small writes from logs, databases and swap wear them out, while power loss mid-write can corrupt the card's internal mapping. Cut the write rate, or move the root filesystem off the card entirely.
solid answer
~50 sTwo distinct failure modes get conflated. **Wear** is gradual: flash cells tolerate a finite number of erase cycles, and a card sized for holiday photos is being asked to absorb continuous journal writes, database commits and swap. **Corruption** is sudden: pulling power during a write can leave the card's internal translation layer inconsistent, and unlike an SSD there is no capacitor to finish the write. The fixes work down two axes. Reduce writes: send logs to RAM instead of the card, turn off swap, mount with `noatime`, and keep chatty state off the root filesystem. Change the medium: boot from a USB SSD or use a Compute Module with soldered eMMC, and where the workload permits, run the root filesystem read-only behind an overlay so a power cut cannot damage anything. Then buy the card deliberately — an endurance-rated or industrial card, not the cheapest one — and plan for clean shutdown or a small UPS.
code
bash · 5 lines# Where is the write traffic actually going?
grep -E 'mmcblk|sda' /proc/diskstats
findmnt -no OPTIONS /
swapon --show
journalctl --disk-usagego deeper
Know that a microSD is consumer flash with limited write endurance, that shutting down cleanly matters, and that constant logging or swapping to the card shortens its life.
Separate the two failure modes and explain them: gradual erase-cycle wear amplified by small frequent writes, and sudden corruption of the card's translation layer when power is lost mid-write. Name concrete write reducers such as RAM-backed logs, no swap, and noatime.
Design the deployment: measure where the writes come from, decide between reducing writes and changing the medium to SSD or eMMC, use a read-only root with an overlay where persistence is not required, and monitor for read-only remounts so a dying card is scheduled maintenance rather than an outage.
Own the fleet economics — endurance-rated media and a supercapacitor or UPS versus the cost of field visits and warranty churn, an update strategy that survives power loss mid-write, and a policy for what state is allowed to live on a device at all.
## Two failure modes, often confused When a fielded Pi stops booting, people say "the SD card died" as if that were one thing. It is two, with different causes and different fixes. **Wear-out.** NAND flash cells are erased in blocks and each block tolerates a finite number of program/erase cycles. A card's controller spreads writes across blocks to extend life, but a cheap microSD controller does this far less capably than an SSD's. Meanwhile the workload is brutal by consumer-card standards: journald writing continuously, an application log, a SQLite or Postgres database committing, atime updates on every read, and swap paging when memory gets tight. Small, frequent, aligned-badly writes trigger write amplification, so the flash sees far more erase traffic than the megabytes your application thinks it wrote. Eventually blocks retire, the spare pool empties, and the card either goes read-only or starts returning bad data. **Power-loss corruption.** This one can kill a brand-new card. Flash is not updated in place; a translation layer maps logical sectors to physical pages and maintains its own metadata. Cut power in the middle of that update and the mapping itself can be left inconsistent — not just your file, the card's own bookkeeping. An enterprise SSD has capacitors to flush in-flight writes; a microSD does not. This is why a Pi that gets unplugged rather than shut down accumulates damage, and why filesystem journalling protects your filesystem but cannot protect the layer beneath it. ## Reducing the write rate Every write you do not issue is endurance you keep. - **Logs to RAM.** The single biggest source of steady writes on an idle appliance is logging. Keeping the journal in memory rather than on disk, or mounting a RAM-backed filesystem over the log directory, removes it. The tradeoff is explicit: logs do not survive a reboot, so ship anything you care about off the device. - **Turn swap off.** On a Raspberry Pi OS install swap is a file on the card, managed by the `dphys-swapfile` service. Swapping to flash is both slow and destructive. Size the workload to fit in RAM and disable it rather than letting the machine grind the card. - **`noatime`.** Without it, reading a file writes back its access time, so a read-mostly workload still produces write traffic. - **Move chatty state.** A database, a metrics store, or a cache belongs on other media or in RAM, not on the boot card. ## Changing the medium Reducing writes buys time; changing the medium changes the category. - **Boot from USB SSD.** On a Pi 4 or 5 the on-board bootloader can boot from USB mass storage, and the Pi 5 additionally exposes a PCIe connection for NVMe. A real SSD brings vastly better endurance, proper wear levelling and, on decent models, power-loss protection. - **eMMC.** A Compute Module with soldered eMMC removes both the endurance problem and the physical failure mode of a card that vibrates loose in a socket. - **Read-only root with an overlay.** For an appliance whose job does not require persistence, mount the root filesystem read-only and put a RAM-backed overlay on top, so all writes land in memory and vanish at reboot. Raspberry Pi OS exposes this as an option in `raspi-config`. It converts "power cut corrupted the card" into "power cut lost this hour's temporary state", which for a kiosk or a sensor node is not a loss at all. ## Buying and operating deliberately Cards are not interchangeable. Speed-class markings describe throughput, not endurance; the relevant products are the high-endurance or industrial ranges, some using pseudo-SLC modes that trade capacity for lifetime. Buy from a supplier you trust — counterfeit cards that misreport capacity are common and fail immediately under sustained write load. Then operate for power loss, because in the field you will get it. Shut down cleanly where you can. Where you cannot, a small UPS or supercapacitor hat that signals impending loss and triggers a shutdown is a cheap fix compared to a truck roll. Design the application so an unexpected restart is uneventful: write state atomically, or accept that the last few seconds are lost. ## Monitoring, so it is not a surprise The last piece is knowing it is coming. Track filesystem errors and I/O errors in the kernel log, watch for the filesystem being remounted read-only — which is what a failing card usually triggers first — and record how many bytes the device is writing per day so you can see a regression when a new release starts logging at debug level. On a fleet, plan for card replacement as scheduled maintenance rather than as an incident.
- Why can power loss corrupt a brand-new card that has plenty of write life left?Because the damage is not wear. Flash is not written in place; the card's translation layer maps logical sectors to physical pages and keeps its own metadata. Losing power mid-update can leave that mapping inconsistent — the card's bookkeeping, beneath any filesystem. A journalling filesystem protects its own structures but has no visibility into that layer, and a microSD has no capacitors to flush in-flight writes the way an enterprise SSD does.
- You mount the root filesystem read-only with a RAM overlay. What have you actually traded away?Persistence. Every write lands in memory and disappears at reboot, so logs, application state and configuration changes do not survive unless you deliberately place them elsewhere — a small writable partition, a network target, or a remote log sink. In exchange, a power cut can no longer corrupt the card, and updates become a deliberate act of remounting read-write. For kiosks and sensor nodes that trade is almost always right.
- Does a faster speed-class card last longer?No — speed class markings describe sustained throughput for video recording, not endurance. The relevant distinction is the high-endurance or industrial product ranges, some of which use pseudo-SLC modes that trade capacity for far more erase cycles. Sourcing matters too: counterfeit cards misreporting capacity are common and fail almost immediately under sustained write load.
saying these in an interview costs you the question
- Just buy a faster class-10 card, it will last
- A journalling filesystem prevents SD corruption
- Pulling the power is fine, Linux handles it
- Cards only fail when they are physically full
- Swap on the card is harmless if there is free space