Data Management
Cloud data patterns for making reads fast, spreading data out, and keeping large payloads out of the wrong places: cache-aside, materialized view, sharding, valet key, claim check and index table.
part ofResilience & cloud-native patternsoverview, primer and where to startread it →on this pageshowhide
explore
- Cache-Aside5 questions
- Materialized View5 questions
- Sharding Pattern6 questions
- Valet Key5 questions
- Claim Check6 questions
- Index Table6 questions
- Event Sourcing & CQRS (bridge)6 questions
- AI Engineerrole
- API Designskill
- Backend Developerrole
- Data Engineerrole
- DevOps / SRE Engineerrole
- Forward Deployed Engineerrole
- Full Stack Developerrole
- Game Developerrole
- Java Backend Developerrole
- Kotlin Backend Developerrole
- Redisskill
- Server-Side Game Developerrole
- Software Architectrole
- Software Design & Architectureskill
- System Designskill
questions
page 2 of 2A hash-sharded cluster running 8 shards needs to grow to 16 because storage per shard is filling up. What makes resharding — moving data to the new shard count — operationally risky, and how do teams typically do it with minimal downtime?
basics
~20 sResharding means moving live data between servers while the system keeps running, so the main risks are reads getting stale or wrong-shard answers mid-move and overloading production during the copy. Teams usually write to both old and new locations, copy the history over in the background, verify it matches, then switch reads over.
In production, a system using pre-signed S3 URLs for client uploads starts seeing a spike in 403 SignatureDoesNotMatch and expired-token errors from clients on flaky mobile networks, even though the URLs were generated correctly seconds before use. What are two distinct root causes worth investigating, and how do they differ?
basics
~20 sEither the client's clock is off so the link looks expired too early, or the upload got retried in a way that changed something about the request (like its size), so it no longer matches what was originally signed. Both look like the same error but need different fixes.
You're reviewing a proposed architecture where every message on an internal event bus, regardless of size, gets routed through claim check by default: producers always write to object storage first and always publish just a reference, even for a 500-byte status update. What's wrong with this default, and when does claim check genuinely not pay for itself?
basics
~20 sUsing claim check for every message, even tiny ones, adds an extra network call and a second system for no benefit -- small payloads fit on the bus fine. It's worth it only when payloads are large.
At what point does introducing a materialized view stop being worth it for a given read path, and what would you reach for instead?
basics
~20 sIf the data changes about as often as it's read, if slightly old data would cause real harm, or if a simple index or better query would already be fast enough, adding a materialized view just adds extra moving parts without enough benefit. Better indexing, caching, or just optimizing the query is often the simpler fix.
You're designing the shard key for a new multi-tenant SaaS platform's core database, where tenants range from a few rows to an enterprise customer with millions of rows. What factors should drive the shard key choice, and what happens in production if you get it wrong?
basics
~20 sPick a shard key that keeps each tenant's data together, since almost every query is scoped to one tenant — but plan for the fact that some tenants will be far bigger than others, or your biggest customer alone can overload a single shard. Getting it wrong means either slow cross-shard queries for a single tenant, or one giant customer crushing whichever shard they land on.
Your organization issues thousands of pre-signed URLs per hour for direct-to-storage access using a single shared account-level signing key. Security now requires the ability to revoke access for a specific issued token within seconds if it's suspected leaked, without breaking every other outstanding URL. How would you redesign the token-issuing architecture to support that?
basics
~20 sStop relying on one shared master key for everything. Instead, make each link individually trackable, for example by making tokens expire very fast, or by adding a lookup list of blocked token IDs that gets checked before the request reaches storage, so you can kill one without touching the rest.
In a cache-aside deployment, an attacker (or a buggy client) repeatedly requests keys that are known not to exist in the database, e.g. random or sequential IDs. Since a cache-aside cache only stores things that were successfully loaded, why does this traffic pattern hit the database on every single request, and what technique addresses it?
basics
~20 sThe cache only remembers things it successfully found. If a key never exists, the database always says 'not found' and nothing ever gets saved to the cache, so every repeat request goes straight to the database again. The fix is to also cache the 'not found' answer for a little while.
A cloud platform's event store holds several hundred million events across many entity streams. After a schema change to how one read projection interprets an old event type, the team needs to rebuild that projection from scratch by replaying the entire log. What operational and architectural considerations does 'replay' raise at this scale?
basics
~20 sRebuilding from hundreds of millions of events isn't instant or free — it takes real time, compute, and can strain the live system if done carelessly. You need a plan for how long it takes, whether old users see stale or half-built data during the rebuild, and how to handle old event formats safely.
A team is about to hand-build and maintain three separate index tables over a dataset stored in DynamoDB, to support queries by three different non-key attributes. What native DynamoDB feature should they evaluate first before committing to hand-rolled index tables, and under what circumstances would hand-rolled index tables still be the right call despite that feature existing?
basics
~20 sDynamoDB has a built-in feature called a Global Secondary Index that does most of this automatically — the database keeps it in sync for you. You'd only build your own index table by hand if you need something the built-in version can't do, like guaranteed instant consistency or very complex derived keys.
showing 31–39 of 39