skip to content

Data Management

Cloud data patterns for making reads fast, spreading data out, and keeping large payloads out of the wrong places: cache-aside, materialized view, sharding, valet key, claim check and index table.

part ofResilience & cloud-native patternsoverview, primer and where to startread it →
on this pageshow

questions

page 2 of 2

A hash-sharded cluster running 8 shards needs to grow to 16 because storage per shard is filling up. What makes resharding — moving data to the new shard count — operationally risky, and how do teams typically do it with minimal downtime?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Resharding means moving live data between servers while the system keeps running, so the main risks are reads getting stale or wrong-shard answers mid-move and overloading production during the copy. Teams usually write to both old and new locations, copy the history over in the background, verify it matches, then switch reads over.

open as a page

In production, a system using pre-signed S3 URLs for client uploads starts seeing a spike in 403 SignatureDoesNotMatch and expired-token errors from clients on flaky mobile networks, even though the URLs were generated correctly seconds before use. What are two distinct root causes worth investigating, and how do they differ?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Either the client's clock is off so the link looks expired too early, or the upload got retried in a way that changed something about the request (like its size), so it no longer matches what was originally signed. Both look like the same error but need different fixes.

open as a page

You're reviewing a proposed architecture where every message on an internal event bus, regardless of size, gets routed through claim check by default: producers always write to object storage first and always publish just a reference, even for a 500-byte status update. What's wrong with this default, and when does claim check genuinely not pay for itself?

level: principalimportance: should knowfreq 38%

basics

~20 s

Using claim check for every message, even tiny ones, adds an extra network call and a second system for no benefit -- small payloads fit on the bus fine. It's worth it only when payloads are large.

open as a page

At what point does introducing a materialized view stop being worth it for a given read path, and what would you reach for instead?

level: principalimportance: should knowfreq 40%

basics

~20 s

If the data changes about as often as it's read, if slightly old data would cause real harm, or if a simple index or better query would already be fast enough, adding a materialized view just adds extra moving parts without enough benefit. Better indexing, caching, or just optimizing the query is often the simpler fix.

open as a page

You're designing the shard key for a new multi-tenant SaaS platform's core database, where tenants range from a few rows to an enterprise customer with millions of rows. What factors should drive the shard key choice, and what happens in production if you get it wrong?

level: principalimportance: should knowfreq 40%

basics

~20 s

Pick a shard key that keeps each tenant's data together, since almost every query is scoped to one tenant — but plan for the fact that some tenants will be far bigger than others, or your biggest customer alone can overload a single shard. Getting it wrong means either slow cross-shard queries for a single tenant, or one giant customer crushing whichever shard they land on.

open as a page

Your organization issues thousands of pre-signed URLs per hour for direct-to-storage access using a single shared account-level signing key. Security now requires the ability to revoke access for a specific issued token within seconds if it's suspected leaked, without breaking every other outstanding URL. How would you redesign the token-issuing architecture to support that?

level: principalimportance: should knowfreq 30%

basics

~20 s

Stop relying on one shared master key for everything. Instead, make each link individually trackable, for example by making tokens expire very fast, or by adding a lookup list of blocked token IDs that gets checked before the request reaches storage, so you can kill one without touching the rest.

open as a page

In a cache-aside deployment, an attacker (or a buggy client) repeatedly requests keys that are known not to exist in the database, e.g. random or sequential IDs. Since a cache-aside cache only stores things that were successfully loaded, why does this traffic pattern hit the database on every single request, and what technique addresses it?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

The cache only remembers things it successfully found. If a key never exists, the database always says 'not found' and nothing ever gets saved to the cache, so every repeat request goes straight to the database again. The fix is to also cache the 'not found' answer for a little while.

open as a page

A cloud platform's event store holds several hundred million events across many entity streams. After a schema change to how one read projection interprets an old event type, the team needs to rebuild that projection from scratch by replaying the entire log. What operational and architectural considerations does 'replay' raise at this scale?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Rebuilding from hundreds of millions of events isn't instant or free — it takes real time, compute, and can strain the live system if done carelessly. You need a plan for how long it takes, whether old users see stale or half-built data during the rebuild, and how to handle old event formats safely.

open as a page

A team is about to hand-build and maintain three separate index tables over a dataset stored in DynamoDB, to support queries by three different non-key attributes. What native DynamoDB feature should they evaluate first before committing to hand-rolled index tables, and under what circumstances would hand-rolled index tables still be the right call despite that feature existing?

level: principalimportance: nice to knowfreq 35%

basics

~20 s

DynamoDB has a built-in feature called a Global Secondary Index that does most of this automatically — the database keeps it in sync for you. You'd only build your own index table by hand if you need something the built-in version can't do, like guaranteed instant consistency or very complex derived keys.

open as a page

showing 31–39 of 39