skip to content

What is the difference between vertically scaling a database (upgrading to a bigger server) and horizontally scaling it (adding more servers such as read replicas), and what is one major limitation of vertical scaling?

level: juniorimportance: must knowfreq 70%

answer

  1. scale up = bigger box, scale out = more boxes
  2. replicas ≈ linear read capacity
  3. vertical hits hardware ceiling
  4. read-your-writes staleness

basics

~20 s

Vertical scaling means giving one database server more CPU/RAM/disk. Horizontal scaling means adding more database servers and spreading the work across them. Vertical scaling is simple but hits a hardware ceiling and stays a single point of failure.

solid answer

~40 s

Vertical scaling (scale-up) buys a bigger box: more CPU, RAM, faster disks/NVMe, more IOPS, for one database instance. It's operationally simple - no application changes - but is bounded by the largest instance a cloud provider offers, gets expensive non-linearly, requires downtime to resize, and leaves a single point of failure. Horizontal scaling (scale-out) adds more machines - most commonly read replicas that serve read traffic in parallel - so read capacity grows roughly linearly and you get failover targets, at the cost of application-level complexity: routing logic, replication lag, and eventual consistency for read replicas. In practice teams vertically scale first because it's cheap and fast, then reach for horizontal techniques once they hit hardware limits or need higher availability.

go deeper

for a junior

Should know the basic definitions - bigger machine vs more machines - and that read replicas exist to handle read traffic. Doesn't need to know replication mechanisms or lag-handling strategies in depth.

for a middle

Should be able to explain read/write splitting at a mechanical level (writes to primary, reads to replicas) and name replication lag as the core trade-off, including at least one concrete symptom like stale reads.

for a senior

Should discuss how to route around replication lag in production (sticky reads, lag-aware routing, pinning post-write reads to primary) and reason about when vertical scaling is still the right first move versus premature horizontal complexity.

for a principal

Should reason about the combined scaling strategy across a system's lifecycle - when to combine vertical and horizontal scaling, how it interacts with cost, failover architecture, and where the line to sharding/write-scaling gets drawn organizationally.

## The two levers Vertical scaling and horizontal scaling are the two **fundamental levers** for growing a database's capacity, and almost every other technique in this space (read replicas, connection pooling, federation) is really a refinement of the horizontal-scaling lever. | Lever | The mechanism | The payoff | |---|---|---| | **Vertical scaling** (scaling up) | increasing the resources of a single database instance | the same connection string, the same single writable dataset, the same transactional guarantees | | **Horizontal scaling** (scaling out) | adding more database servers rather than making one server bigger | read capacity scales out much further than a single box could vertically | ## Scaling up **Vertical scaling**, also called scaling up, means increasing the resources of a single database instance: - more **CPU cores**; - more **RAM**, so more of the working set fits in the buffer pool/cache and fewer disk reads are needed; - **faster storage** - NVMe SSDs instead of spinning disks, or higher-IOPS cloud volumes; - more **network bandwidth**. Mechanically this is often as simple as changing an instance type in a cloud console and restarting the database, or migrating to bigger bare-metal hardware. The appeal is that nothing about the application changes: the same connection string, the same single writable dataset, the same transactional guarantees (a single-node relational database gives you full ACID semantics without any cross-node coordination). For many systems, especially early-stage products, vertical scaling alone can absorb years of growth cheaply and with almost no engineering risk. ## Where scaling up stops The limitation is that it has a ceiling and non-linear economics. - **The ceiling.** Cloud providers sell instance types up to a maximum size (e.g., a few terabytes of RAM, a few hundred vCPUs); once you're on the largest SKU, there is nowhere further to scale up. - **The economics.** Even before hitting that ceiling, price tends to grow faster than capacity - doubling RAM or cores on a top-tier instance frequently costs far more than double, because you're paying for exclusivity of high-end hardware. - **Availability.** Vertical scaling also does nothing for availability: it is still one machine, one point of failure, and resizing typically requires a restart (brief downtime) or a costly blue-green cutover. Crucially, it doesn't address a very common bottleneck shape: most OLTP workloads are read-heavy (often 80-95% reads), and a single writer node handling all reads and writes will eventually be CPU- or IO-bound on read traffic long before write volume is the problem. ## Scaling out **Horizontal scaling**, or scaling out, addresses that shape directly by adding more database servers rather than making one server bigger. The most common first step is **read replicas**: additional database instances that continuously receive a stream of changes from the primary (via the database's native replication mechanism - e.g., MySQL binlog replication, PostgreSQL streaming/WAL replication) and stay in near-real-time sync. The application then splits traffic - **read/write splitting** - - sending all writes (INSERT/UPDATE/DELETE) to the single primary/writer, - and routing read-only queries (`SELECT`) across one or more replicas, often behind a load balancer or a smart driver/proxy (e.g., ProxySQL, pgpool, or an ORM-level read/write router). Because you can add replicas roughly linearly, read capacity scales out much further than a single box could vertically, and replicas double as failover candidates if the primary goes down (one can be promoted). ## The cost of scaling out The trade-off horizontal scaling introduces is complexity, and specifically **consistency complexity**. Replication is asynchronous in almost every mainstream setup (synchronous replication exists but costs write latency because the primary waits for replica acknowledgment), which means there is always some **replication lag** - milliseconds under normal load, but seconds or more under replica load spikes, long-running queries, or network hiccups. Any read routed to a replica during that lag window can return stale data. This produces a very real production failure mode: a user submits a form (write goes to primary), the page immediately reloads and reads from a replica that hasn't caught up yet, and the user sees their own change appear to have been lost or reverted - a classic 'read-your-own-writes' violation. Teams handle this by 1. pinning reads to the primary for a short window after a write, 2. using session/'sticky' read affinity, 3. monitoring replica lag and routing away from lagging replicas, 4. or accepting eventual consistency for genuinely tolerant read paths (e.g., an activity feed). ## How it plays out in practice A concrete real-world shape: an e-commerce site vertically scales its primary Postgres instance for the first couple of years, then adds two read replicas as product-listing and search-page traffic grows 10x faster than checkout traffic. Product pages and category browsing route to replicas; checkout, cart mutations, and 'my orders' immediately after purchase route to the primary to guarantee freshness. Vertical scaling of the primary continues in parallel for write throughput, while horizontal read replicas absorb the read fan-out - the two techniques are complementary, not either/or, and most production databases end up using both.

  • Why do read replicas usually help more than vertical scaling for a typical CRUD web app?
    Because most CRUD workloads are read-heavy (often 80-95% reads), so the bottleneck is usually read throughput, not write throughput. Adding replicas parallelizes exactly that bottleneck across multiple machines, whereas vertical scaling just gives one machine more headroom on both reads and writes without addressing the fan-out problem.
  • What happens to a read replica if the primary database goes down?
    The replica keeps serving reads with whatever data it last received, but it stops getting new updates, so it drifts increasingly stale until either the primary recovers or a replica is promoted to become the new primary. Promotion usually requires orchestration (automatic failover tooling or manual intervention) and a brief write-unavailability window.
  • Can you scale writes horizontally the same way you scale reads with replicas?
    Not with plain read replicas - there is still exactly one writable primary, so write throughput is still bounded by that one machine. Scaling writes horizontally requires a different technique such as sharding/partitioning the data across multiple writable nodes, which is a separate topic from read replication.

Vertical scaling is like replacing a single cashier with a faster, more expensive cashier; horizontal scaling is like opening more checkout lanes - each lane serves customers in parallel, but now you need someone directing shoppers to the right lane.

saying these in an interview costs you the question

  • Claims read replicas make writes faster
  • Doesn't mention replication lag as a cost of horizontal scaling
  • Thinks vertical scaling has no ceiling
  • Confuses horizontal scaling with sharding without distinguishing them
  • Assumes replicas are always perfectly in sync

context