When would you store a dataset in HBase rather than as a Hive table on HDFS?
answer
- One row versus a billion rows
- Sorted keys, ranges served per server
- Write-ahead log, memory, then flushed files
- Milliseconds by key, seconds by scan
- Both write onto the same filesystem
basics
~20 sChoose HBase when you need millisecond random reads and writes of individual rows by key, and updates in place. Choose a Hive table on HDFS when the workload is large sequential scans and aggregations over immutable files.
solid answer
~50 sThey solve opposite access patterns over the same filesystem. **HBase** is a sorted, column-family key-value store: rows are ordered by row key, split into regions served by RegionServers, written first to a write-ahead log and an in-memory MemStore, then flushed to HFiles on HDFS and merged by compaction. That machinery buys you point `get`s, short range scans and single-row updates in milliseconds, at high write rates. A **Hive table on HDFS** is directories of large files, usually ORC or Parquet, read by a batch engine. There is no index by key; the fast path is pruning partitions and streaming whole files. It is excellent at aggregating a billion rows and terrible at fetching one. So: serving layer, entity-by-key lookups, frequent mutations, high-cardinality keys → HBase. Analytics, joins, ad-hoc SQL over history → Hive. Note that HBase's own files live *on* HDFS, so this is not a storage-system choice but an access-path choice.
go deeper
Know the one-line split: HBase for fast lookups and updates by key, Hive tables for large scans and SQL aggregation. Be able to say that HBase's files still live on HDFS.
Explain the mechanics that create the difference — sorted row keys, regions and RegionServers, WAL plus MemStore plus HFiles and compaction — versus columnar files read sequentially with partition pruning.
Show the architecture judgment: batch-compute in the warehouse and publish to HBase for serving, plan compactions, and treat row-key design and pre-splitting as day-one decisions rather than later tuning.
Own whether a keyed serving store belongs in the platform at all, given cloud key-value alternatives, the operational cost of running HBase, and what the serving contract with product teams should be.
## Same storage, opposite access paths Both options ultimately put bytes on HDFS. The difference is what sits between the query and those bytes. ## What a Hive table on HDFS is A directory tree of large files — typically ORC or Parquet — described by a metastore entry. A query engine prunes partitions by directory, then reads files sequentially, exploiting columnar layout and predicate pushdown. There is no per-row index and no random write path: files are written whole, and "updating a row" classically means rewriting a partition. Hive 3 adds ACID `UPDATE`/`DELETE`/`MERGE` on managed transactional tables, but the implementation writes delta files that compaction later merges — it is designed for batch-scale corrections, not for thousands of point updates per second. Strengths: scans, aggregations, joins, SQL, cheap storage, many engines reading the same files. Weaknesses: single-row lookup means scanning; latency is seconds at best; frequent small mutations produce file explosions. ## What HBase is HBase is a distributed sorted map: `(row key, column family, column qualifier, timestamp) → value`. The design consequences matter more than the data model: - **Row keys are sorted bytes.** All rows are kept in lexicographic order, which makes point lookups and contiguous range scans on the key cheap, and makes anything *not* keyed by the row key a full scan. - **Regions and RegionServers.** The key space is cut into contiguous ranges called regions; each region is served by exactly one RegionServer, and regions split as they grow. HMaster handles assignment and balancing; ZooKeeper tracks live servers and the location of the meta table. - **Write path.** A write appends to the write-ahead log (WAL) for durability, then lands in an in-memory MemStore. When the MemStore fills it flushes to an immutable HFile on HDFS. Reads merge MemStore and HFiles, aided by block cache and Bloom filters. - **Compaction.** Minor compaction merges some HFiles; major compaction rewrites all files in a region and physically drops deleted and expired cells. This is the price of the write-optimized design and needs scheduling on busy clusters. - **Column families** are the physical grouping — each family is stored separately, so families should be few and should share access patterns. Strengths: millisecond point reads and writes, very high write throughput, sparse wide rows, versioning by timestamp, in-place updates, TTL per family. Weaknesses: no SQL and no joins natively (Apache Phoenix layers SQL on top), no secondary indexes out of the box, a full table scan is far slower than reading ORC, and everything hinges on the row key you chose on day one. ## The decision Ask what the *dominant* query is. **HBase fits** when requests arrive as "give me the record for this key" or "give me this key's events between two timestamps": user profiles, device or sensor state keyed by device id, a serving store behind an API, message or event stores accessed by entity, and workloads with heavy random updates. **A Hive table fits** when requests are "aggregate everything for last quarter grouped by region": reporting, ad-hoc analytics, feature computation over history, anything expressed naturally in SQL joins. **Very often the answer is both.** The classic Hadoop-era architecture computes in batch over Hive tables and publishes the results into HBase for low-latency serving. Choosing one to do both jobs is the mistake: HBase as an analytics store means scanning a key-value engine, and a Hive table as a serving store means seconds-long lookups. ## The row key is the whole design Because regions are contiguous key ranges, a monotonically increasing row key — a timestamp, a sequence number — sends every new write to the single region holding the highest range, so one RegionServer saturates while the rest idle. The standard remedies are salting (prefix with a hash bucket), hashing or reversing the leading component, and pre-splitting the table into regions at creation. The cost is that a scan then has to hit every bucket. This tradeoff is why "design the row key" is the first question in any real HBase design discussion, and why migrating later is painful: the key is baked into every write. ## Interview framing A weak answer says "HBase is NoSQL, Hive is SQL". The strong answer names the access path — sorted key ranges served from memory plus HFiles versus full sequential scans of columnar files — explains that both sit on HDFS, and ends with the combined architecture plus the row-key caveat.
- An HBase table keyed by an incrementing timestamp has one hot RegionServer. Why, and what do you change?Regions are contiguous row-key ranges, so every new key lands in the one region holding the top of the range and a single RegionServer takes all the writes. Fix it by salting the key with a hash bucket prefix, hashing or reversing the leading component, and pre-splitting the table at creation. The tradeoff is that time-range scans must now read every bucket.
- How do teams run SQL over HBase when they need it?Apache Phoenix layers a SQL engine and JDBC driver over HBase, mapping tables and secondary indexes onto HBase primitives and pushing predicates into region-server coprocessors. It works well for keyed and short-range queries but does not turn HBase into an analytics engine — large aggregations still scan a key-value store and belong on columnar files.
- Why is a full table scan in HBase slower than reading the same data as ORC?HBase stores every cell with its row key, family, qualifier and timestamp, and a read merges MemStore with multiple HFiles. ORC stores columns contiguously with lightweight compression, min/max indexes and predicate pushdown, so a scan reads far fewer bytes and skips whole stripes. The per-cell overhead that makes HBase good at point access is exactly what makes scans expensive.
A Hive table is a warehouse of pallets: unbeatable when you need to count everything, hopeless for fetching one screw. HBase is the parts drawer, labelled and sorted, where any single item is at hand but counting the whole stock takes all day.
saying these in an interview costs you the question
- Says HBase stores data outside HDFS entirely
- Treats HBase as a general analytics or reporting store
- Expects joins and ad-hoc SQL natively from HBase
- Ignores row-key design and then blames the cluster
- Thinks Hive ACID tables give OLTP point-update performance