skip to content

When should you read a Couchbase document with the KV API instead of a SQL++ query?

level: middleimportance: should knowfreq 52%

answer

  1. One path needs no index at all
  2. The other involves three separate services
  3. Latency is a single hop to one node
  4. Deterministic key naming keeps it available
  5. A SQL++ clause names keys directly

basics

~20 s

Whenever the document key is known. A KV read hashes the key to the node that owns it and is served from that node's managed cache, with no index lookup, no query planning and no hop through the Query service.

solid answer

~50 s

Couchbase is a key-value store with a query layer bolted alongside, not on top. A KV `get` sends the key straight to the data node that owns it and is answered from the memory-first managed cache, so it is the cheapest read the cluster offers and it needs no index at all. A SQL++ query, by contrast, goes to a Query node, is parsed and planned, consults the Index service for an access path, then fetches the qualifying documents from data nodes — several more hops and far more CPU per result. So: known key, or a whole document by key, use KV; unknown keys, predicates, joins, aggregation or ad-hoc reporting, use SQL++. Design deterministic keys (`order::{id}`) precisely so the hot paths stay KV. If you must express a key lookup in SQL++, `USE KEYS` skips the index and goes directly to the data nodes.

code

javascript · 3 lines
javascript
// Node SDK: direct key-value read, no index required
const coll = cluster.bucket('app').scope('sales').collection('orders');
const res = await coll.get('order::1001');

go deeper

for a junior

Know that Couchbase can fetch a document directly by its key without any query or index, and that this is the normal way an application reads a document it already identifies.

for a middle

Explain the mechanical difference — one hop to the owning node from the managed cache, versus planning plus index lookup plus fetch — and give the rule for when each is appropriate.

for a senior

Show that you design key schemes so hot paths stay on the KV route, and that you reach for USE KEYS or the sub-document API rather than adding an index to work around a lookup you could have done directly.

for a principal

Own the split at system level: what fraction of the workload should be KV, what the query tier is sized for, and how key design decisions made in the data model determine the cluster's cost and latency profile years later.

## Two access paths to the same documents A Couchbase cluster exposes the same documents through two very different doors. The **KV (key-value) door** is the Data service. The client SDK knows the cluster topology, hashes the document key to find the node that owns it, and sends the operation to that node only. The Data service is memory-first: the working set lives in a managed cache in RAM, and a read that hits the cache never touches disk. No planning, no index, one network round trip, one node involved. The **query door** is the Query service. A statement is parsed and planned on a query node, which asks the Index service for qualifying document keys, then fetches those documents from the owning data nodes, evaluates remaining predicates, and assembles the projection. That is at minimum three services and several round trips, plus per-statement CPU on the query node. The implication is blunt: if you already know the key, going through the query door pays for machinery you do not need. ## When KV is the right call - Reading or writing a single document whose key the caller already holds — a session, a user profile, a cart, a config blob. - Any high-throughput hot path where latency budget is tight; the KV path is the lowest-latency operation the cluster offers. - Writes generally. Upserts, replaces and counters go through KV; SQL++ DML exists but the KV path is the natural one for a service mutating one document at a time. - Partial reads and writes, via the sub-document API, which fetches or mutates a path inside a document without transferring the whole thing — valuable for large documents where only one field is needed. ## When SQL++ is the right call - The key is unknown and must be found from field values. - Filtering, sorting, grouping, aggregation, joins across collections, `UNNEST`/`NEST` over embedded arrays. - Ad-hoc reporting and operational investigation, where writing a query beats writing a program. - Bulk mutations expressed as a set operation rather than a loop of KV calls. ## Key design is the lever Because KV requires a key, the way you mint keys determines how much of your traffic can take the fast path. Deterministic, composable keys — `user::{uuid}`, `order::{orderId}`, `cart::{userId}` — mean the application can construct the key from what it already has and never needs a lookup. Random keys (or keys the caller cannot reconstruct) force every read through the query layer and an index, which is a self-inflicted cost. This is one of the more consequential modelling decisions in a Couchbase application and a frequent interview probe. ## The middle ground: USE KEYS Sometimes a statement needs SQL++ machinery — a join, a projection, an aggregation — but its starting set is a known list of keys. `USE KEYS` expresses that: ```sql SELECT o.status, o.total FROM `app`.sales.orders AS o USE KEYS ["order::1001", "order::1002"]; ``` The Query service resolves those keys directly against the Data service and never consults the Index service, so this works on a collection with no indexes at all. It is also the standard escape hatch when a query fails for want of an index but the caller genuinely knows what it wants. Do not confuse it with `USE INDEX`, which is a hint naming which index the planner should use — a completely different clause. ## What a strong answer sounds like It states the architectural fact (KV is one hop to one node from memory; the query path involves three services), gives the decision rule (known key → KV), mentions key design as the way to keep the fast path available, and cites `USE KEYS` and the sub-document API as the two refinements. A weak answer treats SQL++ as the primary interface and KV as a legacy detail, which inverts how the product is actually built and how successful applications are actually written.

  • What does the USE KEYS clause do in a Couchbase SQL++ statement?
    It restricts the statement to a listed set of document keys, which the Query service resolves directly against the Data service without consulting the Index service. That makes it usable on a collection with no indexes, and it is the natural way to combine SQL++ projection or joins with a caller that already knows its keys.
  • How does key design affect how much Couchbase traffic can use the KV path?
    The KV path requires a key, so deterministic composable keys such as `order::{orderId}` let the application reconstruct a key from data it already holds and read without any lookup. Opaque or random keys force reads through the Query and Index services instead, adding hops and requiring indexes to be maintained for what should have been a direct get.
  • What does the Couchbase sub-document API add over a plain KV get?
    It reads or mutates a specific path inside a document rather than transferring the whole thing, so fetching one field from a large document moves far less data over the network and mutating one field avoids a read-modify-write of the entire document. It still runs on the KV path, so it keeps the single-hop latency profile.

saying these in an interview costs you the question

  • Treats SQL++ as the only way to read documents
  • Runs a WHERE query on the key field instead of a get
  • Thinks a KV get needs an index like a query does
  • Confuses USE KEYS with the USE INDEX hint
  • Uses opaque random keys and then indexes around them

context