skip to content

Buckets, Collections & N1QL

How Couchbase organises data and lets you query it with something that looks like SQL over JSON. The N1QL/SQL++ angle is what interviewers use to contrast it with MongoDB's operator-based API.

on this pageshow

questions

6

In Couchbase 7.x, how do buckets, scopes and collections relate to each other?

level: juniorimportance: must knowfreq 70%

answer

  1. Three nesting levels, one of them physical
  2. Only the outermost level owns a quota
  3. Documents live in the innermost container
  4. Path written bucket dot scope dot collection
  5. Every bucket ships with _default._default

basics

~10 s

A Couchbase bucket is the top-level container that owns the memory quota, replica count and persistence settings. Inside it, scopes group collections, and collections hold the documents. The full keyspace path is bucket.scope.collection.

solid answer

~40 s

Couchbase has three nesting levels. The **bucket** is the physical container: it owns the RAM quota, the number of replicas, the ejection policy and whether data is persisted at all. Inside a bucket, a **scope** is a logical namespace, and inside a scope a **collection** is the keyspace that actually stores documents. Every bucket ships with a `_default` scope containing a `_default` collection, which is where documents written by clients that name only a bucket land. In SQL++ you address a collection by the three-part path `bucket`.`scope`.`collection`, backtick-quoting any name that is not a plain identifier. Document keys are unique per collection, not per bucket, so two collections may both hold a document with key `1234`. Indexes and RBAC roles can be granted at the scope or collection level.

code

sql · 2 lines
sql
CREATE SCOPE `app`.sales;
CREATE COLLECTION `app`.sales.orders;

go deeper

for a junior

Be able to name the three levels in order and write a three-part keyspace path in a query. Know that a bucket is what an administrator creates and a collection is where documents actually live.

for a middle

Explain which settings belong to the bucket (quota, replicas, ejection, persistence) and why scopes and collections carry none of them. Be ready to describe the _default scope and collection and when data ends up there.

for a senior

Show that you can design a collection layout for a real application: which document types get their own collection, how RBAC is granted per scope or collection, and how indexes narrow once a query is confined to a collection.

for a principal

Own the trade-off between bucket count and collection count across a shared cluster: memory quota allocation, tenant isolation, blast radius of a bucket-level operation, and what a bucket-per-tenant design costs you at scale.

## The three levels Couchbase organises data in a strict three-level hierarchy: **bucket → scope → collection → document**. Only the bucket is a physical, cluster-configured object; scopes and collections are logical containers created inside it. A **bucket** is what you create in the cluster UI or with the CLI, and it is the unit at which resources are configured. It owns the memory quota reserved on every data node, the number of replica copies, the ejection policy, and — for the persistent bucket type — the on-disk files. The two bucket types you will meet in practice are the persistent *Couchbase* bucket, which caches in memory and writes to disk, and the *Ephemeral* bucket, which keeps data in memory only. Because the quota is per bucket and every bucket costs metadata overhead on every node, clusters run a small number of buckets, not one per entity type. A **scope** is a named namespace inside a bucket. It carries no quota and no replication setting of its own; it exists so that collections can be grouped and permissions granted wholesale. Typical uses are one scope per application, per tenant, or per microservice sharing a bucket. A **collection** is the keyspace documents actually live in. It is the closest thing Couchbase has to a relational table: you write documents into a collection, index a collection, and query a collection. Document keys are unique *within a collection*, so `order::1` in `sales.orders` and `order::1` in `archive.orders` are two different documents. ## What lives at which level A useful mental checklist for interviews: - Memory quota, replica count, ejection policy, persistence, bucket-level cross-cluster replication: **bucket**. - Logical grouping and coarse access control: **scope**. - Documents, indexes, key uniqueness, fine-grained RBAC: **collection**. RBAC in 7.x can be scoped down: a user can be given read access to one collection rather than a whole bucket, which is what makes multi-tenant or multi-service bucket sharing practical. ## Addressing a collection in SQL++ Queries name the full path: ```sql SELECT a.name FROM `travel-sample`.inventory.airline AS a WHERE a.country = "United States"; ``` The backticks around `travel-sample` are required because the name contains a hyphen; plain identifiers need none. You can also set a default keyspace for a session so short names resolve, but production statements normally spell the path out. Collections and scopes are created with DDL — `CREATE SCOPE` and `CREATE COLLECTION` — or through the UI and CLI. ## The `_default` scope and collection Every bucket is created with a `_default` scope holding a `_default` collection. This is the compatibility path: an older client, or any code that connects to a bucket without naming a scope and collection, reads and writes `_default._default`. In SQL++ a two-part path such as `` `app`.orders `` is interpreted as bucket plus collection in the default scope, and a bare bucket name refers to `_default._default`. ## Why collections were added Before 7.0, a bucket was the only container, so teams either mixed every document type into one bucket and disambiguated with a `type` field and a key prefix, or created a bucket per type and hit the practical ceiling on bucket count and the memory overhead each bucket carries. Collections removed that trade-off: one bucket can now hold many collections that are cheap to create, separately indexable, and separately permissioned. Migrating an old design usually means keeping the key-prefix convention while moving each document type into its own collection, then dropping the `type` predicate from queries and the corresponding leading key from indexes — a query restricted to a collection no longer needs to filter by type at all. ## What interviewers are checking They want to know that you can place the resource knobs at the right level (quota and replicas are bucket-wide, so per-collection tuning does not exist), that you know keys are unique per collection, and that you can write a correctly quoted three-part keyspace path. A candidate who says a scope has its own memory quota, or that document keys must be unique across the whole bucket, is describing a system Couchbase is not.

  • Where do documents go when a client connects to a Couchbase bucket without naming a scope or collection?
    They land in the `_default` collection inside the `_default` scope, which every bucket is created with. That is the compatibility path for pre-7.0 clients and for code that only configures a bucket name. In SQL++, a bare bucket name resolves to that same `_default._default` keyspace.
  • Why do Couchbase clusters run few buckets but many collections?
    Each bucket is a physical container with its own memory quota, replica configuration and per-node metadata overhead, so buckets are expensive and their practical count is limited. Collections are logical, cheap to create, and separately indexable and permissioned, so they are the right unit for splitting document types or tenants.
  • Can two Couchbase collections in the same bucket hold documents with the same key?
    Yes. Key uniqueness is scoped to the collection, so `order::1` can exist independently in `sales.orders` and in `archive.orders`. That is a real behavioural change from pre-7.0 designs, where a single bucket-wide keyspace forced teams to prefix keys with a type marker to avoid collisions.

Think of the bucket as the building with the power and water hookups, scopes as the floors, and collections as the individual rooms where the furniture actually sits.

saying these in an interview costs you the question

  • Says each scope has its own memory quota
  • Claims document keys must be unique bucket-wide
  • Treats collections as just a key prefix convention
  • Configures replica count per collection
  • Forgets backticks around bucket names containing hyphens

context

open as a page

Why does a Couchbase SQL++ query fail with "No index available" when the documents exist?

level: middleimportance: must knowfreq 68%

basics

~20 s

Couchbase's Query service reaches documents only through an index. With no primary index and no secondary GSI whose leading key matches a predicate, there is no access path, so the statement is rejected rather than scanning the whole collection.

open as a page

When should you read a Couchbase document with the KV API instead of a SQL++ query?

level: middleimportance: should knowfreq 52%

basics

~20 s

Whenever the document key is known. A KV read hashes the key to the node that owns it and is served from that node's managed cache, with no index lookup, no query planning and no hop through the Query service.

open as a page

What does UNNEST do in a Couchbase SQL++ query over a document's array field?

level: middleimportance: should knowfreq 62%

basics

~10 s

UNNEST flattens an array embedded in a document into one result row per element, joined back to its parent document, so array elements can be filtered, grouped and projected exactly like ordinary rows.

open as a page

How do you tell whether a Couchbase GSI fully covers a SQL++ query?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A GSI covers a query when every field the query references appears among the index keys, so the Query service answers from the index alone. Run EXPLAIN: a covered plan shows an index scan with a covers list and no Fetch operator.

open as a page

How would you place Couchbase's data, query, index and search services across nodes for a mixed workload?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Couchbase lets each node run a chosen subset of services, so isolate data, query, index and search onto separate node groups once they contend, and size each group to its own bottleneck: RAM for data, CPU for query, memory plus disk for index.

open as a page