skip to content

Weaviate

An open-source vector database with an actual schema: typed collections with properties, per-collection index configuration, and modules that can embed your data for you. Interviews often use it as the contrast to a bare vector index, since the schema decides your filtering and multi-tenancy options later.

on this pageshow

explore

questions

23

In Weaviate, what is a collection and what does collections.create define?

level: juniorimportance: must knowfreq 70%

answer

  1. Think tables, not a bare vector store
  2. Named container with typed properties
  3. Property plus DataType enum
  4. One vector index per collection definition
  5. Also pins replication and tenancy

basics

~20 s

A collection is Weaviate's typed container for objects of one kind, much like a table. collections.create names it and declares its properties with data types, plus per-collection settings such as the vector index, replication and multi-tenancy.

solid answer

~50 s

A Weaviate collection holds objects that share one definition: the same properties, the same vector index, the same replication and tenancy settings. Every object is a UUID plus typed properties plus one or more vectors, so the collection definition tells the server both how to index the payload for filtering and keyword search and how to index the vector for similarity search. With the v4 Python client you call `client.collections.create(name="Article", properties=[Property(name="title", data_type=DataType.TEXT), ...])`. `Property` and the `DataType` enum (`TEXT`, `INT`, `NUMBER`, `BOOL`, `DATE`, `UUID`, `GEO_COORDINATES`, `BLOB`, `OBJECT`, and `_ARRAY` variants) come from `weaviate.classes.config`. The same call also takes the vector index config, inverted-index config, references, replication and multi-tenancy config. Afterwards you work through a handle: `client.collections.get("Article")`, with `collections.exists`, `collections.list_all` and `collections.delete` for lifecycle. This typed schema is the main thing that separates Weaviate from a bare vector index.

code

python · 16 lines
python
import weaviate
import weaviate.classes.config as wvcc

client = weaviate.connect_to_local()
client.collections.create(
    name="Article",
    properties=[
        wvcc.Property(name="title", data_type=wvcc.DataType.TEXT),
        wvcc.Property(name="word_count", data_type=wvcc.DataType.INT),
        wvcc.Property(name="published_at", data_type=wvcc.DataType.DATE),
    ],
)

articles = client.collections.get("Article")
articles.data.insert({"title": "Vector search", "word_count": 1200})
client.close()

go deeper

for a junior

Be able to say plainly that a collection is a named, typed container of objects, and that collections.create declares its properties with data types. Know the create, get, exists and delete calls.

for a middle

Explain what the data type buys you in the inverted index — range filters on numbers and dates, tokenised matching on text — and that the create call also pins the vector index and tenancy configuration.

for a senior

Show that you treat the collection as the unit of blast radius: deletion is total, several settings are fixed at creation, and definitions belong in version-controlled setup code rather than being created ad hoc by application paths.

for a principal

Own the modelling call — how many collections a domain needs, what belongs as a property versus a reference versus a separate collection, and how that shape constrains filtering, tenancy and future re-import cost.

## What a collection is Weaviate's top-level container is a **collection** (the REST schema and older material call it a *class*). A collection holds objects of one kind and every object in it shares one definition: the same property declarations, the same vector index and distance metric, the same inverted-index settings, the same replication factor and the same multi-tenancy setting. That makes it closer to a relational table than to a bare vector index — the schema is a server-side artifact you can read back, not something your application keeps in its head. An object inside a collection is three things at once: a UUID, a set of typed properties (the JSON payload), and one or more vectors. Weaviate stores the properties in an inverted index so they can be filtered and keyword-searched, and the vector in a vector index so it can be searched by similarity. The collection definition is what configures both halves. ## Creating one with the v4 Python client `client.collections.create(...)` is the single entry point. The minimum is a name; in practice you pass `properties=[...]`, each one a `Property(name=..., data_type=...)` from `weaviate.classes.config`. Two naming rules bite newcomers: Weaviate capitalises the first letter of a collection name, so `article` and `Article` are the same collection, and property names must start with a lowercase letter and stay GraphQL-safe. ## Properties and data types The `DataType` enum covers `TEXT`, `INT`, `NUMBER`, `BOOL`, `DATE`, `UUID`, `GEO_COORDINATES`, `BLOB`, `OBJECT` and the corresponding `_ARRAY` variants. The type is not decoration: it decides what the inverted index can do with the value. A `TEXT` property is tokenised and can be BM25-searched; an `INT`, `NUMBER` or `DATE` property supports range comparisons; a `UUID` is stored compactly. Getting `INT` versus `NUMBER` wrong, or storing a timestamp as `TEXT`, quietly costs you range filtering later. `Property` also carries indexing switches — notably `index_filterable` and `index_searchable`, and `tokenization` for text — that let you skip index building for properties you never query on, trading query ability for faster imports and less disk. Vectorisation flags on a property control whether its text feeds the embedding, which is the vectoriser module's concern rather than the schema's. ## What else the definition holds Beyond properties, `collections.create` accepts: the vector index configuration (`Configure.VectorIndex.hnsw()`, `.flat()` or `.dynamic()`) and its distance metric; inverted-index configuration such as BM25 parameters, stopwords and whether timestamps and null states are indexed; `references=[ReferenceProperty(...)]` for links to other collections; `replication_config=Configure.replication(factor=...)`; and `multi_tenancy_config=Configure.multi_tenancy(...)`. All of these are per-collection, which is exactly why the create call carries so much weight — several of these choices cannot be changed afterwards. ## Lifecycle `client.collections.exists("Article")` checks existence, `client.collections.list_all()` enumerates definitions, `client.collections.get("Article")` returns the handle you use for data and queries, and `client.collections.delete("Article")` destroys the definition *and* all of its objects with no undo. On the handle, `collection.config.get()` reads the current definition back, `collection.config.add_property(...)` appends a property, and `collection.config.update(...)` changes the mutable settings. ## Why interviewers start here Because the collection is the unit of configuration, it is also the unit of blast radius. Filtering, hybrid search, tenant isolation, memory footprint and re-import cost are all decided by what you wrote in `collections.create`. A candidate who can describe a collection as "a typed table whose definition also pins the vector index and tenancy model" is set up for every harder Weaviate question that follows.

  • How do you read an existing collection's definition back from the server?
    Take the handle with `client.collections.get("Article")` and call `collection.config.get()`. It returns the full definition — properties and their data types and index flags, the vector index configuration, inverted-index settings, replication and multi-tenancy config. It is the first thing to check when a filter or keyword search behaves unexpectedly, because it shows what the server actually stored rather than what you meant to declare.
  • What happens to the objects when you delete a collection?
    They go with it. `client.collections.delete("Article")` removes the definition and every object, vector and index belonging to it, and there is no undo or soft delete. Recovery means restoring from a backup or re-importing from your source of truth, so in shared environments deletion is worth gating behind a deliberate step rather than leaving it in an idempotent setup script.
  • Why does the data type matter if you are only doing vector search?
    Because pure vector search is rarely the whole query. The data type decides what the inverted index supports: only numeric and date properties give you range comparisons, only text properties are tokenised for keyword matching. If you store a price or a timestamp as text, you can match it exactly but never filter a range without re-importing under a corrected definition.

A collection is a table definition that happens to own an index for meaning as well as indexes for values.

saying these in an interview costs you the question

  • Thinks Weaviate is schemaless like a plain vector index
  • Says collections are created per object insert batch
  • Believes data types are cosmetic and do not affect indexing
  • Assumes deleting a collection keeps the objects somewhere
  • Confuses a collection with a single vector index shard

context

open as a page

In Weaviate's hybrid() query, what does the alpha parameter control?

level: juniorimportance: must knowfreq 78%

basics

~20 s

alpha weights the two halves of a Weaviate hybrid search: alpha=0 is pure BM25 keyword search, alpha=1 is pure vector search, and values in between blend them. The v4 Python client's default of 0.7 leans toward the vector side.

open as a page

In Weaviate, when do you use near_text, near_vector, or near_object?

level: juniorimportance: must knowfreq 74%

basics

~10 s

near_text sends a raw string and lets the collection's configured vectorizer embed it server-side. near_vector takes an embedding you computed yourself. near_object finds the neighbours of an object already stored, referenced by its UUID.

open as a page

In Weaviate, what does setting a text2vec vectorizer on a collection do?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A configured vectorizer makes Weaviate call the embedding model itself: you insert plain properties and the server produces the vector, and a text query is embedded server-side too. Without one you must supply every vector yourself.

open as a page

Why is relying on Weaviate's auto-schema risky for a production collection?

level: middleimportance: must knowfreq 62%

basics

~20 s

Auto-schema is on by default and invents property definitions from whatever object arrives first, so types are guessed from one sample, stray keys become permanent indexed properties, and the resulting schema differs between environments. Property types cannot be changed afterwards.

open as a page

How do Weaviate's HybridFusion.RANKED and RELATIVE_SCORE differ in hybrid results?

level: middleimportance: must knowfreq 62%

basics

~20 s

HybridFusion.RANKED merges the keyword and vector lists using only each object's position, so score magnitude is discarded. HybridFusion.RELATIVE_SCORE normalizes each list's raw scores into a comparable range first, so how far ahead a hit is changes the final order.

open as a page

In Weaviate, how do the distance and certainty thresholds on near_text differ?

level: middleimportance: must knowfreq 62%

basics

~20 s

distance sets a maximum: results further than the value are dropped, and lower is closer. certainty sets a minimum on a normalised 0-to-1 similarity where higher is closer, and it is only defined when the collection uses the cosine metric. Set one, not both.

open as a page

How do you build and combine where filters in the Weaviate v4 Python client?

level: middleimportance: must knowfreq 68%

basics

~10 s

Build conditions with the Filter class — Filter.by_property("category").equal("news") — and combine them with the & and | operators or with Filter.all_of([...]) and Filter.any_of([...]). Pass the result as the filters argument to any query method.

open as a page

In Weaviate, when do you configure no vectorizer and supply vectors yourself?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Configure no vectorizer when you must control the embedding model yourself — pinning a version, reusing embeddings computed elsewhere, or keeping an external provider out of the query path. You then supply a vector on every insert and on every search.

open as a page

How do cross-references work in a Weaviate schema, and what are their limits?

level: middleimportance: should knowfreq 36%

basics

~20 s

A cross-reference is a declared link property pointing at another collection, added with ReferenceProperty in collections.create. It records a pointer to a target object's id. Weaviate does not enforce referential integrity, cascade deletes, or let references span tenants.

open as a page

When would you configure a Weaviate collection with a flat vector index instead of HNSW?

level: middleimportance: should knowfreq 48%

basics

~20 s

Choose flat for small collections, typically a few thousand objects, and for per-tenant shards. It scans every vector, so recall is exact and there is no graph to build or hold in memory, but latency grows linearly as the collection grows.

open as a page

In Weaviate's hybrid(), what does query_properties=['title^2','body'] change?

level: middleimportance: should knowfreq 45%

basics

~20 s

query_properties configures only the keyword half of a Weaviate hybrid query: it restricts BM25 matching to the listed properties and the ^ suffix boosts a property's weight. The vector half is unaffected — it always uses the object's stored vector.

open as a page

In Weaviate, what does the auto_limit (autocut) parameter do to a result list?

level: middleimportance: should knowfreq 38%

basics

~20 s

auto_limit cuts the ranked results where similarity drops off sharply. auto_limit=1 keeps only the first cluster of closely-scoring objects, auto_limit=2 keeps the first two clusters, and so on — so the number of results varies with each query instead of being fixed.

open as a page

How do you control which Weaviate properties a text2vec module embeds?

level: middleimportance: should knowfreq 55%

basics

~10 s

By default a text2vec module embeds every text property, including the property names. Set skip_vectorization=True on a property to exclude it entirely, and vectorize_property_name=False to embed the value without its name.

open as a page

Which Weaviate collection settings can you change after creation, and which are fixed?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Search-time and maintenance settings are mutable through collection.config.update, and properties can be added. The structural choices are fixed: vector index type, distance metric, whether multi-tenancy is enabled, and any existing property's data type. Changing those means a new collection and a re-import.

open as a page

A Weaviate hybrid query ranks an exact keyword match below a fuzzy one — how do you debug it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Ask Weaviate to explain itself: request score and explain_score metadata on the hybrid call to see what each half contributed per object. Then run the query at alpha=0 and alpha=1 separately to find out which half is misbehaving before changing any tuning parameter.

open as a page

Why does a highly selective filter change Weaviate's vector search cost and recall?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Weaviate resolves the filter against its inverted index first and searches only the matching set, so limit is honoured against filtered objects rather than being trimmed afterwards. When that set is small, the engine switches to an exact scan of it, which changes latency and eliminates approximation error.

open as a page

What problem do Weaviate's named vectors solve, and what do they cost?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Named vectors let one object carry several independent embeddings — different properties, different models, or different modalities — each with its own configuration, chosen at query time by name. The cost is one stored vector and one ANN index per name.

open as a page

How do you decide between Weaviate tenants, one collection per customer, and a tenant property?

level: principalimportance: should knowfreq 55%

basics

~20 s

Use built-in multi-tenancy for many isolated customers: each tenant gets its own shard and index, cheap deletion, and an activity status that offloads idle tenants. A collection per customer does not scale in schema metadata; a tenant filter shares one index and gives no isolation.

open as a page

How do you choose a Weaviate vectorization strategy for a production workload?

level: principalimportance: should knowfreq 38%

basics

~20 s

Decide who owns the embedding model. A hosted module is fastest to ship but puts a third party in the query path; a self-hosted inference container keeps data and cost in-house but adds an operated service; self-provided vectors give full model control at the price of owning re-embedding.

open as a page

In Weaviate, what does a generative-* module add to a collection?

level: middleimportance: nice to knowfreq 40%

basics

~20 s

A generative module attaches an LLM to a collection so retrieval and generation happen in one round-trip: the search runs, the retrieved objects are fed into your prompt server-side, and the model's text comes back with the results.

open as a page

How does group_by change the result shape of a Weaviate near_text query?

level: seniorimportance: nice to knowfreq 28%

basics

~10 s

Passing group_by=GroupBy(prop=..., number_of_groups=..., objects_per_group=...) folds the ranked hits into groups keyed by a property value. The response then exposes groups alongside the flat object list, and each object carries the group it belongs to.

open as a page

Weaviate sets hybrid alpha per query, not per collection — how do you manage that?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Because alpha and fusion_type are request arguments, the ranking policy lives in your application, not in Weaviate. Centralize it in one retrieval layer, drive it from configuration, route by query class, and validate changes against a labelled set instead of tuning per call site.

open as a page