skip to content

How do you control which Weaviate properties a text2vec module embeds?

level: middleimportance: should knowfreq 55%

answer

  1. defaults are inclusive, not selective
  2. identifiers pollute the embedded string
  3. two flags: one drops value, one drops label
  4. non-text properties never vectorize
  5. no content means no vector at all

basics

~10 s

By default a text2vec module embeds every text property, including the property names. Set skip_vectorization=True on a property to exclude it entirely, and vectorize_property_name=False to embed the value without its name.

solid answer

~50 s

A `text2vec-*` module builds one string per object out of that object's vectorizable text properties and sends it to the model. The defaults are inclusive: all text-typed properties participate, the property *name* is included alongside its value, and the collection name is prepended unless you turn that off. Two per-property flags in the v4 client narrow it — `Property(..., skip_vectorization=True)` drops the property from the embedded text completely, and `vectorize_property_name=False)` keeps the value but drops the label. Numeric, boolean and date properties are not vectorized at all. This matters because identifiers, URLs, SKUs and status codes are pure noise in an embedding: leaving them in dilutes the semantic signal, and in the worst case an object whose only vectorizable content is an opaque id gets a vector that means nothing. On recent clients you can also target properties positively with `source_properties`.

code

python · 18 lines
python
from weaviate.classes.config import Configure, Property, DataType

client.collections.create(
    name="Product",
    properties=[
        Property(name="sku", data_type=DataType.TEXT, skip_vectorization=True),
        Property(name="source_url", data_type=DataType.TEXT, skip_vectorization=True),
        Property(
            name="description",
            data_type=DataType.TEXT,
            vectorize_property_name=False,
        ),
        Property(name="price", data_type=DataType.NUMBER),
    ],
    vectorizer_config=Configure.Vectorizer.text2vec_openai(
        vectorize_collection_name=False,
    ),
)

go deeper

for a junior

Know that a text2vec module embeds all text properties by default and that you can exclude one with skip_vectorization on the Property. Know numeric fields are not embedded.

for a middle

Explain how the embedded string is assembled — property names plus values, plus the collection name — and what each of the two flags removes. Say why identifiers should be skipped.

for a senior

Show the diagnostic instinct: read an object back with include_vector to confirm it has one, and recognise the 'fetchable but never retrieved' symptom as an object with no vectorizable content.

for a principal

Argue for the positive form. Prefer stating the inclusion set explicitly so a future property added by someone else cannot silently change what every object embeds, and treat the vectorizable set as a documented contract.

## The default is inclusive, and that is the trap When Weaviate vectorizes an object it does not embed "the document" — it embeds a string it constructs from the object. For a `text2vec-*` module the construction is deliberately naive: take the vectorizable properties, pair each property name with its value, and concatenate. The collection name is included too by default. Nothing about this is tuned for your data; it is a reasonable default that assumes your properties are prose. The consequence is that a collection whose schema mixes prose with bookkeeping fields embeds the bookkeeping. A `sku` of `"BX-99417-K"`, a `source_url`, an `internal_status` — each of these contributes tokens to the vector while contributing nothing to meaning. Individually harmless; on an object whose real text is one short title, they can be a large fraction of the embedded string, and similarity starts reflecting id-string surface patterns instead of topic. ## The two flags **`skip_vectorization=True`** removes the property from the embedded text entirely. The property is still stored, still returned, still filterable — it simply does not reach the model. This is what you want for identifiers, foreign keys, URLs, timestamps stored as text, and any field a user would never phrase a query about. **`vectorize_property_name=False`** is the softer knob: the value is embedded but its label is not. Useful when the property name is an internal noun that would drag the vector around (`raw_body_v2`, `field_17`) while the value is genuinely the content you want searchable. Leave it on when the name carries real semantics — a property called `ingredients` genuinely tells the model what the following text is. There is a collection-level analogue as well: `vectorize_collection_name` on the vectorizer config controls whether the collection's own name is folded into every object's text. Turning it off is common when the collection name is an internal label rather than a description of the content. ## The positive form Rather than subtracting property by property, recent clients let you state the inclusion set directly with `source_properties` on the vector configuration — "embed exactly these". For a wide schema this is both shorter and more robust: adding a new bookkeeping property later cannot silently pollute the embedding, because it was never in the list. Subtractive flags have the opposite failure mode — every new property is opted in by default, and the person adding it may not think about vectors at all. ## Types that never vectorize Non-text properties — `INT`, `NUMBER`, `BOOLEAN`, `DATE`, `BLOB` for `text2vec-*` — are not embedded. This surprises people who expect a price or a rating to influence semantic similarity. It will not; numeric relevance belongs in filters or in scoring, not in the embedding. (Image-aware modules such as `multi2vec-clip` are the exception in the other direction: they read designated image fields, which text modules ignore.) ## The pathological case If you skip enough properties, an object can end up with *no* vectorizable content. Weaviate then stores the object without a vector, and it becomes invisible to vector search while remaining perfectly visible to filters and to direct fetches. This produces one of the more confusing bug reports on Weaviate: "the object is in the collection, I can fetch it by id, but semantic search never returns it." The fix is to check what the schema actually feeds the vectorizer. ## How to verify Do not reason about it from the schema alone. Insert a representative object, read it back with `include_vector=True`, and confirm a vector exists and has the expected dimension. Then run two near-text queries — one phrased against the content you meant to embed, one against a field you meant to exclude — and see which one retrieves the object. If a query about an internal SKU retrieves it strongly, your skip flags are not doing what you think. ## Practical policy Treat the vectorizable set as an explicit design decision made when the collection is created, and write it down. Embed the fields a user would phrase a question about; skip everything that exists for joins, auditing, or plumbing. Because the choice is baked into the collection at creation and existing vectors are not regenerated when you change the schema, getting it wrong means a re-import, not a config tweak.

  • What happens if every text property on an object is marked skip_vectorization?
    The object has no vectorizable content, so Weaviate stores it without a vector. It stays fetchable by id and matchable by filters, but vector search can never return it — there is nothing in the ANN index to match. It is a silent failure: no error at insert, just an object that never appears in semantic results.
  • Why does Weaviate include the property name in the embedded text by default?
    Because a bare value is often ambiguous and the label supplies context — embedding "ingredients: flour, butter" carries more meaning than "flour, butter" alone, especially for short values. It is a sensible default for descriptive schemas and a liability for internal field names, which is exactly why the per-property override exists.
  • You add a new text property to an existing collection. Do existing objects get re-embedded to include it?
    No. Vectors are produced once at import; adding a property does not regenerate them. New objects will embed the new property while old ones will not, leaving the collection in two subtly different embedding regimes. If the new property matters for search, re-import the affected objects.

saying these in an interview costs you the question

  • Assuming only properties you explicitly list get embedded
  • Thinking numeric properties like price influence semantic similarity
  • Believing skip_vectorization also hides the property from filters and results
  • Not realising the property name itself is embedded by default
  • Expecting a schema change to re-embed existing objects

context