How do cross-references work in a Weaviate schema, and what are their limits?
answer
- Declared alongside properties, not as a data type
- A pointer to another object's id
- Directional — the inverse is a second property
- Nothing cascades when the target is deleted
- Denormalise what you filter on
basics
~20 sA cross-reference is a declared link property pointing at another collection, added with ReferenceProperty in collections.create. It records a pointer to a target object's id. Weaviate does not enforce referential integrity, cascade deletes, or let references span tenants.
solid answer
~50 sYou declare links in the collection definition: `references=[ReferenceProperty(name="hasAuthor", target_collection="Author")]`, with `ReferenceProperty.MultiTarget` when the link may point at more than one collection. Links are set at insert time or afterwards with `collection.data.reference_add(from_uuid=..., from_property=..., to=...)`, and a reference is directional — the reverse direction is a second property you declare and maintain yourself. The limits matter more than the mechanics. A reference is a stored pointer, not a join: resolving it costs an extra lookup per referenced object, so wide fan-out is expensive. Weaviate does not enforce that the target exists, does not cascade on delete, and does not maintain cardinality — delete the target and the reference dangles. References also cannot cross tenants in a multi-tenant collection. So the practical guidance is to denormalise the few fields you filter or display onto the object, and reserve references for genuinely relational structure you traverse rarely.
code
python · 17 linesimport weaviate
import weaviate.classes.config as wvcc
client = weaviate.connect_to_local()
client.collections.create(
name="Author",
properties=[wvcc.Property(name="name", data_type=wvcc.DataType.TEXT)],
)
client.collections.create(
name="Article",
properties=[
wvcc.Property(name="title", data_type=wvcc.DataType.TEXT),
wvcc.Property(name="author_name", data_type=wvcc.DataType.TEXT),
],
references=[wvcc.ReferenceProperty(name="hasAuthor", target_collection="Author")],
)
client.close()go deeper
Know that Weaviate can link objects across collections, that the link is declared with ReferenceProperty in the collection definition, and that it stores a pointer to the target object.
Explain that references are directional pointers set at insert or via reference_add, that resolution costs a lookup per referenced object, and that no integrity or cascade behaviour is enforced.
Show the modelling judgment: denormalise the fields you filter and display, keep references for genuine traversal, and plan for dangling links and for the fact that references cannot cross tenants.
Own the data-model policy — where the source of truth for duplicated fields lives, how re-denormalisation runs when it changes, and when a relationship belongs in Weaviate at all rather than in the system of record.
## Declaring a reference Weaviate is the vector store with an actual data model, and part of that model is the ability to link objects across collections. In the v4 Python client, links are a separate argument to `collections.create`: `references=[ReferenceProperty(name="hasAuthor", target_collection="Author")]`. `ReferenceProperty` lives in `weaviate.classes.config` alongside `Property`, and `ReferenceProperty.MultiTarget(name=..., target_collections=[...])` declares a link whose target may be one of several collections. A reference property is not a data property. It has no `DataType`; its value is a pointer to another object's UUID, and Weaviate stores it as such. You can also add a reference property to an existing collection later, the same way you add a data property. ## Writing and reading links A reference can be supplied when the object is created, alongside its properties, or attached afterwards with `collection.data.reference_add(from_uuid=..., from_property="hasAuthor", to=target_uuid)`; there are corresponding replace and delete operations for maintaining them. Because a reference lives on the source object, links are directional: if you want to walk from Author to Article as well, that is a second reference property on the Author collection that you populate and keep in sync yourself. Nothing derives the inverse for you. ## What Weaviate does not do for you This is where interviews go, because the gap between "it has references" and "it is a relational database" is exactly where designs go wrong. **No referential integrity.** Nothing checks that the target UUID exists when you write the reference. You can point at an object you never created. **No cascading delete.** Delete the target object and the reference on the source remains, now pointing at nothing. Your application sees a link that resolves to no object, and cleaning that up is your job — either by sweeping references or by treating unresolvable links as expected. **No cardinality or constraints.** There is no unique constraint, no not-null, no foreign-key enforcement. Multiple references from one property are simply a list of pointers. **No tenant crossing.** In a multi-tenant collection, references cannot link objects belonging to different tenants; the isolation model is stronger than the reference model. **Cost on read.** Resolving a reference means fetching the referenced object, so a query that returns fifty results each resolving three references does many more lookups than a plain query. Deep chains multiply this. A reference traversal is not a set-based join executed by a planner; it is pointer-chasing. ## Denormalise first, reference second The pattern that works is to put onto the object whatever you need for filtering, ranking or display, even if it duplicates data. If you filter articles by author name, store `author_name` as a `TEXT` property on Article — that is one indexed property and a fast filter, versus a reference resolution that cannot be filtered nearly as cheaply. Keep the reference for the cases where you genuinely need identity and traversal: fetching the full author record on a detail view, or expressing a real graph you occasionally walk. The cost of denormalising is the familiar one: when the author is renamed you must update every article that copied the name. That is a real burden, but in a retrieval store where reads vastly outnumber writes it is usually the right side of the trade — and it is a burden you can schedule, unlike per-query traversal cost you pay forever. ## Modelling questions worth asking Before adding a reference, ask: do I ever query *through* this link, or do I only need a field from the other side? If it is the latter, copy the field. Ask: how many objects will one source link to? Fan-out of thousands makes traversal a bad idea regardless of how clean the model looks. Ask: does this relationship cross a tenant boundary? If so, references cannot express it and you need a different design, such as a shared non-tenant collection plus identifiers you resolve in the application. ## What a strong answer sounds like "References are declared pointers, not joins. I declare them with ReferenceProperty, set them with reference_add, and I keep them for real traversal. Anything I filter on gets denormalised onto the object, because Weaviate does not enforce integrity, does not cascade deletes, cannot link across tenants, and charges a lookup per resolution."
- How do you attach a reference to an object that already exists?Call `collection.data.reference_add(from_uuid=article_id, from_property="hasAuthor", to=author_id)` on the source collection's handle. There are matching replace and delete operations for maintaining links. Because the reference lives on the source object, adding the inverse direction means a second reference property on the target collection that you populate in the same code path, since Weaviate does not derive it.
- When would you model something as a separate collection with references rather than nested properties?When the linked entity has its own identity and lifecycle — you search it directly, update it independently, or many sources share it. If it is just structured detail belonging to the parent and never queried on its own, an object property or denormalised fields are simpler and avoid resolution cost. Shared, independently searchable entities justify the second collection.
- What breaks if you rely on references inside a multi-tenant collection?References cannot span tenants, so any relationship that crosses a tenant boundary simply cannot be expressed. Within a tenant they work, but the moment your model needs a link to shared, cross-tenant reference data you have to hold identifiers as ordinary properties and resolve them in the application, or keep the shared data in a non-tenant collection queried separately.
A cross-reference is a bookmark slipped into a page, not a foreign key the librarian enforces — remove the target book and the bookmark still points there.
saying these in an interview costs you the question
- Calls a cross-reference a foreign key with integrity checks
- Expects deleting the target to clean up references
- Assumes the inverse direction is created automatically
- Thinks reference resolution is a planner-optimised join
- Models heavy filtering through references instead of denormalising