skip to content

Tables and the Lance Format

You connect to a directory or S3 bucket and create typed tables backed by Lance files, which gives you versioning and time travel for free. Understanding that storage model explains most of LanceDB's behaviour.

on this pageshow

questions

5

In LanceDB, how do you create a table and where can its schema come from?

level: juniorimportance: must knowfreq 72%

answer

  1. two ways: from data or from schema
  2. Arrow types under the hood
  3. empty table needs an explicit schema
  4. pydantic LanceModel with Vector(dim)
  5. fixed_size_list float32 for vectors

basics

~20 s

db.create_table(name, data=...) infers an Arrow schema from a pandas DataFrame, PyArrow table or list of dicts. Passing schema= with a PyArrow schema or a lancedb.pydantic LanceModel defines it explicitly and allows creating an empty table.

solid answer

~40 s

After `db = lancedb.connect("./data")`, the normal path is `db.create_table("docs", data=df)`, where `data` can be a pandas DataFrame, a PyArrow Table or RecordBatch iterator, or a plain list of dicts. LanceDB converts it to Arrow and stores the inferred Arrow schema with the table, so the table is strongly typed from the first write. If you want the schema fixed up front — or an empty table you will populate later — pass `schema=` instead: either a `pyarrow.schema([...])` or a `LanceModel` subclass from `lancedb.pydantic`, where a vector column is declared as `Vector(768)`. The vector column must be a fixed-size list of float32, which is what `Vector(n)` and `pa.list_(pa.float32(), n)` produce. `mode="create"` is the default and errors if the table exists; use `exist_ok=True` or `mode="overwrite"` deliberately.

go deeper

for a junior

Be able to write the three lines from memory: connect, create_table with a DataFrame, then add rows. Know that the schema is fixed at creation and that a vector column is a fixed-size list of float32.

for a middle

Explain the difference between inferring the schema from data and declaring it with a PyArrow schema or LanceModel, why an empty table needs the latter, and what mode and exist_ok do when the table already exists.

for a senior

Show the ingestion discipline: explicit schemas in code, idempotent creation, streaming RecordBatch inputs for large loads, and handling of bad or wrong-length vectors before they reach storage.

for a principal

Own the schema as a contract between the embedding pipeline and the retrieval service — where it is declared, how it is versioned and evolved with add or alter columns, and what a dimension change costs across every table and job that depends on it.

## What a LanceDB table actually is A LanceDB table is a Lance dataset: a directory on disk or in object storage holding columnar data files, plus manifests describing the schema and the current version. There is no server and no schemaless document store — the table carries a strict Apache Arrow schema, and every write is validated against it. That is why table creation is really schema creation: whatever you hand to `create_table` determines the column types the table will enforce for its whole life. ## Creating from data (schema inference) The common form is: `tbl = db.create_table("docs", data=df)` `data` accepts a pandas DataFrame, a PyArrow `Table` or `RecordBatch`, a Polars DataFrame, a list of Python dicts, or an iterator/generator of RecordBatches for streaming large ingests without materialising everything in memory. LanceDB converts the input to Arrow and persists the resulting schema. Inference is convenient but it is the source of most beginner surprises: a column of Python `None`s infers as null type, integers can land as `int64` when you wanted `int32`, and — most importantly — a column of Python lists of floats infers as a *variable-length* list, not the fixed-size list a vector column needs. ## Creating from an explicit schema Two explicit forms exist, and both let you create a table with zero rows: 1. **PyArrow schema.** `schema = pa.schema([pa.field("vector", pa.list_(pa.float32(), 768)), pa.field("text", pa.string())])` then `db.create_table("docs", schema=schema)`. 2. **Pydantic model.** `from lancedb.pydantic import LanceModel, Vector`, then a class with `vector: Vector(768)` and ordinary typed fields, passed as `db.create_table("docs", schema=Doc)`. The same model can be used to validate and to read rows back as typed objects. An empty table is genuinely useful: you create it at deploy time with the exact schema, then stream data in with `add()` later, instead of depending on whatever the first batch happened to look like. ## Why the vector column type matters Vector search and index building require a fixed-size list of float32 (or float16) — the dimension has to be known and identical for every row. `Vector(768)` and `pa.list_(pa.float32(), 768)` both produce that. If inference gives you `list<double>` of variable length, the column looks fine in `to_pandas()` but is not a usable vector column, and you discover it only when you try to search or build an index. NumPy arrays infer better than Python lists, but the reliable fix is to state the schema. `create_table` also takes `on_bad_vectors` with a `fill_value`, controlling what happens to rows whose vector contains NaN or has the wrong length — erroring, dropping or filling them — which matters when embeddings come from a flaky upstream job. ## mode and exist_ok `mode="create"` is the default and raises if the name is taken; `exist_ok=True` makes that a no-op returning the existing table, which is what you want in idempotent startup code. `mode="overwrite"` replaces the contents, including the schema — and because Lance is versioned, the old data still exists as an earlier version rather than being erased immediately. ## Living with the schema afterwards The schema is not frozen forever. `table.schema` shows the current Arrow schema, and `add_columns`, `alter_columns` and `drop_columns` evolve it; because Lance is columnar, adding a column does not rewrite the existing columns' data. But `add()` will reject batches that do not match the table's schema, so the practical discipline is to define the schema once, explicitly, and evolve it with the dedicated methods rather than by re-creating the table. ## Practical guidance For a throwaway notebook, `create_table(name, data=df)` is fine. For anything that runs more than once, declare a `LanceModel` or a PyArrow schema, create with `exist_ok=True`, and let inference nowhere near your vector column. It costs three lines and removes an entire class of bug where the table types drift with the shape of the first batch you happened to load.

  • How would you create a table now but load the rows later that night?
    Create it empty with an explicit schema — a PyArrow schema or a `LanceModel` — via `db.create_table("docs", schema=Doc, exist_ok=True)`. Passing `data` is optional; without it the table exists with zero rows and a fixed schema. The nightly job then calls `tbl.add(batch)`, and every batch is validated against that schema instead of the schema being decided by whichever batch arrives first.
  • What breaks if the vector column ends up as a variable-length list?
    Everything downstream of storage. Vector search and index building require a fixed-size list of float32 so the dimension is known and uniform, so a variable-length list column is not treated as a usable vector column — the data reads back fine in pandas, which is why the problem surfaces late. Declaring `Vector(768)` or `pa.list_(pa.float32(), 768)` avoids it.
  • What does mode="overwrite" do to the data that was already there?
    It replaces the table's contents and schema with the new data, but Lance is versioned, so the previous state remains as an earlier version of the same table rather than being deleted on the spot. The space is reclaimed only when old versions are pruned. Until then you can still read what was overwritten, which makes `overwrite` recoverable but not free.

saying these in an interview costs you the question

  • Claiming LanceDB tables are schemaless like a document store
  • Thinking a list of Python floats gives a proper vector column
  • Believing an empty table can be created without a schema
  • Assuming create_table on an existing name silently succeeds
  • Thinking the schema can never change after creation

context

open as a page

How does LanceDB versioning work, and how do you read an older table version?

level: middleimportance: must knowfreq 65%

basics

~20 s

Every write to a LanceDB table — append, update, delete, overwrite, index build — commits a new immutable version instead of mutating data in place. table.list_versions() enumerates them, table.checkout(n) pins the table object to version n for reading, and table.checkout_latest() returns to the newest.

open as a page

What does lancedb.connect() do for a local path versus an s3:// URI?

level: middleimportance: should knowfreq 60%

basics

~20 s

lancedb.connect(uri) opens a database rooted at a directory — a local path it creates if missing, or an object-storage prefix such as s3://bucket/prefix reached with storage_options for region and credentials. No server or connection pool is involved; the calling process is the database.

open as a page

In LanceDB, what does table.merge_insert() do that table.add() cannot?

level: middleimportance: should knowfreq 48%

basics

~20 s

merge_insert() matches incoming rows against existing ones on a key column and applies update-or-insert semantics in a single commit. add() only appends, so re-running it with the same records silently creates duplicates — LanceDB tables have no primary-key constraint to stop it.

open as a page

Why does a LanceDB table keep growing on disk after deletes, and what fixes it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Deletes commit a new version recording removed rows rather than rewriting data, and every earlier version still references the original files, so nothing is reclaimed. table.optimize() compacts fragments and prunes versions older than an age you pass, which is what actually returns the space.

open as a page