What does lancedb.connect() do for a local path versus an s3:// URI?
answer
- the URI picks storage, not a protocol
- no server process is started
- credentials go in a keyword argument
- object storage changes latency, not the API
- open tables can lag other writers
basics
~20 slancedb.connect(uri) opens a database rooted at a directory — a local path it creates if missing, or an object-storage prefix such as s3://bucket/prefix reached with storage_options for region and credentials. No server or connection pool is involved; the calling process is the database.
solid answer
~50 s`lancedb.connect("./lance")` treats the path as the database root and creates it if it does not exist; each table becomes a `<name>.lance` directory of data files and manifests underneath. Point it at `s3://bucket/prefix` (or `gs://`, `az://`) instead and the same layout lives in object storage, with credentials, region and endpoint passed through `storage_options={...}` rather than a connection string. Either way the call is cheap and local: it does not start a server, open a socket, or hold a pool — LanceDB is embedded, so reads and writes happen in your process against files. The practical consequences are what interviewers want: many readers are fine, concurrent writers to one table can conflict on commit, and on object storage every query pays network latency, so you tune for fewer, larger range reads. `read_consistency_interval` controls how often an open table checks for versions written by other processes.
go deeper
Know that connect takes a directory path or a bucket URI, creates the local directory if needed, and returns a handle you call create_table and open_table on. There is no server to install or start.
Explain that LanceDB is embedded, that storage_options carries region and credentials for object storage, and that the URI changes where bytes live rather than how the API works.
Bring the operational consequences: object-storage latency and co-locating compute, the single-writer-many-readers model and commit conflicts, and read_consistency_interval as the fix for stale readers.
Own the deployment choice — local disk for a single service versus a shared bucket behind many stateless workers — and the tradeoff you accept by having no coordinating server: ingest topology, retry policy, and where consistency requirements are actually enforced.
## connect is an open, not a connection `lancedb.connect(uri)` returns a `DBConnection` object, but almost nothing about it resembles a client-server database connection. There is no handshake, no authentication round trip against a LanceDB daemon, no pool to size, and no port to expose — LanceDB is embedded, so the library running inside your process *is* the database. The URI names a root location; the connection is a handle for listing, opening and creating tables under it. That single fact explains most of LanceDB's operational character. ## A local path Given `./lance` or an absolute directory, LanceDB creates the directory if it is missing and returns immediately. Tables you create appear as subdirectories named `<table>.lance`, each a Lance dataset: columnar data files, a versions/manifest area recording each committed version, index files, and deletion files. `db.table_names()` is a directory listing; `db.open_table(name)` opens the dataset's current manifest. You can copy or rsync the whole directory to move the database, and you can inspect it with ordinary file tools — there is no opaque server state anywhere else. ## An object-storage URI Pass `s3://bucket/prefix` (also `gs://` and `az://`) and the identical layout is written to object storage. Credentials and region do not go in the URI; they go in `storage_options`, for example `lancedb.connect("s3://bucket/prefix", storage_options={"region": "us-east-1"})`, which also carries things like a custom endpoint for S3-compatible stores. Anything not supplied there falls back to the environment's usual credential chain. The behaviour changes even though the API does not. Object storage has high per-request latency and no cheap random small reads, so the format's design — large columnar files read with ranged GETs, metadata in manifests — is what keeps queries viable. Warm local disk gives millisecond queries; the same table on S3 from a laptop can be an order of magnitude slower, and running compute in the same region as the bucket matters more than any tuning parameter. Also note the deployment shape this unlocks: many stateless workers all pointing at one bucket, no cluster to operate. ## Consistency between processes Because each process reads the manifest it opened, a `Table` object does not automatically see versions committed by someone else. `lancedb.connect(uri, read_consistency_interval=timedelta(seconds=5))` sets how often an open table re-checks for newer versions; `timedelta(0)` checks on every operation (strongest, most metadata reads), and the default leaves the table on the version it opened until you explicitly refresh it with `checkout_latest()`. Getting this wrong produces the classic "the writer added rows but my reader still returns the old count" report. ## Writers Lance commits are versioned, and concurrent writers to the same table race to commit the next version; conflicting commits fail and must be retried, so the comfortable model is one writer per table with many readers. If you need multiple ingest processes, partition them across tables, funnel them through a single writer, or accept and handle commit retries. This is the main thing you give up by not running a server. ## Async and cloud `lancedb.connect_async(uri)` gives the async client with the same URI semantics for use inside an event loop. A `db://` URI targets LanceDB Cloud rather than storage you own, in which case an API key is supplied instead of storage credentials — the table API you write afterwards stays the same, which is the point of the uniform URI. ## What to say in an interview Summarise it as: the URI chooses where bytes live, not how you talk to the database. Local paths for development and single-node services, object storage for shared, elastic, serverless-style deployments; `storage_options` for credentials; latency and consistency, not API differences, are what actually change between them.
- A writer process appended rows but a long-running reader still sees the old count. Why?The reader's `Table` object is pinned to the version it opened. LanceDB does not push updates to it; by default it keeps serving that manifest until told otherwise. Either call `checkout_latest()` to jump to the newest version, or set `read_consistency_interval` on `connect` so the table re-checks for new versions on that cadence — `timedelta(0)` for check-every-operation at the cost of extra metadata reads.
- Two ingest jobs write to the same LanceDB table concurrently. What do you expect?Commit conflicts. Each write appends a new version, and two processes trying to commit the next version race; the loser's commit fails and has to be retried against the newer state. LanceDB has no lock server to arbitrate. The safe designs are a single writer per table, sharding writers across separate tables, or wrapping writes in retry logic that tolerates conflicts.
- How do you pass a region and a custom endpoint for an S3-compatible store?Through `storage_options` on `lancedb.connect`, for example `lancedb.connect("s3://bucket/prefix", storage_options={"region": "us-east-1", "endpoint": "..."})`. Credentials can go in the same dict or be left to the environment's normal credential chain. Nothing goes in the URI beyond bucket and prefix, and no separate configuration file is involved.
saying these in an interview costs you the question
- Saying connect() opens a socket to a LanceDB server
- Putting credentials in the s3:// URI instead of storage_options
- Assuming an open table auto-refreshes when another process writes
- Expecting the same latency from S3 as from local disk
- Believing concurrent writers to one table are safely serialised