skip to content

How do you make documents searchable by OpenAI's hosted file_search tool?

level: middleimportance: must knowfreq 56%

answer

  1. upload, collect, wait, reference
  2. the collection has an id you point at
  3. empty results often means not ready yet
  4. storage is rented by the day, not bought
  5. you tune the edges, not the embeddings

basics

~20 s

Upload the files, add them to a vector store, and wait for indexing to finish, then reference that store's id from the file_search tool on your request. OpenAI handles chunking, embedding, and ranking; you pay for storage per day.

solid answer

~50 s

The flow is upload → vector store → attach → query. You upload each document as a file, add it to a vector store (the SDK's file-batch upload-and-poll helper does both and waits), and only then reference `vector_store_ids` from the `file_search` tool — in the Responses API on the tool object itself, in the Assistants API through the assistant's or thread's file-search tool resources. Indexing is **asynchronous**: a vector-store file sits `in_progress` until it is parsed, chunked and embedded, and searching before that returns nothing, which is the number one cause of "my RAG returns empty". OpenAI owns the pipeline — default chunking is roughly 800-token chunks with 400-token overlap, adjustable via a static chunking strategy — and you tune retrieval with `max_num_results`, a score threshold, and metadata filters. Billing is per gigabyte per day of vector-store storage beyond a free allowance, plus a per-call charge for the tool, so set an expiration policy on stores you do not need forever.

code

python · 22 lines
python
store = client.vector_stores.create(
    name="handbook",
    expires_after={"anchor": "last_active_at", "days": 7},
)

with open("handbook.pdf", "rb") as f:
    batch = client.vector_stores.file_batches.upload_and_poll(
        vector_store_id=store.id, files=[f]
    )
print(batch.status, batch.file_counts)

resp = client.responses.create(
    model="gpt-5",
    input="What is the refund window?",
    tools=[{
        "type": "file_search",
        "vector_store_ids": [store.id],
        "max_num_results": 5,
    }],
    include=["file_search_call.results"],
)
print(resp.output_text)

go deeper

for a junior

Know the sequence: upload the document, add it to a vector store, then point the file_search tool at that store's id so the model can quote from it.

for a middle

Explain that indexing is asynchronous and that a file must reach completed before it is searchable, and name the tuning knobs: chunk size and overlap, result count, score threshold, metadata filters.

for a senior

Show operational ownership — check file_counts for failures, inspect returned chunks when answers are wrong, isolate tenants with separate stores, and set expiration policies so idle storage stops billing.

for a principal

Be ready to argue the build-versus-rent line: hosted retrieval buys weeks of work but forfeits embedding choice, hybrid search, reranking and provider portability, which is the deciding factor once retrieval quality becomes the product.

## What the hosted tool actually gives you `file_search` is a managed retrieval pipeline. You hand over documents; OpenAI parses them, splits them into chunks, embeds the chunks, stores the vectors, and at query time rewrites the query, searches, ranks and injects the top passages into the model's context — then cites which chunks it used. You never see an embedding vector or a similarity score computation. That is the whole value proposition and also the whole limitation. ## The ingestion path, step by step 1. **Upload the file.** A document is uploaded as a file object with an assistants-oriented purpose. Individual files can be large — up to 512 MB — with a token ceiling per file, and supported formats cover the usual text and document types. 2. **Create a vector store.** This is the searchable collection. It carries a name, an optional `expires_after` policy, and a `file_counts` object reporting how many of its files are `in_progress`, `completed`, `failed`, or `cancelled`. 3. **Add files to the store.** Either one at a time or as a batch. The SDK's batch upload-and-poll helper uploads and blocks until every file reaches a terminal state, which is what you want in a script; in a service, kick the batch off and track status asynchronously. 4. **Wait for `completed`.** Parsing and embedding take time proportional to document size. Until a file is `completed` its content is not retrievable. A file can also end `failed` — unreadable PDF, unsupported encoding — and a store with silent failures looks identical to one that simply has no relevant content, so check `file_counts` rather than assuming success. 5. **Attach and query.** In the Responses API you pass the tool as `{"type": "file_search", "vector_store_ids": [...]}`, optionally with `max_num_results` and filters. In the Assistants API the store is attached through the file-search tool resources on the assistant (shared corpus) or on the thread (per-conversation documents). ## Tuning the black box You cannot choose the embedding model, but several knobs are exposed: - **Chunking strategy.** The default is automatic — around 800 tokens per chunk with 400 tokens of overlap. A static strategy lets you set both explicitly. Smaller chunks sharpen precision on fact lookups; larger chunks preserve argument structure in prose. Chunking is decided at ingestion, so changing it means re-adding the files. - **Result count.** `max_num_results` caps how many chunks are injected. More results improve recall and inflate your input-token bill on every call, since retrieved text is prompt text. - **Ranking threshold.** A score threshold suppresses weak matches, which is how you stop the model from confabulating around irrelevant passages when the corpus genuinely lacks an answer. - **Attributes / metadata filters.** Files can carry key-value attributes, and a query can filter on them. This is the mechanism for tenant isolation, document-type scoping, and recency windows — one store per tenant is the alternative, and often the safer one. ## Reading the citations The model's answer carries annotations pointing at the file and chunk each claim came from, so you can render sources. In the Responses API you can additionally ask for the raw retrieved chunks in the response by including the file-search call results, which is indispensable for debugging: it tells you whether a wrong answer came from bad retrieval or bad reasoning over good retrieval. ## Cost and lifecycle Vector-store storage is billed **per gigabyte per day** past a free allowance, and in the Responses API each file-search tool call carries its own per-call charge on top of tokens. This makes lifecycle a real design concern: a per-conversation store created for one uploaded PDF and never deleted accrues charges indefinitely. Set `expires_after` with a last-active anchor on ephemeral stores, keep long-lived shared corpora explicit and few, and remember that deleting a file object does not by itself remove the already-indexed chunks from a store — remove the vector-store file too. ## When hosted retrieval is the wrong tool Choose it when the corpus is modest, the documents are ordinary, and shipping fast matters more than control. Reach for your own pipeline when you need a specific embedding model, hybrid keyword-plus-vector search, reranking, chunk-level access control, deterministic reproducibility of retrieval for evaluation, or the ability to move providers. Hosted retrieval trades exactly those things for the several weeks of work it saves.

  • A user complains the assistant says it cannot find something that is definitely in the uploaded PDF. How do you diagnose it?
    Check three layers in order. First, the vector store's file_counts — a file stuck in_progress or ended failed is invisible to search. Second, the retrieved chunks: ask the response to include the file-search results and see whether the right passage was returned at all. If it was returned and the model still said no, that is a reasoning or instruction problem; if it was not, tune chunking, raise max_num_results, or lower the score threshold.
  • How would you isolate documents per customer in a multi-tenant product?
    Prefer one vector store per tenant and pass only that tenant's store id on the request — isolation is then structural and a filter bug cannot leak data. Attribute-based filtering within a shared store is cheaper to operate and viable when tenants are numerous and small, but it makes every query's correctness a security control. Whichever you choose, never let the tenant id reach the request from client-controlled input.
  • Why can more retrieved chunks make answers worse as well as more expensive?
    Every retrieved chunk is injected as prompt text, so raising max_num_results raises input tokens on every single call. Beyond cost, weak matches dilute the context: the model has to pick the relevant passage out of mostly irrelevant ones, and a confidently wrong nearby passage is a common source of confabulation. Pair a modest result count with a score threshold so a genuinely unanswerable question returns nothing rather than noise.
  • Does deleting the uploaded file remove it from the vector store?
    Not reliably — the store holds derived chunks, and the file object and the vector-store file are separate resources. Removing content properly means deleting the vector-store file entry as well as the underlying file, and for data-deletion requests you should verify against the store's file listing rather than assuming. Storage charges follow what the store still holds, so an incomplete cleanup keeps billing too.

saying these in an interview costs you the question

  • Querying immediately after upload and blaming the model for empty results
  • Assuming you can choose the embedding model for file_search
  • Ignoring failed files because the batch call returned
  • Leaving per-conversation vector stores with no expiration policy
  • Thinking retrieved chunks are free rather than billed as input tokens

context