skip to content

What does Mistral's /v1/ocr endpoint return for a multi-page PDF?

level: middleimportance: should knowfreq 32%

answer

  1. One entry per page, not per document
  2. Structure preserved, not flat text
  3. Markdown you can chunk directly
  4. Figures come with boxes, bytes on request
  5. Priced by page, not by token

basics

~20 s

Mistral's document OCR route returns a pages array — one entry per page, each with the page's extracted content as Markdown, its index, page dimensions, and optionally the embedded figures as base64 images. Usage is reported and billed per page processed, not per token.

solid answer

~50 s

You POST to `/v1/ocr` with an OCR model id and a `document` object — either `{"type": "document_url", "document_url": ...}` for a PDF or `{"type": "image_url", "image_url": ...}` for a picture — where the URL may be a public link or a signed URL for a file you uploaded through the Files API. The response is a `pages` array with one entry per page: `index`, `markdown` holding the page's text with structure preserved (headings, lists and tables come back as Markdown), `dimensions`, and an `images` array describing embedded figures with their bounding boxes. Figure bytes are only included if you set `include_image_base64: true`. A usage block reports how many pages were processed, which is also the billing unit — OCR is priced per page, unlike the token-priced text endpoints. The Markdown output is deliberately RAG-friendly: it is what you chunk and embed.

code

bash · 11 lines
bash
curl https://api.mistral.ai/v1/ocr \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $MISTRAL_API_KEY" \
  -d '{
    "model": "mistral-ocr-latest",
    "document": {
      "type": "document_url",
      "document_url": "https://example.com/report.pdf"
    },
    "include_image_base64": false
  }'

go deeper

for a junior

Know that you send a document reference and get back a pages array, and that each page's text arrives as Markdown rather than as one undifferentiated string.

for a middle

Explain the request's document object (document_url vs image_url), the per-page response fields, and the include_image_base64 switch. State clearly that billing counts pages, not tokens.

for a senior

Show how it slots into ingestion: upload and signed URLs for private files, page index carried into chunk metadata for citations, splitting oversized documents with a page offset, and batch submission for archive-scale runs.

for a principal

Own the cost model. Per-page OCR plus per-token embedding gives a two-curve budget for any archive backfill, and quality sampling on scanned material is a policy decision — decide up front what accuracy you must verify before the extracted text becomes a source of truth.

## What the endpoint is for `POST https://api.mistral.ai/v1/ocr` turns a document into structured text. It is not a chat call with a picture attached — it is a dedicated document-understanding route whose job is to give you back the *whole* document as machine-readable text, page by page, in one call. The typical use is the front of an ingestion pipeline: PDF in, Markdown out, chunk, embed, index. ## The request Two fields carry the work: - `model` — an OCR model id, e.g. `mistral-ocr-latest`. - `document` — a tagged object. `{"type": "document_url", "document_url": "https://…"}` for PDFs, `{"type": "image_url", "image_url": "…"}` for images. Data URIs are accepted for the image form, and for private files you upload via the Files API and pass a signed URL rather than a public link. One useful option: `include_image_base64: true` asks for the embedded figures' bytes to be returned inline. Leave it off and you still get each figure's identifier and bounding box, but not the pixels — which keeps responses small when you only want text. ## The response ``` { "pages": [ { "index": 0, "markdown": "# Quarterly Report\n\n| Region | Revenue |\n|---|---|\n...", "images": [{"id": "img-0.jpeg", "top_left_x": 44, "top_left_y": 120, ...}], "dimensions": {"dpi": 200, "height": 2200, "width": 1700} }, { "index": 1, ... } ], "model": "mistral-ocr-latest", "usage_info": {"pages_processed": 2, "doc_size_bytes": 183221} } ``` The important design decisions to be able to talk about: **Per-page granularity.** The unit is the page, not the document. That is what makes citations possible — you can record "page 7" against a chunk and link a retrieved answer back to a location a human can verify. **Markdown, not plain text.** Headings stay headings, tables come back as Markdown tables, lists stay lists. That structure survives into your chunks, and it matters: a table flattened into a wall of words retrieves badly and reads worse when a model has to answer from it. **Figures are described, not inlined by default.** Each image gets an id and a bounding box on the page, and a placeholder reference inside the Markdown, so you can re-associate a figure with its position in the text. Bytes arrive only with `include_image_base64`. ## Billing and limits OCR is billed **per page processed**, which is a genuinely different pricing model from the token-billed chat and embedding routes and is the fact most often missed. `usage_info.pages_processed` is the number you are charged on, so a 400-page PDF is a 400-page charge whether the pages are dense or nearly blank. There are documented ceilings on the size and page count of a single document (on the order of tens of megabytes and around a thousand pages as of mid-2026). Beyond those you split the document client-side and stitch the page arrays back together, keeping a running page offset so your citations stay correct. ## Where it sits in a pipeline A straightforward document-RAG ingestion looks like: upload the PDF through the Files API → get a signed URL → call `/v1/ocr` → walk `pages`, chunking the Markdown while recording `index` as page metadata → embed each chunk → index. Because OCR is per-page priced and embeddings are per-token priced, the two halves of that pipeline have completely different cost curves; sizing them separately is the sane way to forecast a backfill. For large offline runs, the same OCR requests can be submitted as an asynchronous batch job rather than one synchronous call at a time, which is the usual approach for a historical archive where nothing needs an answer in the next second. ## Failure modes to expect - A URL your caller can reach but Mistral cannot — signed URLs expire, and internal hosts are not reachable at all. Upload the file instead of linking to it. - Scanned pages of poor quality: OCR output is text-shaped but wrong. Sample-check pages rather than assuming a 200 means the content is right. - Enormous responses when `include_image_base64` is on for an image-heavy report — turn it off unless you actually need the figures.

  • How does OCR billing differ from the chat and embeddings endpoints?
    Chat and embeddings are billed per token; OCR is billed per page processed, and the response's usage information reports the page count you are charged for. That changes how you forecast a backfill — a large scanned archive costs a function of page count and is completely insensitive to how much text is on each page.
  • You need the figures from a report, not just the text. What changes in the request?
    Set `include_image_base64: true`. Without it each page's `images` entries still give you an id and a bounding box on the page, plus a placeholder reference in the Markdown, but no pixels. With it the encoded bytes come back inline, which can inflate the response substantially on image-heavy documents — so turn it on only for the documents that need it.
  • Your PDF is bigger than the per-document limit. How do you handle it?
    Split it client-side into chunks that fit, OCR each, then concatenate the page arrays while adding a running offset to each `index` so page numbers stay true to the original document. Keeping that offset right matters more than it sounds: it is what makes "see page 412" citations trustworthy downstream.

saying these in an interview costs you the question

  • Expecting one flat text blob for the whole document
  • Assuming OCR is billed per token like chat
  • Thinking figure bytes are always included
  • Believing an internal or expired URL is fetchable by the API
  • Treating a 200 response as proof the extraction is accurate

context