Why does copying text from a scanned PDF return nothing, but not from a born-digital PDF?
answer
- one file, two very different payloads
- characters versus pixels
- test each page, not each file
- scanner and fax pages carry no characters
- OCR creates text that was never there
basics
~20 sA born-digital PDF stores real text objects — character codes with fonts and positions — so extraction just reads them out. A scanned page stores only a photograph of the paper, so no characters exist until OCR creates them.
solid answer
~50 sPDF is a container for page-painting instructions, not a text format. When a program produces the file — an ERP printing an invoice, a word processor exporting a report — the page carries a **text layer**: character codes, their fonts, and the coordinates where each run is drawn. Extracting from that is a parse: instant, free, and exact. When a page arrives through a scanner, a fax gateway or a phone camera, the PDF holds a single raster image; the characters exist only as pixels, so any text has to be *recognised* by OCR, which costs money, adds latency and introduces errors. Real corpora mix both inside one file — a born-digital cover page stapled to scanned annexes — so the working rule is to test each **page** for a usable text layer and send only the pixel pages down the recognition path.
code
python · 11 linesimport fitz # PyMuPDF
doc = fitz.open("bill_of_lading.pdf")
for page in doc:
text = page.get_text("text").strip()
has_full_page_image = any(
page.get_image_bbox(img).get_area() > 0.7 * page.rect.get_area()
for img in page.get_images(full=True)
)
route = "parse" if len(text) > 50 and not has_full_page_image else "ocr"
print(page.number, route)go deeper
Be able to say plainly that a scanned page is an image with no characters in it, and that OCR is what turns those pixels into text. Knowing to check for a text layer before reaching for OCR is the expected answer.
Explain the mechanics: a text layer holds character codes with fonts and coordinates, so extraction is a parse, while a scan needs recognition that costs money and has an error rate. Describe a concrete per-page detection heuristic.
Show that you route per page, not per file, and that you instrument the split because it drives cost. Mention the searchable-PDF trap — an embedded text layer that is really untrusted OCR output — and how you would validate it against a labelled sample.
Own the economics: the parsed-versus-recognised ratio is a budget line that moves when customers change how they submit documents. Be ready to argue for pushing born-digital submission upstream, and to set the policy that no page is recognised twice.
## What a PDF actually stores A PDF describes how to paint a page. For a file produced by software, that description includes text-showing operations: character codes, the font each code belongs to, and the coordinates where each run of text is placed. Collectively this is the *text layer*. Pulling text out of such a page is a parse of data that is already present — it is exact, costs nothing but CPU, and comes with per-run positions you can reuse for layout work. A page that entered the file through a scanner, a fax gateway, or a phone camera contains none of that. The whole page is one large raster image — a grid of pixels — wrapped in a PDF container. There are no character codes anywhere in the file. Selecting text in a viewer selects nothing because there is nothing to select. ## Two paths with very different economics The text-layer path is a parse: microseconds per page, no per-page fee, no recognition errors, and deterministic across reruns. The raster path is recognition: you render the page to an image (200–300 DPI is the usual working range), then run an OCR engine or a document vision-language model over it. That path costs real money per page, adds latency, and has a non-zero character error rate — the standard metric is CER, character error rate, and it climbs sharply with skew, low contrast, fax compression artefacts, and handwriting. The practical consequence: in a freight forwarder's pipeline that ingests bills of lading and arrival notices, the born-digital half is essentially free and perfectly accurate, while the faxed half carries the entire extraction budget and nearly all of the review workload. Knowing which half a page belongs to *before* you spend anything is the first design decision in any document pipeline. ## Both kinds live in the same corpus, and often the same file Documents get merged. A clerk prints a system-generated cover sheet, staples the carrier's faxed annex behind it, and scans the staple line. The result is one PDF where page 1 has a perfect text layer and pages 2–7 are images. Deciding per file therefore produces one of two failures: you OCR pages that did not need it (paying and adding errors), or you skip pages that had no text at all (silently dropping half the shipment's line items). Detection belongs at page granularity. ## How to detect a usable text layer The cheap heuristic is to extract text per page and look at how much came back relative to how much ink is on the page: a page with a handful of characters but a full-page image object is a scan. Refinements worth having are checking whether a single image covers most of the page area, and guarding against legitimately sparse pages (a separator sheet, a signature page) so they are not misrouted as failures. One important special case: **searchable PDFs**. When someone has already run OCR, the tool typically writes the recognised text back into the file as an invisible text layer positioned under the image. Those pages *do* extract text — but that text is somebody else's OCR output, inherited errors and all. Treat it as untrusted: sample a labelled subset and measure its accuracy before deciding whether to trust it or re-recognise the pixels. ## Traps inside a genuine text layer A text layer being present does not make it clean. Fonts embedded without a Unicode mapping extract as garbage or private-use codepoints even though the page renders correctly on screen. Kerning and justification split a single word into several runs. Ligatures and soft hyphens survive into your output. And crucially, the order of drawing operations is whatever the producing software emitted — it is not a reading order, which is why multi-column pages can extract as interleaved nonsense even when every character is exact. ## What this means for the pipeline Route per page into two lanes. The text-layer lane is cheap and exact but still needs layout work to recover structure. The raster lane renders and recognises. A hybrid is legitimate and common: you sometimes rasterise a *born-digital* page on purpose, because a layout-aware model reading the rendered image reconstructs a dense table better than the raw run order does — you are trading a little cost for structure you could not otherwise get. Whatever you choose, instrument the split. The ratio of parsed pages to recognised pages is your cost model, and it drifts as customers change how they send documents.
- A born-digital page extracts every character correctly, yet the words come out jumbled. What is happening?The text layer records drawing operations in whatever order the producing software emitted them, not in reading order. On a multi-column or boxed layout, that order interleaves regions. The characters are exact; the sequence is meaningless. Recovering order requires layout analysis over the run positions rather than trusting the extraction sequence.
- How would you decide a page needs OCR without misclassifying a legitimately near-empty page?Combine signals instead of thresholding character count alone: how many characters came back, whether a single image object covers most of the page area, and how much of the page is non-white ink. A separator sheet has little text and little ink; a scan has little text and a full-page image. That pair separates them.
- A vendor sends 'searchable PDFs' that already contain a text layer. Do you trust it?Not by default. That layer is usually a previous OCR pass, so it carries someone else's recognition errors and may be misaligned with the pixels. Sample a labelled subset, measure character and field-level accuracy, and only then decide between trusting it, re-recognising the image, or using it as a second opinion.
A born-digital PDF is a typed document; a scanned PDF is a photograph of one. You can search the first immediately; the second needs someone to read it back into words first.
saying these in an interview costs you the question
- Assuming every PDF contains extractable text
- Running OCR on born-digital pages that already parse exactly
- Believing OCR improves text that was already exact
- Deciding text-layer vs scan per file instead of per page
- Treating the extraction sequence as the reading order
- Trusting a searchable PDF's embedded OCR without measuring it