A two-column customs form extracts as interleaved gibberish — how do you fix the reading order?
answer
- the characters are right, the sequence is wrong
- drawing order is not reading order
- segment before you serialise
- find the gutter, then read column by column
- keep block boxes to prove the order
basics
~20 sExtraction emitted characters in drawing or naive scan order, not reading order. Fix it upstream with layout analysis: detect columns and blocks first, order text within each region, then read regions in sequence — never by post-processing a flat text dump.
solid answer
~50 sThe characters are usually fine; the *sequence* is wrong. Two causes dominate: a born-digital page extracts in the order the producing software drew the runs, and a naive extractor sorts words purely top-to-bottom then left-to-right, which zips two columns together line by line. Neither is repairable from the flat string, because the geometry that would tell you where the column boundary sits has already been thrown away. The fix is to do segmentation before serialisation — either a geometric split (a recursive horizontal/vertical projection cut that finds the whitespace gutter) or a learned layout detector that returns blocks with types and boxes, then read each block internally and emit blocks in a computed order. Rendering the page and passing it to a layout-aware document model achieves the same thing in one step. Keep the per-block boxes: they are how you verify the order was recovered, and how a reviewer checks a suspicious field.
go deeper
Know that a page has no built-in reading order and that a two-column layout can extract with the columns interleaved. Recognise the symptom — sentences cut in half and spliced together — rather than assuming the OCR misread characters.
Explain both causes: content-stream draw order on born-digital pages, and naive y-then-x sorting of OCR word boxes. Describe segmenting into columns or blocks before serialising, and why the coordinates must still exist when you do it.
Show production judgement: deskew and denoise first, choose between projection cuts and a learned layout detector by page raggedness, merge page-spanning tables by column geometry, and use invariants like totals and column-count stability as runtime alarms.
Own the argument that scrambled order is a correctness risk, not a cosmetic one — it silently rebinds values to the wrong labels. Set the policy that structure recovery is measured on a labelled set and that documents failing invariants are held rather than posted.
## Why the order breaks in the first place A page has no intrinsic linear text. Reading order is a human convention that a machine must reconstruct, and there are two distinct ways it goes wrong. On a **born-digital** page, extraction walks the content stream and returns text runs in the order the producing software emitted them. That order is an artefact of how the generator laid out the page — often frame by frame, sometimes field by field in form-fill order. It has no obligation to match reading order, and on a two-column customs declaration it frequently does not. On a **recognised** page, an OCR engine gives you words with boxes, and something has to serialise them. The naive serialiser sorts by y, then by x. On a single-column page that is correct. On a two-column page it produces exactly the reported symptom: the first line of the left column, then the first line of the right column, then the second line of the left column, and so on — every sentence cut in half and welded to an unrelated one. The damage is worse than it looks. Downstream, a value ends up attached to the wrong label, a description gets stitched onto another commodity's line, and the extraction is not merely ugly — it is confidently wrong. ## Segment before you serialise The repair has to happen while the geometry still exists. **Geometric segmentation.** Project ink onto the x axis and look for a tall, narrow band of whitespace running the height of the text area: that is the gutter. Split there and recurse, alternating horizontal and vertical cuts — the classical recursive projection-cut approach. It is fast, needs no model, and works well on clean rectangular layouts. It degrades on pages where a figure or a wide header straddles both columns, and on noisy scans where speckle fills the gutter, so deskew and denoise first. **Learned layout detection.** A layout model returns regions with types — title, paragraph, table, figure, caption, header, footer — each with a bounding box. You then sort words inside each region and order the regions themselves by a rule (column-major within the body, headers first, footers dropped). This copes with the ragged real-world pages that break projection cuts, at the cost of a model to run and evaluate. **Let a document model do it.** The purpose-built document VLM tier consumes the page image and emits reading-ordered, layout-aware output directly — markdown or HTML with blocks and tables in sequence. It collapses segmentation and serialisation into one step. The catch is that you are trusting an opaque ordering decision, so you still want blocks with coordinates back, and you still need a way to detect when the order is wrong. ## Tables that continue across a page break The same class of problem, one level up. A line-item table with a header repeated on every sheet must become one logical table, and three things have to happen: detect the header on each page, drop the repeats instead of ingesting them as data rows, and merge the remaining rows in page order while keeping column alignment. Column geometry is the reliable join key — the x-ranges of the columns are nearly identical page to page, so you can match a continuation page's columns to the first page's header even when the header itself is missing on later sheets. Watch for rows split across the break, where a description wraps and the second half opens the next page, and for footer subtotals that must not be mistaken for line items. Both are best handled with explicit validation: the number of rows the summary block claims, or a total that must equal the sum of the extracted amounts. ## Key-value pairs on forms A form is not prose at all. "Consignee", "Container No.", "Gross Weight" are labels whose values sit in a spatial relationship — to the right, below, or inside a ruled box. Pairing is a geometry problem: nearest text to the right within the same horizontal band, or the cell adjacent within a detected ruling. Ruling lines, where present, are the strongest signal available and worth detecting explicitly. This is also why a reading-order bug is so damaging on forms: once order is lost, the label-to-value association is lost with it. ## Knowing that you fixed it Do not verify by eye on three documents. Build a small labelled set with the correct block sequence and score against it; for tables, a tree-edit-distance structure metric such as TEDS scores the reconstructed table against the reference rather than just the cell text. Cheap runtime guards catch the rest in production: a sudden spike in mid-sentence breaks, a table whose column count varies page to page, a document where the row count disagrees with a stated total. Each of those routes the document to review instead of letting a plausibly-ordered, actually-scrambled extraction reach the downstream system.
- Why can't you repair the interleaving by post-processing the flat extracted text?Because the flat string has already discarded the coordinates that define where one column ends and the next begins. Any heuristic over the text alone — sentence boundaries, capitalisation — is guessing at a boundary the geometry knew exactly. The repair must happen while the word boxes still exist, which usually means re-extracting rather than cleaning up.
- A full-width header spans both columns and breaks your projection cut. How do you handle it?Cut horizontally before cutting vertically: a recursive projection approach alternates axes, so a full-width band is separated as its own region first, and column splitting is applied only to the body below it. If alternating cuts still fail on ragged layouts, that is the signal to move to a learned layout detector that classifies regions rather than assuming rectangles.
- How would you catch a reading-order regression in production, without labelled data for every document?Use cheap invariants as alarms: a jump in mid-sentence block breaks, column counts that vary between pages of the same table, label-value pairs where the value fails its type or checksum, and extracted totals that disagree with a stated total. Any of those routes the document to human review and flags the batch for inspection.
saying these in an interview costs you the question
- Trying to reorder a flat text dump instead of re-extracting with geometry
- Assuming extraction order equals reading order
- Sorting words purely top-to-bottom then left-to-right on any layout
- Treating a repeated table header on each page as data rows
- Verifying reading order by eyeballing a couple of documents