What is pseudo-localization, and which defects does it catch before any translation exists?
answer
- No translator is involved at all
- The untouched strings are the finding
- Padding stands in for expansion
- Markers expose assembled sentences
- Green means ready to receive translations
basics
~20 sPseudo-localization replaces every extracted string with a mechanically accented, padded, bracketed version of itself and runs the product against that fake locale. Anything still in plain source text was never extracted, and padded text exposes truncation.
solid answer
~50 sThe transform accents or script-shifts each letter, pads the string by roughly a third to a half to simulate expansion, and wraps it in visible boundary markers. Running the product in that pseudo-locale exposes an entire family of internationalization defects with no translator involved: strings hardcoded past extraction, which stay plain; containers that clip or wrap badly once text expands; sentences assembled from fragments, which show as several bracketed pieces in a row; encoding failures, where characters outside the source alphabet arrive as replacement glyphs; and truncation counted in bytes, which splits a character. It is deterministic and cheap, so it belongs in the pipeline from the first sprint. It proves nothing about wording quality, real plural rules, regional formats or mixed-direction text, so a green run means ready to receive translations, not ready for a market.
code
pseudocode · 7 linesfunction pseudoLocalize(source):
accented = mapEachLetter(source, accentTable) // "Save" -> "Savé" style shift
padding = repeat("~", ceil(length(source) * 0.40))
return "[[" + accented + padding + "]]"
// rendered screen: "[[Savé~~]]" extracted
// "Duplicate charge reversed" never extractedgo deeper
Be able to say what a pseudo-localized screen looks like and why plain untouched text on it is a finding. Recognising that no translation is needed to run it is the main point at this level.
Explain each part of the transform and map it to the defect it exposes: accents to extraction, padding to truncation, markers to concatenation, non-source characters to encoding. Expect to be asked what the technique cannot detect.
Show how you turn it into an automated gate with machine-checkable assertions, where it sits relative to translation spend, and how you argue for fixing structural findings before a vendor contract is signed.
Own the sequencing and the economics: readiness gate before translation spend, an expansion budget derived from measured data rather than folklore, and a policy on which findings block a locale launch versus which are absorbed later.
### What pseudo-localization is Pseudo-localization is a build-time transformation that replaces every extracted source string with an algorithmically mangled version of itself, then runs the product against that fake locale. Nothing is translated; the point is that nothing *needs* to be. The transformation is deterministic, costs nothing per run, and can therefore sit in the pipeline from the first sprint — long before a translation budget exists. A conventional transform does four things at once: - **Accents or script-shifts every letter** ("Save" becomes something like "Šàvé"), so text that has gone through the catalogue is visually obvious and text that has *not* stands out as plain source language. - **Pads the string** by a fixed percentage, to simulate expansion. Padding of roughly a third to a half is the common rule of thumb; the real expansion depends heavily on string length — short labels expand proportionally far more than paragraphs — and any single percentage is a heuristic, not a measured law. - **Wraps the string in visible boundary markers**, so a sentence assembled from several fragments shows up as several bracketed pieces in a row. - **Optionally mirrors** the pseudo-locale's declared direction, to get an early look at a right-to-left rendering path. ### What it actually catches Because the transform touches only strings that came from the catalogue, the *absence* of the transform is the signal. A pseudo-localized build reveals, cheaply and repeatably: - **Strings never extracted.** Anything rendering in plain source text is hardcoded past the extraction step, or built at run time from code, or arriving unwrapped from a downstream service. - **Expansion and truncation.** Containers that were sized to the source string clip, ellipsize, wrap badly or push controls out of reach once every string is 40% longer. - **Concatenation.** Boundary markers appearing mid-sentence prove a sentence was assembled from parts, which is unfixable by translation alone. - **Encoding and font path failures.** Characters outside the source alphabet arriving as replacement glyphs, question marks or mojibake point at a byte-level assumption somewhere in storage, transport, logging or rendering. - **Byte-based truncation.** A string cut at a byte count rather than a character count splits a multi-byte character and produces a broken final glyph. - **Fixed-width and overflow assumptions** in exports, receipts and printed documents, which are usually the last surfaces anyone checks. ### What it cannot tell you Pseudo-localization is an internationalization test, not a localization test. It says nothing about whether real wording reads naturally, whether the plural forms of a genuine locale are correct, whether the term matches the glossary, whether a date order or currency convention is right for a region, or whether the tone suits the audience. It also gives only a rough hint about a right-to-left rendering: mirrored pseudo-text exercises the direction plumbing, but not the mixed-direction ordering of real text with embedded numbers and identifiers. Treat a green pseudo-localized run as "the product is ready to receive translations", never as "the product is ready for that market". ### Running it as a gate The pattern that works is a selectable pseudo-locale in the same build the team already runs — chosen by a setting, not a separate branch — plus a scripted walk through the critical flows capturing every screen, and a small set of assertions that can be automated rather than eyeballed: no visible string lacking boundary markers, no clipped text node, no replacement glyph in the captured text, no message containing two closing markers followed by an opening one. The visual review that remains is a human pass over the captured screens. ### A worked example A utility billing portal ran its first pseudo-localized pass over 2,600 catalogue entries and a scripted walk of 41 screens. Seventeen strings rendered in plain source text: fourteen were toasts and error messages built in code, two came unwrapped from a downstream statement service, and one — "Duplicate charge reversed" — was generated inside the payment path, which is exactly the message a customer sees during the incident they are most upset about. A further nine containers clipped under padding, all of them table headers on the account statement. None of that required a translator, a vendor contract or a single translated word; it required one pipeline job and an afternoon of screenshot review, and it was found roughly four months before the first locale shipped. ### Where it sits in the order of work The sequence that avoids rework is: extract strings, run pseudo-localization until it is clean, fix the internationalization defects it exposes, *then* buy translation. Reversing that order means paying for translation of a build that cannot display it, and re-testing every locale after each structural fix.
- How much padding should a pseudo-locale add, and how defensible is that number?Roughly a third to a half is the common working figure, and it is a heuristic rather than a measured law. Real expansion depends strongly on string length: a short button label can more than double, while a paragraph may grow only slightly. If the product has real translated content in any locale, measure the actual expansion distribution from it and set the padding from that instead of quoting a rule of thumb.
- Which locale-readiness defects will a clean pseudo-localized run still miss?Everything that depends on real content or a real region: wording quality and tone, correct plural and gender forms for an actual locale, glossary consistency, regional date, number and currency conventions, calendar choice, collation order, and the ordering of mixed-direction text with embedded numbers and identifiers. It is an internationalization check, not a localization check.
- How would you make a pseudo-localized run a pipeline gate rather than a manual pass?Select the pseudo-locale by configuration in the normal build, script a walk through the critical flows capturing each screen, and assert mechanically: no visible string lacking boundary markers, no clipped text node, no replacement glyph in the captured text, and no message containing a closing marker followed by an opening one. Leave only the visual review of captured screens to a human.
It is a dress rehearsal in costume: none of the text is real yet, but you find out immediately which doorways the costume will not fit through.
saying these in an interview costs you the question
- Thinks it needs real translations to run
- Calls a clean run proof of market readiness
- Skips it until the first locale is contracted
- Treats padding percentage as a measured constant
- Runs it once manually instead of every build