In Ruby, why do ljust, center and format("%-18s") misalign receipt columns when item names contain accented or CJK characters?
answer
- characters, not screen columns
- combining accent counts as one
- wide CJK glyph takes two columns
- width never truncates
- unicode_normalize(:nfc) helps accents only
basics
~10 sljust, rjust, center and format widths count characters, not terminal columns. A decomposed accent adds an invisible character and a CJK character fills two columns, so padding by character count leaves rows misaligned.
solid answer
~40 sRuby's padding methods measure a string's **length in characters**: `ljust`, `rjust`, `center` and `format` widths all count characters, and so does `%.10s` precision. A receipt printer or terminal shows **columns**. Two cases break the match. A **decomposed** accent, `"e\u0301"`, is two characters but one visible glyph, so the row comes out one column short; `unicode_normalize(:nfc)` composes it into one character first. An East Asian **wide** character such as `"抹"` is one character but two columns, so `"抹茶ラテ".ljust(18)` adds 14 spaces to text that already fills 8 columns. Core Ruby has no display-width method, so exact alignment needs a width table, usually from a library, and padding computed from display width instead of `length`. And because widths never truncate, over-long names push later columns right as well.
code
ruby · 8 linesrows = ["Latte", "Cafe\u0301", "抹茶ラテ"]
rows.each { |n| puts "#{n.ljust(10)}|" }
# Latte |
# Café | <- one column short
# 抹茶ラテ | <- four columns too far
rows.map { |n| n.unicode_normalize(:nfc).length }
# => [5, 4, 4] NFC fixes the accent, not the widthgo deeper
Recall that ljust, rjust and center pad to a character count and never shorten a string.
Explain why length can differ from what the screen shows, using a combining accent and a wide CJK character as the two examples.
Diagnose a misaligned receipt or report: inspect codepoints, normalise decomposed input, and pad by display width for wide characters.
Decide how far a product goes for multilingual fixed-width output, from normalising input to adopting a width library and testing on the real output device.
## What Ruby's padding counts Every padding tool on `String` works in **characters**: - `ljust(width)`, `rjust(width)` and `center(width)` pad until `length` reaches `width`. - `format("%-18s", name)` pads to a minimum width in characters; `%.18s` cuts to at most 18 characters. - `length` counts characters in the string's encoding, not bytes. For ASCII text on a fixed-width display, one character is one column, so all of this lines up. Two kinds of text break that assumption. | Text | `length` | Columns shown | Effect of `ljust(10)` | |---|---|---|---| | `"Latte"` | 5 | 5 | aligned | | `"Cafe\u0301"` (e + combining accent) | 5 | 4 | one column short | | `"Café"` (precomposed é) | 4 | 4 | aligned | | `"抹茶"` | 2 | 4 | two columns too long | ## Combining characters Unicode can write é in two ways: one precomposed character, or `e` followed by the **combining** acute accent U+0301. Both render identically, but the second has one extra character that occupies no column. Menu names pasted from a design tool or a web form are often decomposed. The fix for this case is in core Ruby: ```ruby name = "Cafe\u0301 au lait" name.length # => 13 name.unicode_normalize(:nfc).length # => 12 ``` `String#unicode_normalize(:nfc)` composes combining sequences where a precomposed character exists, after which `ljust` counts correctly. ## Wide characters Many East Asian characters are **wide**: fixed-width fonts give them two columns. `"抹茶ラテ"` is four characters and eight columns, so: 1. `"抹茶ラテ".ljust(18)` adds 14 spaces, reaching 18 characters but 22 columns. 2. Every following column on that receipt row shifts four places right. 3. Normalisation does not help; the characters are already single code points. Ruby's core library has no method that returns display width. The usual approach: - compute each string's **display width** with a width table, typically from a library that implements the Unicode East Asian Width property; - pad with `" " * (target - display_width)` instead of `ljust`; - truncate by accumulated display width, not by character count, when a name is too long. ## Over-long values Independently of Unicode, widths are **minimums**: - `"Extra large caramel macchiato".ljust(18)` returns the full 29-character string. - `format("%-18s", same)` does the same. Only a precision (`%-18.18s`) or an explicit slice caps the field. On a receipt, decide whether to truncate, wrap onto a second line, or abbreviate, and apply that before padding. ## Padding by display width Once a display-width function is available, the padding itself is short: ```ruby def pad_columns(text, columns, width_of) text + " " * [columns - width_of.call(text), 0].max end ``` - `width_of` is whatever returns the column count for a string: a library call, or a small table for the scripts your menu uses. - `[..., 0].max` guards against negative counts, because `String#*` raises `ArgumentError` for a negative argument. - The helper never truncates; long names still need a truncation rule of their own, applied by display width. ## Diagnosing a misaligned receipt - Print `name.length` and `name.codepoints` for the bad row. Extra code points after a letter point to combining marks; code points in the CJK ranges point to wide characters. - Compare `name.unicode_normalize(:nfc).length` with `name.length`; a difference means decomposed input. - Check whether the row is simply longer than the field, the most common cause of all. - Remember the output device: a thermal printer's font may render wide characters differently from a terminal, so test on the real target.
- Does unicode_normalize(:nfc) fix the alignment of CJK item names?No. NFC only composes combining sequences into precomposed characters. CJK characters are already single code points; they are misaligned because each occupies two columns, which needs a display-width calculation.
- How do you cap a column at 18 characters with format?Add a precision to the string specification: `%-18.18s` pads to at least 18 and cuts to at most 18 characters. For wide characters the cut must be computed by display width instead.
saying these in an interview costs you the question
- ljust counts bytes, so UTF-8 names get extra padding
- ljust and format widths truncate names that are too long
- unicode_normalize(:nfc) fixes the width of CJK characters
- String#length returns the number of screen columns