skip to content

What does it mean to subset a webfont file, and why does subsetting usually cut its size so dramatically?

level: juniorimportance: should knowfreq 45%

answer

  1. most glyphs are never rendered
  2. rebuild the binary, not compress it
  3. codepoint ranges chosen at build time
  4. fonttools pyftsubset writes woff2 directly
  5. dropped characters fall back per character

basics

~20 s

Subsetting rebuilds a font file so it contains only the glyphs a site actually renders. Retail fonts ship thousands of glyphs covering scripts, symbols and typographic alternates most pages never use, so dropping them commonly removes more than half the bytes.

solid answer

~50 s

A font binary is a container of tables: a character-to-glyph map, the outlines themselves, advance widths, OpenType layout tables for kerning and ligatures, and hinting. A commercial or open text family usually covers Latin, Latin Extended, Greek, Cyrillic, currency signs, arrows and stylistic alternates — and a typical English-language site renders maybe a couple of hundred of those glyphs. Subsetting runs the original through a tool such as `pyftsubset` from the `fonttools` package, keeping a chosen codepoint set and pruning the tables that referenced the removed glyphs, then writing the result straight out as `woff2`. The saving scales with how much you drop; cutting a multi-script family down to Latin plus punctuation is routinely a large majority of the file. The cost is that any character you did not keep now renders from a fallback font, so subsets have to be chosen against real content, not guessed.

code

bash · 5 lines
bash
pyftsubset Inter.ttf \
  --unicodes="U+0000-00FF,U+2013-2014,U+2018-201D,U+20AC" \
  --layout-features="kern,liga" \
  --flavor=woff2 \
  --output-file=inter-latin.woff2

go deeper

for a junior

Be able to say that a font file carries far more glyphs than a page renders, and that subsetting rebuilds it with only the ones you keep. Name that the saving comes from dropping whole scripts and symbol sets.

for a middle

Explain the mechanics: which tables get pruned, that fonttools' pyftsubset takes codepoint ranges and can emit woff2 directly, and that fallback for a missing character happens per character rather than per element.

for a senior

Show judgment about which text is safe to subset aggressively. Talk about deriving the codepoint set from real content versus declaring locale ranges, and about what you would monitor to catch fallback characters appearing in production copy.

for a principal

Own it as a pipeline concern: where subsetting lives in the build, how locale expansion changes the required ranges, how font licences constrain the approach, and the tradeoff between one broad subset that caches well and several tight ones that ship fewer bytes.

## What is actually inside a font file A `woff2` or `ttf` file is not a picture of an alphabet; it is a small binary database of tables. The `cmap` table maps Unicode codepoints to internal glyph ids. The outline table (`glyf` for TrueType outlines, `CFF ` for PostScript ones) holds the actual vector shape of every glyph. `hmtx` holds each glyph's advance width. `GSUB` and `GPOS` are the OpenType layout tables that implement kerning pairs, ligatures, small caps, tabular figures and script-specific shaping. There is also hinting bytecode, plus name and metric metadata. A well-made open-source text family may carry well over a thousand glyphs: Basic Latin, Latin Extended-A and -B for European diacritics, Greek, Cyrillic, Vietnamese, currency symbols, arrows, mathematical operators, and several stylistic sets. A marketing site in English renders a few hundred distinct characters at most — usually far fewer. ## What subsetting does Subsetting produces a **new font file** that keeps only a chosen set of glyphs. It is not compression and it is not a CSS trick: the work happens ahead of time, on the binary, in your build. The tool walks the requested codepoints, keeps their glyphs and anything they depend on (composite accented glyphs reference base glyphs, ligature substitutions reference their components), rewrites `cmap` and `hmtx`, prunes the layout tables to the features you asked to keep, and emits the result. The canonical tool is `pyftsubset`, shipped with the Python `fonttools` package. You give it a codepoint list or ranges, the layout features to preserve, and an output flavour: ```bash pyftsubset Inter.ttf \ --unicodes="U+0000-00FF,U+2018-201D,U+2013-2014,U+20AC" \ --layout-features="kern,liga" \ --flavor=woff2 \ --output-file=inter-latin.woff2 ``` Because subsetting and format conversion are separate steps conceptually, it helps to keep them separate mentally too: `woff2` compression shrinks whatever glyphs are present, subsetting decides which glyphs are present at all. You want both. ## Choosing the character set Two strategies exist. **Static**: declare the ranges the product supports — Basic Latin, the punctuation you typeset with (curly quotes, en and em dashes), the currency symbols, and any diacritics your locales need. **Content-derived**: scan the built site's rendered text and keep exactly the codepoints that appear. Tools such as `glyphhanger` automate the scan and hand the resulting codepoint list to `pyftsubset`. Content-derived subsets are the smallest and the most fragile — they are only safe for text you control, such as a fixed display headline. ## What subsetting can break - **Missing characters fall back.** Font fallback is per character, not per element: a Cyrillic name in a user comment renders from the next family in the stack (or as a `.notdef` box) while the surrounding Latin still uses your webfont. On user-generated content that looks broken. - **Layout features vanish if pruned.** Strip `GPOS`/`GSUB` too eagerly and kerning pairs and ligatures go with them; text still renders, but spacing changes subtly. - **Hinting removal** is a size win but changes rendering at small sizes on platforms that still use hinting. - **Licensing.** Some commercial licences restrict modifying or re-hosting the binary; subsetting is a modification. Check before shipping. ## Where it sits in the wider strategy Subsetting is the cheapest byte saving in font strategy because it removes data no one will ever see, with no rendering tradeoff for the characters you keep. It composes with everything else: subset first, ship as `woff2`, split the subsets across `@font-face` rules with `unicode-range` if you genuinely serve multiple scripts, and pick weights deliberately. A team that skips it often ships several hundred kilobytes of Cyrillic and Greek outlines to an audience that reads only English — bytes on the critical path, competing with the content that actually gets rendered.

  • If subsetting already removes the unused glyphs, why still ship the file as woff2?
    They solve different problems. Subsetting decides which glyphs exist in the file; woff2 compresses whatever glyphs remain, using a font-aware transform on top of Brotli. A subset TrueType file is still far larger than the same subset as woff2, so you do both — subset in the build, emit woff2 as the output flavour.
  • How do you decide a subset is safe for a given block of text on the page?
    Split by content ownership. Text you author and control — nav labels, a display headline, marketing copy — can take an aggressive content-derived subset. Anything user-generated or translated needs a range-based subset broad enough for its locales, or it will show fallback characters mid-sentence. Many teams use a tight subset for the display face and a broader one for body text.
  • What breaks if the subset drops the OpenType layout tables?
    Kerning and ligatures stop applying, so pairs like "AV" or "Ta" lose their tightened spacing and "fi" renders as two separate glyphs. Nothing errors and the text is still legible — which is exactly why it slips through review. Keep at least the kern and liga features unless you have measured that dropping them matters.

saying these in an interview costs you the question

  • Thinks subsetting is just gzipping the font file
  • Believes CSS can hide unused glyphs at runtime
  • Assumes a missing glyph makes the whole paragraph fall back
  • Subsets user-generated content down to ASCII
  • Ignores that some font licences forbid modifying the binary

context