skip to content

Subsetting and Fallback Matching

Most of a font file is glyphs your page never renders. Subsetting plus carefully matched fallback metrics cuts both the bytes and the layout shift the swap causes.

on this pageshow

questions

5

What does it mean to subset a webfont file, and why does subsetting usually cut its size so dramatically?

level: juniorimportance: should knowfreq 45%

answer

  1. most glyphs are never rendered
  2. rebuild the binary, not compress it
  3. codepoint ranges chosen at build time
  4. fonttools pyftsubset writes woff2 directly
  5. dropped characters fall back per character

basics

~20 s

Subsetting rebuilds a font file so it contains only the glyphs a site actually renders. Retail fonts ship thousands of glyphs covering scripts, symbols and typographic alternates most pages never use, so dropping them commonly removes more than half the bytes.

solid answer

~50 s

A font binary is a container of tables: a character-to-glyph map, the outlines themselves, advance widths, OpenType layout tables for kerning and ligatures, and hinting. A commercial or open text family usually covers Latin, Latin Extended, Greek, Cyrillic, currency signs, arrows and stylistic alternates — and a typical English-language site renders maybe a couple of hundred of those glyphs. Subsetting runs the original through a tool such as `pyftsubset` from the `fonttools` package, keeping a chosen codepoint set and pruning the tables that referenced the removed glyphs, then writing the result straight out as `woff2`. The saving scales with how much you drop; cutting a multi-script family down to Latin plus punctuation is routinely a large majority of the file. The cost is that any character you did not keep now renders from a fallback font, so subsets have to be chosen against real content, not guessed.

code

bash · 5 lines
bash
pyftsubset Inter.ttf \
  --unicodes="U+0000-00FF,U+2013-2014,U+2018-201D,U+20AC" \
  --layout-features="kern,liga" \
  --flavor=woff2 \
  --output-file=inter-latin.woff2

go deeper

for a junior

Be able to say that a font file carries far more glyphs than a page renders, and that subsetting rebuilds it with only the ones you keep. Name that the saving comes from dropping whole scripts and symbol sets.

for a middle

Explain the mechanics: which tables get pruned, that fonttools' pyftsubset takes codepoint ranges and can emit woff2 directly, and that fallback for a missing character happens per character rather than per element.

for a senior

Show judgment about which text is safe to subset aggressively. Talk about deriving the codepoint set from real content versus declaring locale ranges, and about what you would monitor to catch fallback characters appearing in production copy.

for a principal

Own it as a pipeline concern: where subsetting lives in the build, how locale expansion changes the required ranges, how font licences constrain the approach, and the tradeoff between one broad subset that caches well and several tight ones that ship fewer bytes.

## What is actually inside a font file A `woff2` or `ttf` file is not a picture of an alphabet; it is a small binary database of tables. The `cmap` table maps Unicode codepoints to internal glyph ids. The outline table (`glyf` for TrueType outlines, `CFF ` for PostScript ones) holds the actual vector shape of every glyph. `hmtx` holds each glyph's advance width. `GSUB` and `GPOS` are the OpenType layout tables that implement kerning pairs, ligatures, small caps, tabular figures and script-specific shaping. There is also hinting bytecode, plus name and metric metadata. A well-made open-source text family may carry well over a thousand glyphs: Basic Latin, Latin Extended-A and -B for European diacritics, Greek, Cyrillic, Vietnamese, currency symbols, arrows, mathematical operators, and several stylistic sets. A marketing site in English renders a few hundred distinct characters at most — usually far fewer. ## What subsetting does Subsetting produces a **new font file** that keeps only a chosen set of glyphs. It is not compression and it is not a CSS trick: the work happens ahead of time, on the binary, in your build. The tool walks the requested codepoints, keeps their glyphs and anything they depend on (composite accented glyphs reference base glyphs, ligature substitutions reference their components), rewrites `cmap` and `hmtx`, prunes the layout tables to the features you asked to keep, and emits the result. The canonical tool is `pyftsubset`, shipped with the Python `fonttools` package. You give it a codepoint list or ranges, the layout features to preserve, and an output flavour: ```bash pyftsubset Inter.ttf \ --unicodes="U+0000-00FF,U+2018-201D,U+2013-2014,U+20AC" \ --layout-features="kern,liga" \ --flavor=woff2 \ --output-file=inter-latin.woff2 ``` Because subsetting and format conversion are separate steps conceptually, it helps to keep them separate mentally too: `woff2` compression shrinks whatever glyphs are present, subsetting decides which glyphs are present at all. You want both. ## Choosing the character set Two strategies exist. **Static**: declare the ranges the product supports — Basic Latin, the punctuation you typeset with (curly quotes, en and em dashes), the currency symbols, and any diacritics your locales need. **Content-derived**: scan the built site's rendered text and keep exactly the codepoints that appear. Tools such as `glyphhanger` automate the scan and hand the resulting codepoint list to `pyftsubset`. Content-derived subsets are the smallest and the most fragile — they are only safe for text you control, such as a fixed display headline. ## What subsetting can break - **Missing characters fall back.** Font fallback is per character, not per element: a Cyrillic name in a user comment renders from the next family in the stack (or as a `.notdef` box) while the surrounding Latin still uses your webfont. On user-generated content that looks broken. - **Layout features vanish if pruned.** Strip `GPOS`/`GSUB` too eagerly and kerning pairs and ligatures go with them; text still renders, but spacing changes subtly. - **Hinting removal** is a size win but changes rendering at small sizes on platforms that still use hinting. - **Licensing.** Some commercial licences restrict modifying or re-hosting the binary; subsetting is a modification. Check before shipping. ## Where it sits in the wider strategy Subsetting is the cheapest byte saving in font strategy because it removes data no one will ever see, with no rendering tradeoff for the characters you keep. It composes with everything else: subset first, ship as `woff2`, split the subsets across `@font-face` rules with `unicode-range` if you genuinely serve multiple scripts, and pick weights deliberately. A team that skips it often ships several hundred kilobytes of Cyrillic and Greek outlines to an audience that reads only English — bytes on the critical path, competing with the content that actually gets rendered.

  • If subsetting already removes the unused glyphs, why still ship the file as woff2?
    They solve different problems. Subsetting decides which glyphs exist in the file; woff2 compresses whatever glyphs remain, using a font-aware transform on top of Brotli. A subset TrueType file is still far larger than the same subset as woff2, so you do both — subset in the build, emit woff2 as the output flavour.
  • How do you decide a subset is safe for a given block of text on the page?
    Split by content ownership. Text you author and control — nav labels, a display headline, marketing copy — can take an aggressive content-derived subset. Anything user-generated or translated needs a range-based subset broad enough for its locales, or it will show fallback characters mid-sentence. Many teams use a tight subset for the display face and a broader one for body text.
  • What breaks if the subset drops the OpenType layout tables?
    Kerning and ligatures stop applying, so pairs like "AV" or "Ta" lose their tightened spacing and "fi" renders as two separate glyphs. Nothing errors and the text is still legible — which is exactly why it slips through review. Keep at least the kern and liga features unless you have measured that dropping them matters.

saying these in an interview costs you the question

  • Thinks subsetting is just gzipping the font file
  • Believes CSS can hide unused glyphs at runtime
  • Assumes a missing glyph makes the whole paragraph fall back
  • Subsets user-generated content down to ASCII
  • Ignores that some font licences forbid modifying the binary

context

open as a page

In a CSS @font-face rule, what does the unicode-range descriptor do, and how does it change which font files the browser downloads?

level: middleimportance: should knowfreq 45%

basics

~20 s

The unicode-range descriptor declares which codepoints a font face covers. The browser downloads that face only if the page actually renders at least one character in the range, so a family split into per-script faces fetches only the scripts the page uses.

open as a page

A site self-hosts four static weights of one family (400, 500, 600, 700). When does replacing them with a single variable font actually reduce bytes, and when does it make things worse?

level: middleimportance: should knowfreq 40%

basics

~20 s

A variable font stores one set of outlines plus delta data describing how they change along an axis, so it beats three or four separate static weight files but loses to one or two. Using few weights, or shipping the full axis range when you need a slice of it, makes it heavier.

open as a page

Text visibly reflows when a webfont replaces the fallback. What do the CSS @font-face descriptors size-adjust, ascent-override and descent-override do about that, and what are you matching them against?

level: seniorimportance: should knowfreq 35%

basics

~20 s

They retune a fallback font's metrics to match the webfont's. size-adjust scales glyph outlines and advance widths so text occupies the same width; ascent-override and descent-override replace the metrics that set line-box height. Matched metrics mean the swap changes glyphs without moving anything.

open as a page

A team argues that loading their webfont from a public font service such as Google Fonts is faster than self-hosting the woff2. Which part of that argument stopped being true, and what does self-hosting change about the critical path?

level: seniorimportance: should knowfreq 45%

basics

~20 s

The shared-cache argument is dead: browsers now partition the HTTP cache by top-level site, so a font fetched on another site is refetched on yours. Self-hosting removes a third-party stylesheet-then-font request chain and two extra connection setups from the critical path.

open as a page