skip to content

What do lazy-pull container image formats such as eStargz and SOCI change about container start-up?

level: seniorimportance: nice to knowfreq 17%

answer

  1. gzip is not seekable
  2. containers read little of what they carry
  3. one approach rewrites the layer, one adds an index
  4. a plugin below the runtime does the fetching
  5. the registry stays on the critical path afterwards

basics

~20 s

They make image layers seekable, so a container starts before its image has finished transferring and file content is fetched on demand. Both need a containerd snapshotter on the node, and they trade a faster start for a live registry dependency while running.

solid answer

~50 s

A normal OCI layer is a gzipped tar: not seekable, so the runtime must fetch and decompress the whole blob before any file inside it exists on disk. Yet a container typically reads only a small fraction of its image before it is serving. **eStargz** rebuilds each layer as a seekable, gzip-compatible archive with a table of contents (and an optional prioritised-file set that is prefetched); **SOCI** leaves the image bytes untouched and publishes a separate index artefact in the registry that maps files to offsets. Either way a compatible **containerd snapshotter** — stargz-snapshotter or the SOCI snapshotter — mounts the layer immediately and pulls file ranges on demand, so start-up overlaps with transfer. The classic dockerd graph driver does not do this. The costs are real: the container now depends on the registry after it has started, first access to a cold file pays latency, and you add a conversion or index-building step to the build pipeline.

code

bash · 3 lines
bash
docker buildx build \
  --output type=image,name=registry.example.com/tileserver:1.9,push=true,oci-mediatypes=true,compression=estargz \
  .

go deeper

for a junior

You are not expected to know these formats. Knowing that a normal pull must complete before the container starts is the useful takeaway at this level.

for a middle

Be able to explain why a gzip layer cannot be read partway through, and that lazy pulling needs support below the CLI rather than a Dockerfile change.

for a senior

Show that you weigh it as a tradeoff, not a free win: which workloads have a small hot set, what the registry dependency does to your failure model, and what the conversion or indexing step costs the pipeline.

for a principal

Own the decision at fleet level: whether the profile justifies a runtime-stack change on every node, how it interacts with registry availability targets, and whether pre-warming solves the same problem with less coupling.

### The problem with a tar.gz layer An ordinary image layer is a tar stream compressed with gzip. Gzip is a single continuous stream, so you cannot jump to the bytes of one file inside it: to read anything, you decompress everything before it. In practice that means a runtime must download an entire layer blob, decompress it, and unpack it before the container can open a single file from that layer. For a 1.87 GB image on a fresh host, that is minutes of work before the first line of application code runs. The waste is that a container reads very little of its image at start-up. Measurements of real workloads have repeatedly put the actively read fraction in the single-digit percentages — a Kotlin geospatial tile server needs its JRE, a handful of native libraries and its own jar, not the several hundred megabytes of baked tiles it will touch lazily, if at all. ### The two approaches **eStargz** (from the stargz-snapshotter project) changes the layer format. The layer is rebuilt so each file is compressed independently and a table of contents is appended, making the archive *seekable* while remaining a valid gzip stream — an ordinary client that knows nothing about eStargz can still pull the image normally; it is simply slightly larger. eStargz also supports a *prioritised files* set, recorded at build time, which the snapshotter prefetches in one contiguous read so the hottest files are local before the process asks for them. **SOCI** (Seekable OCI) leaves the image completely alone. It builds a separate index artefact — stored in the registry alongside the image and referring to it — that records where each file lives inside the existing gzip layers. Nothing about the image digest changes, which means an image you did not build, and cannot rebuild, can still be lazily pulled. ### What actually does the work Neither format is self-executing. Both need a **containerd snapshotter plugin** on the node: stargz-snapshotter for eStargz, the SOCI snapshotter for SOCI. The snapshotter presents the layer as a filesystem immediately, backed by a network-backed FUSE mount, and fetches byte ranges from the registry — using HTTP range requests — the first time a file region is read, caching what it fetched locally. The container therefore starts while transfer is still in progress. This is a runtime-stack property, not a Dockerfile property. The classic dockerd storage path materialises complete layers, so building an eStargz image and pulling it with an ordinary `docker pull` gets you a normal, complete pull. BuildKit can produce the format: ``` docker buildx build \ --output type=image,name=registry.example.com/tileserver:1.9,push=true,oci-mediatypes=true,compression=estargz \ . ``` ### What you give up - **A live registry dependency.** With a conventional pull, once the container is running the registry can disappear and nothing breaks. With lazy pulling, a file first touched twenty minutes in still has to be fetched. A registry outage or a network partition then reaches into running workloads, which is a materially different failure model. - **Tail latency.** The first read of a cold region pays a network round trip. A request that happens to touch a cold code path can be dramatically slower than its neighbours, which shows up as a start-up latency tail rather than a clean improvement. - **Pipeline cost.** eStargz requires producing the format (and its images are somewhat larger); SOCI requires generating and pushing an index for every image. Either is a step someone has to own and keep working. - **Node prerequisites.** A snapshotter plugin must be installed and configured on every node, and the fleet must agree — a mixed fleet where half the nodes lack it silently falls back to a full pull, which is fine but makes benchmarks confusing. - **Limited upside for some workloads.** If the process genuinely reads most of the image at start — a runtime that memory-maps a large model or dataset, for example — lazy pulling only moves the wait, it does not remove it. ### When it is worth it The profile that pays is a large image with a small hot set, on hosts created frequently enough that cold starts dominate, where the registry is close and reliable. That is a real profile — machine-learning serving images and fat JVM images are the classic cases — but it is a narrower one than the headline benchmark numbers suggest. For most fleets, getting the image onto the node before it is needed is a simpler answer with no runtime coupling, and lazy pulling is the option you reach for when that is not possible.

  • Can an ordinary client pull an eStargz image if the node has no lazy-pull support?
    Yes. eStargz layers are still valid gzip tar streams, so a client that knows nothing about the format pulls and extracts them normally; the only penalty is a slightly larger image. That backwards compatibility is a deliberate design goal and is why converting an image to eStargz is comparatively low risk for a mixed fleet.
  • What is the main operational risk of lazy pulling that a conventional pull does not have?
    The running container keeps a dependency on the registry. With a conventional pull all bytes are local once the pull finishes, so the registry can fail without touching running workloads. With lazy pulling, a file first read long after start still triggers a fetch, so registry or network trouble becomes a runtime fault, and cold reads add a latency tail.
  • How does SOCI differ from eStargz in what it asks of your build pipeline?
    eStargz changes the layer format, so the image must be produced or converted in that form and its digest differs from the original. SOCI leaves the image bytes untouched and publishes a separate index artefact in the registry that maps files to offsets in the existing layers, so it can be applied to images you did not build and cannot rebuild.

It is the difference between shipping a whole library and shipping a catalogue: the reader starts with the catalogue and the couriers keep delivering individual books as pages are turned — fast to begin, but the courier must stay reachable.

saying these in an interview costs you the question

  • Thinks lazy pulling works with any runtime out of the box
  • Believes a normal gzip layer can be read at an offset
  • Assumes eStargz images break ordinary clients
  • Ignores the registry dependency after container start
  • Expects a win for workloads that read the whole image
  • Confuses it with simply making the image smaller

context