A service scans every class in its deployment unit at startup to find marked components — what does a build-time index of those markers change?
answer
- cost scales with declarations, not matches
- paid on every process start
- move the search into the build
- read a list, then resolve entries
- an index can disagree with the code
basics
~20 sIt moves the search off the startup path. The build records which declarations carry the marker, and the booting process reads that list instead of inspecting every class. Startup gets cheaper, and a stale index becomes a new failure mode.
solid answer
~50 sThe cost of a startup scan is proportional to the number of declarations in the deployment unit, not to the number that turn out to be marked: every declaration must be inspected before it is known to be uninteresting, and that grows with every dependency pulled in. A build-time index shifts that work to a moment where much of it is already being done — the build walks the declarations anyway — and records the marked ones as plain data inside the artifact. At boot the consumer reads a short list and resolves only those entries. What it buys is a startup cost proportional to the marked set. What it costs is a second source of truth: the index must be regenerated by the same build that compiled the code, and anything contributed after the build is not in it.
code
json · 24 lines{
"indexVersion": 1,
"builtFrom": "unit-2f19ac",
"marked": [
{
"marker": "Scheduled",
"targetKind": "type",
"target": "reports.ReportRefresher",
"members": { "everySeconds": 30 }
},
{
"marker": "Scheduled",
"targetKind": "type",
"target": "billing.InvoiceSweeper",
"members": { "everySeconds": 300 }
},
{
"marker": "Validated",
"targetKind": "member",
"target": "billing.InvoiceSweeper.sweep",
"members": {}
}
]
}go deeper
Hold on to the shape: finding marked components at boot means inspecting everything to find a few, and that work can instead be recorded once by the build.
Explain why the cost tracks the total number of declarations rather than the marked ones, and what the build has already done that makes recording them nearly free.
Show the operational judgment: measure the startup walk, decide whether short-lived instances make it matter, and name the staleness failure the index introduces and how you fail loudly on it.
The decision is whether the platform standardises on a build-produced index for every service — buying startup latency at the price of a second source of truth and a build step every team must keep correct.
## What the startup scan actually costs Suppose a deployment unit contains forty thousand declarations and sixty of them carry the marker a consumer cares about. The scan's cost tracks the **forty thousand**, not the sixty, because a declaration must be opened and inspected before anyone can know it is unmarked. The work per declaration is small, but it is repeated across everything the unit contains — including all the code pulled in by dependencies that the team never reads. Three properties make this worse than it first sounds: - It is paid on **every process start**, not once per deployment. A service running many short-lived instances, or scaled up and down frequently, pays it constantly. - It grows with **dependencies**, which grow on their own over time and are nobody's line item. - It happens at the worst moment: before the process can serve anything, so the cost lands directly on startup latency and on how fast a deployment or a scale-up can absorb traffic. ## What an index replaces The insight is that the search has already been performed once, by the build. The build parses every declaration to compile it, so it can record the marked ones for almost nothing extra, and write that record into the artifact as data: 1. The build walks the declarations it compiles, as it must anyway. 2. For each marked declaration it records what a consumer would need: which marker, which kind of declaration, how to name the target, and the member values. 3. It writes those entries as an ordered data file placed inside the artifact. 4. At boot, the consumer reads the file and resolves only the entries listed. | | Startup scan | Build-time index | |---|---|---| | When the search happens | every process start | once, during the build | | Cost scales with | every declaration in the unit | only the marked declarations | | Sees declarations added after the build | yes, if they are in its candidate set | no | | Failure when wrong | slower boot, or a missed candidate set | entries that do not match the code | The change is therefore not "startup becomes free" — the consumer still resolves every entry it reads, and that cost scales with the marked set. The change is that the *search* is gone from the startup path. ## What the index cannot cover An index is a statement about the code as it was at build time, and that is exactly its limit: - **Anything contributed later is invisible.** A component supplied by a plug-in dropped in after the build, or generated while the process runs, is not in a list written before it existed. - **Anything outside what the build saw** is missing for the same reason as before — an index inherits the candidate-set problem rather than curing it. - **Conditional markers.** If what counts as marked depends on configuration resolved at boot, the index can only record the raw facts and leave the decision to the consumer. Systems that need both usually keep the index as the fast path and allow an explicit, narrow scan of a declared extension area, so the expensive walk covers a small region instead of the whole unit. ## Keeping the index honest The index introduces a failure the scan did not have: it can **disagree with the code it claims to describe**. A scan can be pointed at the wrong place, but it cannot be out of date, because it reads what is actually there. Practices that keep an index trustworthy: - Generate it **in the same build step** that compiles the code, never as a separate job and never by hand. - Treat it as a **build output, not a source file** — an index that is edited or committed will eventually describe code that no longer exists. - **Fail loudly on an entry that does not resolve** at boot. An index naming a declaration that is not present means the artifact and the index came from different states, and continuing quietly turns a build problem into a mystery in production. - Record **what the index was built from**, so a mismatch can be diagnosed rather than guessed at. ## When scanning is still the right answer A scan is simpler, self-correcting and always current, and for a small unit, a long-lived process, or a service that starts a handful of times a day, it costs nothing anyone can measure. The index earns its keep where startup latency is a real constraint — short-lived instances, aggressive scaling, cold starts on the request path — or where the unit has grown large enough that the walk is measured in seconds rather than milliseconds. Reach for it when the number is actually on a graph, not on the assumption that a scan must be slow.
- Why does the scan cost hurt more as a service runs on more, shorter-lived instances?Because the scan is paid per process start, not per deployment. One long-lived instance amortises it over days; many instances that start, serve briefly and exit pay the full walk every time, and the cost lands on how quickly new capacity can take traffic.
- What stops a build-time index from drifting away from the code it describes?Generating it inside the same build that compiles the code, never editing or committing it, and failing at boot on any entry that does not resolve. Recording what the index was built from turns a silent mismatch into a diagnosable one.
- Can a system use an index and still support components added after the build?Yes, by keeping the index as the fast path for everything known at build time and scanning only a small, explicitly declared extension area at boot. The expensive walk then covers a narrow region instead of the whole unit.
saying these in an interview costs you the question
- Thinks the scan cost scales with the number of marked classes
- Treats startup scanning as free because it happens once per start
- Believes a committed index stays valid as the code changes
- Says an index makes late-contributed components discoverable
- Ignores that the index and the code can describe different states
- Reaches for an index before startup time is actually measured