What creates instrumentation lock-in across a service fleet, and what would you standardise to reduce it?
answer
- The emitting code is the cheap part
- Count what your queries hardcode
- Dashboards and alert rules outlive backends
- One spelling per concept, fleet-wide
- Short lists get followed, long ones do not
basics
~20 sLock-in lives in what was built on top of the telemetry, not in the code that emits it. Changing the emitting mechanism is a rollout; re-authoring every dashboard and alert rule that hardcodes an attribute name is the real cost.
solid answer
~50 sAn instrumentation estate has three layers with wildly different migration costs. **Emission** - the agent or the calls in your code - is a bounded redeploy. **Transport and pipeline** is a re-point or a translation, owned by one team. **Derived assets** - dashboards, alert rules, saved searches, runbooks - are unbounded, spread across every team, and nobody has an inventory of them. Each of those assets is a query, and every query hardcodes span names, metric names and attribute keys. So the portable asset is the **vocabulary**, not the mechanism: if three services spell the same outcome three ways, no fleet-wide rule can cover them and any move must be renegotiated service by service. Standardise service identity, operation naming, outcome, the two or three business dimensions everyone slices by, and units. Leave mechanism choice free - it buys far less portability than agreeing on names.
code
text · 5 linesservice-a span: HTTP GET /menus/{id} attrs: http.status_code=200 tenant.id=...
service-b span: get_menu attrs: httpStatus=200 school=...
service-c span: MenusController.get attrs: resp_code=200 org_id=...
no single alert rule or fleet-wide query can cover all threego deeper
Recall that dashboards and alert rules are written against the names your telemetry emits, so changing a span name or an attribute key can quietly break somebody's alert. Names are shared, not private to your service.
Explain why a consistent vocabulary matters: one query can only cover services that spell a concept the same way. Be able to give an example of three spellings of the same outcome and what that costs.
Show that you cost a migration by its derived assets rather than by its redeploys, and that you can name the handful of keys worth standardising first - identity, operation, outcome, the shared business dimensions, units.
Own the strategy: make the standard the default path via shared defaults and drift checks, keep the list short enough to be followed, decide deliberately where lock-in is an acceptable trade, and be able to state what evidence you could still produce after a backend change.
## Three layers, three very different costs "Lock-in" is usually discussed as though it lived in the code that emits telemetry. It rarely does. An instrumentation estate has three layers, and their migration costs differ by orders of magnitude. | Layer | What it is | Cost to move | | --- | --- | --- | | Emission | The agent attached, or the calls in your code | A redeploy per service — bounded, mechanical | | Transport and pipeline | The wire format and whatever sits between apps and storage | A translation or a re-point — bounded, one team | | Derived assets | Dashboards, alert rules, saved searches, runbook links, reports | Unbounded, spread across every team, and nobody has an inventory | Almost everyone underestimates the third row, because it is the only one not represented in a repository anybody owns. The dashboards were built over four years by people who have moved on. The alert rules encode judgement nobody wrote down. Each of them hardcodes span names, metric names and attribute keys — and that, not the agent binary, is what actually holds you in place. ## Why the vocabulary is the portable asset Every derived asset is a query, and every query names things. Which means the property that makes an estate movable is not which mechanism produced the telemetry but whether the same concept is spelled the same way everywhere. Consider three services in one fleet describing the same HTTP outcome as `http.status_code`, `httpStatus` and `resp_code`, and the same customer as `tenant.id`, `school` and `org_id`. Nothing is broken. Each service's own dashboard works. But no single alert rule covers all three, no fleet-wide query is possible, and a migration must be negotiated service by service because there is no shared thing to port. Now invert it. If the fleet spells each concept one way, the derived assets are portable in the only sense that matters: they must be re-authored in the new query surface, but they are re-authored against the same names, mechanically, by one team, in a known amount of time. **A consistent vocabulary converts an unbounded organisational migration into a bounded engineering task.** That is why naming discipline is worth more than mechanism choice, and it is the answer interviewers are listening for. ## What to standardise — and what to leave alone Standardise: 1. **Service identity** — one spelling of what a service is called, and it must match what deploys and pages call it. 2. **Operation naming** — low-cardinality names for the units of work, so a name is a groupable thing rather than a unique string per request. 3. **Outcome** — one key for status and one convention for what counts as an error, because every alert rule in the estate reads it. 4. **The two or three business dimensions everyone slices by** — tenant, region, plan — spelled identically in every service. 5. **Units** — pick milliseconds or seconds and never mix them within one concept. Leave alone: which mechanism each service uses, and the internals of any one service's own spans. A batch job in one language and a request service in another can be instrumented completely differently and still be perfectly portable, as long as they agree on the five things above. Standardising mechanism is the seductive wrong answer — it is highly visible, it generates a long migration, and it buys much less portability than agreeing on names. ## Making the vocabulary hold Publishing a naming document changes nothing on its own. What works is making the standard the path of least resistance and then checking it: - Ship a shared library or base configuration that emits the agreed keys by default, so the standard is what you get by doing nothing. - Do the rewriting in one place where possible, so a service that got a key wrong is corrected without a fleet-wide code change. - Check emitted names against the agreed list automatically and report drift to the owning team, rather than discovering it mid-incident. - Keep the list short. Twelve keys get followed; two hundred is a document. ## The evidence test A useful way to size real portability: suppose a regulator asks for six months of evidence about a service's behaviour, and the backend changed three months ago. What has to be true? The **names** must be identical either side of the change, or the two halves cannot be compared at all. The **data** must be exportable in a form readable without the vendor's own interface. And **retention** must be a decision you own rather than one you inherit from whatever tier you happen to be on. If any of those three is false, the answer is that the evidence does not exist — and that is a lock-in cost measured in obligations, not in engineering days. ## Where accepting lock-in is rational Not all of it is worth fighting. The query and visualisation surface is where a product's value actually lives, and re-authoring queries is bounded work once names are stable, so binding to one is a reasonable trade. What is not reasonable is letting the emission vocabulary fragment: that cost is unbounded, paid by every team rather than one, and it grows every quarter you defer it. Standardise the names, stay relaxed about the mechanism, and treat the derived assets as an inventory you must be able to count.
- How would you measure, today, how locked in you actually are?Inventory the derived assets - dashboards, alert rules, saved searches, runbook links - and count, per concept, how many distinct spellings appear across them. That number is the real portability metric: it predicts how much re-authoring a backend change costs and how much of it can be done mechanically rather than negotiated team by team.
- A regulator asks for six months of evidence spanning a backend change three months ago. What has to be true?The names must be identical either side of the change, or the two halves cannot be compared. The data must be exportable in a form readable outside the vendor's interface. And retention must be a decision you own rather than one inherited from whatever tier you are on. If any fails, the evidence effectively does not exist.
- Where is it rational to accept lock-in rather than fight it?In the query and visualisation surface - that is where a product's value actually lives, and re-authoring queries is bounded work once names are stable. What is never rational is letting the emission vocabulary fragment, because that cost is unbounded, is paid by every team rather than one, and grows every quarter it is deferred.
Changing telemetry backends is like changing filing cabinets. The cabinet is cheap; what hurts is that every index card, cross-reference and instruction sheet you ever wrote refers to the old labels.
saying these in an interview costs you the question
- Equates lock-in with the vendor's agent binary alone
- Thinks a shared wire format by itself makes telemetry portable
- Omits dashboards and alert rules when costing a migration
- Lets each team invent its own attribute spellings
- Assumes renaming a key is free because queries can be edited
- Standardises the mechanism instead of the vocabulary