Your organisation runs a commercial APM vendor's proprietary agent across its services and is considering moving instrumentation to OpenTelemetry. Argue both sides honestly, and describe how you would actually run the migration.
answer
- the real question: who owns the instrumentation contract
- OTel = portability, one vocabulary, pipeline-side redaction
- agent = depth out of the box, one supported thing, coupled features
- vendors ingest OTLP now → keep UI, own instrumentation
- fan out from the collector; time-boxed dual run with an exit date
basics
~20 sOpenTelemetry buys instrumentation you own and can point anywhere, with a shared semantic vocabulary; a proprietary agent buys depth, defaults and support that the vendor tunes end to end. Migrate incrementally per service through a collector that can fan out to both backends, with an agreed dual-run window, a comparison method, and a date to stop paying twice.
solid answer
~60 s**For OpenTelemetry**: instrumentation becomes an asset you own rather than a vendor artefact, so changing backend is a routing change, not a re-instrumentation project. One vocabulary (semantic conventions) across languages and signals. Community coverage of libraries the vendor may never prioritise. Negotiating leverage, and the ability to route different signals to different destinations. **For the proprietary agent**: it is usually deeper out of the box — code-level profiling, automatic baselining, database and error analytics wired to the vendor's UI — and it is one supported thing rather than an assembled pipeline you now operate. Vendor features often expect vendor-specific data; some will degrade on OTLP ingest. **Migration**: pick a representative service, run OTel alongside the agent, and export through a collector that fans out to the incumbent and the candidate. Compare on real questions — can I still answer the incident I answered last month — not on feature checklists. Watch for double instrumentation (two agents in one process), the cost of dual ingest, and semantic-convention churn. Set an exit date; a permanent dual-run is the expensive failure mode.
code
text · 10 linesservices (OTel SDK, OTLP out)
|
collector --fan-out--> incumbent vendor backend (existing alerts)
| \
| -------> candidate backend (being validated)
|
one place to change routing; applications untouched
exit condition: candidate answers the replayed incidents
-> stop the incumbent leg -> stop paying twicego deeper
Say OpenTelemetry is vendor-neutral instrumentation you own, while a proprietary agent is deeper out of the box but ties you to one vendor.
Add that most vendors accept OTLP now, so the realistic choice is owning instrumentation while keeping a vendor UI, and that migration means running both for a while.
Give a concrete plan: pilot service, collector fan-out, no stacked agents, comparison by replaying real incidents, and the artefact migration cost.
Frame it as ownership of the instrumentation contract, quantify dual-run and pipeline operating cost, name exit criteria and a date, and set a stable internal naming layer to absorb future convention churn.
## Frame the decision correctly The question is not 'which telemetry is better'. It is **where the instrumentation contract lives**. With a proprietary agent, the contract belongs to the vendor: the data shape, the field names, the extension points and the release cadence are theirs, and your dashboards, alerts and runbooks are written against their model. With OpenTelemetry, the contract is a public specification and the data is portable by construction. Everything else follows from that. ## The honest case for OpenTelemetry **Portability and leverage.** Applications emit OTLP; a collector routes it. Changing backend becomes a configuration change plus a dashboard/alert migration, rather than touching every service. That materially changes contract negotiations, and it changes what happens when a vendor's pricing model or product direction shifts. **One vocabulary.** Semantic conventions give the same attribute names for HTTP, database, messaging and infrastructure attributes across every language and every signal. Cross-language correlation stops being a per-team convention. **Coverage on your terms.** Community instrumentation covers a long tail the vendor may never prioritise, and where it does not, you write it yourself against a stable API — an option a closed agent does not offer. **Routing freedom.** Traces to one system, metrics to a metrics store you already run, logs to a cheaper archive tier, plus filtering and redaction in the pipeline before anything leaves your network. That last point — PII scrubbing under your control — is often the argument that lands with legal and security. ## The honest case for the proprietary agent **Depth per unit of effort.** Mature agents do things the open ecosystem does unevenly: continuous profiling correlated to spans, automatic dependency and baseline detection, error grouping, database statement analysis, real-user monitoring stitched to backend traces. Much of this exists in the OTel ecosystem, but assembled from parts you integrate and operate. **One supported thing.** With an agent you have a vendor to call. With OTel you own a collector fleet — its capacity, its failure modes, its upgrades. That is real, ongoing engineering cost and it must be counted, not waved away. **Feature coupling.** Some vendor features are driven by vendor-specific data the agent produces. Ingesting OTLP into the same vendor may light up most of the UI but leave specific capabilities degraded. Establish this empirically, per feature you actually use, before committing. **Spec churn.** Semantic conventions have moved — attribute renames and stabilisation waves are real, and they break dashboards and alerts written against older names. The ecosystem provides migration windows and dual-emission options, but you are now exposed to upstream change on a cadence you do not control. ## What has changed the calculus Most commercial vendors now ingest OTLP natively, either at their edge or through their own agent acting as an OTLP receiver. That collapses the question from 'OTel or vendor' into 'who owns the instrumentation' — you can keep the vendor's UI and analytics while owning the instrumentation. That is usually the pragmatic landing spot and it should be stated explicitly, because candidates often frame this as all-or-nothing. One technical wrinkle: temporality preferences differ. Several commercial backends prefer delta metrics while Prometheus-style stores require cumulative, so a fan-out to both may need conversion or separate pipelines — a concrete cost of dual-running. ## Running the migration **1. Decide the target state on paper.** Which signals move, which backend each lands in, who operates the pipeline, and what the exit criteria are. Without this the effort becomes an unbounded dual-run. **2. Pick a representative pilot** — a service with real traffic, real dependencies and a real on-call rotation. Not a toy. **3. Run both, fan out from one point.** Instrument with OTel, export OTLP to a collector, and have the collector send to both the incumbent and the candidate. One place to change routing; the application stays out of it. **4. Beware double instrumentation.** Two agents in one process is a genuine hazard: duplicated spans, conflicting context propagation, doubled overhead, occasionally instability. Check whether the vendor's agent has a supported coexistence mode, and if not, alternate rather than overlap in-process. **5. Compare on questions, not features.** Replay a past incident: can you answer it in the new stack, at the same speed, with the same confidence? Check the specific things people rely on — a saved query, an alert that catches a known failure mode, a drill-down someone uses weekly. **6. Budget the transitional cost.** Dual ingest means paying twice for the pilot's volume, plus egress and collector capacity. Make it visible and time-boxed. **7. Migrate the artefacts, not just the data.** Dashboards, alerts, runbooks and saved queries are where the effort actually lives, and attribute renames make it worse. Plan for it explicitly, and consider defining alerts against a stable internal naming layer so future churn is absorbed in one place. **8. Roll out in waves with a stop condition,** and name the date the dual-run ends. Permanent dual-running is the classic failure: all the cost of the migration and none of the savings. ## Interview framing The grader is looking for whether you can argue against your own preference. Give the vendor's case with real substance, name the collapse of the dichotomy (vendors accept OTLP now), then a migration plan with a fan-out point, a dual-run budget, an artefact-migration line item, and an exit date.
- What is the strongest argument for keeping the proprietary agent?Depth per unit of effort plus a single support relationship. Mature agents ship correlated profiling, automatic baselining and error analytics tuned end to end, and the vendor operates the pipeline. Adopting OpenTelemetry means you now run a collector fleet with its own capacity planning, failure modes and upgrade cadence — a real, permanent engineering cost that has to be counted against the portability benefit.
- Why is running both agents in the same process dangerous?Two instrumentation layers can each wrap the same libraries, producing duplicated spans and inconsistent parenting, and they may disagree about context propagation formats so downstream services see conflicting headers. Overhead roughly doubles on the hot paths. Unless the vendor documents a supported coexistence mode, compare by alternating deployments or by fanning out from the collector rather than by stacking agents.
- Where does most of the migration effort actually go?Not into instrumenting services — auto-instrumentation covers much of that — but into the artefacts built on top of the old data model: dashboards, alert rules, saved queries and runbooks, all written against vendor field names. Semantic-convention renames compound it. Defining alerts against a stable internal naming layer limits how often that work has to be repeated.
- How do you know the migration is finished?When the new stack answers the questions the old one answered — verified by replaying past incidents and by on-call using it during a real one — and when the incumbent has been switched off so you stop paying twice. An explicit exit date in the plan is what prevents the permanent dual-run, which costs more than either option alone.
saying these in an interview costs you the question
- Presenting it as all-or-nothing when most vendors now ingest OTLP directly.
- Ignoring the operating cost of a collector fleet you now own.
- Running two in-process agents simultaneously and expecting clean data.
- Budgeting only the instrumentation work and not the dashboard/alert/runbook migration.
- Starting a dual-run with no exit date, so the organisation pays for both indefinitely.
- Assuming semantic conventions are frozen — attribute renames have broken dashboards before.