A vendor voids support if you install the EDR agent and the platform team refuses it on CPU grounds. How do you design around that?
answer
- the refusal is a constraint, not an argument to win
- make the device forward its own logs
- watch the path in, not the host
- sometimes remove trust instead of adding visibility
- named owner, review date, procurement criterion
basics
~20 sStop arguing about the agent and buy visibility elsewhere: mandatory log forwarding, network chokepoints, jump-host-only administration with recorded sessions, and segmentation. Then record the gap as an owned, dated risk acceptance and make agent support a procurement requirement.
solid answer
~50 sBoth refusals are usually legitimate, so treat them as constraints rather than obstruction. Reframe the ask from "install the agent" to "what signal do we get, and how small is the blast radius". Require the device to forward its own logs to your collector; an appliance that cannot is a procurement red flag. Force administration through a jump host with MFA and recorded sessions, so authentication to the device is observable even when the device is not. Place it where you own the flow records, and make sure it holds no reusable credential that reaches the rest of the estate. On the CPU objection, measure the real cost rather than accepting the number. Finally, write the residual gap down with a named owner and a review date, and make agent supportability a criterion at renewal, because that is where this actually gets fixed.
go deeper
Know that some hosts cannot take an agent for contractual or performance reasons, and that the answer is other sources of signal rather than sneaking software onto them.
Be able to list the compensating controls concretely: forwarded device logs, jump-host administration, flow collection on the segment, and no reusable credentials on the host.
Show you would test the performance claim rather than accept or dismiss it, and that you can state exactly what a flow record or a forwarded appliance log can and cannot establish later.
Own the whole call: who accepts the residual risk by name, how exceptions expire, how you stop one exemption becoming the fleet default, and how agent supportability and log export get into procurement and renewal criteria.
## Why this is a judgment question, not a technical one The technical answer to "can we put a sensor on it" is often simply no, and no amount of engineering changes a support contract, a certified image, or a latency budget on a trading or clinical platform. What separates a strong answer here is that it stops trying to win the argument and starts designing for the world in which the argument is lost. ## Take the objections seriously first A vendor support clause is a commercial reality: if you install anyway and the appliance misbehaves, you own the outage and the vendor walks away. A platform owner citing a CPU budget is often right, and sometimes measured. But "the agent costs four per cent" deserves a test rather than a shrug: measure on a representative host, with the policy you would actually deploy, and see whether a scoped profile or a narrow, well-owned exclusion brings it into range. Sometimes it does, and then you are trading a known blind spot for a smaller one, deliberately. Sometimes it does not, and you have a measurement instead of an assertion, which changes the conversation with the risk owner. ## What to buy instead of the agent Work through four layers, because no single one replaces endpoint telemetry: **Make the device talk.** Almost every appliance and hypervisor can forward its own logs: authentication to the management interface, configuration changes, shell or SSH being enabled, service restarts, account creation. Insist on this at deployment, and treat inability to forward structured logs as a procurement defect. This is the single highest-value substitute, because it is the only source that sees the device's own actions. **Own the path in.** Administration only from a jump host, with MFA at the identity provider and recorded sessions. Now the *authentication* and the *session* are observable even though the host is not, and you have an identity-side trail for every human who touched it. **Own the network around it.** Put it on a segment where you collect flow, so you can establish contact, direction and volume, and place egress through a proxy or firewall you log. Remember the limits you will be asked about: a flow record proves bytes moved between endpoints and carries no payload; TLS SNI and certificate subjects are visible to a passive observer while the body is not. **Shrink the blast radius.** The unagented host should hold no credential that is reusable elsewhere, should not be reachable from general user networks, and should not be able to reach the crown jewels directly. If you cannot see it, at least ensure that what happens on it does not immediately become an estate-wide problem. ## Some refusals are best answered by removing trust, not adding visibility For a contractor or BYOD laptop, chasing telemetry on a machine you do not own is usually the wrong fight. The better design removes the device from the trust path: access through VDI or an isolated browser, conditional access that requires a managed device for anything sensitive, and no direct network reachability. You accept that you cannot see the endpoint, and you make that acceptable by ensuring nothing important depends on it. ## The organisational half, which is the part people skip A gap you cannot close becomes defensible only if it is explicit. That means a written risk acceptance naming the specific hosts, the compensating controls actually built (not proposed), the person accountable, and a date when it is looked at again. It means agreeing who is paged when the compensating signal fires, since the SOC's usual endpoint playbook does not apply. It means keeping an exception register so that one appliance's exemption does not become the default answer for the next forty. And it means agent supportability, or at minimum structured log export, appearing as a scored requirement in procurement and renewal, which is where the population of unagentable devices is actually determined. ## Why this matters more than it looks The realistic outcome on these hosts is that nothing is ever found there first. An intruder resident on a hypervisor or an appliance tends to be discovered only when they reach back onto a monitored endpoint, and the earlier residence is then reconstructed after the fact. Whether that reconstruction is possible at all, and whether it takes an afternoon or is simply impossible, is decided long before the incident, by whether somebody made the device forward its logs and put its management path behind something you record. That is the argument to make to the risk owner: you are not asking for the agent any more, you are asking for the ability to answer the question later. ## What a weak answer sounds like Installing the agent covertly, escalating to an executive as the first move rather than after a design attempt, or accepting the refusal silently and letting the coverage story stay comfortable. All three end the same way: an unmonitored host with no compensating signal, and nobody who agreed to own it.
- The platform team quotes four per cent CPU. How do you respond without escalating?Measure it. Deploy the exact policy you intend on a representative host under production load and compare, then tune: a reduced profile, or one narrow and well-owned exclusion for the hot path, often brings it inside the budget. If it genuinely does not, you now have a measurement to take to the risk owner instead of two opinions.
- What makes an accepted coverage gap defensible a year later, in an audit or a post-incident review?That it was written down with the specific hosts named, that the compensating controls listed were actually built and verified, that a person was accountable by name, and that a review date came and was honoured. An acceptance with no owner, no expiry and aspirational controls is indistinguishable from having ignored it.
- Your appliance vendor cannot forward structured logs at all. What now?Treat it as a procurement defect and record it as such. In the meantime, compensate outside the box: log every authentication to its management interface at the jump host and identity provider, collect flow to and from its management segment, restrict administration to a tiny named group, and raise log export as a scored requirement at renewal.
saying these in an interview costs you the question
- Installs the agent anyway and hides the support risk
- Escalates to an executive before attempting a design
- Accepts the refusal with no compensating controls
- Treats a risk acceptance with no owner or date as closure
- Chases endpoint telemetry on a device the company does not own