Your logistics partner's webhook is trusted on a shared secret you cannot change — how do you model and bound a compromised partner?
answer
- assume the attacker holds the credential
- the envelope is not the claim
- their staging reaches your production
- cap and reconcile before anything pays
- an off switch you can pull alone
basics
~20 sPut the secret in the attacker's hands and re-read the diagram. A shared secret and an address allowlist authenticate a channel, not the truth of what it asserts, so treat partner events as claims and bound their financial damage.
solid answer
~50 sI model the partner as an adversary holding every credential the integration issued, because that is what compromised means, and because their staging reaching my production endpoint says a lower-assurance system already has it. The secret and the allowlist prove an event came through the partner's channel; they prove nothing about whether the delivery or damage claim it describes happened. So the impact is money leaving through claims and refunds, and audit truth corrupted with events my records will treat as fact. Since I cannot change the partner, I bound it on my side: partner-asserted events are claims, not facts; idempotency keys and a replay window; caps and reconciliation against an independent signal before anything pays out; alerting on volume and shape; separate credentials for their non-production traffic; and a rotation and kill path I can use without them. The residual risk is then written down and owned by name.
go deeper
Know that authenticating who sent a message is not the same as verifying that what it says is true, and that a shared secret can leak from either side of an integration.
Explain the mechanics you would add: timestamps and a replay window, idempotency keys, separate credentials per partner environment, and validation of the payload against your own state.
Show that you bound impact rather than chase assurance: pending states, value caps, reconciliation against an independent signal, and shape-based alerting, plus a rotation and disable path you can execute alone.
Own the decision. Argue what residual risk remains once your side is bounded, who accepts it by name, and what triggers a re-model — renewal, volume growth, or a disclosed incident at the partner.
## The move this question is testing Most integration threat models quietly assume the far side is honest and only ask whether an outsider can get in. The compromised-vendor scenario removes that assumption: you grant the partner exactly the trust the design gives them, then ask what an attacker in possession of that trust can accomplish. It is the vendor equivalent of assuming breach, and it is the difference between a model that produces reassurance and one that produces decisions. ## What the current design actually proves A shared secret plus an IP allowlist establishes, at best, that a request arrived through a channel the partner controls. It does not establish: - that the *event described* occurred — the payload is an assertion, and authentication of the sender never validates the claim; - that the sender was the partner's production system, when their staging shares the secret and the egress range; - that the message is fresh, unless you enforce timestamps and a replay window; - that the partner's own environment is uncompromised, which is precisely the thing under test. The first and second points are where teams lose the most. "The webhook is signed" is a statement about the envelope; the business logic behind it is what converts the envelope into a refund. ## Naming the impact in the right units Delivery and damage-claim events drive two things: **money** (claims paid, refunds issued, shipping costs credited) and **audit truth** (your record of what happened to a shipment, which downstream disputes and reporting rely on). In STRIDE terms the dominant categories are spoofing at the boundary and tampering with the state those events write, with repudiation close behind — once forged events land in the ledger, distinguishing real from fabricated after the fact may be impossible. Availability barely features, and that is worth saying out loud: a threat model that only produces "they could DoS us" has not looked at what the integration is *for*. An attack tree helps make the branches explicit: ``` Goal: extract money via forged claim events OR |-- obtain the shared secret | OR | |-- compromise partner production | |-- compromise partner staging (same secret) | +-- obtain it from a partner employee or config store +-- reach the endpoint from an allowed address AND |-- egress from the partner's address range +-- send a well-formed claim payload ``` At an OR node the attacker takes the cheapest child; at an AND node every child must be satisfied. Reading it that way shows the allowlist as one conjunct in a branch the attacker mostly already satisfies, not as an independent defence. ## Bounding it when you have no leverage The honest constraint in this scenario is that you cannot make the partner better. You can only change your side, and that is where the interesting principal judgment sits. The moves worth arguing for: - **Demote the event to a claim.** Nothing partner-asserted should mutate financial state directly. Land it in a pending state that a reconciliation or a second signal promotes. - **Reconcile against something the partner does not control** — your own scan data, carrier tracking you fetch yourself, or customer confirmation. Independence is the whole value; a second feed from the same partner is not one. - **Cap the blast radius numerically.** Per-event, per-hour and per-account limits on value, with anything above them queued for human review. A cap turns an unbounded loss into a bounded one and does not require the partner's cooperation. - **Enforce freshness and idempotency.** Timestamped, single-use identifiers with a narrow replay window remove trivially cheap branches of the tree. - **Separate their environments.** Insist on a distinct endpoint and credential for non-production traffic even if you cannot audit them; if they refuse, treat production as reachable from their staging and say so in the model. - **Own the off switch.** Rotation and disablement must be executable by you, alone, within minutes, and rehearsed. A control you can only exercise by filing a ticket with the party you are worried about is not a control. - **Detect on shape, not just volume.** Sudden changes in claim rate, value distribution, or the set of accounts touched are what a forged-event campaign looks like from your side. ## Where this stops Two adjacent activities are not this question. Assessing the partner against a control framework or sending them a questionnaire is procurement's risk process and does not change your diagram. Writing down "the partner protects the secret" as an inherited assumption is an artifact-hygiene practice worth doing, but the point here is the opposite instinct: model what happens when the assumption is false rather than filing it. ## Closing the loop as a lead The deliverable is a decision, not a list. Say which mitigations you are implementing, what residual risk remains after them, who accepts it by name, and what would trigger revisiting — a contract renewal, a change in claim volume, a security incident disclosed by the partner, or a material change in the integration. A threat model of a vendor integration that ends without an owner for the leftover risk has produced a document; one that ends with an owner has produced a decision.
- Would upgrading the shared secret to a per-message signature close this?It improves sender authentication and stops some tampering in transit, but the modeled attacker holds the signing key, so signed forged events are still forged. It is worth doing and it does not touch the dominant threat. The controls that do are the ones that decide whether a claimed event becomes money: reconciliation, caps and pending states.
- How do you decide how much of this to build when the partner is small and the integration is low volume?Size the controls to the loss, not the partner. Low volume with high per-claim value still deserves caps and reconciliation; low value and low volume may justify only idempotency, alerting and a rehearsed off switch. Write the reasoning down so the decision can be revisited when volume or value changes, because it will.
- The partner asks you to allowlist a broader address range for a migration. What do you say?That the allowlist was already a weak conjunct rather than a real defence, so widening it changes little in the model — but it is a signal to check the rest. I would take the change alongside a separate credential for the new environment, confirm non-production traffic is not on that range, and use the moment to get the rotation path tested.
- What does 'accepting' this residual risk actually require?A named owner with the authority to absorb the loss, a written statement of what is being accepted and how large it could be after the mitigations, and a trigger for revisiting — renewal, volume growth, or a disclosed incident. Acceptance without a name and a number is just an unexamined assumption wearing a process.
saying these in an interview costs you the question
- Says a shared secret plus IP allowlist makes the events trustworthy
- Treats a signed webhook as proof the event occurred
- Assumes the partner's staging cannot reach production
- Answers only with 'ask the vendor to fix it'
- Confuses contractual liability with a technical control
- Ends with a risk list and no named owner or trigger