skip to content

A support vendor's chat and session-replay script runs on your logged-in pages — what data actually leaves your boundary?

level: seniorimportance: should knowfreq 45%

answer

  1. the script runs in your origin
  2. not the payload you designed
  3. who at the vendor can open a replay
  4. it can write to the page, not only read
  5. iframe on their origin changes the grant

basics

~10 s

Whatever the authenticated page contains: form fields, tokens in the DOM, personal data, anything typed. Script you include in your page runs with your page's privileges, so the vendor receives flows you never designed.

solid answer

~50 s

The mistake is modeling this as "an API call to the vendor". A script included first-party executes inside your origin with full access to the DOM, storage and input events on an authenticated page, so the outbound flow is not the payload you specified — it is everything the page holds while a customer uses it, including card fields, tokens rendered into the page and free text they type. I model two elements: a flow from the browser to the vendor carrying uncontrolled content, and a vendor-side store whose console lets any of their agents open any replay, which puts people I do not employ in the reader set. Because the script writes as well as reads, a compromised vendor gets tampering and credential capture inside my authenticated pages. So: keep it off the highest-value pages, deny by default at capture, require named-agent access to replays.

go deeper

for a junior

Know that a script tag you add to your page runs with the same access as your own code, so a vendor widget on a logged-in page can see what the user sees and types.

for a middle

Explain the mechanics: DOM and input access inside your origin, capture sent to the vendor, replays stored and viewable by their staff. Name the categories — disclosure, tampering, credential capture.

for a senior

Demonstrate the operating call: keep DOM-reading vendor code off authentication, payment and settings pages, insist on deny-by-default capture, and ask who at the vendor can open an arbitrary customer's replay.

for a principal

Own the trade: a support or analytics capability bought at the price of a standing read-and-write grant inside authenticated sessions. Set the policy for which page classes may host third-party code at all, and make exceptions decisions with named owners.

## Why this integration is different from an API integration Most vendor integrations are server-to-server: you decide the fields, you build the request, and the diagram shows one flow with a known payload. A widget delivered as a script tag inverts that. You are not sending the vendor data; you are inviting the vendor's code into the page where your customer is authenticated, and it decides what to collect. On a data-flow diagram the honest depiction is not one labelled arrow — it is a process you do not control, running inside the client, with a flow out of it whose contents you have not enumerated. ## What "first-party inclusion" actually grants A script loaded into your page's origin can, by design of the web platform: - read and modify the entire DOM, including values rendered into forms and pages; - observe keystrokes, clicks, focus and form input as the customer types; - read storage and any token the page keeps where script can reach it; - issue requests as the page, and change what the page displays. Session replay makes deliberate use of most of that: reconstructing a session means capturing the DOM and the input stream. So the capability you granted and the capability the product needs are close to identical, which is why "we only use it for support" is not a boundary — it is a usage intention on top of an unrestricted grant. ## The two flows worth drawing **Flow one: browser to vendor.** Content is whatever the page held. On a settings or checkout page that includes address, payment fields, and sometimes values that were never meant to leave, such as an identifier a customer pasted in. Modeling it as "session metadata" is where teams go wrong. **Flow two: vendor store to vendor staff.** The replay is not just stored, it is *watchable*, and typically by any agent in the vendor's console. A threat model that stops at "the vendor holds the data" misses that the practical adversary here is often a single rogue or socially engineered agent, not a breach of the vendor. Ask who at the vendor can open a replay of an arbitrary customer, whether that access is per-agent identified, and whether it is logged in a way you can ever review. ## Threat categories in play - **Information disclosure**: the headline. Customer data and any credential-like material visible in the page cross a boundary continuously, at the volume of your traffic. - **Tampering**: the script can alter the rendered page. A compromised vendor script can change displayed values, inject fields, or silently modify a form before submission. - **Spoofing / credential capture**: input capture on an authenticated page reaches passwords typed into re-authentication prompts and one-time codes, which is a path to account takeover rather than mere data leakage. - **Elevation of privilege**: not in your servers, but inside the customer's session, which is the privilege that matters to them. ## The same pattern, a different vendor The managed observability vendor is this problem with the arrow pointing out of the backend instead of the browser. Nobody decided to send request bodies and authorization headers to a third party; the logging call sites simply dumped what they had, and the pipeline forwarded it. In both cases the defining property is the same: **data leaves your boundary because of how a mechanism works, not because anyone specified that flow.** That is the class of threat this leaf exists to make you look for — and the question to ask of any vendor integration is not only "what did we agree to send?" but "what does this mechanism carry by default?" ## Controls that follow from the model - **Placement.** Do not load a DOM-reading vendor script on authentication, payment or account-settings pages. This costs product capability and is usually the single largest risk reduction available. - **Deny by default at capture.** Configure the capture to record nothing sensitive unless explicitly allowed, rather than masking a list of known-bad selectors. Understand that this configuration runs inside a page the vendor's code also controls, so it reduces accidental collection well and reduces deliberate collection much less. - **Isolate where the feature allows.** If the vendor supports rendering in an iframe on their own origin, the same-origin policy removes their DOM access outright, and the flow becomes only what you explicitly send in. Many chat products work this way; replay products generally cannot, which is itself a useful signal about the trust each one requires. - **Constrain the reader set.** Require per-agent identity, need-based access to replays, retention limits, and access logs you can request. - **Content controls on your side.** Keep tokens out of DOM-reachable places, and stop rendering full values you do not need on screen. ## The answer that lands in an interview State the grant plainly — first-party script means full access to the authenticated page — then name both flows, then explain that the compromised or rogue-agent case is the *modeled* threat rather than a footnote. Finish on the control that a strong candidate reaches for first: not a longer masking list, but keeping the vendor's code off the pages where the value is.

  • The vendor says replays are masked by default. What do you check before accepting that?
    What the masking actually covers — typed input as well as rendered values, and network payloads if the product captures them — and whether it is deny-by-default or a list of known-sensitive selectors that new pages will not match. Then check who can change it: masking configured in a console the vendor controls, executing in code the vendor ships, is a control you do not hold.
  • Same question for a managed observability vendor receiving every log line. What is the analogous threat?
    Request bodies and authorization headers ride out unscrubbed because nobody decided they should — you end up operating an offsite copy of production data and credentials. Model the vendor as a reader with staff access, scrub at the call site rather than in the pipeline, and treat access to that console as access to production.
  • How does the model change if the widget loads in an iframe from the vendor's own origin?
    The same-origin policy removes their direct read and write access to your DOM, so the flow shrinks to the messages you deliberately pass in. The uncontrolled outbound flow becomes a designed one. The trade-off is capability: a vendor that insists on first-party inclusion is asking for a far larger trust grant, and that should be an explicit decision.

Embedding a vendor's script is less like mailing them a report and more like handing a contractor a badge that opens every room your customer is standing in.

saying these in an interview costs you the question

  • Assumes the vendor only receives what your API sends them
  • Calls an embedded widget out of scope for the threat model
  • Treats vendor-configured masking as a control you own
  • Ignores that the script can modify the page and capture input
  • Confuses a retention or deletion promise with a technical restriction

context