skip to content

How would you scope a Playwright recordHar capture before those archives are committed to a repository?

level: principalimportance: should knowfreq 28%

answer

  1. Filter at capture, not afterwards
  2. Smallest file that still replays
  3. One artefact, not scattered attachments
  4. Recordings carry headers and keys
  5. Every archive needs a named owner

basics

~20 s

Decide before recording what belongs in the file: narrow urlFilter to the endpoints under test, pick minimal mode and a zip path so archives stay small and self-contained, and treat everything captured as permanently readable repository data.

solid answer

~50 s

Recording is cheap, so the discipline has to come from the capture settings rather than from cleanup afterwards. Narrow `recordHar.urlFilter` to the endpoints the tests actually exercise — `'**/api/forecast**'` rather than the whole dashboard — which keeps images, fonts and analytics out of the file. Use `mode: 'minimal'` when the archive exists only to be replayed, and a `.zip` path so the HAR and its attachments travel as one reviewable artefact. Then treat the file as published data: it records full request URLs, headers and cookies, so an API key in a query string or an `Authorization` header for the forecast provider is now in your history. Record against a sandbox credential, keep authenticated third-party calls out of the capture, and name an owner for each archive so refreshing it is somebody's job rather than nobody's.

code

typescript · 15 lines
typescript
import { defineConfig } from '@playwright/test';

export default defineConfig({
  use: {
    baseURL: 'https://dashboard.example',
    contextOptions: {
      recordHar: {
        path: 'har/weather.har.zip',
        urlFilter: '**/api/forecast**',
        mode: 'minimal',
        content: 'attach',
      },
    },
  },
});

go deeper

for a junior

Know that recordHar takes a urlFilter and that you should point it at the API under test rather than recording the whole site. Do not commit an archive you have not opened.

for a middle

Explain what each capture setting changes in the file: what urlFilter excludes, what minimal mode drops, and why a zip path keeps the archive and its attachments together.

for a senior

Weigh the review and secret-handling consequences of a capture, and set up the refresh path so archives stay small, self-contained and readable in a diff.

for a principal

Own the standard across teams: which credentials may be recorded, how large an archive may be, who refreshes it, and what review a new or updated archive must pass before it lands.

## Decide what goes in before you record An archive is easy to produce and hard to unpublish. Every decision worth making happens at capture time, because after the fact you are editing a JSON blob by hand or re-recording anyway. Three questions settle the shape of the file: 1. **Which endpoints does the suite actually replay?** Only those belong in the archive. For a weather dashboard that is the forecast API, not the page HTML, the map tiles or the analytics beacon. 2. **Will anything but Playwright read this file?** If not, the archive can drop everything replay does not consult. 3. **What does the traffic carry?** Credentials, tokens and personal data in a recording become repository content with a very long life. ## The knobs and what each one buys | setting | choice | why | |---|---|---| | `recordHar.urlFilter` | `'**/api/forecast**'` | keeps assets, fonts and telemetry out of the file entirely | | `recordHar.mode` | `'minimal'` | drops timings, sizes, cookies and the page section replay never reads | | `recordHar.content` | `'attach'` | bodies as attachments rather than a wall of inline base64 | | `recordHar.path` | `har/weather.har.zip` | archive and attachments travel as one file that cannot be half-committed | | capture route | `--save-har` from the CLI | one-off recording by hand, with `--save-har-glob` as the filter | The CLI equivalent of a scoped capture is `npx playwright open --save-har=har/weather.har.zip --save-har-glob='**/api/**' https://dashboard.example`, which is often the fastest way to seed the first archive before any test exists. ## Secrets end up in the file A HAR entry stores the full request line and headers. That means, by default: - **Query parameters**, including the forecast provider's API key if it travels in the URL. - **Request headers**, including `Authorization` and any custom key header. - **Cookies**, unless `mode: 'minimal'` has dropped that section. - **Response bodies**, including whatever personal data the account you recorded with could see. So the policy is: record with a throwaway or sandbox credential, keep authenticated third-party calls out of the capture with `urlFilter` where you can, and read the archive once before the first commit. Secret scanning that watches source files will not necessarily look inside a `.zip` archive, so the review is a human step, not a tooling assumption. ## Size and reviewability An unfiltered capture of a dashboard is megabytes of images and fonts, and no reviewer will read it. That has real consequences: nobody can tell a legitimate refresh from an accidental one, the repository grows on every re-record, and checkout times drift upward across every job that clones it. - Filter at capture time; deleting entries afterwards is manual work nobody repeats. - Prefer `minimal` so a refresh diff shows changed payloads rather than changed timings. - Keep one archive per scenario when the scenarios need different answers for the same endpoint, since a given URL is answered one way from a single archive — a storm-warning run and a clear-day run want their own files. - Keep archives out of any bundle that ships, and near the specs that use them so ownership is obvious. ## Ownership and refresh An archive with no owner is a liability. Name one, and make the refresh path mechanical: 1. One helper performs every `routeFromHAR` call, with the update flag read from the environment. 2. Refreshing is a documented command, run against a known environment, not an ad-hoc local edit. 3. The refreshed archive is reviewed like code — which is only possible because of the filtering decisions above. 4. The credential used to record is one you can rotate without touching anything else. ## What the archive is not for Recording is a way to capture a real payload faithfully; it is not a substitute for deciding what your suite should assert. Capture the endpoints under test, keep the file small enough that a human reads it, and make sure everything inside it is something you are content to publish.

  • What stops a recorded archive from leaking the forecast API key?
    Nothing automatic — the archive stores full URLs and request headers as they were sent. Record with a sandbox credential you can rotate, use urlFilter to leave authenticated third-party calls out of the capture where the tests do not need them, and read the file once before the first commit.
  • When is one archive per test better than a single shared archive?
    When scenarios need different answers for the same endpoint. A single archive answers a given URL one way, so a storm-warning case and a clear-day case each want their own file. Shared archives suit stable reference data and keep the repository smaller.
  • How do you keep archives from bloating the repository over time?
    Filter at capture with urlFilter so assets never enter the file, record in minimal mode, and keep each archive scoped to one scenario. A refresh then rewrites a small file rather than adding another few megabytes of images to the history.

saying these in an interview costs you the question

  • Records the whole site when one API is under test
  • Assumes a HAR file never contains credentials
  • Commits multi-megabyte archives nobody can review
  • Treats recording as a one-off nobody has to own
  • Keeps full mode when only replay data is needed
  • Plans to strip sensitive entries by hand later