skip to content

You swap a plan's spiderAjax job for spiderClient and change nothing else. Which defaults move under you?

level: seniorimportance: should knowfreq 46%

answer

  1. Same parameter names, different defaults
  2. Scope handling and logout behaviour both move
  3. Depth and browser count move too
  4. Write the values into the job

basics

~20 s

The two browser-crawler jobs share parameter names but not their defaults: scopeCheck flips from Strict to Flexible, logoutAvoidance from off to on, and crawl depth, browser count and duration all change. Pin the dials in the plan.

solid answer

~40 s

The two browser-crawler jobs take many of the same parameter names, which makes the swap look like a one-word edit. It is not. `scopeCheck` defaults to `Strict` on `spiderAjax` and `Flexible` on `spiderClient`, so out-of-scope resources go from blocked to loaded. `logoutAvoidance` defaults off on `spiderAjax` and on on `spiderClient`, so the crawl goes from happily clicking a sign-out link to trying not to. `maxCrawlDepth` and `numberOfBrowsers` also differ — the client spider's defaults are the shallower and the smaller. A few do agree, including the browser id. The fix is not to memorise which is which: **write the values you depend on into the job**, so a crawler swap changes the crawler and nothing else.

code

yaml · 7 lines
yaml
# pin the dials, then the job type is the only thing the swap changes
- type: spiderClient
  parameters:
    context: myapp
    url: https://example.com
    scopeCheck: Flexible
    logoutAvoidance: true

go deeper

for a junior

The lesson to take away is simple: if a setting matters to you, write it into the job rather than letting the job supply it, because two jobs can spell a parameter the same way and mean different things by leaving it out.

for a middle

Be able to name the settings that move — scope check, logout avoidance, depth, browser count, duration — and say which direction each moves in when you go from the older crawler to the newer one.

for a senior

Show how you would attribute a change in results: hold the dials fixed, change only the engine, and resist reporting a default effect as a capability difference.

for a principal

Decide what your organisation pins centrally versus leaves to each team, knowing that an unpinned default is a dependency on somebody else's release notes that no review will catch.

## Why this bites An automation plan names a job by `type` and then lists `parameters`. Both browser crawlers accept a similar set of parameter names — `context`, `user`, `url`, `maxDuration`, `maxCrawlDepth`, `numberOfBrowsers`, `browserId`, `scopeCheck`, `logoutAvoidance` — so replacing `type: spiderAjax` with `type: spiderClient` looks like a safe, one-token change. Every parameter you did **not** write down is then supplied by the new job's own defaults, and the two add-ons chose differently. ## What actually moves | parameter | `spiderAjax` default | `spiderClient` default | what changes | |---|---|---|---| | `scopeCheck` | `Strict` | `Flexible` | whether out-of-scope resources are blocked at the crawl's proxy or allowed to load | | `logoutAvoidance` | off | on | whether the crawl tries not to click sign-out controls | | `maxCrawlDepth` | the deeper value | the shallower value | how far from the start URL the crawl goes | | `numberOfBrowsers` | about one per core | about half that, capped | memory and wall-clock cost of the crawl | | `maxDuration` | a bounded run | unbounded | whether the job stops itself | | `browserId` | headless, same on both | headless, same on both | nothing | The `scopeCheck` and `logoutAvoidance` rows are the ones that change what the crawl *finds*. The depth and browser-count rows change what it *costs*. The `maxDuration` row is the one that most often surprises a pipeline, and it is worth dwelling on: **both shipped job templates carry the same comment for that parameter, and the comment is only right for one of them.** The AJAX spider job initialises it to a positive number of minutes and enforces a deadline; the client spider job leaves it unset, and an unset value means no deadline at all. A crawl that used to stop on its own can start running until something else kills it. ## The right lesson is not "memorise the table" The table will drift. The habit that survives is: 1. **Write down every parameter whose value you depend on**, even when it matches today's default. A plan is a record of intent; an omitted key records nothing. 2. **Treat a template comment as a comment.** These templates are what the plan generator writes for you to copy, and at least one of their default claims does not match the code. The validator and the job class are the authority. 3. **Diff the effective configuration, not the plan.** If two runs disagree and the plan text is identical except for a job type, the difference is in the defaults you did not write. 4. **Re-check after a crawler swap, not before.** The point of the swap is usually more coverage; if the numbers move a lot, work out how much of that is the engine and how much is `scopeCheck` flipping to `Flexible`. ## A worked case Take a single-page app whose navigation never appears to a link parser, served as a shell page with its bundle and its API on other hostnames. Run it under the AJAX spider job with nothing pinned: `scopeCheck` is `Strict`, the bundle's host is not in scope, the crawl's own proxy refuses the request, the app never boots, and the crawl reports the shell and little else. Swap the job type to `spiderClient` with nothing pinned: `scopeCheck` is now `Flexible`, the bundle loads, the router runs, and the crawl explores the app. It is tempting to write that up as "the client spider is much better on single-page apps". Part of it is real — the DOM reporting genuinely finds more. But a meaningful part of that particular jump is a **default**, not a capability, and you can get most of it by setting `scopeCheck: Flexible` on the AJAX spider job. If you do not separate the two effects, you will draw the wrong conclusion about your tooling, and you will be surprised again the next time a default moves. ## And the direction nobody expects The swap also turns `logoutAvoidance` **on**, which sounds like an unambiguous improvement and mostly is — but it means the newer crawler will decline to click some elements the older one clicked. If a comparison run shows the client spider missing a page the AJAX spider found, an avoided element whose visible text happened to look like a sign-out control is a live explanation, and it is not one you would ever guess from the job type alone.

  • Which shared parameters do agree between the two jobs?
    `context`, `user` and `url` behave the same way, and `browserId` defaults to the same headless browser on both. The ones to watch are `scopeCheck`, `logoutAvoidance`, `maxCrawlDepth`, `numberOfBrowsers` and `maxDuration` — every one of those is supplied differently when you leave it out.
  • What happens if you write an unrecognised value for scopeCheck?
    It is not an error. The value is parsed case-insensitively and anything that does not match falls back to that add-on's own default — which is `Strict` for `spiderAjax` and `Flexible` for `spiderClient`. So a typo does not fail the plan; it silently hands you the default you were trying to override, and the two crawlers give you different ones.
  • How would you attribute a coverage jump after swapping crawlers?
    Change one thing at a time. Run the old job with the new job's scope setting pinned, and compare that against both originals. Whatever coverage the old crawler gains from the setting alone is a default effect; what remains is the engine.

saying these in an interview costs you the question

  • Assumes identically named parameters carry identical defaults
  • Credits a coverage jump entirely to the newer engine
  • Trusts a template comment over the job's code
  • Says an unrecognised scopeCheck value fails the plan
  • Leaves scope and duration unset because the defaults look sane