How do you decide which Cucumber implementations a polyglot organisation should support?
answer
- Two layers with different coupling
- Ask what the glue is bound to
- Uniformity is not itself a goal
- Standardise conventions, not the runtime
- Write down the inputs, set a review trigger
basics
~20 sMatch each service to the implementation native to its language rather than mandating one estate-wide. Glue calls production code, so an out-of-stack choice buys shared tooling at the cost of a language boundary. Standardise conventions and reporting instead.
solid answer
~50 sStart from what a single implementation actually buys: one glue idiom to review and train on, one CI recipe, one report pipeline, and engineers who can move between suites. Then weigh it against what it costs, which is usually decisive — **the step body calls the system under test**, so mandating one implementation forces some teams to write acceptance tests outside their service's language. They lose the service's own domain types, test doubles and fixtures, gain a process or client boundary, and debug across two toolchains. A second implementation is the right answer whenever a team's service is in a different stack and its scenarios exercise that service's code. What you standardise instead is the layer that genuinely is shared: Gherkin conventions and step phrasing, the tag vocabulary, the abstraction level of scenarios, a normalised report format, and a per-pipeline time budget every suite must meet.
go deeper
Know that each language stack has its own Gherkin implementation and that step definitions are written in the service's language. You are not expected to make the organisational call at this level.
Be able to name what differs between implementations and what does not, and to explain why sharing feature files is easier than sharing step definitions.
Argue the tradeoff concretely for a specific estate: where scenarios execute, whether the suite can reuse a service's own test assets, and what an out-of-stack suite costs in pipeline time and debugging loop.
Own the decision and its review trigger. Separate what is standardised — conventions, tag vocabulary, report shape, pipeline budget — from what is chosen per stack, and be ready to defend a mixed answer against a request for uniformity.
This is a judgement question with no single right answer, and the interviewer is listening for whether you reason from **what the tool is coupled to** rather than from a preference for uniformity. ## Frame the decision correctly The question is usually posed as "should we all use the same BDD tool?", which smuggles in an assumption. A Gherkin suite has two layers with completely different coupling: - The **feature file** is coupled to the business language. It is portable across the family and is the layer where consistency genuinely pays. - The **glue** is coupled to the system under test. A step body constructs your client, calls your service, asserts on your types. It is not portable, and no amount of standardisation makes it so. So the real decision is not one question but two: *which implementation does each team run?* and *what do all teams share regardless?* Conflating them is the failure mode. ## What a single implementation actually buys 1. **One idiom to learn and review.** Reviewers hold one mental model of hooks, state and matching instead of three, and internal guidance has one worked example. 2. **One CI recipe.** A single container image, cache strategy, parallel-execution setting and artefact layout. 3. **One report pipeline.** Output arrives in one machine-readable shape, so coverage and trend dashboards are built once. 4. **Mobility.** An engineer moving between squads is productive in the test suite immediately. 5. **One upgrade to do.** Version bumps, security advisories and deprecations are handled once. These are real, and on an estate of near-identical services in one language the answer is obvious. ## What it costs when the stacks genuinely differ Mandating, say, the JVM implementation across a Node service team means their acceptance tests live outside their service's language. The concrete costs: - **No reuse of the service's own test assets.** Fixtures, factories, in-memory doubles and domain types built for the service's unit tests are unusable from the out-of-stack suite, so they are rebuilt and then drift. - **A boundary appears.** The suite must reach the service over HTTP or a spawned process even where an in-process test would have been faster and more precise, which pushes the suite toward end-to-end shape and pushes pipeline time up. - **Debugging crosses toolchains.** A failing scenario is diagnosed in one language and fixed in another, which lengthens the loop and quietly transfers ownership of the tests away from the team that owns the code. - **A skills tax.** The team maintains competence in a language they otherwise do not use, and the suite becomes the thing nobody wants to touch. ## The decision inputs I would actually weigh | Input | Pushes toward one implementation | Pushes toward several | | --- | --- | --- | | Number of production languages | one or two | three or more, each with a real owning team | | Where scenarios execute | a shared deployed environment | in-process against each service | | Feature-file sharing | rules genuinely shared across services | each service has its own rules | | Team ownership | a central quality-engineering group | squads own their own tests | | Pipeline budget | generous | tight enough that in-process speed matters | ## A worked call A ferry-timetable booking product has three squads — a Node booking API, a Python pricing service and a .NET back-office — sharing a 63-scenario feature set describing fare, timetable and cancellation rules, with an eleven-minute budget per pipeline. The decision I would defend: **three implementations, one specification.** Each squad runs the Gherkin implementation native to its stack, so scenarios execute in-process and fit the budget. The 63 shared scenarios ship as a versioned feature-file package that all three consume, with undefined steps failing the build so drift is loud. Three things are standardised organisation-wide: step phrasing conventions and the abstraction level scenarios are written at; the tag vocabulary and what each tag obliges a pipeline to do; and a normalised report shape so one dashboard answers "which rules are covered, and did they pass?". The call flips when the inputs flip. If the same organisation ran a single deployed environment and one quality-engineering group owned all acceptance tests, the glue would not be reaching into anyone's process, the coupling argument would evaporate, and one implementation would be the obviously cheaper answer. ## How I would keep the decision honest - Write the decision down with the inputs that produced it, so it can be revisited when a stack is retired rather than defended by habit. - Set a review trigger — a new production language, or a squad asking for an exception — instead of a calendar date. - Measure the thing you actually care about: pipeline time, flake rate and rule coverage per squad. If a second implementation is not showing up as worse on those, uniformity was never the goal.
- A squad asks to adopt a fourth implementation for one service. What would make you say yes?That the service is in a language none of the current implementations serve, that the squad owns and will maintain the suite, and that the scenarios genuinely exercise that service's code rather than a shared environment. I would also require that it consume the same shared feature-file package, honour the tag vocabulary, emit the normalised report shape and meet the same pipeline budget — the standards are on the interfaces, not the runtime.
- How would you tell whether standardising had actually paid off a year later?By the measures that motivated it, not by tool count: pipeline time per squad, flake rate, how long a new engineer takes to add a scenario, and how much shared tooling was genuinely reused rather than forked. If reviewers still hold one mental model and dashboards still work, uniformity earned its keep; if squads have quietly built adapters around the mandated tool, it did not.
- What is the strongest argument you would expect from someone advocating one implementation, and how do you answer it?That three implementations mean three sets of upgrades, three CI recipes and three ways to be wrong, which is real and I would not dismiss it. The answer is to cut the duplication where it is actually duplicated — one shared feature-file package, one tag vocabulary, one report shape, one pipeline template parameterised by stack — so what remains distinct is only the layer that was always going to be written per language anyway.
saying these in an interview costs you the question
- Argues for one implementation purely because uniformity is tidier
- Ignores that step bodies call the system under test in its own language
- Assumes a shared feature-file package removes the need for local glue
- Treats tool count as the metric rather than pipeline time and coverage
- Mandates a stack without asking who will maintain the resulting suite