In Cucumber-JVM, how does a doc string's content type reach a @DocStringType transformer?
answer
- a doc string can be typed
- content type sits after the opening delimiter
- the pair is content type plus parameter
- no transformer means a plain String
- tables and doc strings use separate registries
basics
~20 sWrite the content type immediately after the doc string's opening delimiter in the feature file. Cucumber-JVM pairs it with the step method's declared parameter type to select a @DocStringType transformer. With none registered, the body arrives as a plain String.
solid answer
~40 sA Gherkin doc string is the block between two triple-quote delimiters under a step, and it may carry a **content type** written directly after the opening delimiter with no space. In Cucumber-JVM that content type is not decoration: it is one half of the key Cucumber uses to pick a converter, the other half being the type the step method declares for the parameter. A `@DocStringType` method takes the raw body as a `String`, returns the target type, and names the content type it serves in the annotation's `contentType` attribute; a transformer declared without one serves doc strings written without one. If the step method simply declares `String`, no conversion happens and the body is passed through with the opening delimiter's indentation already stripped from every line.
code
java · 9 lines@DocStringType(contentType = "json")
public TariffSchedule tariffJson(String docString) throws IOException {
return MAPPER.readValue(docString, TariffSchedule.class);
}
@Given("the district-heating tariff schedule is:")
public void loadSchedule(TariffSchedule schedule) {
billing.load(schedule);
}go deeper
Know that a doc string is a step argument written between triple-quote delimiters, that it arrives as a string by default, and that a word may follow the opening delimiter to say what the body is.
Explain that Cucumber-JVM selects a @DocStringType from the content type together with the declared parameter type, what happens with no transformer, and how indentation is normalised relative to the opening delimiter.
Show when typing a doc string pays: several steps parsing the same payload with duplicated error handling, versus a single step where a transformer is indirection for its own sake. Recognise a conversion failure as distinct from an assertion failure.
Own whether feature files may steer conversion at all. Deciding that a content type in the Gherkin selects glue behaviour puts parsing policy into files non-engineers edit, and that tradeoff should be a deliberate call rather than an accident.
## The doc string, mechanically A doc string is the second kind of step argument Gherkin offers: a block of free text written under a step line between two triple-quote delimiters. Unlike a data table it has no rows and no columns — it is one string, and it exists for payloads whose shape is not tabular. Three mechanical facts matter before any conversion is discussed: - **Indentation is relative.** Gherkin strips from every line the indentation of the *opening* delimiter, so the block can be indented to sit under its step without that indentation reaching your code. Anything indented further keeps its extra spaces, which matters when the body is whitespace-sensitive. - **The delimiter can be escaped.** A literal triple quote inside the body is written with backslashes so the parser does not read it as the closing delimiter. - **The opening delimiter can carry a content type.** Written immediately after it, with no space: a media-type-style word such as `json` or `xml`. ```gherkin Given the district-heating tariff schedule is: """json { "standing": 0.4137, "tiers": [ { "upTo": 1200, "rate": 0.0912 } ] } """ ``` Gherkin also accepts three backticks as an alternative delimiter, which is convenient when the body itself contains triple quotes. ## The content type is half of a key In Cucumber-JVM that content type is not decoration and it is not a comment. It is one half of the key Cucumber uses to select a converter; the other half is the **type the step method declares for the parameter**. `@DocStringType` marks a glue method as such a converter. The method takes the raw body as a `String` and returns the target type; the annotation's `contentType` attribute names the content type it serves. | Doc string in the feature file | Step method parameter | What runs | |---|---|---| | Carries a content type | A domain type | The `@DocStringType` registered for that pair | | Carries no content type | A domain type | A transformer registered without a content type | | Anything | `String` | No conversion; the raw body is passed through | | Carries a content type | A domain type with no matching transformer | The step fails at conversion time | The practical consequence is that **one step text can accept differently-typed payloads**. Two steps whose bodies look almost identical can be parsed into two different objects purely because their opening delimiters say `json` and `xml`, with no change to the step wording and no extra parameter in the step text. ## When nothing matches If the step method declares `String`, nothing has to match: the content type is simply unused and the body arrives verbatim, indentation already normalised. That is the default behaviour and the reason many suites never discover doc-string typing at all. If the step method declares a domain type and no transformer produces it, there is no fallback to the string. The step fails at conversion time, before the method body executes, with a message naming the type. As with data tables, this is deliberate loudness: the run stops on the argument rather than on an assertion further down. ## Doc strings and data tables are separate machinery It is worth being explicit, because candidates conflate them. The two argument kinds have independent registries: - A `@DataTableType` never converts a doc string, and a `@DocStringType` never converts a table. - Table conversion is driven by the *declared parameter type alone*; doc-string conversion is driven by the parameter type **together with** the content type written in the feature file. - The default transformers that catch unregistered types on the table side have no doc-string equivalent to lean on. That second bullet is the interesting asymmetry: the doc string is the only step argument whose conversion can be steered from the feature file itself, without touching the glue. ## Across the implementation family | Implementation | What the step receives | |---|---| | Cucumber-JVM | The converted type, or the raw body as a `String` | | cucumber-js | The doc string as a plain string argument | | Behave | The block on the context object's text attribute | | SpecFlow/Reqnroll | A multiline text argument as a string parameter | Only Cucumber-JVM turns the content type into a dispatch key, so this is squarely a Cucumber-JVM question rather than a Gherkin one. ## Where it earns its keep A district-heating billing team carried tariff schedules as JSON payloads in about a dozen scenarios of its 148-scenario nightly pack. Before typing them, every step parsed the string itself, so `ObjectMapper` calls and their error handling were duplicated across nine step methods and one of them silently swallowed a malformed body. One `@DocStringType(contentType = "json")` method returning the schema type deleted all nine, moved parse failures into a single place that fails the step loudly, and — three weeks from a contractor handover — left one obvious hook for anyone adding a new payload format. The counter-case is worth saying too: if only one step in a suite ever takes a JSON body, registering a transformer buys indirection for nothing. Take the `String` and parse it in the step.
- The feature file gives the doc string a content type but no matching transformer is registered. What does the step method receive?If it declares `String`, the raw body with every line intact and the content type simply unused. If it declares a domain type instead, there is nothing to build it with, so the step fails at conversion time naming that type rather than falling back to handing over the string.
- Two steps need the same doc-string body parsed into different objects. How do you arrange that?Give the two doc strings different content types and register a `@DocStringType` for each, returning a different type. Selection is by content type together with the declared parameter type, so the two steps can carry identical body text and still receive different objects with no change to the step wording.
- How is the doc string's indentation handled before your transformer sees it?Gherkin strips from every line the indentation of the opening delimiter, so the block can be indented to line up under its step without those spaces reaching your code. Lines indented further than the delimiter keep the extra spaces, which matters when the payload is whitespace-sensitive.
saying these in an interview costs you the question
- Thinks the content type is only a comment for readers
- Expects a @DataTableType to convert a doc string too
- Believes all leading whitespace is stripped from every line
- Cannot say what a step gets when no transformer matches
- Confuses a doc string with an Examples table