skip to content

In a feature file, how does a data table differ from a doc string as a step argument?

level: juniorimportance: nice to knowfreq 26%

answer

  1. Both belong to one step
  2. Grid of fields versus block of text
  3. Line breaks survive in one form
  4. Runs once, not once per row
  5. Do not confuse with an Examples table

basics

~20 s

Both attach data to the single step above them. A data table is pipe-delimited rows and columns, read as structured records; a doc string is a delimited block of free text passed through as one string, line breaks intact.

solid answer

~50 s

Both forms are arguments to **one step**, not to the scenario. A **data table** is a grid of pipe-delimited rows, usually with a header row, and it arrives as structured data — a list of records, a list of rows, or a set of key/value pairs depending on its shape. A **doc string** is a delimited block of free text, conventionally triple-quoted, that arrives as a single string with its line breaks and relative indentation preserved; many dialects allow a content-type hint on the opening delimiter. Use a table when the data has fields, and a doc string when the value genuinely is a block of text — a message body, a payload, an expected document. The distinction people get wrong is with an **Examples** table: a data table feeds one step and the scenario still runs once, while an Examples table under an outline runs the scenario once per row.

code

pseudocode · 14 lines
pseudocode
Scenario: Nightly rates are published to a booking channel
  Given the channel manager holds these rates
    | date       | room type | nightly rate | currency |
    | 2026-04-11 | standard  | 138.50       | EUR      |
    | 2026-04-12 | standard  | 146.00       | EUR      |
    | 2026-04-11 | suite     | 297.25       | EUR      |
  When the rates are published to the wholesale channel
  Then the channel receives the message
    """
    RATE UPDATE
    property: 4471
    nights: 2
    note: prices exclude city tax
    """

go deeper

for a junior

Know that both forms hang off the single step directly above them, that one is a pipe-delimited grid and the other a delimited block of free text, and that neither multiplies the number of times the scenario runs.

for a middle

Explain what each form reaches the implementation as — structured rows or one preserved string — and be crisp about the contrast with an Examples table, which multiplies scenarios instead of feeding a step.

for a senior

Show judgment about size. Be ready to say when a large payload or a wide table has turned a readable example into an unreviewed fixture, and where that exact-value comparison belongs instead.

for a principal

Own the house rule for how much data may live in business-readable text at all, and the argument for it: incidental detail in a specification is what erodes the readability that justified the format in the first place.

### Two ways to hang extra data off a single step Some steps need more than a sentence. A feature file offers two argument forms for that, and both attach to the **one step immediately above them** — they are part of that step, not of the scenario. **A data table** is a block of pipe-delimited rows written under the step. It is structured: rows and columns, usually with a header row, and it reaches the step's implementation as a table object that can be read as a list of records, a list of rows, or a set of key/value pairs depending on its shape. **A doc string** is a block of free text, conventionally delimited by triple quotes (some dialects also accept a triple-backtick form) and indented to the step. It is unstructured: it reaches the implementation as a **single string**, with its internal line breaks and relative indentation preserved. Many dialects allow a content-type hint on the opening delimiter, which is a note to the reader and to the implementation about how to parse the text. ``` Scenario: A channel's nightly rates are published for a date range Given the channel manager holds these rates | date | room type | nightly rate | currency | | 2026-04-11 | standard | 138.50 | EUR | | 2026-04-12 | standard | 146.00 | EUR | | 2026-04-11 | suite | 297.25 | EUR | When the rates are published to the wholesale channel Then the channel receives the message """ RATE UPDATE property: 4471 nights: 2 note: prices exclude city tax """ ``` ### Choosing between them Use a **data table** when the data has fields: several records of the same shape, a set of named properties, a small matrix. The table is readable to a domain expert, the columns name the fields, and the implementation gets something it can iterate without parsing. Use a **doc string** when the value **is** a block of text and its shape matters: a message body, a document payload, a rendered notification, an expected multi-line output. Flattening that into a table would misrepresent it, and putting it on the step line would make the step unreadable. Whitespace and line breaks survive, which is precisely why it suits payloads and precisely why a stray trailing space can make an equality assertion fail in a way that is invisible on screen. Both are equally valid on a Given, a When or a Then — a table of input records on a Given, a payload on a When, an expected document on a Then. ### The confusion worth being crisp about The near-universal mix-up is between a **data table** and an **Examples table**, because both are pipe-delimited grids. * A **data table** belongs to one step, inside one scenario, and is passed to that step as an argument. The scenario runs **once**. * An **Examples table** belongs to a Scenario Outline and supplies substitution values. The scenario runs **once per row**. So a three-row data table means one scenario handling three records; a three-row Examples table means three scenarios. Getting this backwards is a reliable signal that someone has read the format but not run it, and it is the reason this question gets asked at all. Note also that they compose: an outline's placeholders are usually substituted inside a step's data table or doc string as well, so a row of an Examples table can vary one cell of a payload. ### Restraint Both forms are easy to abuse. A doc string holding a two-hundred-line document turns a specification into a fixture file that nobody reviews; if the exact bytes matter more than the behaviour, the comparison probably belongs in a lower-level check and the scenario should assert the property that matters ("the message names the property and both nights"). A data table with fourteen columns has the same problem: the reader cannot see which column the scenario is actually about. A useful rule is that a step argument should still fit on a screen and every column in it should be one a domain expert would ask about. When it does not, the detail is incidental to the behaviour, and incidental detail is what pushes a business-readable file back into being a script.

  • Which of the two forms would you use for an expected multi-line message, and why?
    A doc string, because the value is text and its shape is part of what is being asserted — line breaks and relative indentation are preserved exactly. A table would impose fields the message does not have. The caution is that invisible whitespace then matters, so a trailing space can fail an equality check for reasons nobody can see on screen.
  • Can a step argument appear under a step inside a Scenario Outline?
    Yes, and placeholders are normally substituted inside the argument too, so a row of the Examples table can vary a single cell of a table or one line of a text block. That is useful, but it compounds quickly — if a reader has to hold both the row and the payload in their head, the scenario is probably carrying incidental detail.
  • When is a large step argument a smell rather than a convenience?
    When it stops being reviewable. A two-hundred-line text block or a fourteen-column table turns a specification into a fixture nobody reads, and it hides which part of the data the scenario is actually about. Assert the property that matters in the scenario and move exact-payload comparison to a lower-level check.

saying these in an interview costs you the question

  • Says a three-row data table makes the scenario run three times
  • Confuses a data table with an Examples table
  • Thinks a step argument belongs to the scenario, not one step
  • Pastes a long document into a scenario as an assertion
  • Assumes a doc string is parsed into fields automatically

context