skip to content

In an ADF Mapping Data Flow, why does Data preview require debug mode to be on?

level: juniorimportance: should knowfreq 48%

answer

  1. the preview is a real execution
  2. something has to run the query
  3. the toggle at the top of the designer
  4. a session cluster, billed while open
  5. previews read, debug runs write

basics

~20 s

Data preview really executes the flow up to that transformation against live data. Debug mode starts a dedicated Spark cluster for your design session; without it there is no compute to run the sampled query on.

solid answer

~50 s

Turning on the data flow debug toggle starts a debug Spark cluster attached to your session. With it running, the **Data preview** tab on any transformation executes the graph from the sources up to that node and returns a limited sample of rows, so you can see what a Derived Column expression produced or which marker an Alter Row assigned before you ever schedule anything. Two details matter in practice. Preview reads sources but does not write to sinks - previewing a Sink transformation shows the rows arriving, not rows landed in the target. A pipeline **Debug run** of the Data Flow activity is different: that executes for real and does write. The debug cluster consumes compute for as long as the session is open, and debug settings let you cap preview rows, point sources at sample files, and supply parameter values.

code

text · 7 lines
text
Debug settings
  Row limit          : 1000
  Source settings
    SourceCustomers  : sample file  adls/dev/customers_sample.csv
    SourceOrders     : use source schema, row limit 500
  Parameters
    p_load_date      : 2026-08-20

go deeper

for a junior

Know that the debug toggle starts a cluster and that Data preview needs it because the preview is a real execution over sampled rows. Remember to switch debug off when you are done.

for a middle

Explain the split precisely: preview reads and displays without writing, while a pipeline debug run executes end to end and writes to sinks. Mention debug settings for row limits, sample sources and parameter values.

for a senior

Use previews as a diagnostic ladder, checking each transformation's output to isolate where rows or values were lost, and be clear that sampled previews say nothing about full-volume performance.

for a principal

Set the team's rules: which sources debug sessions may touch, that debug runs never point at production sinks, and how session sprawl across developers is kept from becoming a standing compute bill.

## What the toggle actually starts A Mapping Data Flow is a design surface, but everything it shows you about your data comes from real execution. The **Data flow debug** toggle at the top of the designer starts a Spark cluster dedicated to your session. It is the same kind of managed compute a triggered run would use, held open interactively so the designer can ask it questions. That is the whole answer to why preview needs it: without a live cluster there is nowhere to run the query that produces the preview. Because a cluster has to be acquired, flipping the toggle takes a couple of minutes before previews become available - the same cold-start cost that a scheduled data flow pays. ## What Data preview gives you Every transformation in the graph has a Data preview tab. Clicking it executes the flow from the sources up to and including that transformation and returns a limited number of rows. That is enormously useful because you can inspect the intermediate state at each hop rather than reasoning about the whole graph at once: - On a Source, confirm the projection and the types the service inferred. - On a Derived Column, see the computed values and catch a null-propagating expression immediately. - On a Join, see how many rows survived and whether the key matched at all. - On an Alter Row, see the marker assigned to each row, which is the fastest way to prove a policy expression is firing. Previews are also where you notice type surprises early - a numeric column read as string from a delimited text source, or a date that failed to parse. ## Debug settings The **Debug settings** panel controls the session. You can cap the number of rows a preview returns so the round trip stays fast, point a source at a sample file or a smaller row limit instead of the full production dataset, and supply values for the flow's parameters so parameterized expressions resolve while you are designing. Setting parameter values here is easy to forget and produces confusing previews when an expression silently evaluates against an empty parameter. ## Preview does not write; a debug run does This distinction catches people out and is worth being precise about in an interview. **Data preview** inside the designer reads from sources and materialises rows for display. It does not write to sinks - previewing the Sink transformation shows what would arrive there. So previewing against production sources is safe from the target's point of view, though it does read real data. A **Debug run of the pipeline** containing the Data Flow activity is a genuine execution. It uses the debug session's compute rather than waiting for a fresh cluster, but the flow runs end to end and the sinks are written with whatever settings they carry, including table actions like truncate or recreate. Pointing a debug run at a production table because "it is only debug" is a real way to lose data. ## Cost and hygiene The debug session holds a cluster, and holding a cluster consumes compute time whether you are actively previewing or reading email. Sessions have an inactivity shutdown, but the disciplined habit is to turn debug off when you finish designing. On a shared factory, several developers each holding a debug session is a quiet, recurring cost that nobody notices until the bill arrives. One more caution: because previews are sampled and often run against smaller inputs, a preview that returns instantly tells you nothing about how the flow will perform on the full dataset. Correctness feedback, yes; performance evidence, no. For that you need a real run and its monitoring detail. ## The short version for an interview Debug mode starts a session cluster; Data preview executes the graph up to a node on that cluster and shows sampled rows; previews never write to sinks but a pipeline debug run does; debug settings cap rows, choose sample sources and supply parameters; and the session costs compute until you switch it off.

  • Is it safe to preview a Sink transformation that points at a production table?
    Yes for the target - preview shows the rows arriving at the sink and does not write them. It does read from the real sources, so it is not free of side effects on load. What is not safe is running a pipeline Debug run against that sink: that executes for real, including any table action such as truncate or recreate.
  • Your preview shows nulls where an expression should produce values. What do you check first?
    Preview the upstream transformation to see whether the input columns are present and typed as you expect - a mistyped or drifted column is a common cause. Then check Debug settings for unset flow parameters, since an expression referencing an empty parameter evaluates to nothing without raising an error.

saying these in an interview costs you the question

  • Thinks Data preview writes rows into the sink
  • Believes previews prove production performance
  • Assumes triggered runs use the debug session's cluster
  • Leaves debug sessions open all day on a shared factory
  • Expects preview to read the entire source dataset

context