skip to content

In Azure Data Factory, how do a linked service, a dataset and an activity differ?

level: juniorimportance: must knowfreq 76%

answer

  1. three layers: where, what, and do
  2. one object stores the connection
  3. another names the table or file
  4. the step that acts is the activity

basics

~20 s

In Azure Data Factory a linked service holds the connection and credentials for a store or compute, a dataset names the data inside it using that linked service, and an activity is the pipeline step that acts on the data.

solid answer

~40 s

They are three layers of the same reference chain. A **linked service** is the connection: the endpoint, the auth (managed identity, or a Key Vault secret reference) and, via `connectVia`, which integration runtime executes against it. A **dataset** is a named, typed view of data reachable through one linked service — a SQL schema/table, a blob folder and file pattern, a Parquet path — and it carries only shape and location, never credentials. An **activity** is a step inside a pipeline: `Copy`, `Lookup`, `GetMetadata`, `ForEach`, `If Condition`, `Execute Pipeline`, `Set Variable`. Data-movement activities bind datasets as inputs and outputs; control activities bind none. A **pipeline** is the ordered group of activities, and activity order comes from `dependsOn` conditions rather than from the datasets.

go deeper

for a junior

Be ready to name the objects and say which one holds credentials. This is the opening question in almost any Azure Data Factory screen, and a wrong answer here ends the section.

for a middle

Explain the reference chain: activity binds datasets, a dataset binds one linked service, and the linked service picks an integration runtime through connectVia. Add why control activities need no dataset.

for a senior

Show how you keep an estate maintainable: parameterised datasets and linked services, secrets in Key Vault or managed identity, a control table plus ForEach instead of hand-built pipelines per source.

for a principal

Own the standard: naming, environment promotion through ARM template parameters and CI/CD, and the call between generic metadata-driven pipelines and per-source pipelines that juniors can read.

## The object model in one sentence Azure Data Factory (ADF) separates *where the data lives* (linked service), *what the data is* (dataset), *what you do to it* (activity), *how those steps are grouped* (pipeline) and *what fires the group* (trigger). Almost every ADF confusion in an interview comes from collapsing two of those layers. ## Linked service — the connection A linked service is essentially a connection string plus an authentication method plus a pointer to the compute that will use it. It has a `type` naming the store or compute (`AzureSqlDatabase`, `AzureBlobStorage`, `SqlServer`, `AzureDatabricks`, `AzureKeyVault`, …), `typeProperties` carrying endpoint details, and an optional `connectVia` block naming the integration runtime that dials the connection. Credentials should not sit inline: a linked service can authenticate with the factory's managed identity or reference a secret in an Azure Key Vault linked service, which is what keeps the JSON safe to check into git. One linked service is typically reused by many datasets. Linked services can also be parameterised, so a single "SQL server" linked service can serve several databases whose names arrive at runtime. ## Dataset — the named data A dataset always references exactly one linked service through `linkedServiceName`, and adds the coordinates within that store: schema and table for a relational source, container/folder/file plus a format for a lake source. It may also declare a `schema` (column list) and `structure`, though ADF can work schema-less and infer at run time for many formats. A dataset holds **no credentials**. It is the object teams most often over-create: one per table quickly becomes hundreds. The idiomatic fix is a *parameterised dataset* — declare `parameters` on the dataset (for example `tableName`, `folderPath`) and reference them in `typeProperties` with `@dataset().tableName`. A single dataset then serves every table, with values supplied per activity, usually driven from a control table read by a `Lookup` and iterated with `ForEach`. This is the backbone of metadata-driven ADF. ## Activity — the unit of work An activity is a step. ADF groups them roughly as: - **Data movement** — the `Copy` activity, which binds `inputs` and `outputs` datasets and moves bytes between them. - **Data transformation** — `Mapping Data Flow`, `Databricks Notebook`, `Stored Procedure`, `Script`, `HDInsight`, `Azure Function`. - **Control flow** — `ForEach`, `If Condition`, `Switch`, `Until`, `Wait`, `Filter`, `Set Variable`, `Append Variable`, `Execute Pipeline`, `Fail`, `Web`, `Webhook`, `Validation`, `Lookup`, `GetMetadata`, `Delete`. Control activities take no dataset at all — a common junior error is claiming every activity needs input and output datasets. Activities expose outputs that later activities read with `@activity('ActivityName').output…`, which is how a `Lookup` result reaches a `ForEach`. ## Pipeline — the grouping and the ordering A pipeline is a named list of activities plus `parameters` (set once when the run starts, immutable thereafter) and `variables` (mutable during the run via `Set Variable`). Order is not implied by datasets; it comes from each activity's `dependsOn` array, where every entry names a prior activity and a dependency condition — `Succeeded`, `Failed`, `Skipped` or `Completed`. Several entries on one activity are ANDed together. The pipeline is also the unit of triggering, monitoring and rerun: a run has one run id, one set of parameter values, and appears as a single row in Monitor. ## Where the integration runtime fits The fifth object is the integration runtime — the compute that actually performs copies, runs data flows and dispatches activities. Datasets never reference it; linked services do, through `connectVia`. So the full chain for a copy is: activity → dataset → linked service → integration runtime → the actual store. ## What interviewers listen for A solid answer states the reference direction (activity → dataset → linked service), puts credentials only in the linked service, notes that control activities have no datasets, and mentions dataset parameterisation as the way to avoid object sprawl. A weak answer describes a dataset as "the connection to the database".

  • Does every activity need input and output datasets?
    No. Copy, Lookup, GetMetadata and Delete bind datasets, but control activities — ForEach, If Condition, Until, Set Variable, Execute Pipeline, Wait — take none, and Web or Azure Function activities call an endpoint directly. Treating datasets as mandatory is a common misread of the model.
  • How do you avoid creating one dataset per source table?
    Parameterise the dataset. Declare parameters such as schema and table (or folder and file), reference them as @dataset().tableName inside typeProperties, and pass values from the activity. Drive the values from a control table read by a Lookup and iterated with ForEach, so one dataset serves hundreds of tables.
  • Where should the credentials for a linked service live?
    Prefer the factory's managed identity where the target supports it. Otherwise store the secret in Azure Key Vault and have the linked service reference it through a Key Vault linked service. Inline connection strings end up in ARM templates and source control, which is exactly what you want to avoid.

saying these in an interview costs you the question

  • Says the dataset stores the connection string and credentials
  • Claims every activity must have input and output datasets
  • Confuses the integration runtime with the linked service
  • Creates one dataset per table instead of parameterising
  • Thinks activity order is derived from dataset relationships

context