In Azure Data Factory, when do you need a self-hosted integration runtime instead of the Azure one?
answer
- compute lives outside the factory's control plane
- three flavours, one you install yourself
- the agent dials out, nothing dials in
- one flavour exists only for legacy packages
basics
~20 sYou need a self-hosted integration runtime when the source or sink is not reachable from the public Azure network — on-premises servers or private-network stores — because it runs as an agent inside your network and connects outbound to Azure Data Factory.
solid answer
~50 sADF has three integration runtime types. The **Azure IR** is fully managed and serverless: it performs copies, runs Mapping Data Flows and dispatches activities from a chosen region, and needs the endpoint to be reachable from Azure. The **self-hosted IR** is software you install on a machine inside your own network; it makes outbound connections to ADF, so no inbound firewall ports are opened, and it is the way to reach an on-premises SQL Server, a private file share or anything behind a corporate boundary. It can be scaled out across multiple nodes for throughput and availability, and shared with other data factories. The **Azure-SSIS IR** is a managed cluster that lifts and shifts existing SSIS packages. A fourth option often mentioned alongside them is an Azure IR with a managed virtual network and managed private endpoints, which reaches private Azure PaaS resources without any self-hosted machine.
code
json · 18 lines{
"name": "SqlServerOnPrem",
"properties": {
"type": "SqlServer",
"connectVia": {
"referenceName": "CorpSelfHostedIR",
"type": "IntegrationRuntimeReference"
},
"typeProperties": {
"connectionString": "Server=sql01;Database=Sales;Integrated Security=False;User ID=svc_adf",
"password": {
"type": "AzureKeyVaultSecret",
"store": { "referenceName": "KeyVaultLS", "type": "LinkedServiceReference" },
"secretName": "sql01-svc-adf"
}
}
}
}go deeper
Recall that the integration runtime is the compute doing the work, that Azure's is managed, and that a self-hosted one is needed to reach on-premises data.
Explain that a linked service selects the runtime via connectVia, describe all three types, and state the outbound-only connectivity model of the self-hosted agent.
Show sizing and availability judgment: multiple nodes, sharing across factories, the bottleneck when one end of a copy is private, and when a managed virtual network removes the need for a self-hosted node entirely.
Own the connectivity standard across the estate — who runs the self-hosted fleet, how it is patched and monitored, and the migration path off on-premises hops toward managed private networking.
## What an integration runtime is The integration runtime (IR) is the compute that ADF actually uses. The factory itself is a control plane: it stores definitions and schedules runs. Every byte a `Copy` activity moves, every Mapping Data Flow, and every dispatch of an activity to an external compute passes through an IR. A dataset never names an IR; a **linked service** does, through its `connectVia` property. If you omit `connectVia`, the factory's default AutoResolve Azure IR is used. ## Azure integration runtime The Azure IR is fully managed and serverless — nothing to patch, nothing to size in advance. It does three jobs: - **Data movement** for the `Copy` activity between publicly reachable cloud stores. - **Data flow execution**, spinning up the managed Spark cluster behind a Mapping Data Flow according to the compute size and time-to-live configured on the IR. - **Activity dispatch** to external computes such as Databricks, HDInsight or a stored procedure on a SQL database. Its main configuration is the region. "AutoResolve" lets ADF pick a region close to the sink, which is convenient, but a compliance requirement to keep data in a named geography is a reason to pin an explicit region instead. For copy throughput on the Azure IR you tune Data Integration Units and parallel copies on the `Copy` activity, not the IR itself. ## Self-hosted integration runtime The self-hosted IR (often abbreviated SHIR) is an agent you install on a Windows machine — a VM in a virtual network, or a physical server on-premises. Its defining property is the direction of connectivity: it dials **out** to ADF, so a corporate firewall does not need inbound rules. That makes it the answer whenever the data lives somewhere Azure cannot reach: an on-premises SQL Server or Oracle instance, a file share, a private network segment, or a cloud store restricted to specific networks. A self-hosted IR also does more than movement. It dispatches activities that must run against on-premises compute, and it is required for a copy where *either* side is unreachable from the public network — the same node handles both ends of that copy, so it needs enough CPU, memory and network to sustain the transfer. Operationally, three points matter. It can be **scaled out to several nodes**, giving both throughput and high availability if one node goes down. It can be **shared** with other data factories, so a single hardened node serves multiple teams. And it is *your* machine: patching, disk, upgrades and the certificate/credential store are your responsibility, which is precisely why teams reach for a managed alternative when one exists. ## Azure-SSIS integration runtime The Azure-SSIS IR is a managed cluster of nodes that runs SQL Server Integration Services packages, so an existing SSIS estate can be lifted into ADF and invoked by the Execute SSIS Package activity. You choose node size and node count, and you can join it to a virtual network to reach private data sources. It exists for migration rather than for new development; nobody builds new SSIS packages to run on ADF today. ## Managed virtual network and private endpoints The modern middle ground is an Azure IR with a **managed virtual network** enabled and **managed private endpoints** to the target resources. This reaches Azure PaaS services locked to private networking — a storage account or SQL database with public access disabled — without any machine you have to own. It does not help with genuinely on-premises sources, which still need a self-hosted IR, but it removes the most common reason teams used to install one. Interviewers like this distinction because it separates candidates who learned ADF years ago from those working with it now. ## Choosing, in one pass 1. Are both ends publicly reachable cloud services? Azure IR, pin the region if compliance requires it. 2. Is the target an Azure resource locked behind private networking? Azure IR with managed VNet and managed private endpoints. 3. Is anything genuinely on-premises or on a private non-Azure network? Self-hosted IR, sized for the transfer and ideally multi-node. 4. Are you running existing SSIS packages? Azure-SSIS IR. ## Pitfalls worth naming - Believing a self-hosted IR needs inbound firewall openings; it does not, and saying so signals unfamiliarity. - Running production transfers on a single SHIR node, then discovering the whole factory's on-prem connectivity is one unpatched VM. - Tuning Data Integration Units for a self-hosted copy — DIUs apply to the Azure IR; self-hosted throughput comes from node resources and concurrent job settings. - Forgetting that a copy with one private end pins the *whole* copy to the self-hosted node, which then becomes the bottleneck. - Assuming the Azure-SSIS IR is needed for ordinary pipelines; it is only for SSIS packages.
- Does a self-hosted integration runtime require inbound firewall rules?No. The agent makes outbound connections to the Data Factory service, so no inbound ports are opened in the corporate firewall. It needs outbound access to the service endpoints, and can be configured to use a proxy. Claiming it needs inbound access is a common giveaway.
- How do you make a self-hosted integration runtime highly available?Register several nodes against the same self-hosted IR. Work is distributed across nodes, and losing one does not take connectivity down. Size each node for the copies it must sustain, and remember it can also be shared with other data factories so one hardened cluster serves several teams.
- When can you reach a private Azure storage account without a self-hosted runtime?When you enable a managed virtual network on the Azure IR and create managed private endpoints to the resource. The factory then reaches the locked-down PaaS service privately with no machine of your own. Genuinely on-premises sources still require a self-hosted runtime.
saying these in an interview costs you the question
- Says the self-hosted runtime needs inbound firewall ports opened
- Thinks datasets reference the integration runtime directly
- Tunes Data Integration Units for a self-hosted copy
- Believes Azure-SSIS runtime is needed for ordinary pipelines
- Runs all production on-prem transfers on one single node