Describe the connector restart endpoint and the includeTasks and onlyFailed query parameters. How would you restart just the failed tasks of a connector?
answer
- POST /restart, no params = connector only, 204
- includeTasks=true -> tasks too
- onlyFailed=true -> only FAILED
- both = heal failed tasks (KIP-745, 3.0)
- tasks/{id}/restart = single task
basics
~10 sPOST /connectors/{name}/restart restarts only the Connector instance by default. Add ?includeTasks=true to also restart tasks, and ?onlyFailed=true to limit the restart to FAILED instances. To restart just failed tasks: POST /connectors/{name}/restart?includeTasks=true&onlyFailed=true.
solid answer
~40 sPOST /connectors/{name}/restart, with no parameters, restarts only the Connector object — not its tasks — and returns 204 No Content. Two query parameters (added in KIP-745, Kafka 3.0) make it useful for bulk recovery: includeTasks=true expands the restart to the connector's tasks, and onlyFailed=true restricts the operation to instances currently in the FAILED state. Combined as ?includeTasks=true&onlyFailed=true, it restarts the connector and/or its tasks only where they have failed — the idiomatic 'heal the failed tasks' call. When parameters are supplied the endpoint returns 200 with a body summarizing the restart plan (which connector/tasks were restarted). For a single task you can instead POST /connectors/{name}/tasks/{taskId}/restart. The default (no params) is a common gotcha: people expect it to restart tasks too, but it only restarts the connector, leaving FAILED tasks down.
go deeper
Know there is a POST restart endpoint to recover a connector.
Know includeTasks and onlyFailed exist and what each does.
Compose ?includeTasks=true&onlyFailed=true as the targeted heal call and pair it with /status polling.
Design self-healing automation (poll status -> targeted restart -> escalate) and tie restarts to errors.tolerance/DLQ to avoid restart loops.
When a Kafka Connect task throws an unrecoverable exception it enters the **FAILED** state and stays there — Connect never auto-restarts it. The restart endpoint is how you recover. **POST /connectors/{name}/restart** — the base behavior, with **no query parameters**, restarts **only the Connector instance** and **returns 204 No Content**. Critically, it does **not** touch the tasks. Since data flows through tasks, restarting only the connector rarely fixes an outage — this is the most common misconception about the endpoint. **KIP-745 (Kafka 3.0)** added two boolean query parameters that make the endpoint genuinely useful for recovery: - **`includeTasks`** (default false): when true, the restart also targets the connector's **tasks**, not just the Connector object. - **`onlyFailed`** (default false): when true, the operation is restricted to instances in the **FAILED** state (and the connector if it is FAILED). RUNNING instances are left untouched. **The recovery recipe** — restart only what's broken: ``` POST /connectors/{name}/restart?includeTasks=true&onlyFailed=true ``` This restarts the connector and any FAILED tasks while leaving healthy tasks running. It's the call you script after polling /status and finding FAILED tasks. **Parameter matrix:** - `()` — connector only, healthy or not. 204. - `includeTasks=true` — connector + all tasks (full bounce). - `onlyFailed=true` — connector only, and only if FAILED. - `includeTasks=true&onlyFailed=true` — connector + tasks, but only the FAILED ones. **Response semantics:** with the KIP-745 parameters present, the endpoint returns **200/202** with a JSON body describing the restart plan (the connector state and the list of tasks that were restarted). The bare parameterless form keeps the legacy **204** response. **Single-task restart:** **POST /connectors/{name}/tasks/{taskId}/restart** restarts exactly one task by index (returns 204). Use this when you know precisely which task to recover and don't want the plan-based bulk endpoint. **Operational note:** restarting a task re-instantiates it; a sink task resumes from its last committed consumer offset and a source task from its last committed source offset, so at-least-once semantics hold and you may see brief reprocessing. If the underlying cause (bad data, unreachable system) persists, the task will simply fail again — restarts are remediation, not a fix for the root cause. Pairing restarts with error-handling configs (errors.tolerance, dead-letter queue via errors.deadletterqueue.topic.name) prevents the failure recurring.
- What does a bare POST /connectors/{name}/restart (no query params) actually restart?Only the Connector instance, not its tasks — returning 204. Failed tasks stay FAILED. You need includeTasks=true to restart tasks.
- After restarting failed tasks they fail again immediately. What's your next move?Restart only remediates; the root cause persists. Read the /status trace, fix the underlying issue (bad record/unreachable system), and configure errors.tolerance with a dead-letter queue so poison records don't kill the task.
saying these in an interview costs you the question
- Believing POST /restart restarts tasks by default — it restarts only the connector unless includeTasks=true.
- Confusing onlyFailed (filters by state) with includeTasks (expands scope to tasks).
- Thinking restart fixes root causes — it just re-instantiates; a persistent error re-fails.
- Assuming there's a single-call 'restart everything failed' before KIP-745 — those params arrived in Kafka 3.0.