skip to content

In A2A, how does a caller track a remote task that runs for minutes?

level: seniorimportance: should knowfreq 38%

answer

  1. a connection cannot hold the work
  2. the work item gets its own identity
  3. states including one that asks you something
  4. stream, webhook, or poll with backoff
  5. persist the id, never blind-resubmit

basics

~20 s

A2A makes the unit of work a task with an id and a lifecycle, not a single request/response. The caller subscribes to streamed status and artifact updates, or registers a webhook for push notifications, and can re-attach to the task by id after a disconnect.

solid answer

~50 s

Holding an HTTP request open for four minutes while a supplier agent prices a bill of materials is a design that fails on the first proxy timeout. A2A models the work as a **task**: the caller submits a message, receives a task id, and the task moves through states — submitted, working, input-required, and terminal states such as completed, failed or canceled. Progress arrives one of two ways: a streamed subscription that pushes status changes and artifacts as they are produced, or push notifications to a caller-supplied webhook for tasks long enough that keeping any connection open is unrealistic. Both are declared as capabilities on the peer's Agent Card, so the caller knows which is available. The design consequence is that the caller must be **resumable**: it holds the task id durably, tolerates a dropped stream by re-attaching or polling, and treats input-required as a normal branch — the remote agent asking for a missing purchase-order number rather than failing.

code

json · 10 lines
json
{
  "task_id": "t-8fd21c",
  "updates": [
    { "state": "submitted", "at": "2026-08-19T09:00:00Z" },
    { "state": "working", "at": "2026-08-19T09:00:04Z", "note": "pricing 40 line items" },
    { "state": "input-required", "at": "2026-08-19T09:01:11Z", "asks_for": "purchase_order_number" },
    { "state": "working", "at": "2026-08-19T09:03:02Z" },
    { "state": "completed", "at": "2026-08-19T09:04:26Z", "artifacts": ["quote-4471.json"] }
  ]
}

go deeper

for a junior

Know that A2A work is a task with an id and a lifecycle, and that a caller gets progress through streamed updates rather than by waiting on one response.

for a middle

Name the states and explain why input-required exists — an autonomous peer can answer with a question, and the caller must be able to supply the missing value and resume.

for a senior

Show the operational discipline: persist the task id, re-attach rather than resubmit, set your own wall-clock timeout for zombie tasks, and authenticate any notification webhook.

for a principal

Own when this machinery is justified at all. Weigh duration, deployment boundary and the cost of losing in-flight work against the complexity of resumable callers and duplicate-suppression at the business layer.

## Why request/response breaks Agent work is not API work. A supplier agent asked to quote a 40-line bill of materials may call its own tools, consult an inventory system, and run several model turns. Minutes, not milliseconds. Every layer between the two parties is hostile to that: load balancers and proxies time out idle connections, clients get redeployed mid-flight, and a dropped socket destroys work that cost real money to produce. So an agent protocol cannot treat a call as one round trip. It has to make the *work item* addressable independently of the connection that started it. ## The task as a first-class object In A2A the unit is a **task**, created when the caller sends its first message and identified by a task id that both sides use from then on. The task carries state, the message history that produced it, and any **artifacts** the remote agent has produced — the actual deliverables, which may be structured data, files, or text. Because the task has an identity separate from the transport, the caller can lose its connection, restart, and come back to the same task. That is the property that makes long-running cross-organization work survivable. ## The lifecycle A task moves through states the caller must handle distinctly: - **submitted** — accepted, not yet started. - **working** — in progress; interim status and artifacts may arrive. - **input-required** — the remote agent needs something from the caller to continue. This is not an error. A supplier agent that discovers your request lacks a purchase-order number should ask, and the caller should be built to answer and resume, not to fail the task. - **completed / failed / canceled** — terminal. Terminal means no further updates; a subsequent change requires a new task. The pattern to internalize is that **an autonomous peer can respond with a question**. Code that assumes every remote call resolves to a value or an exception will mishandle the most useful state in the list. ## Getting updates: three shapes 1. **Streaming subscription.** The caller opens a stream and the peer pushes status and artifact updates as they occur (server-sent events being the usual carrier). Best for interactive work where a human is watching a progress indicator and the session lives minutes, not hours. 2. **Push notifications.** The caller registers a webhook; the peer calls it on state changes. This is the option for work measured in tens of minutes or longer, or where the caller is a serverless function that will not exist when the answer arrives. It brings its own obligations: the webhook is a public endpoint, so it must authenticate the caller and verify that the notification actually corresponds to a task you started. 3. **Polling by task id.** Always available as a fallback, and the honest answer when a peer declares no streaming capability. Poll with backoff, not in a tight loop. Which options exist is declared as a capability on the peer's Agent Card, so a caller can check before designing around one. ## What the caller owes - **Durability of the task id.** If the id lives only in memory, a restart orphans work you have already paid for. Persist it with the work item it belongs to. - **Re-attachment logic.** A dropped stream is normal, not exceptional. Reconnect or fall back to polling; do not resubmit, which risks a duplicate task and, in a procurement setting, a duplicate order. - **Idempotency at the business layer.** Because retries and reconnects are routine, a resubmitted request must not create a second obligation. - **A timeout of your own.** The protocol does not promise the peer will ever reach a terminal state. The caller needs a wall-clock budget and a decision about what to do when it expires. - **Treating artifacts as untrusted.** Content from an external agent is input, not instruction, and validating it against the expected shape is the boundary discipline of any cross-organization exchange. ## Failure modes worth naming - **The zombie task**: peer stops updating, never reaches terminal state. Only the caller's own timeout resolves this. - **Duplicate submission**: caller treats a lost connection as a lost task and resubmits, and now two supplier agents are quoting the same line. - **Silent partial success**: task fails after emitting artifacts. Partial artifacts are real output; decide explicitly whether you keep them. - **Webhook spoofing**: an unauthenticated notification endpoint accepting a fabricated "completed" callback with attacker-chosen artifacts. - **Blocking on input-required**: nobody is watching, so a task that asked a question sits forever. Route these to a human queue with a deadline. ## When you do not need any of this If the remote step finishes in two seconds and both ends deploy from your repository, a plain synchronous call is correct and a task lifecycle is overhead. The machinery earns its place when duration exceeds what a connection can hold, when the peer is outside your deployment, or when the work is valuable enough that losing it to a socket reset is unacceptable.

  • Your stream drops at minute three. What do you do, and what must you not do?
    Re-attach using the task id, or fall back to polling with backoff — the task is still running on the peer regardless of your connection. What you must not do is resubmit the original message: that creates a second task, and in a procurement or payment setting a second obligation. This is why the task id has to be persisted with the work item rather than held in memory.
  • How is input-required different from a failure, and why does that distinction matter architecturally?
    It means the peer is blocked on you, not broken. Architecturally it forces the caller to be a participant rather than a caller: it needs a path to supply the missing value, from cached context, from a lookup, or from a human queue with a deadline. Systems that map every non-success state to an exception silently abandon tasks that were one field away from completing.
  • What has to be true of a webhook you register for push notifications?
    It must authenticate the notifier and bind the notification to a task you actually started, or anyone who learns the URL can assert that a task completed with artifacts of their choosing. Treat it as a public, attacker-reachable endpoint: verify the sender, check the task id against your own records, and validate artifact payloads against the expected schema before acting on them.

saying these in an interview costs you the question

  • Holds an HTTP request open until the agent finishes
  • Treats input-required as a task failure
  • Resubmits the message when the stream drops
  • Keeps the task id only in process memory
  • Trusts an unauthenticated completion webhook

context