Concepts

Jobs and durability

Why a job survives crashes, restarts and lost nodes, and what that means for your agents.

A job is one run of an agent. Under the hood it is a Temporal workflow, AgentWorkflow, and its job ID is the workflow ID. Nothing else stores the job: Temporal holds its status, input, answer and full history.

What durable means here

  • The agent's loop is workflow code. Temporal records every step in the job's history.
  • Every model call is an activity. It streams, has a ten-minute timeout and up to five attempts, and heartbeats while it runs. A crashed node loses at most the call in flight. Temporal is its only retry layer: the provider SDKs do not retry on their own, so an attempt is one call to the provider, and it waits as long as the provider's retry-after asks. An error no retry can fix, such as a bad request, a spent cap or budget, or a missing key, fails the job at once.
  • Every tool call is an activity on the home node's queue. Its result is recorded, so it is not called again on replay. Tools can have side effects, so a failed tool call is not retried; the error goes back to the model, which decides what to do.
  • Stopping the client does not stop the job. orchestrator-zero run waits for the answer, but if you press Ctrl+C the job keeps running. Follow it with job get.

If the server restarts, Temporal resumes the job where it was. If a node crashes in the middle of a step, another node with the same agent takes the step over. Tool calls stay on the home node, so a job whose home node is gone waits for it when it needs a tool; moving the home to another node is planned.

IDs and trees

Top-level jobs get an ID such as job-bb270cf1-b194-4bf1-a0e3-267e061e416a, or the one you pass with --id. Starting a job with an ID that is still running follows that job instead of starting a new one, which makes retries from apps safe.

Children carry their parent's ID: <parent>/<agent>-<8 hex>. That is how job get finds a job's tree and how cost is added up per branch.

Limits and contracts

The platform enforces limits and contracts in workflow code, so a model cannot talk its way around them:

  • max_steps, max_tokens and timeout end a job when reached.
  • An input that does not match the agent's input_schema fails the job before it starts.
  • An answer that does not match output_schema goes back to the model with the reason, at most twice.

See Contracts and limits.

History

Every job has a full history in Temporal: each model call with its request and response, each tool call and each child. The edge reads it as the job's steps, with what each one cost: orchestrator-zero job trace, and the job's page in the web UI (Job history). In dev mode the Temporal UI shows the history itself in a browser. History stays as long as the namespace keeps finished workflows; usage rows for cost stay longer.

Payloads larger than 256 KiB, such as a large input or a long conversation that every model call carries, do not go into the history itself: the node stores them on the edge, which keeps them per tenant in the database, and the history holds a reference (ADR 0035). So a job can take inputs of many megabytes, up to 32 MiB a payload, where Temporal alone would refuse anything over 2 MB. The edge deletes a payload 90 days after a run last stored it. The Temporal UI shows such a payload as its reference.

Copyright © 2026