Jobs and durability
A job is one run of an agent. Under the hood it is a Temporal workflow, AgentWorkflow, and its job ID is the workflow ID. Nothing else stores the job: Temporal holds its status, input, answer and full history.
What durable means here
- The agent's loop is workflow code. Temporal records every step in the job's history.
- Every model call is an activity. It streams, has a ten-minute timeout and up to five attempts, and heartbeats while it runs. A crashed node loses at most the call in flight. Temporal is its only retry layer: the provider SDKs do not retry on their own, so an attempt is one call to the provider, and it waits as long as the provider's
retry-afterasks. An error no retry can fix, such as a bad request, a spent cap or budget, or a missing key, fails the job at once. - Every tool call is an activity on the home node's queue. Its result is recorded, so it is not called again on replay. Tools can have side effects, so a failed tool call is not retried; the error goes back to the model, which decides what to do.
- Stopping the client does not stop the job.
orchestrator-zero runwaits for the answer, but if you press Ctrl+C the job keeps running. Follow it withjob get.
If the server restarts, Temporal resumes the job where it was. If a node crashes in the middle of a step, another node with the same agent takes the step over. Tool calls stay on the home node, so a job whose home node is gone waits for it when it needs a tool; moving the home to another node is planned.
IDs and trees
Top-level jobs get an ID such as job-bb270cf1-b194-4bf1-a0e3-267e061e416a, or the one you pass with --id. Starting a job with an ID that is still running follows that job instead of starting a new one, which makes retries from apps safe.
Children carry their parent's ID: <parent>/<agent>-<8 hex>. That is how job get finds a job's tree and how cost is added up per branch.
Limits and contracts
The platform enforces limits and contracts in workflow code, so a model cannot talk its way around them:
max_steps,max_tokensandtimeoutend a job when reached.- An input that does not match the agent's
input_schemafails the job before it starts. - An answer that does not match
output_schemagoes back to the model with the reason, at most twice.
See Contracts and limits.
History
Every job has a full history in Temporal: each model call with its request and response, each tool call and each child. The edge reads it as the job's steps, with what each one cost: orchestrator-zero job trace, and the job's page in the web UI (Job history). In dev mode the Temporal UI shows the history itself in a browser. History stays as long as the namespace keeps finished workflows; usage rows for cost stay longer.
Payloads larger than 256 KiB, such as a large input or a long conversation that every model call carries, do not go into the history itself: the node stores them on the edge, which keeps them per tenant in the database, and the history holds a reference (ADR 0035). So a job can take inputs of many megabytes, up to 32 MiB a payload, where Temporal alone would refuse anything over 2 MB. The edge deletes a payload 90 days after a run last stored it. The Temporal UI shows such a payload as its reference.