Monitor

Job history

Every step of every job, as a trace you can read from the CLI and the web UI, with findings about what looks wrong.

Every job is a Temporal workflow, so its full history is recorded: each model call, each tool call with its arguments and result, each hook and approval, each child job, and the error when something fails. The edge reads that history and shows it as the job's steps.

A job's steps

Terminal
orchestrator-zero job trace <job-id>
Output
Job:     job-f60d03cb-c410-4352-a3b4-8ee042b805aa (lead), completed after 1.5s, $0.0024
Trace:   2e156de40610647ba330ff84c01601f4 in your tracing backend

AT      TOOK   KIND   STEP                                     NODE               TOKENS  COST     STATUS  ID
+197ms  33ms   start  job start                                node-d5fba02a1365  -       -        ok      activity-1
+324ms  105ms  model  anthropic/claude-sonnet-5-5              node-d5fba02a1365  196/20  $0.0006  ok      activity-2
+555ms  652ms  child  helper (job …/helper-e2c82af1)           node-8eda3997fd54  307/44  $0.0010  ok      child-23
+1.3s   101ms  model  anthropic/claude-sonnet-5-5              node-d5fba02a1365  269/25  $0.0008  ok      activity-3

Each row is one step, in the order they began:

KindWhat it is
queueThe job waited at least a second for a node to take it, or to take its next step
startThe job's start on its home node, which reads the hooks and approvals that apply to it
modelOne model request, with the tokens and the cost of all its attempts, as the gateway metered them
toolA tool call, on the node it ran on, with what its plugin reported for the services it called. A harness session's tool calls are listed inside the session
hookA hook and what it decided: allow, deny, modify or escalate
approvalA wait for a person, who decided and how, or timeout when nobody did
childA child job, with what its whole branch cost. job trace on its ID shows its own steps
harnessA harness session, with every model call it made

The time a step waited for a node is the gap between when it was asked for and when it started. A model call that was tried again shows its last attempt and why the one before failed. While a job runs, job trace shows what runs now, and how long it has run.

job trace --json prints the same for scripts.

What looks wrong

Below the steps come findings: rules over the trace that point at what is worth a look.

FindingWhen
idleNothing ran for ten minutes
waiting_for_nodeThe job or a step waited more than a minute for a node
repeated_toolOne tool was called ten times or more
possible_loopOne tool was called three times with the same input
repeated_errorThree steps or more failed the same way
retriesThree steps or more needed another attempt, or one step needed three
budgetThe job has spent 80% of its budget
workflow_failuresThe job's workflow code failed on a node, for example after a change that broke a running job
long_runThe job has run for more than half an hour
slow_stepA model call, tool call or hook took more than five minutes

Findings only report. A job keeps running whatever they say.

Across jobs

The edge also keeps a summary of every job in the management database: its agent, status, times, steps, calls, tokens, its own cost, its error and its findings. It writes the summary when the job ends, and every 30 seconds while it runs, so a looping job shows up before it is done. Summaries stay after Temporal has forgotten the job.

Terminal
orchestrator-zero job findings                       # jobs with a warning, the last day
orchestrator-zero job findings --code possible_loop --since 168h
orchestrator-zero job findings --code all --agent lead --json

The web UI's Findings page, under Quality, shows the same with filters for the rule, the agent and the period. Open a job there to see all its findings and go to its steps.

What a step read and wrote

A trace shows the shape of a job: names, times, tokens and cost. What a step read and wrote can be your customers' data, so it is shown one step at a time, to those allowed to see it:

Terminal
orchestrator-zero job step <job-id> activity-2

That prints the step's input and output: a model call's messages and its response, a tool call's arguments and result, what a person was asked and what they answered. Payloads kept in the edge's storage are read back for it. It needs the permission to see the tenant's content, which operators, admins and owners have and readers do not, unless an owner says otherwise (roles). Each look is recorded in the audit log, as job.step.read:

Terminal
orchestrator-zero audit --action job.step.read

Without that permission, the answer, the live stream's text, tool arguments and results, and approvals' details are left out everywhere: in the web UI, job get, job watch and the runtime API.

In your tracing backend

When the edge exports traces, job trace names the job's OpenTelemetry trace. The deep dive and the trace show the same job: the deep dive knows the waits, the approvals and the cost, and your tracing backend keeps it as long as you keep traces.

In the web UI

A job's page in the web UI shows the same steps as a waterfall, with the findings above it. Each bar starts when its step was asked for: the dashed part is time spent waiting for a node, and the solid part the step itself. Click a step to inspect it: when it began, how long it waited and ran, its node, attempts, tokens and cost, and what it decided. Admins can open its input and output there, which the audit log records too.

The page also shows the answer or the error, the live stream while the job runs, the tree of child jobs with the cost of each branch, and a link to the job's history in the Temporal UI.

Other ways in

Terminal
orchestrator-zero job list
orchestrator-zero job watch <job-id>        # follow a running job live
orchestrator-zero job get <job-id>          # status, answer or error, children, cost per branch
orchestrator-zero job get <job-id> --json   # everything, for scripts

In the Temporal UI

Every server serves Temporal's web UI at https://<server>:8443/temporal/, behind the web UI's sign-in and for admins only, since it can change any job. A job's page in the web UI links straight to its history there. In dev mode it also runs on port 8233 without a sign-in, for the local machine (orchestrator-zero dev prints its address).

Open the tenant's namespace, default unless you created others, and find the job by its ID. Each activity's summary says what it is: tool github.merge, hook pre_tool_call guard.check, job start, harness claude-agent-sdk, and request model: … for model calls. Child workflows are delegation.

How long history stays

History stays as long as the namespace keeps finished workflows, and so do traces and steps. job get can show a job's answer only while Temporal still has it. Usage rows for cost and job summaries stay in the management database longer.

Copyright © 2026