Job history
Every job is a Temporal workflow, so its full history is recorded: each model call, each tool call with its arguments and result, each hook and approval, each child job, and the error when something fails. The edge reads that history and shows it as the job's steps.
A job's steps
orchestrator-zero job trace <job-id>
Job: job-f60d03cb-c410-4352-a3b4-8ee042b805aa (lead), completed after 1.5s, $0.0024
Trace: 2e156de40610647ba330ff84c01601f4 in your tracing backend
AT TOOK KIND STEP NODE TOKENS COST STATUS ID
+197ms 33ms start job start node-d5fba02a1365 - - ok activity-1
+324ms 105ms model anthropic/claude-sonnet-5-5 node-d5fba02a1365 196/20 $0.0006 ok activity-2
+555ms 652ms child helper (job …/helper-e2c82af1) node-8eda3997fd54 307/44 $0.0010 ok child-23
+1.3s 101ms model anthropic/claude-sonnet-5-5 node-d5fba02a1365 269/25 $0.0008 ok activity-3
Each row is one step, in the order they began:
| Kind | What it is |
|---|---|
queue | The job waited at least a second for a node to take it, or to take its next step |
start | The job's start on its home node, which reads the hooks and approvals that apply to it |
model | One model request, with the tokens and the cost of all its attempts, as the gateway metered them |
tool | A tool call, on the node it ran on, with what its plugin reported for the services it called. A harness session's tool calls are listed inside the session |
hook | A hook and what it decided: allow, deny, modify or escalate |
approval | A wait for a person, who decided and how, or timeout when nobody did |
child | A child job, with what its whole branch cost. job trace on its ID shows its own steps |
harness | A harness session, with every model call it made |
The time a step waited for a node is the gap between when it was asked for and when it started. A model call that was tried again shows its last attempt and why the one before failed. While a job runs, job trace shows what runs now, and how long it has run.
job trace --json prints the same for scripts.
What looks wrong
Below the steps come findings: rules over the trace that point at what is worth a look.
| Finding | When |
|---|---|
idle | Nothing ran for ten minutes |
waiting_for_node | The job or a step waited more than a minute for a node |
repeated_tool | One tool was called ten times or more |
possible_loop | One tool was called three times with the same input |
repeated_error | Three steps or more failed the same way |
retries | Three steps or more needed another attempt, or one step needed three |
budget | The job has spent 80% of its budget |
workflow_failures | The job's workflow code failed on a node, for example after a change that broke a running job |
long_run | The job has run for more than half an hour |
slow_step | A model call, tool call or hook took more than five minutes |
Findings only report. A job keeps running whatever they say.
Across jobs
The edge also keeps a summary of every job in the management database: its agent, status, times, steps, calls, tokens, its own cost, its error and its findings. It writes the summary when the job ends, and every 30 seconds while it runs, so a looping job shows up before it is done. Summaries stay after Temporal has forgotten the job.
orchestrator-zero job findings # jobs with a warning, the last day
orchestrator-zero job findings --code possible_loop --since 168h
orchestrator-zero job findings --code all --agent lead --json
The web UI's Findings page, under Quality, shows the same with filters for the rule, the agent and the period. Open a job there to see all its findings and go to its steps.
What a step read and wrote
A trace shows the shape of a job: names, times, tokens and cost. What a step read and wrote can be your customers' data, so it is shown one step at a time, to those allowed to see it:
orchestrator-zero job step <job-id> activity-2
That prints the step's input and output: a model call's messages and its response, a tool call's arguments and result, what a person was asked and what they answered. Payloads kept in the edge's storage are read back for it. It needs the permission to see the tenant's content, which operators, admins and owners have and readers do not, unless an owner says otherwise (roles). Each look is recorded in the audit log, as job.step.read:
orchestrator-zero audit --action job.step.read
Without that permission, the answer, the live stream's text, tool arguments and results, and approvals' details are left out everywhere: in the web UI, job get, job watch and the runtime API.
In your tracing backend
When the edge exports traces, job trace names the job's OpenTelemetry trace. The deep dive and the trace show the same job: the deep dive knows the waits, the approvals and the cost, and your tracing backend keeps it as long as you keep traces.
In the web UI
A job's page in the web UI shows the same steps as a waterfall, with the findings above it. Each bar starts when its step was asked for: the dashed part is time spent waiting for a node, and the solid part the step itself. Click a step to inspect it: when it began, how long it waited and ran, its node, attempts, tokens and cost, and what it decided. Admins can open its input and output there, which the audit log records too.
The page also shows the answer or the error, the live stream while the job runs, the tree of child jobs with the cost of each branch, and a link to the job's history in the Temporal UI.
Other ways in
orchestrator-zero job list
orchestrator-zero job watch <job-id> # follow a running job live
orchestrator-zero job get <job-id> # status, answer or error, children, cost per branch
orchestrator-zero job get <job-id> --json # everything, for scripts
In the Temporal UI
Every server serves Temporal's web UI at https://<server>:8443/temporal/, behind the web UI's sign-in and for admins only, since it can change any job. A job's page in the web UI links straight to its history there. In dev mode it also runs on port 8233 without a sign-in, for the local machine (orchestrator-zero dev prints its address).
Open the tenant's namespace, default unless you created others, and find the job by its ID. Each activity's summary says what it is: tool github.merge, hook pre_tool_call guard.check, job start, harness claude-agent-sdk, and request model: … for model calls. Child workflows are delegation.
How long history stays
History stays as long as the namespace keeps finished workflows, and so do traces and steps. job get can show a job's answer only while Temporal still has it. Usage rows for cost and job summaries stay in the management database longer.