Monitor

Tracing and metrics

Metrics for Prometheus, and an OpenTelemetry trace for every job.

Metrics

Each server serves /metrics for Prometheus when you give it an address:

Terminal
orchestrator-zero server start --metrics-listen 127.0.0.1:9464 ...
curl -s http://127.0.0.1:9464/metrics | grep '^oz0_'

The endpoint is plain HTTP and not authenticated, so keep it on a private interface or behind your scraper's network. Each server counts what it saw itself; let Prometheus add them up across a cluster.

MetricLabelsWhat it counts
oz0_llm_calls_totaltenant, provider, model, statusModel calls through the gateway. status is 2xx, 4xx, 5xx, or the exact code for 429, 500, 502, 503 and 529
oz0_llm_tokens_totaltenant, provider, model, kindTokens: input, output, cache_read, cache_write
oz0_llm_cost_usd_totaltenant, provider, modelWhat the calls cost, in dollars
oz0_reported_cost_usd_totaltenant, plugin, serviceWhat plugins reported for their own services (Tools)
oz0_jobs_started_totaltenant, agent, byJobs started: by is api (the CLI, the web UI and apps), webhook or gate
oz0_nodes_connectedtenantNodes with a control stream open to this server
oz0_node_health_scoretenant, nodeEach connected node's health score
oz0_node_health_statetenant, node, state1 for the node's state: healthy, degraded or unhealthy
oz0_gate_verdicts_totaltenant, plugin, verdictQuality gates: passed or stopped
oz0_rollouts_totaltenant, kind, verdictRollouts that ended, of a runtime or a plugin (kind): done or rolled-back

Go's and the process's own metrics (go_*, process_*) are there too. A useful alert to start from:

alerts.yaml
- alert: NodeUnhealthy
  expr: oz0_node_health_state{state="unhealthy"} == 1
  for: 2m

Traces

Every job is one OpenTelemetry trace. It starts on the edge, which starts the job's workflow, and continues on the nodes that run it:

  • the job's workflow and every child job, on whichever nodes they run;
  • each activity: model requests, tool calls, hooks, the job's start;
  • each model call, with its model, provider and tokens, in OpenTelemetry's GenAI conventions (chat claude-sonnet-5-5, gen_ai.usage.input_tokens);
  • each agent's run and each of its tool calls (invoke_agent, execute_tool).

Send them to any backend that takes OTLP over HTTP, such as an OpenTelemetry Collector, Jaeger, Grafana Tempo, Honeycomb or Langfuse:

Terminal
orchestrator-zero server start --otlp-endpoint http://collector:4318
orchestrator-zero server start --otlp-endpoint https://otlp.example.com --otlp-header "Authorization=Bearer $OTLP_TOKEN"

Nodes send their spans to the edge, never to the backend: the edge tells them at connect whether it relays spans, and sets oz0.tenant and oz0.node on every span from the node's certificate, so a node cannot pass its spans off as another's. The edge's own spans carry service.name orchestrator-zero, the nodes' oz0-runtime.

Prompts and answers stay out of the spans, because they may hold your customers' data: a span shows the shape of each message, not its text. Without --otlp-endpoint, nothing is traced and nodes export nothing.

The edge does not trace the live stream's polls, so watching a job adds no spans of its own. Harness sessions are one activity span today; the Claude Agent SDK's own spans inside a session do not join the job's trace yet.

Live view

The live stream works from the CLI, the web UI and the runtime API's WatchJob, on Temporal's Workflow Streams. Each viewer follows the stream on its own today; the edge will subscribe once per job and fan it out, so many viewers do not load Temporal.

Copyright © 2026