Monitor

Node health

See which nodes are connected, what they run and what state they are in.

What you can see today

Terminal
orchestrator-zero node list                   # every node with its state
orchestrator-zero node show <node-id>         # one node in detail
orchestrator-zero plugin list                 # each node's state with each plugin
orchestrator-zero catalog                     # which nodes are ready to run each agent
orchestrator-zero cluster info                # the servers that are up

Nodes keep a control stream open to the edge and send heartbeats over it. A node that is blocked, rejected or deleted has its stream closed within two seconds.

Health scores

The edge that holds a node's control stream scores the node at every heartbeat, every ten seconds, from 0 to 100:

score = 100 − resources − failures − latency − model calls − restarts
PenaltyUp toGrows with
Resources25The busiest of CPU, memory, disk and GPU memory over 80%, smoothed over a few heartbeats
Failures30The share of the node's activities in the last 15 minutes that failed because of the node or a plugin, over 2%. A tool that answers with an error, a model that is overloaded or a job out of budget is not the node's failure.
Latency15How much slower the node runs a type of activity than the cluster's median, from 1.5 times; at least 20 runs
Model calls10The share of its model calls in 15 minutes that got 429 or 5xx, or never reached the provider, over 5%
Restarts20Restarts of its runtime in 15 minutes; three cost everything
StateScoreWhat happens
Healthy80–100Nothing
Degraded50–79The node takes half as many activities at once until it recovers. The audit log records node.degraded.
Unhealthy0–49The node gets no new tasks and drains: its polls come back empty, and what it runs finishes as usual. The audit log records node.unhealthy.
Offline—No heartbeat for 30 seconds. Its activities time out and run elsewhere.

A disk over 95% full, or a runtime that has been down for 30 seconds, makes a node unhealthy whatever its score. A state changes after two heartbeats in a row agree, and back to healthy takes three.

Terminal
orchestrator-zero node list                   # HEALTH: state and score
orchestrator-zero node show <node-id>         # the reason and every penalty
orchestrator-zero audit --action node.        # when nodes changed state, and why
Health      unhealthy 75: disk 97% full
Penalties   resources 25.0, failures 0.0, latency 0.0, model calls 0.0, restarts 0.0 (scored 3s ago)

The web UI's Nodes page shows the same, with the penalties on the badge's tooltip. A tool with affinity: node runs only on its job's home node, so its calls wait while that node is unhealthy.

Each server answers GET /healthz on its edge port (7443) and its management port (8443), for load balancers and monitoring.

Logs

Servers and nodes log to stderr. Raise the level when you need detail:

Terminal
orchestrator-zero server start --log-level debug ...
orchestrator-zero-node run --log-level debug ...

Temporal's own logs are quieter by default; set --temporal-log-level on the server to see them.

The edge tells the node its verdict on the control stream, and the node's log says what it does about it (health verdict state=degraded ...).

Planned (M6): logs that stream from nodes to the edge.
Copyright © 2026