Node health
What you can see today
orchestrator-zero node list # every node with its state
orchestrator-zero node show <node-id> # one node in detail
orchestrator-zero plugin list # each node's state with each plugin
orchestrator-zero catalog # which nodes are ready to run each agent
orchestrator-zero cluster info # the servers that are up
Nodes keep a control stream open to the edge and send heartbeats over it. A node that is blocked, rejected or deleted has its stream closed within two seconds.
Health scores
The edge that holds a node's control stream scores the node at every heartbeat, every ten seconds, from 0 to 100:
score = 100 − resources − failures − latency − model calls − restarts
| Penalty | Up to | Grows with |
|---|---|---|
| Resources | 25 | The busiest of CPU, memory, disk and GPU memory over 80%, smoothed over a few heartbeats |
| Failures | 30 | The share of the node's activities in the last 15 minutes that failed because of the node or a plugin, over 2%. A tool that answers with an error, a model that is overloaded or a job out of budget is not the node's failure. |
| Latency | 15 | How much slower the node runs a type of activity than the cluster's median, from 1.5 times; at least 20 runs |
| Model calls | 10 | The share of its model calls in 15 minutes that got 429 or 5xx, or never reached the provider, over 5% |
| Restarts | 20 | Restarts of its runtime in 15 minutes; three cost everything |
| State | Score | What happens |
|---|---|---|
| Healthy | 80–100 | Nothing |
| Degraded | 50–79 | The node takes half as many activities at once until it recovers. The audit log records node.degraded. |
| Unhealthy | 0–49 | The node gets no new tasks and drains: its polls come back empty, and what it runs finishes as usual. The audit log records node.unhealthy. |
| Offline | — | No heartbeat for 30 seconds. Its activities time out and run elsewhere. |
A disk over 95% full, or a runtime that has been down for 30 seconds, makes a node unhealthy whatever its score. A state changes after two heartbeats in a row agree, and back to healthy takes three.
orchestrator-zero node list # HEALTH: state and score
orchestrator-zero node show <node-id> # the reason and every penalty
orchestrator-zero audit --action node. # when nodes changed state, and why
Health unhealthy 75: disk 97% full
Penalties resources 25.0, failures 0.0, latency 0.0, model calls 0.0, restarts 0.0 (scored 3s ago)
The web UI's Nodes page shows the same, with the penalties on the badge's tooltip. A tool with affinity: node runs only on its job's home node, so its calls wait while that node is unhealthy.
Each server answers GET /healthz on its edge port (7443) and its management port (8443), for load balancers and monitoring.
Logs
Servers and nodes log to stderr. Raise the level when you need detail:
orchestrator-zero server start --log-level debug ...
orchestrator-zero-node run --log-level debug ...
Temporal's own logs are quieter by default; set --temporal-log-level on the server to see them.
The edge tells the node its verdict on the control stream, and the node's log says what it does about it (health verdict state=degraded ...).