Run a fleet
High availability
Run two or more servers so nodes and jobs keep going when one stops.
Run at least two servers. Every server is a full member: management, edge and Temporal. When one stops, its nodes move to another server's edge within seconds and jobs carry on.
How nodes fail over
- Every edge certificate carries the extra name
edge.oz0, so one TLS configuration fits every edge of the cluster. - When a node connects, the edge sends it the list of every server's edge that reported in the last 30 seconds, and sends it again when the list changes.
- The node keeps that list and always tries the address it joined through first, then the others.
- A server that shuts down closes its control streams at once and gives Temporal calls in flight two seconds, so its nodes move quickly.
Behind a load balancer
If nodes reach the edges through a load balancer or a DNS name, give that name to every server with --san, so the edge's certificate is valid for it:
Terminal
orchestrator-zero server start --master-key-file /etc/orchestrator-zero/master.key \
--advertise 10.0.0.11 --san edge.example.com
Join nodes with --server https://edge.example.com:7443. They keep trying that address first, so the load balancer stays in charge.
What still depends on one thing
- PostgreSQL. Every server uses it. Run it with your usual high-availability setup. If it is down, management stops; the edges keep authorizing nodes from their cached registry, and with embedded Temporal, jobs pause.
- The master key. Every server needs a copy. Store it like any other critical key.