Upgrades, backups and restores
Upgrade servers
Upgrade one server at a time. server start applies pending schema migrations, ours and Temporal's, under a PostgreSQL advisory lock, so the first upgraded server migrates and the others wait. A binary refuses to start on a database whose schema is newer than it knows.
Temporal must move one minor version at a time. Read the release notes before you skip a release.
Upgrade nodes
Every job stays on the runtime version it started on. Each tenant's Temporal namespace has a Worker Deployment, oz0-node, and a runtime's version is its Build ID: the runtime's workers for agents, flows and gates are versioned, and jobs are pinned to the version that started them, so an upgrade never replays a job's history with other workflow code. The node's own queue, node.<id>, which carries its tool calls, is not versioned.
The edge points each tenant's deployment at the version its nodes are meant to run: the version most of the tenant's online nodes are told to run, or run when told nothing, becomes current as soon as one of them runs it, and new jobs go there. While a rollout runs, it moves the deployment instead.
Upgrade the runtime
The edge serves every runtime version in its dist directory: make dist adds the one you build to dist/runtime/, and the newest is what new nodes install. Tell nodes to run another one:
orchestrator-zero node runtimes # the versions the edge serves, and how many nodes run each
orchestrator-zero node upgrade node-7f3a01b2c4d5 --runtime 0.5.0+1a2b3c4d5e6f
orchestrator-zero node upgrade --all-in default --runtime 0.5.0+1a2b3c4d5e6f
orchestrator-zero node list # RUNTIME: what each node runs, and what it moves to
A node told to run a version downloads it from the edge, checks its digest, installs it and starts it next to the one it runs. New jobs go to the new version once most of the tenant's nodes run it, and jobs that started on the old version finish there. When no job is pinned to the old version any more, which Temporal reports within half a minute of the last one ending, the nodes stop it:
NODE NAME TENANT STATE STATUS HEALTH RUNTIME
node-40f20944d825 node-a default accepted online on dev healthy 100 0.5.0+1a2b3c4d5e6f, 0.4.0+9f8e7d6c5b4a
job get shows the runtime a job runs on. A version that cannot be downloaded or installed leaves the node on what it runs, and node show says why. node upgrade --runtime "" lets nodes keep what they run.
Upgrade the node binary
Nodes replace their own binary with the one the edge serves:
orchestrator-zero node upgrade node-7f3a01b2c4d5 --binary
orchestrator-zero node upgrade --all-in default --binary
orchestrator-zero node show node-7f3a01b2c4d5 # Supervisor: the version it runs
A node told to run another binary downloads the edge's release, and checks it before it changes anything:
- The release's
manifest.jsonmust be signed with a key the node's own binary trusts. - The binary for its platform must match the digest in that manifest.
- The new binary must start and say it is the version asked for.
Then the node keeps its old binary next to the new one as orchestrator-zero-node.previous, stops its runtimes and starts again as the new binary, in the same process. A release the node does not trust leaves it on what it runs, and its log says why.
The new binary has two minutes to reach an edge (OZ0_UPDATE_GRACE changes it). If it does not, or it stops before then, the node goes back:
- the previous binary is put back in place, and the new one is kept as
orchestrator-zero-node.failed; - the node does not install that version again. Upgrade it to another one, or delete
node-update.jsonfrom the node's data directory and restart the node to try again.
The node records the update in node-update.json before the new binary starts. A new binary that crashes, and is started again by a service manager, still finds it and goes back.
Runtimes are checked the same way: a node takes a runtime's digest only from the signed manifest. A node that starts while no edge offers a runtime it trusts, because the edge is down or serves a release the node does not trust, starts the runtime it already has.
Releases and their signatures
make dist signs dist/manifest.json with a release key, and every binary it builds trusts that key. The first time, it makes a key of its own in ~/.config/orchestrator-zero/release.key. Keep that file safe: whoever has it can sign node binaries your nodes will run. Set RELEASE_KEY to keep it somewhere else.
Mirror a release that is served elsewhere, such as a release's download folder, into the edge's dist directory:
orchestrator-zero server dist pull https://releases.example.com/orchestrator-zero/0.7.0 --dist-dir /srv/oz0/dist
The server checks the release's signature against the keys its own binary trusts, downloads every file it names and checks each digest. The manifest goes in place last, so nodes see the new release whole or not at all. Node binaries of releases older than the one before are removed from the dist directory; runtimes stay until you delete them.
A signature proves that a release is yours, not that it is the newest: someone who controls an edge can serve an older release you signed. To retire a release for good, sign the next ones with a new key and upgrade the nodes to binaries that trust only it.
Roll a runtime out in steps
node upgrade moves the nodes you name at once. A rollout moves all of a tenant's nodes in steps, and stops by itself when the new version does worse:
orchestrator-zero rollout start --runtime 0.5.0+1a2b3c4d5e6f # tenant default: a canary, then the rest in two steps
orchestrator-zero rollout start --runtime 0.5.0+1a2b3c4d5e6f --tenant acme --canary 2 --batch 5 --soak 10m
orchestrator-zero rollout show rollout-3f9c2a1b
orchestrator-zero rollout list
orchestrator-zero rollout abort rollout-3f9c2a1b # roll back by hand
In the web UI, Upgrade on the Nodes page starts one: choose a runtime the edge serves, the canary, the nodes per later step, the soak and the start timeout. The panel above the nodes follows each step to its verdict, and Roll back aborts it.
A rollout plans every online node of the tenant that does not already run only the new version: a canary of one node (--canary), then the others in steps of half of them each (--batch). Each step:
- tells its nodes to run the new version, and waits until they run it, for at most
--start-timeout(5 minutes); - sends the new version its share of the tenant's new jobs: with one node of four upgraded, a quarter of them;
- watches the upgraded nodes for
--soak(1 minute), and then judges the step.
A step passes two gates:
- Health: every upgraded node stays online, keeps running the new version and stays healthy, all through the soak. A node that goes bad fails the step at once.
- Quality: of the jobs on the new version that ended during the soak, the share that failed is at most 10 points worse than on the old version over the same time. The gate needs five such jobs; a quiet tenant passes on health alone, and the step says so.
When the last step passes, the new version becomes current and the old one drains, as above. When a step fails, or you abort, the rollout rolls back by itself: the upgraded nodes are told to run the old version again, the old version becomes current again, and jobs that started on the new version finish there. A runtime that does not start never gets past its canary. Here the canary's new runtime exited as soon as it started, with --start-timeout 20s:
Rollout rollout-0ca52a09: tenant default, runtime from 0.5.0+1a2b3c4d5e6f to 0.5.1+7c6d5e4f3a2b, rolled-back
canary failed node-07ecbf417308: step 1: node node-07ecbf417308 did not run 0.5.1+7c6d5e4f3a2b (RUNTIME_STATE_CRASHED after 4 restarts: exited with status 0) within 20s
step 2 pending node-5cd237748ea5
step 3 pending node-e8db051fed0d
Reason: step 1: node node-07ecbf417308 did not run 0.5.1+7c6d5e4f3a2b (RUNTIME_STATE_CRASHED after 4 restarts: exited with status 0) within 20s
In make e2e-rollout, which does this while jobs run back to back, none of the 39 jobs that ran meanwhile failed: the other nodes, and the canary's old runtime, took them.
A tenant has one rollout at a time: rollout start refuses while another runs. A server with the edge role drives it and holds a lease on it in the database; if that server stops, another one carries on from the same step within half a minute. The audit log has every rollout (orchestrator-zero audit --action rollout.: create, start, step, done, rolled-back and abort), and the metric oz0_rollouts_total counts how they ended.
Plugin versions roll out the same way, canary first, when a plugin that runs on two or more nodes gets a new version: see New versions roll out to canaries first. Such a rollout waits while a runtime rollout runs, and rollout list shows both kinds.
Back up
server backup writes one archive of the management database while the servers keep running. It reads every table in one snapshot, and for a dev server it copies the embedded Temporal's database too:
orchestrator-zero server backup --out oz0-backup.tar.gz # a cluster: --db or OZ0_DB
orchestrator-zero server backup --dev --data-dir .oz0/dev --out oz0-backup.tar.gz
The archive holds tenants, nodes and join tokens, plugins and their artifacts, sealed secrets, accounts and memberships, usage, the audit log, job summaries, and the large payloads that jobs keep out of their histories. It leaves out:
- the master key. Keep it apart, and give it to
server restore; - what describes the moment: live servers, nodes' connections and sign-in sessions;
- the dist directory, which
make distorserver dist pullfills again.
| What | Why | How |
|---|---|---|
| The master key file | Without it, the CAs and secrets cannot be unsealed and the cluster is lost | Copy it somewhere safe and offline once, after server init |
| The management database | Everything in the list above | server backup on a schedule, or PostgreSQL's own backups for point-in-time recovery |
| Temporal's two databases (a cluster) | Running jobs and their recent histories | PostgreSQL's own backups: the archive does not hold them |
| Operator contexts | Operator certificates | Renew them with operator renew. If one is lost or expired, issue new credentials on a server with server operator issue. Keep them private |
Nodes hold nothing you need to back up: their plugins come from the edge, and a lost node can simply join again.
Restore
orchestrator-zero server restore oz0-backup.tar.gz --master-key-file /etc/orchestrator-zero/master.key \
--db postgres://oz0@db.example.com:5432/oz0
orchestrator-zero server restore oz0-backup.tar.gz --master-key-file .oz0/dev/master.key --dev --data-dir .oz0/restored
A restore goes into an empty database: a new PostgreSQL database, or a new dev data directory. Before it writes anything, it checks every file in the archive against its digest, and checks that the master key unseals the backup's CAs. Then it inserts every row in one transaction and migrates to the binary's schema. A restore that fails leaves a dev data directory as it found it.
After a restore:
- start the servers with the same master key;
- nodes keep their certificates and reconnect when they reach the edge at the address they know;
- operators get new credentials with
server operator issue. A dev server issues its own when it starts; - accounts keep their passwords, and sign in again.
A cluster restored without Temporal's databases has everything but the jobs that were running and their recent histories. A dev backup has those too. make e2e-backup backs up a dev server while a job waits for a person, restores it into a new data directory, and the job finishes there once it is approved.