Evals and quality gates
Evals are test cases in your plugin: an input for an agent or a flow, and what a good answer looks like. orchestrator-zero plugin test runs them on a dev server and a node of its own, and model answers can be recorded once and replayed after, so tests in CI are fast, free and repeatable.
agent: greeter # or flow: <name>
threshold: 0.8 # the share of cases that must pass; all of them by default
cases:
- name: greets by name
input: Please greet Ada
expect:
contains: ["Hello, Ada"]
cel: "calls <= 3 && cost_usd < 0.02"
- name: refuses an empty name
input: { name: "" }
expect:
status: failed
error_contains: InvalidInput
List eval files or folders in the manifest with evals: [evals/]. plugin lint checks them, CEL included.
What a case can expect
Every expectation that is set must hold:
| Expectation | Holds when |
|---|---|
status | The job ends completed (the default) or failed |
error_contains | The failed job's error contains this text (with status: failed) |
contains, not_contains | The answer, as text (JSON for structured answers), contains each, or none, of these |
equals | The answer is exactly this value |
cel | This CEL expression is true; it sees output (the answer), cost_usd, calls and seconds |
max_cost_usd, max_seconds | The job cost and took no more than this |
An eval passes when the share of passing cases reaches its threshold.
Run them
orchestrator-zero plugin test ./greeter # every eval of the plugin
orchestrator-zero plugin test ./greeter --eval greetings
PASS greetings / greets by name (replayed, 2 calls, $0.0011, 0.5 s)
FAIL greetings / says goodbye (replayed, 2 calls, $0.0011, 0.3 s): the answer does not contain "Goodbye"
greetings (greeter): 2 of 3 passed (67%), threshold 60%: ok; $0.0032, 1.2 s
The command starts a dev server and a node, installs the plugin from its working tree, and runs the cases one at a time as jobs. It exits non-zero when an eval is below its threshold, so it can gate a pull request. --json prints the report for machines. It needs what orchestrator-zero dev needs: the node binary next to the CLI and a runtime bundle.
Record once, replay after
Every model call of a case goes through the test server's gateway, and from there through a recorder in the CLI:
--recordcalls the real API for every case, withANTHROPIC_API_KEYorOPENAI_API_KEYfrom your environment, and keeps each answer inevals/cassettes/<eval>/<case>.jsonl.- Without flags, a case that has a recording is answered from it, in order, without calling anyone and without a key; a case without one calls the real API.
--liveignores recordings.
Commit the recordings with the plugin: they hold what the model answered, never your key or the request's headers. When a prompt changes, a replay still answers in order and warns that the call differs from the recording; record again to bring it up to date. A case that makes more calls than its recording holds fails and says so. Costs in a replay are what the recorded calls cost.
Quality gates
When you install a version of a plugin that has evals, the edge runs them before the fleet gets the version, against your tenant's real models and keys:
The version waits as the plugin's candidate
The fleet keeps the version it has. A plugin's first version waits too: its agents run once it passes.
One node tests it
The edge picks an online node of the tenant that matches the plugin's selector and the labels its agents need, and prefers nodes labelled oz0.gate=true. That node runs the new version instead of the installed one, on queues of its own: no fleet job reaches it, and no eval case reaches the fleet. While no such node is online, the gate waits for one.
The evals run there
Every case runs as a job through the gateway, as in production: real models, real tools, real cost. Recordings are not used.
The verdict
The version passes when every eval reaches its threshold and passes at least the share of cases that the installed version passed at its own gate. It then becomes the installed version and rolls out, and its results become the bar for the next version. Otherwise it is stopped: the fleet keeps the installed version and the test node goes back to it.
orchestrator-zero plugin install https://github.com/acme/team --ref v0.2.0
Testing team 0.2.0 (commit ff4d7eae2c6d…) in tenant default before the fleet gets it: its evals run on a test node.
Until it passes, the fleet keeps team 0.1.0.
Follow with: orchestrator-zero plugin list (--no-gate installs without the evals)
plugin list and the Plugins page of the web UI show the gate: waiting, testing on a node, or how it ended and why, with each eval's result. A version that does worse reads like this:
Gate of team 0.2.0: stopped: eval greetings: 0 of 1 cases passed (0%), below its threshold of 100%
The audit log records every verdict. plugin install --no-gate installs a version at once, without running its evals.
Good to know:
- A dedicated test node keeps gates away from the nodes that serve jobs: give it the label
oz0.gate=truewith its join token (token create --label oz0.gate=true) or when you accept it. While a gate runs, its node serves no fleet jobs of that plugin, so a tenant with a single node waits for the gate before those agents run again. - A version that cannot start on the test node, or does not start within 10 minutes, is stopped. A case may run for 10 minutes.
- Children stay with the candidate: a flow's agent steps and an agent's delegations to agents of the same plugin run the new version on the test node; agents of other plugins answer from the fleet, in their installed versions.
- Installing again during a gate: the version under test again (a
plugin updatewhile it runs) lets its gate go on; a newer version replaces it and its gate starts over; the installed version again, or--no-gate, cancels it. A version that was stopped is tested again when you install it again. orchestrator-zero devandplugin testinstall your working tree without a gate.- Servers share the work: any server can drive a gate. If it stops, another one runs the gate again from the start within 30 seconds.