Build plugins

Harnesses

Run a whole session of an existing agent harness, such as the Claude Agent SDK, as a job.

Some agents are better run by an existing harness than by a framework. The Claude Agent SDK brings its own tools (reading and editing files, searching, running commands), skills and subagents. A harness plugin adapts a harness to Orchestrator Zero: each job runs as one session of it, in its own process on the node.

Use the Claude Agent SDK

The official harness plugin is harness-claude. Install it on the nodes that should run harness agents, and set the tenant's Anthropic key once:

Terminal
orchestrator-zero plugin install <git url of the harness-claude plugin>
printf %s "$ANTHROPIC_API_KEY" | orchestrator-zero secret set ANTHROPIC_API_KEY

Then give an agent, in any plugin, runtime: harness:claude-agent-sdk:

agents/coder.yaml
name: coder
description: Makes small changes to code and runs the tests.
runtime: harness:claude-agent-sdk
model:
  primary: anthropic/claude-sonnet-5-5
instructions: prompts/coder.md        # added to the SDK's own system prompt
limits:
  max_steps: 30                       # turns
  timeout: 30m
harness:                              # options for the harness, passed on as they are
  tools: [Read, Write, Edit, Glob, Grep, Bash]   # the default; anything else is refused
  permission_mode: acceptEdits        # default, acceptEdits, plan or bypassPermissions
  system_prompt: append               # or replace, to use only the instructions
  max_budget_usd: 2.0                 # the SDK stops the session at this cost

Run it like any agent: orchestrator-zero run coder "Fix the failing test in tests/test_dates.py". Its text and tool calls appear in the live stream, and its model calls are metered under the job.

What a session gets

  • An empty working directory of its own, removed when the session ends, and a private state directory for the harness's configuration.
  • No user, project or local settings, no memory between sessions, no telemetry and no auto-update.
  • The node's LLM proxy as its API, so the tenant's key stays in the edge, with the job and agent on every call so the gateway meters the session.
  • A minimal environment: never the Temporal proxy's token or a provider key.

Hooks and approvals for the session's tools

Every tool call in a Claude harness session asks first (ADR 0031). The session can use only the tools in harness.tools, and each call runs only when the job allows it:

  • Hooks see it as a pre_tool_call named <harness>.<tool>, such as claude-agent-sdk.Bash, with the tool's input as arguments. They allow it, deny it, change its input or ask a person, as for any tool call.
  • Tools in harness.requires_approval wait for a person's approval of each call, once the hooks have let it through:
agents/coder.yaml
harness:
  tools: [Read, Glob, Grep, Bash]
  requires_approval: [Bash]          # someone approves every command

A call that no hook matches and that needs no approval runs at once. The others go to the job's workflow, through the edge, so a call can wait days for a person without the session losing its place. A refused call is not made: the agent reads why, as the SDK reports a denied tool.

A sandbox for sessions comes later. Hooks and approvals decide whether a call runs, not what it can reach once it does. Give an agent only the tools its task needs, and keep agents with Bash, Write or Edit away from input you do not trust: plugin install refuses a flow a webhook starts that can reach one (ADR 0030).

A session that fails fails the job; it is not retried. Inside the session, the Claude Agent SDK retries a failed model call itself, up to four times, so a call reaches the provider at most five times, as for Pydantic AI agents. Resuming a session on another node needs the harness's session store and comes later.

Write a harness for another tool

A harness is a command in the plugin's manifest:

oz0-plugin.yaml
harnesses:
  - name: my-harness
    command: ["python", "-m", "my_harness"]

Agents use it with runtime: harness:my-harness. The node starts the command once per session and talks to it in JSON lines over stdin and stdout:

  1. The runtime writes one session request to stdin: the job and agent, the model, the instructions, the input, the working and state directories, the turn limit and the agent's harness: options.
  2. The harness writes session events to stdout, one JSON object per line: text, tool calls and tool results, and a result last. Logs go to stderr.
  3. The harness sends the request's job and agent as the headers Oz0-Job and Oz0-Agent on its model calls.
  4. A harness with permissions: true in the manifest asks before each tool call. When the request has permissions set, the harness writes a permission request event (id, tool, input) and waits: the runtime answers on stdin with a permission answer (id, allow, input, reason), and the call runs only if it is allowed, with the answer's input when it has one. The runtime closes stdin when the session's result is out.
  5. Cancelling a job sends SIGTERM to the session's process group, then SIGKILL after ten seconds.

The messages are defined in proto/oz0/harness/v1/harness.proto and written in their proto3 JSON form. A harness can be written in any language and tested with a pipe; node/runtime/tests/fake_harness.py in the repository is a complete example in about 50 lines, permission requests included.

One harness for the Agent Client Protocol (ACP) is planned for v1.1, in place of separate Codex, OpenCode and goose harnesses: it reaches them and Gemini CLI, Cline and others through one protocol, and sends their permission requests to your hooks and approvals.

Copyright © 2026