Contracts and limits
Users build the agents; the platform makes sure they stay inside the lines. None of this relies on the agent behaving well: it is enforced in workflow code.
Input contract
input_schema: schemas/review-input.json
{
"type": "object",
"required": ["repo", "pr"],
"properties": {
"repo": { "type": "string" },
"pr": { "type": "integer" }
}
}
A job whose input does not match fails before the agent starts, with InvalidInput, and is not retried. With an input schema, start jobs with JSON: orchestrator-zero run reviewer '{"repo": "acme/api", "pr": 42}'.
Output contract
output_schema: schemas/review.json
The model is asked to answer in that structure. An answer that does not match goes back to the model with the reason, at most twice; after that the job fails with a clear error. The job's answer is then JSON that matches the schema, ready for an app to use.
Limits
limits:
max_steps: 20 # model calls per job (default 50)
max_tokens: 200000 # input plus output tokens per job
timeout: 10m # wall-clock time, such as 90s, 10m or 1h30m
max_cost_usd: 2.0 # the job's default budget, for it and every job it delegates to
A job that reaches a limit ends with the reason. run --timeout adds a limit on top for one job.
Budgets
A budget caps what a job and every job below it may spend on model calls, in US dollars. limits.max_cost_usd sets the default for an agent's jobs, and run --budget (or budget_usd in the runtime API) sets it for one job:
orchestrator-zero run reviewer --budget 0.50 "Review PR 42 in acme/api"
The LLM gateway checks every call against what the whole tree has spent, costs that plugins report for their own services included. Once the budget is gone it refuses the next call with 402 oz0_budget_exceeded, and the job fails with BudgetExceeded without retrying. The check happens before each call, so a job can end a call's worth over its budget. Calls that are still running count too: each one holds an estimate of what it may cost, from the size of its input and its output limit (max_tokens, at most 4,096 tokens), until its real cost is recorded. So fifty children that start at once cannot all pass a budget that only one of them fits. Each edge counts the calls it carries; calls through another edge show up once they are recorded. job get shows the budget next to the cost.
A budget per tenant and month
A tenant can have a budget for each calendar month (UTC), which covers every model call of the tenant and what its plugins report:
orchestrator-zero tenant budget acme 250 # $250 a month
orchestrator-zero tenant list # this month's spending and the budget, per tenant
orchestrator-zero tenant budget acme 0 # no monthly budget
Once the month's budget is spent, the gateway refuses the tenant's model calls with 402 oz0_budget_exceeded until the month ends or the budget is raised, and jobs fail with BudgetExceeded:
BudgetExceeded: tenant acme has spent $250.0012 of its $250.00 budget for October 2026
The gateway reads what the tenant has spent at most every two seconds, and counts the calls it is carrying with their estimates, so the month can end a few calls over its budget. A new budget bites within five seconds. The web UI shows the month on the Cost page and sets budgets under Settings, Tenants.
Tools and delegation
An agent can call only the tools it lists and delegate only to the agents in delegation.allow, within max_depth and max_children.
requires_approval wait for a person before they run; see Approvals. Your own hooks can allow, deny or change calls, and evals stop a worse version at install. Planned later: budgets per tenant and month (M6), an outcome check by a second model, and shadow runs of a new version on real jobs.