Run agents

Models and keys

How model calls reach providers through the edge's LLM gateway, and why nodes never hold a key.

Every model call goes through the LLM gateway in the edge. It adds the tenant's API key, passes the call to the provider unchanged, streaming included, and records tokens and cost.

Name the model

In an agent, models are <provider>/<model>:

agents/reviewer.yaml
model:
  primary: anthropic/claude-sonnet-5-5
  fallback: [anthropic/claude-haiku-4-5]
ProviderGateway pathKey secret
anthropic/llm/anthropic/...: Anthropic's Messages API, unchangedANTHROPIC_API_KEY
openai/llm/openai/...: an OpenAI-compatible API, chat completions and the Responses APIOPENAI_API_KEY

Because the gateway passes calls through unchanged, provider features such as prompt caching and extended thinking work without translation. Its anthropic-* request headers, betas included, reach the provider, and the provider's rate-limit and retry headers come back: retry-after, retry-after-ms, x-should-retry, anthropic-ratelimit-* and x-ratelimit-*.

Store the key once

Terminal
printf %s "$ANTHROPIC_API_KEY" | orchestrator-zero secret set ANTHROPIC_API_KEY

The key is sealed with the master key and unsealed only in the edge. If a tenant has no key for a provider, the gateway answers 412 with the command that sets it. An agent with a fallback model on another provider falls back to it; when every model was refused, the job fails at once with MissingKey and that command, without retrying.

Why nodes never see keys

Each node runs a local LLM proxy. The agent runtime and its plugins see ANTHROPIC_BASE_URL and OPENAI_BASE_URL pointing at it, with a per-start token as their API key. Unmodified SDKs work, the proxy forwards calls to the edge with the node's certificate, and a stolen laptop holds no provider key. When an edge cannot be reached, the proxy tries the next one; a call that an edge has already taken is never sent to a second edge, so a provider never runs it twice, and the activity's own retry decides what happens next.

Point the gateway somewhere else

Terminal
orchestrator-zero server start --llm-anthropic-url https://llm-proxy.internal \
  --llm-openai-url https://vllm.internal:8000

--llm-anthropic-url and --llm-openai-url point the gateway at a corporate proxy, a mirror, a test server, or an OpenAI-compatible server such as vLLM.

Cost

The gateway reads usage from every response and stream and prices it with genai-prices: Anthropic's message_start and message_delta events, the last chunk of a chat completion streamed with stream_options.include_usage, and the Responses API's response.completed. Tokens read from the cache are counted apart from new input tokens, so they are priced as cache reads. Each call is recorded with its job, agent, node and model. See Cost and usage.

Server tools

The gateway passes Anthropic's Messages API through as it is, so the provider's own server tools (web search, web fetch, code execution) work for the agents and harnesses that use them. Web searches are counted from the response's usage.server_tool_use and priced at $10 per 1,000 on top of tokens. Server tools work only with that provider's models, and the data they fetch leaves your environment; tools from plugins work with every model.

Budgets per job tree and per tenant and month are in Contracts and limits. Planned: local models a node calls directly with the usage reported, and more providers.
Copyright © 2026