Models and keys
Every model call goes through the LLM gateway in the edge. It adds the tenant's API key, passes the call to the provider unchanged, streaming included, and records tokens and cost.
Name the model
In an agent, models are <provider>/<model>:
model:
primary: anthropic/claude-sonnet-5-5
fallback: [anthropic/claude-haiku-4-5]
| Provider | Gateway path | Key secret |
|---|---|---|
anthropic | /llm/anthropic/...: Anthropic's Messages API, unchanged | ANTHROPIC_API_KEY |
openai | /llm/openai/...: an OpenAI-compatible API, chat completions and the Responses API | OPENAI_API_KEY |
Because the gateway passes calls through unchanged, provider features such as prompt caching and extended thinking work without translation. Its anthropic-* request headers, betas included, reach the provider, and the provider's rate-limit and retry headers come back: retry-after, retry-after-ms, x-should-retry, anthropic-ratelimit-* and x-ratelimit-*.
Store the key once
printf %s "$ANTHROPIC_API_KEY" | orchestrator-zero secret set ANTHROPIC_API_KEY
The key is sealed with the master key and unsealed only in the edge. If a tenant has no key for a provider, the gateway answers 412 with the command that sets it. An agent with a fallback model on another provider falls back to it; when every model was refused, the job fails at once with MissingKey and that command, without retrying.
Why nodes never see keys
Each node runs a local LLM proxy. The agent runtime and its plugins see ANTHROPIC_BASE_URL and OPENAI_BASE_URL pointing at it, with a per-start token as their API key. Unmodified SDKs work, the proxy forwards calls to the edge with the node's certificate, and a stolen laptop holds no provider key. When an edge cannot be reached, the proxy tries the next one; a call that an edge has already taken is never sent to a second edge, so a provider never runs it twice, and the activity's own retry decides what happens next.
Point the gateway somewhere else
orchestrator-zero server start --llm-anthropic-url https://llm-proxy.internal \
--llm-openai-url https://vllm.internal:8000
--llm-anthropic-url and --llm-openai-url point the gateway at a corporate proxy, a mirror, a test server, or an OpenAI-compatible server such as vLLM.
Cost
The gateway reads usage from every response and stream and prices it with genai-prices: Anthropic's message_start and message_delta events, the last chunk of a chat completion streamed with stream_options.include_usage, and the Responses API's response.completed. Tokens read from the cache are counted apart from new input tokens, so they are priced as cache reads. Each call is recorded with its job, agent, node and model. See Cost and usage.
Server tools
The gateway passes Anthropic's Messages API through as it is, so the provider's own server tools (web search, web fetch, code execution) work for the agents and harnesses that use them. Web searches are counted from the response's usage.server_tool_use and priced at $10 per 1,000 on top of tokens. Server tools work only with that provider's models, and the data they fetch leaves your environment; tools from plugins work with every model.