diff --git a/docs/integrations/oh_my_pi.md b/docs/integrations/oh_my_pi.md index 0e0d9f89d..2fadf0fa4 100644 --- a/docs/integrations/oh_my_pi.md +++ b/docs/integrations/oh_my_pi.md @@ -29,7 +29,8 @@ providers: - `auth: none` marks the provider as keyless. Switchyard ignores client keys unless an LLM client sets `forward_auth = true`. Without `auth: none`, `omp` refuses to send a - request. + request. To send your gateway key through a route that forwards it, see + [Forwarded keys](#forwarded-keys). - `models[].id` must equal a route `id` from your TOML file. `contextWindow` and `maxTokens` set `omp`'s compaction limit and output cap. `reasoning: true` turns on the `--thinking` flag. @@ -68,8 +69,8 @@ the default model. ## Check the routing The checks in [Use Switchyard with pi](pi.md#check-the-routing) work the same way for -`omp`. On the Chat Completions API, `omp` sends no session header for a custom provider, -so the routing log records `"session_id": null`. Routes with +`omp`. On the Chat Completions and Responses APIs, `omp` sends no session header for a +custom provider, so the routing log records `"session_id": null`. Routes with `classify_trigger = "user_turn"` or `"new_session"`, advisor budgets, and the stage router's `capable_hold_turns` then treat each request as its own session. If you need per-session routing, use `anthropic-messages`. On that API `omp` sends the @@ -94,3 +95,44 @@ Port 4000 is also the default port for `omp`'s `litellm` provider and for as a LiteLLM proxy. In that case, run Switchyard on another port or set `LITELLM_BASE_URL`. Set `cost` on the model entry if you want `omp` to show a non-zero cost. + +## Claude through an LLM gateway + +The pi guide's [Claude through an LLM gateway](pi.md#claude-through-an-llm-gateway) +section applies to `omp`: give Claude targets an LLM client with +`format = "anthropic_messages"`. It shows how to check prompt caching and estimate the +cost. This section covers what differs for `omp`, tested with Oh My Pi 18.2.11. + +### Thinking on `anthropic-messages` + +On `openai-completions` and `openai-responses`, Switchyard turns `omp`'s thinking level +into adaptive thinking for an `anthropic_messages` target. On `anthropic-messages`, +Switchyard sends `omp`'s own `thinking` object unchanged. `omp` does not +recognize a route ID such as `switchyard` as a Claude model. With thinking on, it sends +`thinking: {type: "enabled"}`, and Claude Opus 5.5 and Sonnet 5 return HTTP 400. Set +adaptive thinking on the model entry: + +```yaml + - id: switchyard + reasoning: true + thinking: + mode: anthropic-adaptive + efforts: [low, medium, high] +``` + +`omp` requires `efforts` next to `mode`. With both set, it sends +`thinking: {type: "adaptive"}` and `output_config.effort`. + +### Forwarded keys + +To forward `omp`'s key, remove `auth: none` and set `apiKey: GATEWAY_API_KEY`, the name of +the environment variable that holds your gateway key. Unlike pi, `omp` reads the name +without a leading `$`. With `auth: none`, `omp` sends no key, and the gateway returns HTTP +401. Only standalone `switchyard-server` forwards keys. The native Nemo Relay plugin +rejects routes that use `forward_auth = true` (see +[Request Handling](nemo_relay.md#request-handling)). + +The pi guide's [Forwarded keys](pi.md#forwarded-keys) table shows which request APIs a +route accepts when it forwards the key. A route that forwards the key to an OpenAI-format +LLM client, such as a classifier's GPT judge, accepts only `openai-completions` and +`openai-responses` requests. diff --git a/docs/integrations/pi.md b/docs/integrations/pi.md index 1560d14a2..8560fa2b5 100644 --- a/docs/integrations/pi.md +++ b/docs/integrations/pi.md @@ -41,7 +41,9 @@ with route id `switchyard`. - `models[].id` must equal a route `id` from your TOML file. Add one entry per route. - `apiKey` is a placeholder. Switchyard ignores client keys unless an LLM client sets - `forward_auth = true`. pi still needs some value here before it lists the model. + `forward_auth = true`. pi still needs some value here before it lists the model. To + send your gateway key through a route that forwards it, see + [Forwarded keys](#forwarded-keys). - `contextWindow` and `maxTokens` set pi's compaction limit and output cap. pi does not read these values from the server. Use the smallest context window among the route's targets. @@ -96,3 +98,156 @@ clients keep the local id `switchyard` instead. Set `cost` on the model entry if you want pi to show a non-zero cost. [`benchmark/run-baseline.sh`](../../benchmark/README.md) runs Terminal-Bench tasks with pi through Switchyard when you pass `--agent pi`. + +## Claude through an LLM gateway + +An LLM gateway, such as a LiteLLM proxy, can serve Claude on three endpoints: +`/v1/chat/completions`, `/v1/responses`, and `/v1/messages`. Switchyard calls the +endpoint that matches the `format` of the Claude target's LLM client, whichever `api` pi +uses. Give Claude targets an LLM client with `format = "anthropic_messages"`: + +```toml +[llm_clients.gateway_claude] +format = "anthropic_messages" +base_url = "https://gateway.example.com" +api_key_env = "GATEWAY_API_KEY" + +[targets.claude] +id = "claude-sonnet-5" # the gateway's model ID +llm_client = "gateway_claude" +``` + +On the tested gateway, both OpenAI formats returned HTTP 400 when pi sent a thinking +level, and `openai_responses` never read the prompt cache: + +| Claude LLM client `format` | Gateway endpoint | Prompt cache | pi's thinking level | +|---|---|---|---| +| `anthropic_messages` | `/v1/messages` | Read on repeated prompts | Works | +| `openai_chat` | `/v1/chat/completions` | Read on repeated prompts | HTTP 400 | +| `openai_responses` | `/v1/responses` | Never read | HTTP 400 | + +[Prompt caching](#prompt-caching) and [Thinking](#thinking) explain both failures. To +use each developer's own gateway key instead of a key that the server holds, see +[Forwarded keys](#forwarded-keys). + +### Prompt caching + +pi sends the whole conversation on every turn. When the gateway reads the repeated part +from Claude's prompt cache, that part costs 0.1 times the input price on Claude Sonnet 5 +and 0.05 times on Claude Opus 5.5, according to Anthropic's +[pricing page](https://platform.claude.com/docs/en/about-claude/pricing). The tested +gateway never read Claude prompts from the cache on `/v1/responses`, so every turn there +paid the full input price for the whole conversation. + +To check your gateway, start the server with `--routing-log-file PATH` and send the same +prompt twice. Claude does not cache short prompts, so use a prompt of at least 5,000 +tokens. Then read the records: + +```bash +jq -c '{route_id, prompt_tokens, cached_tokens, cache_creation_tokens}' PATH +``` + +When caching works, the first record shows the prompt in `cache_creation_tokens`, and the +second shows `cached_tokens` close to `prompt_tokens`. If the second record shows +`"cached_tokens": 0`, the gateway read nothing from the cache. + +#### Estimate the cost + +The routing log records token counts, not prices. To estimate what a request cost, +multiply each count in its record by the matching price and add the results: + +| Tokens in the record | Price | +|---|---| +| `prompt_tokens - cached_tokens - cache_creation_tokens` | Input | +| `cached_tokens` | Cache read | +| `cache_creation_tokens` | Cache write | +| `completion_tokens` | Output | + +Write the prices to `prices.json` in USD per million tokens, and key each entry by the +record's `model` value. This example uses Anthropic's list prices for Claude Sonnet 5 on +2026-10-02: + +```json +{ + "claude-sonnet-5": {"input": 2.00, "cache_read": 0.20, "cache_write": 2.50, "output": 10.00} +} +``` + +The example's `cache_write` price is for a 5-minute cache. A 1-hour cache costs 2 times +the input price instead of 1.25 times. The routing log does not say which cache the +gateway used. A gateway may also charge its own prices, so the result is an estimate, not +the gateway's bill. This command prints one estimated cost per record, or a warning for a +model that has no entry in `prices.json`: + +```bash +jq -r --slurpfile prices prices.json ' + . as $r + | ($prices[0][$r.model // ""]) as $p + | if $p == null then + "warning: no price for model \($r.model); add it to prices.json" + else + ((($r.prompt_tokens // 0) - ($r.cached_tokens // 0) - ($r.cache_creation_tokens // 0)) * $p.input + + ($r.cached_tokens // 0) * $p.cache_read + + ($r.cache_creation_tokens // 0) * $p.cache_write + + ($r.completion_tokens // 0) * $p.output) / 1000000 + | "\($r.route_id) \($r.model) estimated $\(. * 1000000 | round / 1000000)" + end' PATH +``` + +### Thinking + +With `reasoning: true` on the model entry, pi sends a thinking level on every request, +even when you do not pass `--thinking`. Switchyard passes that level to an OpenAI-format +target in OpenAI form: `reasoning_effort` on `openai_chat` and `reasoning.effort` on +`openai_responses`. The tested gateway turned either field into Anthropic's older +`thinking: {type: "enabled"}`. Claude Opus 5.5 and Sonnet 5 refuse that form, so the +gateway returned HTTP 400: + +```text +"thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior. +``` + +For an `anthropic_messages` target, Switchyard sends the level in the form that Claude +accepts: `thinking: {type: "adaptive"}` with `output_config.effort`. pi's thinking level +then takes effect. + +If a Claude target must stay on an OpenAI format, remove the effort field from its +requests with [`omit_body_fields`](../reference/toml_schema.md#targetsname). Claude then +thinks at its default effort, and pi's thinking level has no effect on that target: + +```toml +[targets.claude] +id = "claude-opus-5-5" +llm_client = "gateway_chat" # format = "openai_chat" +omit_body_fields = ["reasoning_effort"] # use "reasoning" on openai_responses +``` + +Switchyard applies the target's `extra_body` and `reasoning_effort` after the removal, so +either can set the field again. + +### Forwarded keys + +With `api_key_env`, every caller's requests use one gateway key that the server holds. +To use each developer's own gateway key instead, set `forward_auth = true` on the LLM +client. Then set pi's `apiKey` to `$` plus the name of the environment variable that +holds your key: `"apiKey": "$GATEWAY_API_KEY"`. Without the `$`, pi sends the name itself +as the key, and the gateway returns HTTP 401. Only standalone `switchyard-server` forwards +keys. The native Nemo Relay plugin rejects routes that use `forward_auth = true` (see +[Request Handling](nemo_relay.md#request-handling)). + +Which request APIs a route accepts depends on its LLM clients that forward the key. +Clients with `api_key_env` do not count: + +| Forwarding LLM clients in the route | Request APIs the route accepts | +|---|---| +| None | All three | +| Only OpenAI formats | `/v1/chat/completions` and `/v1/responses` | +| OpenAI formats and `anthropic_messages`, all on the same scheme, host, and port | `/v1/chat/completions` and `/v1/responses` | +| Only `anthropic_messages` | `/v1/messages`, which pi should not use (see [Which request API](#which-request-api)) | + +For example, a classifier route can forward pi's key to its GPT judge on +`openai_responses` and to its Claude targets on `anthropic_messages` when both LLM clients +point at the same gateway. If they use different hosts, ports, or schemes, the server does +not start. If a route forwards the key only to Claude targets on `anthropic_messages`, pi +cannot use it. Put those targets on `openai_chat` with `omit_body_fields` (see +[Thinking](#thinking)), or let the server hold the key with `api_key_env`. diff --git a/docs/reference/toml_schema.md b/docs/reference/toml_schema.md index 9f820e72f..ce95015ee 100644 --- a/docs/reference/toml_schema.md +++ b/docs/reference/toml_schema.md @@ -149,6 +149,7 @@ such clients. | `llm_client` | Yes | — | Key under `[llm_clients]`. | | `system_prompt` | No | unset | System prompt prepended when this target serves a completion. | | `extra_body` | No | `{}` | Values merged into the upstream request when the request does not already set that key. | +| `omit_body_fields` | No | `[]` | Top-level fields removed from every request body that Switchyard sends to this target. Switchyard removes them after it translates the request to the LLM client's `format`, so use that format's field names, for example `reasoning_effort` on `openai_chat` or `reasoning` on `openai_responses`. Switchyard applies `extra_body` and `reasoning_effort` after the removal, so either can set a removed field again. | | `reasoning_effort` | No | unset | Reasoning effort forced on every request to this target, replacing the value the caller sent (`reasoning.effort` on `openai_responses`, `reasoning_effort` on `openai_chat`). Rejected on `anthropic_messages` clients. Use it to run one target at a different effort than the client asked for, for example a strong tier at `max` behind a client that sends `high`. Targets with different effort settings need distinct model IDs when used within one route. Separate routes may use the same model ID with separate `llm_clients` entries (same endpoint, different name). | Within one route, callable targets with the same model ID must use the same `llm_client`.