From 6e7b22d343bb5914cd04f36fd1ba7c7e6a87607e Mon Sep 17 00:00:00 2001 From: Elyas Mehtabuddin Date: Tue, 29 Sep 2026 23:13:25 -0700 Subject: [PATCH 1/7] docs(integrations): add Claude gateway caching and thinking notes Signed-off-by: Elyas Mehtabuddin --- docs/integrations/oh_my_pi.md | 22 +++++++++++++ docs/integrations/pi.md | 59 +++++++++++++++++++++++++++++++++++ 2 files changed, 81 insertions(+) diff --git a/docs/integrations/oh_my_pi.md b/docs/integrations/oh_my_pi.md index 0e0d9f89d..1ba5674e5 100644 --- a/docs/integrations/oh_my_pi.md +++ b/docs/integrations/oh_my_pi.md @@ -94,3 +94,25 @@ Port 4000 is also the default port for `omp`'s `litellm` provider and for as a LiteLLM proxy. In that case, run Switchyard on another port or set `LITELLM_BASE_URL`. Set `cost` on the model entry if you want `omp` to show a non-zero cost. + +### Claude targets behind an OpenAI-compatible gateway + +The [pi guide's notes on Claude targets behind a gateway](pi.md#claude-targets-behind-an-openai-compatible-gateway) +apply to `omp` too. The target's LLM client `format` decides prompt caching and +thinking, not the `api` that `omp` uses. Prefer `format = "openai_chat"` or +`"anthropic_messages"` for Claude targets, because a gateway may not cache Claude +prompts on `/v1/responses`. + +With thinking on, `omp` sends a reasoning effort on `openai-completions` and +`openai-responses`. Switchyard passes it to an `openai_chat` target as +`reasoning_effort`, which some gateways reject for Claude Opus 5.5 and Sonnet 5 with +HTTP 400. Switchyard turns the effort into adaptive thinking only when it translates a +Chat Completions or Responses request for an `anthropic_messages` target. It forwards an +`anthropic-messages` request to that target with `omp`'s own `thinking` settings. + +A route that forwards the caller's key to an `anthropic_messages` client accepts +requests only on `/v1/messages`, and it cannot also forward the key to an +`openai_chat` or `openai_responses` client. If such a route also needs an OpenAI-format +target, such as a GPT judge on `openai_responses`, keep the Claude targets on +`openai_chat` with `omit_body_fields = ["reasoning_effort"]`. The models then think at +their default effort, and `--thinking` has no effect on them. diff --git a/docs/integrations/pi.md b/docs/integrations/pi.md index 1560d14a2..3af94fa95 100644 --- a/docs/integrations/pi.md +++ b/docs/integrations/pi.md @@ -96,3 +96,62 @@ clients keep the local id `switchyard` instead. Set `cost` on the model entry if you want pi to show a non-zero cost. [`benchmark/run-baseline.sh`](../../benchmark/README.md) runs Terminal-Bench tasks with pi through Switchyard when you pass `--agent pi`. + +### Claude targets behind an OpenAI-compatible gateway + +Some gateways, such as a LiteLLM proxy, serve Claude models on `/v1/chat/completions`, +`/v1/responses`, and `/v1/messages` with one API key. The target's LLM client `format` +decides which endpoint Switchyard calls. That choice can change prompt caching and +thinking for Claude. + +**Prompt caching.** On the LiteLLM gateway tested for this page, a repeated Claude prompt +was read from the cache on `/v1/chat/completions` and `/v1/messages`, but never on +`/v1/responses`. Every request to `/v1/responses` paid for the whole prompt again. Use +`format = "openai_chat"` or `"anthropic_messages"` for Claude targets. To check your +gateway, start the server with `--routing-log-file PATH` and send the same long prompt +twice. Claude does not cache short prompts, so use one of at least 5,000 tokens. Then +read the records: + +```bash +jq -c '{route_id, prompt_tokens, cached_tokens, cache_creation_tokens}' PATH +``` + +If the gateway caches the prompt, the first record shows it in `cache_creation_tokens` +and the second shows `cached_tokens` close to `prompt_tokens`. Writing the cache makes +the first request cost more, and each later request that reuses the prompt costs much +less. If the second record shows `"cached_tokens": 0`, the gateway billed the whole +prompt again. + +**Thinking.** Claude Opus 5.5 and Sonnet 5 accept only adaptive thinking. Some gateways +turn the Chat Completions `reasoning_effort` field into Anthropic's older +`thinking: {type: "enabled"}`, and the model then returns HTTP 400: + +```text +"thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior. +``` + +With `reasoning: true`, pi sends `reasoning_effort`, and Switchyard passes it to an +`openai_chat` target. You have two options: + +- Use an `anthropic_messages` target. Switchyard turns the requested effort into + `thinking: {type: "adaptive"}` and `output_config.effort`. +- Keep the `openai_chat` target and drop the field: + + ```toml + [targets.claude] + id = "claude-opus-5-5" + llm_client = "gateway_chat" # format = "openai_chat" + omit_body_fields = ["reasoning_effort"] + ``` + + The request then succeeds, and the model thinks at its default effort. pi's + `--thinking` level has no effect on this target. + +**Forwarded keys.** Every LLM client in one route that sets `forward_auth = true` must +use the same API family: `openai_chat` and `openai_responses`, or `anthropic_messages`. +Otherwise the server does not start and prints +`route cannot forward both Anthropic and OpenAI caller credentials`. A route that +forwards the caller's key to an `anthropic_messages` client also accepts requests only +on `/v1/messages`, and pi should not use that API. So when Switchyard forwards pi's key, +use `openai_chat` Claude targets with `omit_body_fields`. If the server holds the key +through `api_key_env`, an `anthropic_messages` target works with every request API. From 0321e77244d3a2518f86864c689ad0ff2a59340c Mon Sep 17 00:00:00 2001 From: Elyas Mehtabuddin Date: Tue, 29 Sep 2026 23:36:20 -0700 Subject: [PATCH 2/7] docs(integrations): clarify when each Claude gateway setup applies Signed-off-by: Elyas Mehtabuddin --- docs/integrations/oh_my_pi.md | 34 ++++++++++++++++++++++------------ docs/integrations/pi.md | 20 ++++++++++++++------ 2 files changed, 36 insertions(+), 18 deletions(-) diff --git a/docs/integrations/oh_my_pi.md b/docs/integrations/oh_my_pi.md index 1ba5674e5..e9a180c00 100644 --- a/docs/integrations/oh_my_pi.md +++ b/docs/integrations/oh_my_pi.md @@ -98,21 +98,31 @@ cost. ### Claude targets behind an OpenAI-compatible gateway The [pi guide's notes on Claude targets behind a gateway](pi.md#claude-targets-behind-an-openai-compatible-gateway) -apply to `omp` too. The target's LLM client `format` decides prompt caching and -thinking, not the `api` that `omp` uses. Prefer `format = "openai_chat"` or -`"anthropic_messages"` for Claude targets, because a gateway may not cache Claude -prompts on `/v1/responses`. +apply to `omp` too. The target's LLM client `format`, not the `api` that `omp` uses, +decides which gateway endpoint Switchyard calls and so whether the gateway caches the +prompt. Prefer `format = "openai_chat"` or `"anthropic_messages"` for Claude targets, +because a gateway may not cache Claude prompts on `/v1/responses`. Thinking depends on +both the `format` and the `api`, as the next paragraph explains. With thinking on, `omp` sends a reasoning effort on `openai-completions` and `openai-responses`. Switchyard passes it to an `openai_chat` target as -`reasoning_effort`, which some gateways reject for Claude Opus 5.5 and Sonnet 5 with -HTTP 400. Switchyard turns the effort into adaptive thinking only when it translates a -Chat Completions or Responses request for an `anthropic_messages` target. It forwards an -`anthropic-messages` request to that target with `omp`'s own `thinking` settings. +`reasoning_effort`. Some gateways turn that field into a thinking setting that Claude +Opus 5.5 and Sonnet 5 reject with HTTP 400. Switchyard turns the effort into adaptive +thinking only when it translates a Chat Completions or Responses request for an +`anthropic_messages` target. When `omp` uses `anthropic-messages`, Switchyard forwards +the request to that target with `omp`'s own `thinking` settings unchanged. A route that forwards the caller's key to an `anthropic_messages` client accepts requests only on `/v1/messages`, and it cannot also forward the key to an -`openai_chat` or `openai_responses` client. If such a route also needs an OpenAI-format -target, such as a GPT judge on `openai_responses`, keep the Claude targets on -`openai_chat` with `omit_body_fields = ["reasoning_effort"]`. The models then think at -their default effort, and `--thinking` has no effect on them. +`openai_chat` or `openai_responses` client. So a route with an OpenAI-format target, +such as a GPT judge on `openai_responses`, cannot forward the caller's key to +both the judge and Claude targets on `anthropic_messages`. Choose one of two setups: + +- Forward the key to every client, and keep the Claude targets on `openai_chat` with + `omit_body_fields = ["reasoning_effort"]`. The Claude models then think at their + default effort, and `--thinking` has no effect on them. +- Forward the key only to the GPT judge, and put the Claude targets on an + `anthropic_messages` client that reads a key held by the server from `api_key_env`. + Switchyard then turns the effort into adaptive thinking. The route accepts only + `/v1/chat/completions` and `/v1/responses`, so use `openai-completions` or + `openai-responses`. diff --git a/docs/integrations/pi.md b/docs/integrations/pi.md index 3af94fa95..b9aed795d 100644 --- a/docs/integrations/pi.md +++ b/docs/integrations/pi.md @@ -109,8 +109,8 @@ was read from the cache on `/v1/chat/completions` and `/v1/messages`, but never `/v1/responses`. Every request to `/v1/responses` paid for the whole prompt again. Use `format = "openai_chat"` or `"anthropic_messages"` for Claude targets. To check your gateway, start the server with `--routing-log-file PATH` and send the same long prompt -twice. Claude does not cache short prompts, so use one of at least 5,000 tokens. Then -read the records: +twice. Claude does not cache short prompts. A prompt of at least 5,000 tokens is a safe +size; it is not Claude's exact minimum. Then read the records: ```bash jq -c '{route_id, prompt_tokens, cached_tokens, cache_creation_tokens}' PATH @@ -150,8 +150,16 @@ With `reasoning: true`, pi sends `reasoning_effort`, and Switchyard passes it to **Forwarded keys.** Every LLM client in one route that sets `forward_auth = true` must use the same API family: `openai_chat` and `openai_responses`, or `anthropic_messages`. Otherwise the server does not start and prints -`route cannot forward both Anthropic and OpenAI caller credentials`. A route that +`route cannot forward both Anthropic and OpenAI caller credentials`, where +`` is the route's `[routes.]` table key, not its `id`. A route that forwards the caller's key to an `anthropic_messages` client also accepts requests only -on `/v1/messages`, and pi should not use that API. So when Switchyard forwards pi's key, -use `openai_chat` Claude targets with `omit_body_fields`. If the server holds the key -through `api_key_env`, an `anthropic_messages` target works with every request API. +on `/v1/messages`, and pi should not use that API. So when Switchyard forwards pi's key +to the Claude targets, use `openai_chat` Claude targets with `omit_body_fields`. + +An `anthropic_messages` client that reads the key from `api_key_env` works with every +request API, as long as no other LLM client in the route sets `forward_auth = true`. If +one does, the forwarded client limits the route's endpoints. For example, a route with a +forwarded `openai_responses` client accepts only `/v1/chat/completions` and +`/v1/responses`, and returns HTTP 400 on `/v1/messages`. pi uses those APIs, so such a +route can forward pi's key to a GPT judge on `openai_responses` while its Claude targets +use `anthropic_messages` with a key that the server holds. From e6fa560ac46a7e1c6c1c3f214c22564cf726ca58 Mon Sep 17 00:00:00 2001 From: Elyas Mehtabuddin Date: Thu, 1 Oct 2026 13:07:10 -0700 Subject: [PATCH 3/7] docs(integrations): estimate Claude cache costs and fix the thinking and key setup steps Signed-off-by: Elyas Mehtabuddin --- docs/integrations/oh_my_pi.md | 96 ++++++++++++------ docs/integrations/pi.md | 181 +++++++++++++++++++++++++--------- docs/reference/toml_schema.md | 1 + 3 files changed, 202 insertions(+), 76 deletions(-) diff --git a/docs/integrations/oh_my_pi.md b/docs/integrations/oh_my_pi.md index e9a180c00..69d9b6a7a 100644 --- a/docs/integrations/oh_my_pi.md +++ b/docs/integrations/oh_my_pi.md @@ -29,7 +29,8 @@ providers: - `auth: none` marks the provider as keyless. Switchyard ignores client keys unless an LLM client sets `forward_auth = true`. Without `auth: none`, `omp` refuses to send a - request. + request. To send your gateway key through a route that forwards it, see + [Forwarded keys](#forwarded-keys). - `models[].id` must equal a route `id` from your TOML file. `contextWindow` and `maxTokens` set `omp`'s compaction limit and output cap. `reasoning: true` turns on the `--thinking` flag. @@ -68,8 +69,8 @@ the default model. ## Check the routing The checks in [Use Switchyard with pi](pi.md#check-the-routing) work the same way for -`omp`. On the Chat Completions API, `omp` sends no session header for a custom provider, -so the routing log records `"session_id": null`. Routes with +`omp`. On the Chat Completions and Responses APIs, `omp` sends no session header for a +custom provider, so the routing log records `"session_id": null`. Routes with `classify_trigger = "user_turn"` or `"new_session"`, advisor budgets, and the stage router's `capable_hold_turns` then treat each request as its own session. If you need per-session routing, use `anthropic-messages`. On that API `omp` sends the @@ -95,34 +96,65 @@ as a LiteLLM proxy. In that case, run Switchyard on another port or set `LITELLM_BASE_URL`. Set `cost` on the model entry if you want `omp` to show a non-zero cost. -### Claude targets behind an OpenAI-compatible gateway - -The [pi guide's notes on Claude targets behind a gateway](pi.md#claude-targets-behind-an-openai-compatible-gateway) -apply to `omp` too. The target's LLM client `format`, not the `api` that `omp` uses, -decides which gateway endpoint Switchyard calls and so whether the gateway caches the -prompt. Prefer `format = "openai_chat"` or `"anthropic_messages"` for Claude targets, -because a gateway may not cache Claude prompts on `/v1/responses`. Thinking depends on -both the `format` and the `api`, as the next paragraph explains. - -With thinking on, `omp` sends a reasoning effort on `openai-completions` and -`openai-responses`. Switchyard passes it to an `openai_chat` target as -`reasoning_effort`. Some gateways turn that field into a thinking setting that Claude -Opus 5.5 and Sonnet 5 reject with HTTP 400. Switchyard turns the effort into adaptive -thinking only when it translates a Chat Completions or Responses request for an -`anthropic_messages` target. When `omp` uses `anthropic-messages`, Switchyard forwards -the request to that target with `omp`'s own `thinking` settings unchanged. - -A route that forwards the caller's key to an `anthropic_messages` client accepts -requests only on `/v1/messages`, and it cannot also forward the key to an -`openai_chat` or `openai_responses` client. So a route with an OpenAI-format target, -such as a GPT judge on `openai_responses`, cannot forward the caller's key to -both the judge and Claude targets on `anthropic_messages`. Choose one of two setups: - -- Forward the key to every client, and keep the Claude targets on `openai_chat` with +## Claude targets behind an OpenAI-compatible gateway + +The pi guide's section +[Claude targets behind an OpenAI-compatible gateway](pi.md#claude-targets-behind-an-openai-compatible-gateway) +applies to `omp` too. Use its table to choose the Claude LLM client by who holds the +gateway key. Switchyard calls the gateway endpoint that matches the target's LLM client +`format`, whatever `api` `omp` uses, so the `format` decides whether the gateway caches +the prompt. Use `format = "openai_chat"` or `"anthropic_messages"` for Claude targets, +because a gateway may not cache Claude prompts on `/v1/responses`. To estimate what the +requests cost from the routing log, see [Estimate the cost](pi.md#estimate-the-cost) in +the pi guide. This section covers what differs for `omp`, checked with Oh My Pi 18.2.11. + +### Thinking + +Thinking depends on both the LLM client `format` and the `api`: + +- On `openai-completions`, `omp` sends `reasoning_effort`, and on `openai-responses` it + sends `reasoning.effort`. Switchyard passes the effort to an `openai_chat` target as + `reasoning_effort` and to an `openai_responses` target as `reasoning.effort`. Some + gateways turn either field into a thinking setting that Claude Opus 5.5 and Sonnet 5 + refuse with HTTP 400. Switchyard turns the effort into adaptive thinking only for an + `anthropic_messages` target. The pi guide's [Thinking](pi.md#thinking) section shows + how to remove the field with `omit_body_fields` instead. +- On `anthropic-messages`, Switchyard sends `omp`'s own `thinking` settings to an + `anthropic_messages` target unchanged. `omp` does not recognize a route id such as + `switchyard` as a Claude model, so with thinking on it sends + `thinking: {type: "enabled"}`, and Claude Opus 5.5 and Sonnet 5 return HTTP 400. Tell + `omp` to use adaptive thinking on the model entry: + + ```yaml + - id: switchyard + reasoning: true + thinking: + mode: anthropic-adaptive + efforts: [low, medium, high] + ``` + + `omp` requires `efforts` next to `mode`. It then sends `thinking: {type: "adaptive"}` + and `output_config.effort`. + +### Forwarded keys + +To forward `omp`'s key, remove `auth: none` and set `apiKey: GATEWAY_API_KEY`, the name +of the environment variable that holds your gateway key. Unlike pi, `omp` reads the name +without a leading `$`. + +A route that forwards the key to an `anthropic_messages` client accepts requests only on +`/v1/messages`, and it cannot also forward the key to an `openai_chat` or +`openai_responses` client. So a route with a GPT judge on `openai_responses` cannot +forward the caller's key to both the judge and Claude targets on `anthropic_messages`. +Choose one of two setups: + +- Forward the key to every LLM client, and keep the Claude targets on `openai_chat` with `omit_body_fields = ["reasoning_effort"]`. The Claude models then think at their default effort, and `--thinking` has no effect on them. -- Forward the key only to the GPT judge, and put the Claude targets on an - `anthropic_messages` client that reads a key held by the server from `api_key_env`. - Switchyard then turns the effort into adaptive thinking. The route accepts only - `/v1/chat/completions` and `/v1/responses`, so use `openai-completions` or - `openai-responses`. +- Forward the key only to the GPT judge, and give the Claude targets an + `anthropic_messages` client with `api_key_env`, so they use a server-owned key. + Switchyard then turns the effort into adaptive thinking. + +Both setups forward the key to an OpenAI-format LLM client, so the route accepts only +`/v1/chat/completions` and `/v1/responses` and returns HTTP 400 on `/v1/messages`. Use +`openai-completions` or `openai-responses`. diff --git a/docs/integrations/pi.md b/docs/integrations/pi.md index b9aed795d..69a990b92 100644 --- a/docs/integrations/pi.md +++ b/docs/integrations/pi.md @@ -41,7 +41,9 @@ with route id `switchyard`. - `models[].id` must equal a route `id` from your TOML file. Add one entry per route. - `apiKey` is a placeholder. Switchyard ignores client keys unless an LLM client sets - `forward_auth = true`. pi still needs some value here before it lists the model. + `forward_auth = true`. pi still needs some value here before it lists the model. To + send your gateway key through a route that forwards it, see + [Forwarded keys](#forwarded-keys). - `contextWindow` and `maxTokens` set pi's compaction limit and output cap. pi does not read these values from the server. Use the smallest context window among the route's targets. @@ -97,45 +99,119 @@ Set `cost` on the model entry if you want pi to show a non-zero cost. [`benchmark/run-baseline.sh`](../../benchmark/README.md) runs Terminal-Bench tasks with pi through Switchyard when you pass `--agent pi`. -### Claude targets behind an OpenAI-compatible gateway +## Claude targets behind an OpenAI-compatible gateway Some gateways, such as a LiteLLM proxy, serve Claude models on `/v1/chat/completions`, -`/v1/responses`, and `/v1/messages` with one API key. The target's LLM client `format` -decides which endpoint Switchyard calls. That choice can change prompt caching and -thinking for Claude. - -**Prompt caching.** On the LiteLLM gateway tested for this page, a repeated Claude prompt -was read from the cache on `/v1/chat/completions` and `/v1/messages`, but never on -`/v1/responses`. Every request to `/v1/responses` paid for the whole prompt again. Use -`format = "openai_chat"` or `"anthropic_messages"` for Claude targets. To check your -gateway, start the server with `--routing-log-file PATH` and send the same long prompt -twice. Claude does not cache short prompts. A prompt of at least 5,000 tokens is a safe -size; it is not Claude's exact minimum. Then read the records: +`/v1/responses`, and `/v1/messages` with one API key. Switchyard calls the endpoint that +matches the target's LLM client `format`, whatever `api` pi uses. On the gateway tested +for this page, the `format` decided whether Claude prompts were cached and whether pi's +`--thinking` level worked. + +Choose the Claude LLM client by who holds the gateway key: + +| Who holds the gateway key | Claude LLM client | Result | +|---|---|---| +| The server, through `api_key_env` | `format = "anthropic_messages"` | Prompt caching and pi's `--thinking` level both work. Every caller's Claude requests use the server-owned key. | +| pi sends it as `apiKey`, and the Claude LLM client forwards it with `forward_auth = true` | `format = "openai_chat"`, with `omit_body_fields = ["reasoning_effort"]` on the target | Prompt caching works. pi's `--thinking` level has no effect, and Claude thinks at its default effort. | + +Do not use `format = "openai_responses"` for Claude targets on such a gateway. The +gateway tested for this page never cached Claude prompts on `/v1/responses`, and it +returned HTTP 400 when thinking was on (see [Thinking](#thinking)). In both setups, a GPT +judge on `openai_responses` can still use pi's forwarded key (see +[Forwarded keys](#forwarded-keys)). + +### Prompt caching + +On the LiteLLM gateway tested for this page, a repeated Claude prompt was read from the +cache on `/v1/chat/completions` and `/v1/messages`, but never on `/v1/responses`. Every +`/v1/responses` request counted the whole prompt as uncached input. + +To check your gateway, start the server with `--routing-log-file PATH` and send the same +prompt twice. Claude does not cache short prompts, so use a prompt of at least 5,000 +tokens. In the tests for this page, prompts of about 5,000 tokens were cached on Claude +Opus 5.5 and Sonnet 5; the tests did not find Claude's exact minimum. Then read the +records: ```bash jq -c '{route_id, prompt_tokens, cached_tokens, cache_creation_tokens}' PATH ``` If the gateway caches the prompt, the first record shows it in `cache_creation_tokens` -and the second shows `cached_tokens` close to `prompt_tokens`. Writing the cache makes -the first request cost more, and each later request that reuses the prompt costs much -less. If the second record shows `"cached_tokens": 0`, the gateway billed the whole -prompt again. +and the second shows `cached_tokens` close to `prompt_tokens`. If the second record shows +`"cached_tokens": 0`, the gateway read nothing from the cache, and the whole prompt +counts as uncached input again. + +#### Estimate the cost + +The routing log records token counts, not prices. To estimate what a request cost, +multiply each token count in its record by the matching price, and add the results: + +| Tokens in the record | Price | +|---|---| +| Uncached input: `prompt_tokens - cached_tokens - cache_creation_tokens` | Input price | +| `cached_tokens` | Cache-read price | +| `cache_creation_tokens` | Cache-write price | +| `completion_tokens` | Output price | + +Anthropic's published pricing sets the cache-read price at 0.1 times the input price. It +sets the cache-write price at 1.25 times the input price for a 5-minute cache, or 2 times +for a 1-hour cache. The routing log does not record which cache lifetime the gateway +used, and a gateway may charge its own prices, so the result is an estimate, not the +gateway's bill. + +Put your prices in a `prices.json` file, in USD per million tokens. Key each entry by the +`model` value from the routing log. The rates below are illustrative: they only follow +Anthropic's published ratios, with cache reads at 0.1 times and 5-minute cache writes at +1.25 times the input price. Replace them with your provider's current prices. + +```json +{ + "claude-opus-5-5": {"input": 10.00, "cache_read": 1.00, "cache_write": 12.50, "output": 50.00} +} +``` + +The command below prints one estimated cost per record. For a record whose model has no +entry in `prices.json`, it prints a warning instead of a cost: + +```bash +jq -r --slurpfile prices prices.json ' + . as $r + | ($prices[0][$r.model // ""]) as $p + | if $p == null then + "warning: no price for model \($r.model); add it to prices.json" + else + ((($r.prompt_tokens // 0) - ($r.cached_tokens // 0) - ($r.cache_creation_tokens // 0)) * $p.input + + ($r.cached_tokens // 0) * $p.cache_read + + ($r.cache_creation_tokens // 0) * $p.cache_write + + ($r.completion_tokens // 0) * $p.output) / 1000000 + | "\($r.route_id) \($r.model) estimated $\(. * 1000000 | round / 1000000)" + end' PATH +``` -**Thinking.** Claude Opus 5.5 and Sonnet 5 accept only adaptive thinking. Some gateways -turn the Chat Completions `reasoning_effort` field into Anthropic's older -`thinking: {type: "enabled"}`, and the model then returns HTTP 400: +### Thinking + +Claude Opus 5.5 and Sonnet 5 accept only adaptive thinking. With `reasoning: true`, pi +sends `reasoning_effort`, even when you do not pass `--thinking`, and Switchyard +passes that field unchanged to an `openai_chat` target. Some gateways, including the one +tested for this page, turn `reasoning_effort` into Anthropic's older +`thinking: {type: "enabled"}` and return HTTP 400: ```text "thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior. ``` -With `reasoning: true`, pi sends `reasoning_effort`, and Switchyard passes it to an -`openai_chat` target. You have two options: +An `openai_responses` target fails the same way. Switchyard sends the effort to it as +`reasoning.effort`, and the gateway returns the same error. + +You have two options: -- Use an `anthropic_messages` target. Switchyard turns the requested effort into - `thinking: {type: "adaptive"}` and `output_config.effort`. -- Keep the `openai_chat` target and drop the field: +- Use an `anthropic_messages` LLM client. Switchyard turns pi's effort into + `thinking: {type: "adaptive"}` and `output_config.effort`, so pi's `--thinking` level + still applies. +- Keep the OpenAI-format LLM client and remove the effort field from requests to the + target with [`omit_body_fields`](../reference/toml_schema.md#targetsname). Use the + field name of the target's format: `reasoning_effort` on `openai_chat`, or `reasoning` + on `openai_responses`. ```toml [targets.claude] @@ -144,22 +220,39 @@ With `reasoning: true`, pi sends `reasoning_effort`, and Switchyard passes it to omit_body_fields = ["reasoning_effort"] ``` - The request then succeeds, and the model thinks at its default effort. pi's - `--thinking` level has no effect on this target. - -**Forwarded keys.** Every LLM client in one route that sets `forward_auth = true` must -use the same API family: `openai_chat` and `openai_responses`, or `anthropic_messages`. -Otherwise the server does not start and prints -`route cannot forward both Anthropic and OpenAI caller credentials`, where -`` is the route's `[routes.]` table key, not its `id`. A route that -forwards the caller's key to an `anthropic_messages` client also accepts requests only -on `/v1/messages`, and pi should not use that API. So when Switchyard forwards pi's key -to the Claude targets, use `openai_chat` Claude targets with `omit_body_fields`. - -An `anthropic_messages` client that reads the key from `api_key_env` works with every -request API, as long as no other LLM client in the route sets `forward_auth = true`. If -one does, the forwarded client limits the route's endpoints. For example, a route with a -forwarded `openai_responses` client accepts only `/v1/chat/completions` and -`/v1/responses`, and returns HTTP 400 on `/v1/messages`. pi uses those APIs, so such a -route can forward pi's key to a GPT judge on `openai_responses` while its Claude targets -use `anthropic_messages` with a key that the server holds. + The request then succeeds, and Claude thinks at its default effort. pi's `--thinking` + level has no effect on this target. + +### Forwarded keys + +An LLM client with `forward_auth = true` sends the caller's key to the gateway. An LLM +client with `api_key_env` sends a server-owned key, which the server reads from an +environment variable (see +[`[llm_clients.]`](../reference/toml_schema.md#llm_clientsname)). To forward pi's +key, replace the `apiKey` placeholder with the name of an environment variable that holds +your gateway key, with a leading `$`: `"apiKey": "$GATEWAY_API_KEY"`. Without the `$`, +pi sends the name itself as the key. + +Two rules limit forwarding in one route: + +- Every LLM client in the route that sets `forward_auth = true` must use the same API + family: `openai_chat` and `openai_responses`, or `anthropic_messages`. Otherwise the + server does not start and prints + `route cannot forward both Anthropic and OpenAI caller credentials`. `` is + the route's `[routes.]` table key, not its `id`. +- A route that forwards the key to an `anthropic_messages` client accepts requests only on + `/v1/messages`, and pi should not use that API (see [Which request API](#which-request-api)). + +So if the route forwards pi's key to its Claude targets, put them on `openai_chat` with +`omit_body_fields`. + +To keep pi's `--thinking` level, give the Claude targets an `anthropic_messages` client +with `api_key_env`. Which endpoints the route accepts then depends on the route's other +LLM clients: + +- If no LLM client in the route sets `forward_auth = true`, the route accepts every + request API. +- If an OpenAI-format LLM client forwards the key, for example a GPT judge on + `openai_responses`, the route accepts only `/v1/chat/completions` and `/v1/responses` + and returns HTTP 400 on `/v1/messages`. pi uses those two APIs, so this setup works + with pi. diff --git a/docs/reference/toml_schema.md b/docs/reference/toml_schema.md index 9f820e72f..ce95015ee 100644 --- a/docs/reference/toml_schema.md +++ b/docs/reference/toml_schema.md @@ -149,6 +149,7 @@ such clients. | `llm_client` | Yes | — | Key under `[llm_clients]`. | | `system_prompt` | No | unset | System prompt prepended when this target serves a completion. | | `extra_body` | No | `{}` | Values merged into the upstream request when the request does not already set that key. | +| `omit_body_fields` | No | `[]` | Top-level fields removed from every request body that Switchyard sends to this target. Switchyard removes them after it translates the request to the LLM client's `format`, so use that format's field names, for example `reasoning_effort` on `openai_chat` or `reasoning` on `openai_responses`. Switchyard applies `extra_body` and `reasoning_effort` after the removal, so either can set a removed field again. | | `reasoning_effort` | No | unset | Reasoning effort forced on every request to this target, replacing the value the caller sent (`reasoning.effort` on `openai_responses`, `reasoning_effort` on `openai_chat`). Rejected on `anthropic_messages` clients. Use it to run one target at a different effort than the client asked for, for example a strong tier at `max` behind a client that sends `high`. Targets with different effort settings need distinct model IDs when used within one route. Separate routes may use the same model ID with separate `llm_clients` entries (same endpoint, different name). | Within one route, callable targets with the same model ID must use the same `llm_client`. From 81740fe5b68345ef706ee50869fd89c2a3262950 Mon Sep 17 00:00:00 2001 From: Elyas Mehtabuddin Date: Fri, 2 Oct 2026 11:09:12 -0700 Subject: [PATCH 4/7] docs(integrations): note when omit_body_fields holds and use Anthropic's published cache prices Signed-off-by: Elyas Mehtabuddin --- docs/integrations/pi.md | 25 ++++++++++++++----------- 1 file changed, 14 insertions(+), 11 deletions(-) diff --git a/docs/integrations/pi.md b/docs/integrations/pi.md index 69a990b92..3a3519bf4 100644 --- a/docs/integrations/pi.md +++ b/docs/integrations/pi.md @@ -112,7 +112,7 @@ Choose the Claude LLM client by who holds the gateway key: | Who holds the gateway key | Claude LLM client | Result | |---|---|---| | The server, through `api_key_env` | `format = "anthropic_messages"` | Prompt caching and pi's `--thinking` level both work. Every caller's Claude requests use the server-owned key. | -| pi sends it as `apiKey`, and the Claude LLM client forwards it with `forward_auth = true` | `format = "openai_chat"`, with `omit_body_fields = ["reasoning_effort"]` on the target | Prompt caching works. pi's `--thinking` level has no effect, and Claude thinks at its default effort. | +| pi sends it as `apiKey`, and the Claude LLM client forwards it with `forward_auth = true` | `format = "openai_chat"`, with `omit_body_fields = ["reasoning_effort"]` on the target | Prompt caching works. pi's `--thinking` level has no effect, and Claude thinks at its default effort. The omitted field stays out only if the target's `extra_body` and `reasoning_effort` do not set it again. | Do not use `format = "openai_responses"` for Claude targets on such a gateway. The gateway tested for this page never cached Claude prompts on `/v1/responses`, and it @@ -153,20 +153,21 @@ multiply each token count in its record by the matching price, and add the resul | `cache_creation_tokens` | Cache-write price | | `completion_tokens` | Output price | -Anthropic's published pricing sets the cache-read price at 0.1 times the input price. It -sets the cache-write price at 1.25 times the input price for a 5-minute cache, or 2 times -for a 1-hour cache. The routing log does not record which cache lifetime the gateway -used, and a gateway may charge its own prices, so the result is an estimate, not the -gateway's bill. +Anthropic's [pricing page](https://platform.claude.com/docs/en/about-claude/pricing) +sets the cache-read price at 0.1 times the input price for most models, including +Claude Sonnet 5, but at 0.05 times for Claude Opus 5.5. It sets the cache-write price at +1.25 times the input price for a 5-minute cache, or 2 times for a 1-hour cache. The +routing log does not record which cache lifetime the gateway used, and a gateway may +charge its own prices, so the result is an estimate, not the gateway's bill. Put your prices in a `prices.json` file, in USD per million tokens. Key each entry by the -`model` value from the routing log. The rates below are illustrative: they only follow -Anthropic's published ratios, with cache reads at 0.1 times and 5-minute cache writes at -1.25 times the input price. Replace them with your provider's current prices. +`model` value from the routing log. The example below uses Anthropic's list prices for +Claude Sonnet 5 on 2026-10-02, with the 5-minute cache-write price. Check the current +pricing page or your gateway's prices before you rely on the result. ```json { - "claude-opus-5-5": {"input": 10.00, "cache_read": 1.00, "cache_write": 12.50, "output": 50.00} + "claude-sonnet-5": {"input": 2.00, "cache_read": 0.20, "cache_write": 2.50, "output": 10.00} } ``` @@ -211,7 +212,9 @@ You have two options: - Keep the OpenAI-format LLM client and remove the effort field from requests to the target with [`omit_body_fields`](../reference/toml_schema.md#targetsname). Use the field name of the target's format: `reasoning_effort` on `openai_chat`, or `reasoning` - on `openai_responses`. + on `openai_responses`. Switchyard applies the target's `extra_body` and + `reasoning_effort` after `omit_body_fields`, so this works only if neither sets the + field again. ```toml [targets.claude] From 57e67f7206ca0d8fcb16c9583434ecb0df5feb5d Mon Sep 17 00:00:00 2001 From: Elyas Mehtabuddin Date: Fri, 2 Oct 2026 15:10:19 -0700 Subject: [PATCH 5/7] docs(integrations): document forwarding one gateway key to OpenAI and Anthropic clients Signed-off-by: Elyas Mehtabuddin --- docs/integrations/oh_my_pi.md | 17 +++++++------ docs/integrations/pi.md | 45 ++++++++++++++++++----------------- 2 files changed, 31 insertions(+), 31 deletions(-) diff --git a/docs/integrations/oh_my_pi.md b/docs/integrations/oh_my_pi.md index 69d9b6a7a..f276fe34c 100644 --- a/docs/integrations/oh_my_pi.md +++ b/docs/integrations/oh_my_pi.md @@ -142,19 +142,18 @@ To forward `omp`'s key, remove `auth: none` and set `apiKey: GATEWAY_API_KEY`, t of the environment variable that holds your gateway key. Unlike pi, `omp` reads the name without a leading `$`. -A route that forwards the key to an `anthropic_messages` client accepts requests only on -`/v1/messages`, and it cannot also forward the key to an `openai_chat` or -`openai_responses` client. So a route with a GPT judge on `openai_responses` cannot -forward the caller's key to both the judge and Claude targets on `anthropic_messages`. -Choose one of two setups: +A route that forwards the key only to `anthropic_messages` clients accepts requests only +on `/v1/messages`. A route that forwards the key to both a GPT judge on `openai_responses` +and Claude targets on `anthropic_messages` works when both LLM clients use the same +scheme, host, and port, as clients on one gateway do. Switchyard then turns the effort +into adaptive thinking, and every request uses `omp`'s key. Two other setups also work: - Forward the key to every LLM client, and keep the Claude targets on `openai_chat` with - `omit_body_fields = ["reasoning_effort"]`. The Claude models then think at their - default effort, and `--thinking` has no effect on them. + `omit_body_fields = ["reasoning_effort"]`. This works across hosts, but the Claude + models then think at their default effort, and `--thinking` has no effect on them. - Forward the key only to the GPT judge, and give the Claude targets an `anthropic_messages` client with `api_key_env`, so they use a server-owned key. - Switchyard then turns the effort into adaptive thinking. -Both setups forward the key to an OpenAI-format LLM client, so the route accepts only +All three setups forward the key to an OpenAI-format LLM client, so the route accepts only `/v1/chat/completions` and `/v1/responses` and returns HTTP 400 on `/v1/messages`. Use `openai-completions` or `openai-responses`. diff --git a/docs/integrations/pi.md b/docs/integrations/pi.md index 3a3519bf4..72903d280 100644 --- a/docs/integrations/pi.md +++ b/docs/integrations/pi.md @@ -112,11 +112,12 @@ Choose the Claude LLM client by who holds the gateway key: | Who holds the gateway key | Claude LLM client | Result | |---|---|---| | The server, through `api_key_env` | `format = "anthropic_messages"` | Prompt caching and pi's `--thinking` level both work. Every caller's Claude requests use the server-owned key. | -| pi sends it as `apiKey`, and the Claude LLM client forwards it with `forward_auth = true` | `format = "openai_chat"`, with `omit_body_fields = ["reasoning_effort"]` on the target | Prompt caching works. pi's `--thinking` level has no effect, and Claude thinks at its default effort. The omitted field stays out only if the target's `extra_body` and `reasoning_effort` do not set it again. | +| pi sends it as `apiKey`, and the route also forwards it to an OpenAI-format LLM client on the same gateway, such as a GPT judge | `format = "anthropic_messages"` with `forward_auth = true`, on the same scheme, host, and port as that client | Prompt caching and pi's `--thinking` level both work, and every request uses pi's own key. | +| pi sends it as `apiKey`, and the route forwards it only to Claude targets | `format = "openai_chat"`, with `omit_body_fields = ["reasoning_effort"]` on the target | Prompt caching works. pi's `--thinking` level has no effect, and Claude thinks at its default effort. The omitted field stays out only if the target's `extra_body` and `reasoning_effort` do not set it again. | Do not use `format = "openai_responses"` for Claude targets on such a gateway. The gateway tested for this page never cached Claude prompts on `/v1/responses`, and it -returned HTTP 400 when thinking was on (see [Thinking](#thinking)). In both setups, a GPT +returned HTTP 400 when thinking was on (see [Thinking](#thinking)). In every setup, a GPT judge on `openai_responses` can still use pi's forwarded key (see [Forwarded keys](#forwarded-keys)). @@ -236,26 +237,26 @@ key, replace the `apiKey` placeholder with the name of an environment variable t your gateway key, with a leading `$`: `"apiKey": "$GATEWAY_API_KEY"`. Without the `$`, pi sends the name itself as the key. -Two rules limit forwarding in one route: +The LLM clients that forward the key decide which request APIs a route accepts: -- Every LLM client in the route that sets `forward_auth = true` must use the same API - family: `openai_chat` and `openai_responses`, or `anthropic_messages`. Otherwise the - server does not start and prints +- If they all use OpenAI formats (`openai_chat` or `openai_responses`), the route accepts + `/v1/chat/completions` and `/v1/responses`, the two APIs pi uses. +- If they all use `anthropic_messages`, the route accepts only `/v1/messages`, and pi + should not use that API (see [Which request API](#which-request-api)). +- If they mix the two, the route accepts `/v1/chat/completions` and `/v1/responses`. The + server starts only if all of them use the same scheme, host, and port in `base_url`, + as clients on one gateway do. Otherwise it prints an error that starts with `route cannot forward both Anthropic and OpenAI caller credentials`. `` is the route's `[routes.]` table key, not its `id`. -- A route that forwards the key to an `anthropic_messages` client accepts requests only on - `/v1/messages`, and pi should not use that API (see [Which request API](#which-request-api)). - -So if the route forwards pi's key to its Claude targets, put them on `openai_chat` with -`omit_body_fields`. - -To keep pi's `--thinking` level, give the Claude targets an `anthropic_messages` client -with `api_key_env`. Which endpoints the route accepts then depends on the route's other -LLM clients: - -- If no LLM client in the route sets `forward_auth = true`, the route accepts every - request API. -- If an OpenAI-format LLM client forwards the key, for example a GPT judge on - `openai_responses`, the route accepts only `/v1/chat/completions` and `/v1/responses` - and returns HTTP 400 on `/v1/messages`. pi uses those two APIs, so this setup works - with pi. + +So to forward pi's key and keep pi's `--thinking` level, the route needs an OpenAI-format +LLM client that also forwards the key on the same gateway, such as a GPT judge on +`openai_responses`. Then put the Claude targets on an `anthropic_messages` client with +`forward_auth = true` (see the example under +[`[llm_clients.]`](../reference/toml_schema.md#llm_clientsname)). + +A route without such a client cannot do both. To forward pi's key to Claude, put the +Claude targets on `openai_chat` with `omit_body_fields`. To keep `--thinking`, give them +an `anthropic_messages` client with `api_key_env`. Clients with `api_key_env` never limit +the request APIs: if no LLM client in the route sets `forward_auth = true`, the route +accepts every request API. From 85aad1145570de8559cc6634154486977e4acda9 Mon Sep 17 00:00:00 2001 From: Elyas Mehtabuddin Date: Fri, 2 Oct 2026 15:21:18 -0700 Subject: [PATCH 6/7] docs(integrations): say only standalone switchyard-server forwards keys Signed-off-by: Elyas Mehtabuddin --- docs/integrations/oh_my_pi.md | 4 ++++ docs/integrations/pi.md | 4 ++++ 2 files changed, 8 insertions(+) diff --git a/docs/integrations/oh_my_pi.md b/docs/integrations/oh_my_pi.md index f276fe34c..f8f094618 100644 --- a/docs/integrations/oh_my_pi.md +++ b/docs/integrations/oh_my_pi.md @@ -142,6 +142,10 @@ To forward `omp`'s key, remove `auth: none` and set `apiKey: GATEWAY_API_KEY`, t of the environment variable that holds your gateway key. Unlike pi, `omp` reads the name without a leading `$`. +Only standalone `switchyard-server` forwards keys. The native Nemo Relay plugin rejects +routes that use `forward_auth = true` (see +[Request Handling](nemo_relay.md#request-handling)). + A route that forwards the key only to `anthropic_messages` clients accepts requests only on `/v1/messages`. A route that forwards the key to both a GPT judge on `openai_responses` and Claude targets on `anthropic_messages` works when both LLM clients use the same diff --git a/docs/integrations/pi.md b/docs/integrations/pi.md index 72903d280..505f6e150 100644 --- a/docs/integrations/pi.md +++ b/docs/integrations/pi.md @@ -237,6 +237,10 @@ key, replace the `apiKey` placeholder with the name of an environment variable t your gateway key, with a leading `$`: `"apiKey": "$GATEWAY_API_KEY"`. Without the `$`, pi sends the name itself as the key. +Only standalone `switchyard-server` forwards keys. The native Nemo Relay plugin rejects +routes that use `forward_auth = true` (see +[Request Handling](nemo_relay.md#request-handling)). + The LLM clients that forward the key decide which request APIs a route accepts: - If they all use OpenAI formats (`openai_chat` or `openai_responses`), the route accepts From d655d29d045d3e8373c0aeedfc56265993ccd314 Mon Sep 17 00:00:00 2001 From: Elyas Mehtabuddin Date: Fri, 2 Oct 2026 16:11:36 -0700 Subject: [PATCH 7/7] docs(integrations): lead the Claude gateway sections with the format to use Signed-off-by: Elyas Mehtabuddin --- docs/integrations/oh_my_pi.md | 95 ++++++---------- docs/integrations/pi.md | 199 ++++++++++++++++------------------ 2 files changed, 128 insertions(+), 166 deletions(-) diff --git a/docs/integrations/oh_my_pi.md b/docs/integrations/oh_my_pi.md index f8f094618..2fadf0fa4 100644 --- a/docs/integrations/oh_my_pi.md +++ b/docs/integrations/oh_my_pi.md @@ -96,68 +96,43 @@ as a LiteLLM proxy. In that case, run Switchyard on another port or set `LITELLM_BASE_URL`. Set `cost` on the model entry if you want `omp` to show a non-zero cost. -## Claude targets behind an OpenAI-compatible gateway - -The pi guide's section -[Claude targets behind an OpenAI-compatible gateway](pi.md#claude-targets-behind-an-openai-compatible-gateway) -applies to `omp` too. Use its table to choose the Claude LLM client by who holds the -gateway key. Switchyard calls the gateway endpoint that matches the target's LLM client -`format`, whatever `api` `omp` uses, so the `format` decides whether the gateway caches -the prompt. Use `format = "openai_chat"` or `"anthropic_messages"` for Claude targets, -because a gateway may not cache Claude prompts on `/v1/responses`. To estimate what the -requests cost from the routing log, see [Estimate the cost](pi.md#estimate-the-cost) in -the pi guide. This section covers what differs for `omp`, checked with Oh My Pi 18.2.11. - -### Thinking - -Thinking depends on both the LLM client `format` and the `api`: - -- On `openai-completions`, `omp` sends `reasoning_effort`, and on `openai-responses` it - sends `reasoning.effort`. Switchyard passes the effort to an `openai_chat` target as - `reasoning_effort` and to an `openai_responses` target as `reasoning.effort`. Some - gateways turn either field into a thinking setting that Claude Opus 5.5 and Sonnet 5 - refuse with HTTP 400. Switchyard turns the effort into adaptive thinking only for an - `anthropic_messages` target. The pi guide's [Thinking](pi.md#thinking) section shows - how to remove the field with `omit_body_fields` instead. -- On `anthropic-messages`, Switchyard sends `omp`'s own `thinking` settings to an - `anthropic_messages` target unchanged. `omp` does not recognize a route id such as - `switchyard` as a Claude model, so with thinking on it sends - `thinking: {type: "enabled"}`, and Claude Opus 5.5 and Sonnet 5 return HTTP 400. Tell - `omp` to use adaptive thinking on the model entry: - - ```yaml - - id: switchyard - reasoning: true - thinking: - mode: anthropic-adaptive - efforts: [low, medium, high] - ``` - - `omp` requires `efforts` next to `mode`. It then sends `thinking: {type: "adaptive"}` - and `output_config.effort`. +## Claude through an LLM gateway -### Forwarded keys +The pi guide's [Claude through an LLM gateway](pi.md#claude-through-an-llm-gateway) +section applies to `omp`: give Claude targets an LLM client with +`format = "anthropic_messages"`. It shows how to check prompt caching and estimate the +cost. This section covers what differs for `omp`, tested with Oh My Pi 18.2.11. + +### Thinking on `anthropic-messages` + +On `openai-completions` and `openai-responses`, Switchyard turns `omp`'s thinking level +into adaptive thinking for an `anthropic_messages` target. On `anthropic-messages`, +Switchyard sends `omp`'s own `thinking` object unchanged. `omp` does not +recognize a route ID such as `switchyard` as a Claude model. With thinking on, it sends +`thinking: {type: "enabled"}`, and Claude Opus 5.5 and Sonnet 5 return HTTP 400. Set +adaptive thinking on the model entry: + +```yaml + - id: switchyard + reasoning: true + thinking: + mode: anthropic-adaptive + efforts: [low, medium, high] +``` -To forward `omp`'s key, remove `auth: none` and set `apiKey: GATEWAY_API_KEY`, the name -of the environment variable that holds your gateway key. Unlike pi, `omp` reads the name -without a leading `$`. +`omp` requires `efforts` next to `mode`. With both set, it sends +`thinking: {type: "adaptive"}` and `output_config.effort`. + +### Forwarded keys -Only standalone `switchyard-server` forwards keys. The native Nemo Relay plugin rejects -routes that use `forward_auth = true` (see +To forward `omp`'s key, remove `auth: none` and set `apiKey: GATEWAY_API_KEY`, the name of +the environment variable that holds your gateway key. Unlike pi, `omp` reads the name +without a leading `$`. With `auth: none`, `omp` sends no key, and the gateway returns HTTP +401. Only standalone `switchyard-server` forwards keys. The native Nemo Relay plugin +rejects routes that use `forward_auth = true` (see [Request Handling](nemo_relay.md#request-handling)). -A route that forwards the key only to `anthropic_messages` clients accepts requests only -on `/v1/messages`. A route that forwards the key to both a GPT judge on `openai_responses` -and Claude targets on `anthropic_messages` works when both LLM clients use the same -scheme, host, and port, as clients on one gateway do. Switchyard then turns the effort -into adaptive thinking, and every request uses `omp`'s key. Two other setups also work: - -- Forward the key to every LLM client, and keep the Claude targets on `openai_chat` with - `omit_body_fields = ["reasoning_effort"]`. This works across hosts, but the Claude - models then think at their default effort, and `--thinking` has no effect on them. -- Forward the key only to the GPT judge, and give the Claude targets an - `anthropic_messages` client with `api_key_env`, so they use a server-owned key. - -All three setups forward the key to an OpenAI-format LLM client, so the route accepts only -`/v1/chat/completions` and `/v1/responses` and returns HTTP 400 on `/v1/messages`. Use -`openai-completions` or `openai-responses`. +The pi guide's [Forwarded keys](pi.md#forwarded-keys) table shows which request APIs a +route accepts when it forwards the key. A route that forwards the key to an OpenAI-format +LLM client, such as a classifier's GPT judge, accepts only `openai-completions` and +`openai-responses` requests. diff --git a/docs/integrations/pi.md b/docs/integrations/pi.md index 505f6e150..8560fa2b5 100644 --- a/docs/integrations/pi.md +++ b/docs/integrations/pi.md @@ -99,72 +99,73 @@ Set `cost` on the model entry if you want pi to show a non-zero cost. [`benchmark/run-baseline.sh`](../../benchmark/README.md) runs Terminal-Bench tasks with pi through Switchyard when you pass `--agent pi`. -## Claude targets behind an OpenAI-compatible gateway - -Some gateways, such as a LiteLLM proxy, serve Claude models on `/v1/chat/completions`, -`/v1/responses`, and `/v1/messages` with one API key. Switchyard calls the endpoint that -matches the target's LLM client `format`, whatever `api` pi uses. On the gateway tested -for this page, the `format` decided whether Claude prompts were cached and whether pi's -`--thinking` level worked. +## Claude through an LLM gateway + +An LLM gateway, such as a LiteLLM proxy, can serve Claude on three endpoints: +`/v1/chat/completions`, `/v1/responses`, and `/v1/messages`. Switchyard calls the +endpoint that matches the `format` of the Claude target's LLM client, whichever `api` pi +uses. Give Claude targets an LLM client with `format = "anthropic_messages"`: + +```toml +[llm_clients.gateway_claude] +format = "anthropic_messages" +base_url = "https://gateway.example.com" +api_key_env = "GATEWAY_API_KEY" + +[targets.claude] +id = "claude-sonnet-5" # the gateway's model ID +llm_client = "gateway_claude" +``` -Choose the Claude LLM client by who holds the gateway key: +On the tested gateway, both OpenAI formats returned HTTP 400 when pi sent a thinking +level, and `openai_responses` never read the prompt cache: -| Who holds the gateway key | Claude LLM client | Result | -|---|---|---| -| The server, through `api_key_env` | `format = "anthropic_messages"` | Prompt caching and pi's `--thinking` level both work. Every caller's Claude requests use the server-owned key. | -| pi sends it as `apiKey`, and the route also forwards it to an OpenAI-format LLM client on the same gateway, such as a GPT judge | `format = "anthropic_messages"` with `forward_auth = true`, on the same scheme, host, and port as that client | Prompt caching and pi's `--thinking` level both work, and every request uses pi's own key. | -| pi sends it as `apiKey`, and the route forwards it only to Claude targets | `format = "openai_chat"`, with `omit_body_fields = ["reasoning_effort"]` on the target | Prompt caching works. pi's `--thinking` level has no effect, and Claude thinks at its default effort. The omitted field stays out only if the target's `extra_body` and `reasoning_effort` do not set it again. | +| Claude LLM client `format` | Gateway endpoint | Prompt cache | pi's thinking level | +|---|---|---|---| +| `anthropic_messages` | `/v1/messages` | Read on repeated prompts | Works | +| `openai_chat` | `/v1/chat/completions` | Read on repeated prompts | HTTP 400 | +| `openai_responses` | `/v1/responses` | Never read | HTTP 400 | -Do not use `format = "openai_responses"` for Claude targets on such a gateway. The -gateway tested for this page never cached Claude prompts on `/v1/responses`, and it -returned HTTP 400 when thinking was on (see [Thinking](#thinking)). In every setup, a GPT -judge on `openai_responses` can still use pi's forwarded key (see -[Forwarded keys](#forwarded-keys)). +[Prompt caching](#prompt-caching) and [Thinking](#thinking) explain both failures. To +use each developer's own gateway key instead of a key that the server holds, see +[Forwarded keys](#forwarded-keys). ### Prompt caching -On the LiteLLM gateway tested for this page, a repeated Claude prompt was read from the -cache on `/v1/chat/completions` and `/v1/messages`, but never on `/v1/responses`. Every -`/v1/responses` request counted the whole prompt as uncached input. +pi sends the whole conversation on every turn. When the gateway reads the repeated part +from Claude's prompt cache, that part costs 0.1 times the input price on Claude Sonnet 5 +and 0.05 times on Claude Opus 5.5, according to Anthropic's +[pricing page](https://platform.claude.com/docs/en/about-claude/pricing). The tested +gateway never read Claude prompts from the cache on `/v1/responses`, so every turn there +paid the full input price for the whole conversation. To check your gateway, start the server with `--routing-log-file PATH` and send the same prompt twice. Claude does not cache short prompts, so use a prompt of at least 5,000 -tokens. In the tests for this page, prompts of about 5,000 tokens were cached on Claude -Opus 5.5 and Sonnet 5; the tests did not find Claude's exact minimum. Then read the -records: +tokens. Then read the records: ```bash jq -c '{route_id, prompt_tokens, cached_tokens, cache_creation_tokens}' PATH ``` -If the gateway caches the prompt, the first record shows it in `cache_creation_tokens` -and the second shows `cached_tokens` close to `prompt_tokens`. If the second record shows -`"cached_tokens": 0`, the gateway read nothing from the cache, and the whole prompt -counts as uncached input again. +When caching works, the first record shows the prompt in `cache_creation_tokens`, and the +second shows `cached_tokens` close to `prompt_tokens`. If the second record shows +`"cached_tokens": 0`, the gateway read nothing from the cache. #### Estimate the cost The routing log records token counts, not prices. To estimate what a request cost, -multiply each token count in its record by the matching price, and add the results: +multiply each count in its record by the matching price and add the results: | Tokens in the record | Price | |---|---| -| Uncached input: `prompt_tokens - cached_tokens - cache_creation_tokens` | Input price | -| `cached_tokens` | Cache-read price | -| `cache_creation_tokens` | Cache-write price | -| `completion_tokens` | Output price | - -Anthropic's [pricing page](https://platform.claude.com/docs/en/about-claude/pricing) -sets the cache-read price at 0.1 times the input price for most models, including -Claude Sonnet 5, but at 0.05 times for Claude Opus 5.5. It sets the cache-write price at -1.25 times the input price for a 5-minute cache, or 2 times for a 1-hour cache. The -routing log does not record which cache lifetime the gateway used, and a gateway may -charge its own prices, so the result is an estimate, not the gateway's bill. - -Put your prices in a `prices.json` file, in USD per million tokens. Key each entry by the -`model` value from the routing log. The example below uses Anthropic's list prices for -Claude Sonnet 5 on 2026-10-02, with the 5-minute cache-write price. Check the current -pricing page or your gateway's prices before you rely on the result. +| `prompt_tokens - cached_tokens - cache_creation_tokens` | Input | +| `cached_tokens` | Cache read | +| `cache_creation_tokens` | Cache write | +| `completion_tokens` | Output | + +Write the prices to `prices.json` in USD per million tokens, and key each entry by the +record's `model` value. This example uses Anthropic's list prices for Claude Sonnet 5 on +2026-10-02: ```json { @@ -172,8 +173,11 @@ pricing page or your gateway's prices before you rely on the result. } ``` -The command below prints one estimated cost per record. For a record whose model has no -entry in `prices.json`, it prints a warning instead of a cost: +The example's `cache_write` price is for a 5-minute cache. A 1-hour cache costs 2 times +the input price instead of 1.25 times. The routing log does not say which cache the +gateway used. A gateway may also charge its own prices, so the result is an estimate, not +the gateway's bill. This command prints one estimated cost per record, or a warning for a +model that has no entry in `prices.json`: ```bash jq -r --slurpfile prices prices.json ' @@ -192,75 +196,58 @@ jq -r --slurpfile prices prices.json ' ### Thinking -Claude Opus 5.5 and Sonnet 5 accept only adaptive thinking. With `reasoning: true`, pi -sends `reasoning_effort`, even when you do not pass `--thinking`, and Switchyard -passes that field unchanged to an `openai_chat` target. Some gateways, including the one -tested for this page, turn `reasoning_effort` into Anthropic's older -`thinking: {type: "enabled"}` and return HTTP 400: +With `reasoning: true` on the model entry, pi sends a thinking level on every request, +even when you do not pass `--thinking`. Switchyard passes that level to an OpenAI-format +target in OpenAI form: `reasoning_effort` on `openai_chat` and `reasoning.effort` on +`openai_responses`. The tested gateway turned either field into Anthropic's older +`thinking: {type: "enabled"}`. Claude Opus 5.5 and Sonnet 5 refuse that form, so the +gateway returned HTTP 400: ```text "thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior. ``` -An `openai_responses` target fails the same way. Switchyard sends the effort to it as -`reasoning.effort`, and the gateway returns the same error. - -You have two options: +For an `anthropic_messages` target, Switchyard sends the level in the form that Claude +accepts: `thinking: {type: "adaptive"}` with `output_config.effort`. pi's thinking level +then takes effect. -- Use an `anthropic_messages` LLM client. Switchyard turns pi's effort into - `thinking: {type: "adaptive"}` and `output_config.effort`, so pi's `--thinking` level - still applies. -- Keep the OpenAI-format LLM client and remove the effort field from requests to the - target with [`omit_body_fields`](../reference/toml_schema.md#targetsname). Use the - field name of the target's format: `reasoning_effort` on `openai_chat`, or `reasoning` - on `openai_responses`. Switchyard applies the target's `extra_body` and - `reasoning_effort` after `omit_body_fields`, so this works only if neither sets the - field again. +If a Claude target must stay on an OpenAI format, remove the effort field from its +requests with [`omit_body_fields`](../reference/toml_schema.md#targetsname). Claude then +thinks at its default effort, and pi's thinking level has no effect on that target: - ```toml - [targets.claude] - id = "claude-opus-5-5" - llm_client = "gateway_chat" # format = "openai_chat" - omit_body_fields = ["reasoning_effort"] - ``` +```toml +[targets.claude] +id = "claude-opus-5-5" +llm_client = "gateway_chat" # format = "openai_chat" +omit_body_fields = ["reasoning_effort"] # use "reasoning" on openai_responses +``` - The request then succeeds, and Claude thinks at its default effort. pi's `--thinking` - level has no effect on this target. +Switchyard applies the target's `extra_body` and `reasoning_effort` after the removal, so +either can set the field again. ### Forwarded keys -An LLM client with `forward_auth = true` sends the caller's key to the gateway. An LLM -client with `api_key_env` sends a server-owned key, which the server reads from an -environment variable (see -[`[llm_clients.]`](../reference/toml_schema.md#llm_clientsname)). To forward pi's -key, replace the `apiKey` placeholder with the name of an environment variable that holds -your gateway key, with a leading `$`: `"apiKey": "$GATEWAY_API_KEY"`. Without the `$`, -pi sends the name itself as the key. - -Only standalone `switchyard-server` forwards keys. The native Nemo Relay plugin rejects -routes that use `forward_auth = true` (see +With `api_key_env`, every caller's requests use one gateway key that the server holds. +To use each developer's own gateway key instead, set `forward_auth = true` on the LLM +client. Then set pi's `apiKey` to `$` plus the name of the environment variable that +holds your key: `"apiKey": "$GATEWAY_API_KEY"`. Without the `$`, pi sends the name itself +as the key, and the gateway returns HTTP 401. Only standalone `switchyard-server` forwards +keys. The native Nemo Relay plugin rejects routes that use `forward_auth = true` (see [Request Handling](nemo_relay.md#request-handling)). -The LLM clients that forward the key decide which request APIs a route accepts: - -- If they all use OpenAI formats (`openai_chat` or `openai_responses`), the route accepts - `/v1/chat/completions` and `/v1/responses`, the two APIs pi uses. -- If they all use `anthropic_messages`, the route accepts only `/v1/messages`, and pi - should not use that API (see [Which request API](#which-request-api)). -- If they mix the two, the route accepts `/v1/chat/completions` and `/v1/responses`. The - server starts only if all of them use the same scheme, host, and port in `base_url`, - as clients on one gateway do. Otherwise it prints an error that starts with - `route cannot forward both Anthropic and OpenAI caller credentials`. `` is - the route's `[routes.]` table key, not its `id`. - -So to forward pi's key and keep pi's `--thinking` level, the route needs an OpenAI-format -LLM client that also forwards the key on the same gateway, such as a GPT judge on -`openai_responses`. Then put the Claude targets on an `anthropic_messages` client with -`forward_auth = true` (see the example under -[`[llm_clients.]`](../reference/toml_schema.md#llm_clientsname)). - -A route without such a client cannot do both. To forward pi's key to Claude, put the -Claude targets on `openai_chat` with `omit_body_fields`. To keep `--thinking`, give them -an `anthropic_messages` client with `api_key_env`. Clients with `api_key_env` never limit -the request APIs: if no LLM client in the route sets `forward_auth = true`, the route -accepts every request API. +Which request APIs a route accepts depends on its LLM clients that forward the key. +Clients with `api_key_env` do not count: + +| Forwarding LLM clients in the route | Request APIs the route accepts | +|---|---| +| None | All three | +| Only OpenAI formats | `/v1/chat/completions` and `/v1/responses` | +| OpenAI formats and `anthropic_messages`, all on the same scheme, host, and port | `/v1/chat/completions` and `/v1/responses` | +| Only `anthropic_messages` | `/v1/messages`, which pi should not use (see [Which request API](#which-request-api)) | + +For example, a classifier route can forward pi's key to its GPT judge on +`openai_responses` and to its Claude targets on `anthropic_messages` when both LLM clients +point at the same gateway. If they use different hosts, ports, or schemes, the server does +not start. If a route forwards the key only to Claude targets on `anthropic_messages`, pi +cannot use it. Put those targets on `openai_chat` with `omit_body_fields` (see +[Thinking](#thinking)), or let the server hold the key with `api_key_env`.