Skip to content

feat(server): let a forwarding route mix OpenAI and Anthropic clients on one host - #873

Open
elyasmnvidian wants to merge 1 commit into
mainfrom
emehtabuddin/switch-1631-forward-auth-across-formats
Open

elyasmnvidian wants to merge 1 commit into
mainfrom
emehtabuddin/switch-1631-forward-auth-across-formats

Conversation

@elyasmnvidian

@elyasmnvidian elyasmnvidian commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

What

A route fails to load when its forward_auth = true clients mix OpenAI formats (openai_chat, openai_responses) and anthropic_messages, even if all of them point at the same gateway. For example, a route with a GPT judge on the gateway's /v1/responses and Claude on its /v1/messages fails with:

route claude_route cannot forward both Anthropic and OpenAI caller credentials

With forward_auth = true, a client sends the caller's own credential upstream. The config check puts each forwarding client in one of two credential families by its format, OpenAI or Anthropic, and refuses a route that uses both. The check exists so that a credential meant for one provider, such as a Codex user's ChatGPT token, never reaches another provider. But a gateway such as a LiteLLM proxy serves GPT and Claude on one host and accepts one key as Authorization: Bearer on all three endpoints. A route on that gateway sends the key to no other host, yet the check refuses it.

This PR lets a forwarding route mix the two credential families when all of its forwarding clients use the same scheme, host, and port in base_url. The path does not count, so https://gateway.example.com and https://gateway.example.com/v1 match. This config now loads:

[llm_clients.gateway_responses]
format = "openai_responses"
base_url = "https://gateway.example.com/v1"
forward_auth = true

[llm_clients.gateway_messages]
format = "anthropic_messages"
base_url = "https://gateway.example.com"
forward_auth = true

A mixed route serves Chat Completions and Responses callers and forwards the caller's bearer token to every client. Its anthropic_messages clients send the authorization header unchanged, as the Anthropic backend already does. A /v1/messages caller gets the existing 400 before any upstream call. Mixing the families across different hosts still fails, and the error now says how to fix it:

route claude_route cannot forward both Anthropic and OpenAI caller credentials to different hosts (https://api.anthropic.com, https://gateway.example.com); point all of its forwarding clients at one host

What stays the same: there is no new config key, and every config that loads today loads the same way. A route that does not mix families behaves as before, so a passthrough route on the anthropic_messages client above still serves /v1/messages callers such as Claude Code.

Why

Claude must be called through /v1/messages to get the reasoning effort the caller asks for. On the gateway I tested, Claude returns 400 when Switchyard sends the effort in OpenAI form (reasoning_effort on /v1/chat/completions, reasoning.effort on /v1/responses): "thinking.type.enabled" is not supported for this model. On /v1/messages, Switchyard sends thinking: {"type": "adaptive"} plus output_config.effort, and the gateway accepts it. Caching is not the reason: /v1/chat/completions caches Claude prompts too.

Today a route can forward the caller's key or send Claude the effort, not both. Claude on openai_chat with omit_body_fields = ["reasoning_effort"] uses the caller's key but never gets the effort. Claude on anthropic_messages with api_key_env gets the effort, but the server must hold the gateway key.

The new rule still keeps the caller's credential on one host. When every forwarding client of a route calls the same scheme, host, and port, the credential can reach only that host, which the operator configured. The mixed route serves only Chat Completions and Responses callers because an anthropic_messages client forwards their bearer token as is. An OpenAI-format client, by contrast, would pass a Messages caller's x-api-key on as an ordinary header.

Notes for reviewers

Start with build_route_clients in crates/switchyard-runner/src/config.rs. The server's per-request check reads the credential family that this function returns, so no server code changed. As with any forward_auth client today, the check cannot tell whether that one host should receive the caller's login; the operator decides that.

#874 says in the pi and Oh My Pi guides that a forwarding route cannot use both formats. Whichever PR merges second should add the one-host exception there.

Tests
  • route_on_one_host_forwards_the_bearer_token_to_responses_and_messages (crates/switchyard-server/tests/server.rs): one stub serves /v1/responses and /v1/messages. A composite route has its judge on an openai_responses client and its tiers on an anthropic_messages client, both forwarding to the stub. The stub records every value of authorization, x-api-key, chatgpt-account-id, and anthropic-version. For a Chat Completions caller and a Responses caller, the test expects exactly one unchanged bearer token on both endpoints, no x-api-key, and anthropic-version only on /v1/messages. A /v1/messages caller on the composite route gets 400 with no stub call. A passthrough route on the same anthropic_messages client serves a /v1/messages caller. On main, the test fails because the config does not load: route agent cannot forward both Anthropic and OpenAI caller credentials.
  • forwarding_route_mixes_credential_families_only_on_one_host (crates/switchyard-runner/src/config.rs): the mixed route loads with the OpenAI family when the Messages client uses the same host, with and without /v1. A passthrough route on the Messages client keeps the Anthropic family. A different host, port, or scheme fails with the full new error, which lists both origins.
Live run

Host names are replaced with gateway.example.com, and the gateway's model IDs with public model IDs. The config held no key. Both clients pointed at a local logging proxy in front of the gateway. The proxy recorded header names and SHA-256 prefixes of credential headers, never their values. Each caller sent Authorization: Bearer <gateway key>.

[llm_clients.gateway_responses]
format = "openai_responses"
base_url = "http://127.0.0.1:18731/v1"
forward_auth = true
max_retries = 0

[llm_clients.gateway_messages]
format = "anthropic_messages"
base_url = "http://127.0.0.1:18731"
forward_auth = true
max_retries = 0

[targets]
capable = { id = "claude-opus-5-5", llm_client = "gateway_messages" }
efficient = { id = "claude-sonnet-5", llm_client = "gateway_messages" }
judge = { id = "gpt-5.6-terra", llm_client = "gateway_responses" }

[routes.claude_route]
id = "claude-route"
type = "composite"
classifier = { target = "judge", base_threshold = 0.5, classify_trigger = "user_turn", message_hash_fallback = true }
stage = { capable_target = "capable", efficient_target = "efficient", confidence_threshold = 0.5 }

[routes.sonnet_route]
id = "sonnet-route"
type = "passthrough"
target = "efficient"

--dry-run with the gateway's own base URLs, then with the Messages client moved to a second host:

$ switchyard-server --config one-host.toml --dry-run
server OK: claude-route, sonnet-route
$ switchyard-server --config second-host.toml --dry-run
invalid server config second-host.toml: route claude_route cannot forward both Anthropic and OpenAI caller credentials to different hosts (https://api.anthropic.com, https://gateway.example.com); point all of its forwarding clients at one host: route claude_route cannot forward both ...

The error appears twice because of the existing error-chain printing. The same config with both clients on /v1 also loads, and two loopback ports fail with the same error.

# Caller endpoint Mode Route Served model Result Upstream calls
A /v1/responses, reasoning.effort: high stream claude-route claude-sonnet-5 200, answer 24, 0 reasoning tokens judge /v1/responses, answer /v1/messages
A2 /v1/responses, reasoning.effort: high stream claude-route claude-sonnet-5 200, reasoning item, 200 reasoning tokens, answer 105 judge /v1/responses, answer /v1/messages
B /v1/chat/completions buffered claude-route claude-sonnet-5 200, 1073, finish_reason stop judge /v1/responses, answer /v1/messages
C /v1/messages buffered claude-route none 400 route claude-route forwards an OpenAI login; call it through /v1/chat/completions or /v1/responses none
D /v1/messages buffered sonnet-route claude-sonnet-5 200, **Hallo!** answer /v1/messages

Across the 7 upstream calls, each call carried exactly one authorization whose hash matched the caller's, and none carried x-api-key. Only the /v1/messages calls carried anthropic-version: 2023-06-01. For A and A2, the /v1/messages body had thinking: {"type": "adaptive"} and output_config: {"effort": "high"}. With adaptive thinking, Claude skipped thinking on A's easy question and thought on A2's.

Summary by CodeRabbit

  • New Features
    • Forwarding routes can now combine OpenAI and Anthropic clients when all clients share the same scheme, host, and port.
    • These mixed routes serve Chat Completions and Responses requests, forwarding the caller’s bearer token to each client.
  • Behavior Changes
    • Requests using an API that the route does not serve return HTTP 400 before being forwarded.
    • Mixed-family routes pointing to different origins are rejected during configuration.
  • Documentation
    • Updated setup and integration guides to describe the routing and compatibility rules.

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.8.1

🚀 View preview at
https://NVIDIA-NeMo.github.io/Switchyard/pr-preview/pr-873/

Built to branch gh-pages at 2026-10-02 17:14 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

@elyasmnvidian
elyasmnvidian force-pushed the emehtabuddin/switch-1631-forward-auth-across-formats branch from ac88535 to 117e8d6 Compare October 1, 2026 20:28
@elyasmnvidian elyasmnvidian changed the title feat(server): add forward_auth_from to send an OpenAI login to anthropic_messages clients feat(server): let a forwarding route mix OpenAI and Anthropic clients on one host Oct 1, 2026
… on one host

Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
@elyasmnvidian
elyasmnvidian force-pushed the emehtabuddin/switch-1631-forward-auth-across-formats branch from 117e8d6 to e33fd8a Compare October 2, 2026 17:13
@elyasmnvidian
elyasmnvidian marked this pull request as ready for review October 2, 2026 17:59
@elyasmnvidian
elyasmnvidian requested a review from a team as a code owner October 2, 2026 17:59
@coderabbitai

coderabbitai Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: NVIDIA-NeMo/Switchyard/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: c7d41bd7-9553-4c3a-a8e0-191175524734

📥 Commits

Reviewing files that changed from the base of the PR and between 16cbe59 and e33fd8a.

📒 Files selected for processing (10)
  • crates/libsy-llm-client/README.md
  • crates/libsy-llm-client/src/backend.rs
  • crates/libsy-llm-client/src/client.rs
  • crates/switchyard-nemo-relay-plugin/README.md
  • crates/switchyard-runner/src/config.rs
  • crates/switchyard-server/README.md
  • crates/switchyard-server/tests/server.rs
  • docs/getting_started.md
  • docs/integrations/nemo_relay.md
  • docs/reference/toml_schema.md

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


Walkthrough

Forwarding routes now allow OpenAI and Anthropic clients together when they share a scheme, host, and port. Tests check configuration, bearer-token forwarding, request headers, and unsupported API responses. Documentation describes the updated rules.

Changes

Credential forwarding

Layer / File(s) Summary
Mixed-family route configuration
crates/switchyard-runner/src/config.rs, crates/libsy-llm-client/src/backend.rs, crates/libsy-llm-client/src/client.rs, docs/reference/toml_schema.md
Configuration accepts mixed credential families when forwarding clients share a scheme, host, and port. Tests cover shared origins with different paths and reject different hosts, ports, or schemes.
Request forwarding and API handling
crates/switchyard-server/tests/server.rs, crates/libsy-llm-client/README.md, crates/switchyard-nemo-relay-plugin/README.md, crates/switchyard-server/README.md, docs/getting_started.md, docs/integrations/nemo_relay.md, docs/reference/toml_schema.md
Integration tests check bearer-token forwarding, format-specific headers, and rejection of unsupported API requests before upstream dispatch. Documentation describes the forwarding behavior and supported APIs.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~20 minutes

Merge Risk: ⚪ Minimal · up to e33fd

The supplied evidence identifies no actionable issue that should block merging.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 77.78% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 9 functions across 4 files. (6 skipped: 6… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: allowing forwarding routes to mix OpenAI and Anthropic clients when they share one host.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Full details: Docstring Coverage

Explanation

Docstring coverage is 77.78% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 9 functions across 4 files. (6 skipped: 6 unsupported.)

  • Fix all pre-merge checks with AI
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


I’m a rabbit with a token to share,
I hop through routes with careful flair.
OpenAI and Anthropic meet,
When hosts and ports align and greet.
A wrong API gets stopped upstream-free,
Then I nibble docs beside my tea.

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant