feat(server): let a forwarding route mix OpenAI and Anthropic clients on one host - #873
elyasmnvidian wants to merge 1 commit into
Conversation
|
ac88535 to
117e8d6
Compare
… on one host Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
117e8d6 to
e33fd8a
Compare
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: NVIDIA-NeMo/Switchyard/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (10)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review. WalkthroughForwarding routes now allow OpenAI and Anthropic clients together when they share a scheme, host, and port. Tests check configuration, bearer-token forwarding, request headers, and unsupported API responses. Documentation describes the updated rules. ChangesCredential forwarding
Priority: ➖ Normal Estimated code review effort: 3 (Moderate) | ~20 minutes Merge Risk: ⚪ Minimal · up to The supplied evidence identifies no actionable issue that should block merging. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 77.78% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 9 functions across 4 files. (6 skipped: 6 unsupported.)
I’m a rabbit with a token to share, Comment |
What
A route fails to load when its
forward_auth = trueclients mix OpenAI formats (openai_chat,openai_responses) andanthropic_messages, even if all of them point at the same gateway. For example, a route with a GPT judge on the gateway's/v1/responsesand Claude on its/v1/messagesfails with:With
forward_auth = true, a client sends the caller's own credential upstream. The config check puts each forwarding client in one of two credential families by itsformat, OpenAI or Anthropic, and refuses a route that uses both. The check exists so that a credential meant for one provider, such as a Codex user's ChatGPT token, never reaches another provider. But a gateway such as a LiteLLM proxy serves GPT and Claude on one host and accepts one key asAuthorization: Beareron all three endpoints. A route on that gateway sends the key to no other host, yet the check refuses it.This PR lets a forwarding route mix the two credential families when all of its forwarding clients use the same scheme, host, and port in
base_url. The path does not count, sohttps://gateway.example.comandhttps://gateway.example.com/v1match. This config now loads:A mixed route serves Chat Completions and Responses callers and forwards the caller's bearer token to every client. Its
anthropic_messagesclients send theauthorizationheader unchanged, as the Anthropic backend already does. A/v1/messagescaller gets the existing 400 before any upstream call. Mixing the families across different hosts still fails, and the error now says how to fix it:What stays the same: there is no new config key, and every config that loads today loads the same way. A route that does not mix families behaves as before, so a passthrough route on the
anthropic_messagesclient above still serves/v1/messagescallers such as Claude Code.Why
Claude must be called through
/v1/messagesto get the reasoning effort the caller asks for. On the gateway I tested, Claude returns 400 when Switchyard sends the effort in OpenAI form (reasoning_efforton/v1/chat/completions,reasoning.efforton/v1/responses):"thinking.type.enabled" is not supported for this model. On/v1/messages, Switchyard sendsthinking: {"type": "adaptive"}plusoutput_config.effort, and the gateway accepts it. Caching is not the reason:/v1/chat/completionscaches Claude prompts too.Today a route can forward the caller's key or send Claude the effort, not both. Claude on
openai_chatwithomit_body_fields = ["reasoning_effort"]uses the caller's key but never gets the effort. Claude onanthropic_messageswithapi_key_envgets the effort, but the server must hold the gateway key.The new rule still keeps the caller's credential on one host. When every forwarding client of a route calls the same scheme, host, and port, the credential can reach only that host, which the operator configured. The mixed route serves only Chat Completions and Responses callers because an
anthropic_messagesclient forwards their bearer token as is. An OpenAI-format client, by contrast, would pass a Messages caller'sx-api-keyon as an ordinary header.Notes for reviewers
Start with
build_route_clientsincrates/switchyard-runner/src/config.rs. The server's per-request check reads the credential family that this function returns, so no server code changed. As with anyforward_authclient today, the check cannot tell whether that one host should receive the caller's login; the operator decides that.#874 says in the pi and Oh My Pi guides that a forwarding route cannot use both formats. Whichever PR merges second should add the one-host exception there.
Tests
route_on_one_host_forwards_the_bearer_token_to_responses_and_messages(crates/switchyard-server/tests/server.rs): one stub serves/v1/responsesand/v1/messages. Acompositeroute has its judge on anopenai_responsesclient and its tiers on ananthropic_messagesclient, both forwarding to the stub. The stub records every value ofauthorization,x-api-key,chatgpt-account-id, andanthropic-version. For a Chat Completions caller and a Responses caller, the test expects exactly one unchanged bearer token on both endpoints, nox-api-key, andanthropic-versiononly on/v1/messages. A/v1/messagescaller on the composite route gets 400 with no stub call. A passthrough route on the sameanthropic_messagesclient serves a/v1/messagescaller. Onmain, the test fails because the config does not load:route agent cannot forward both Anthropic and OpenAI caller credentials.forwarding_route_mixes_credential_families_only_on_one_host(crates/switchyard-runner/src/config.rs): the mixed route loads with the OpenAI family when the Messages client uses the same host, with and without/v1. A passthrough route on the Messages client keeps the Anthropic family. A different host, port, or scheme fails with the full new error, which lists both origins.Live run
Host names are replaced with
gateway.example.com, and the gateway's model IDs with public model IDs. The config held no key. Both clients pointed at a local logging proxy in front of the gateway. The proxy recorded header names and SHA-256 prefixes of credential headers, never their values. Each caller sentAuthorization: Bearer <gateway key>.--dry-runwith the gateway's own base URLs, then with the Messages client moved to a second host:The error appears twice because of the existing error-chain printing. The same config with both clients on
/v1also loads, and two loopback ports fail with the same error./v1/responses,reasoning.effort: highclaude-route24, 0 reasoning tokens/v1/responses, answer/v1/messages/v1/responses,reasoning.effort: highclaude-routereasoningitem, 200 reasoning tokens, answer105/v1/responses, answer/v1/messages/v1/chat/completionsclaude-route1073,finish_reasonstop/v1/responses, answer/v1/messages/v1/messagesclaude-routeroute claude-route forwards an OpenAI login; call it through /v1/chat/completions or /v1/responses/v1/messagessonnet-route**Hallo!**/v1/messagesAcross the 7 upstream calls, each call carried exactly one
authorizationwhose hash matched the caller's, and none carriedx-api-key. Only the/v1/messagescalls carriedanthropic-version: 2023-06-01. For A and A2, the/v1/messagesbody hadthinking: {"type": "adaptive"}andoutput_config: {"effort": "high"}. With adaptive thinking, Claude skipped thinking on A's easy question and thought on A2's.Summary by CodeRabbit