diff --git a/crates/switchyard-server/README.md b/crates/switchyard-server/README.md index 152dbfe02..8bb64b91f 100644 --- a/crates/switchyard-server/README.md +++ b/crates/switchyard-server/README.md @@ -176,6 +176,88 @@ To add instructions for a target, set `system_prompt` on its `[targets.]` Switchyard prepends that text when the selected target serves a completion and retains the caller's instructions. Omit the setting to add no target instructions. +## Codex "Approve for me" + +With "Approve for me" on, Codex asks a reviewer model to check each action that +needs approval. Codex sends each review request to its own model provider, so +when Codex points at Switchyard, the review goes to Switchyard. Unless Codex is +logged in with an OpenAI API key, the request names the model ID +`codex-auto-review`. If no route has that `id`, the server returns HTTP 404 +`model_not_found`. Codex treats a failed review as a denial, so it declines the +action. + +You need the route if you use Codex through Switchyard with "Approve for me" on +in the Default or Read Only mode. This includes Codex logged in with ChatGPT. +Full Access mode never asks for approval, so it never sends review requests. + +Add a route whose `id` is `codex-auto-review`. `examples/run_codex.sh` and +`scripts/config/composite.toml` forward the review unchanged to Codex's own +reviewer model on the ChatGPT backend: + +```toml +[llm_clients.chatgpt_backend] +format = "openai_responses" +base_url = "https://chatgpt.com/backend-api/codex" +forward_auth = true + +[targets.reviewer] +id = "codex-auto-review" +llm_client = "chatgpt_backend" + +[routes.codex_auto_review] +id = "codex-auto-review" +type = "passthrough" +target = "reviewer" +``` + +To review with another model, point `target` at any other target. Use a small, +fast model at low reasoning effort. Codex waits for each review before it runs +the action, so a slow reviewer slows down every step that needs approval. +Codex already asks for `low` effort in each review request, so don't set a +higher `reasoning_effort` on the reviewer target. A larger model costs more and +takes longer per review, but OpenAI found that stronger models catch risky +actions more reliably +([Auto-review](https://alignment.openai.com/auto-review/)). The model must also +follow a JSON output schema, because Codex reads the reviewer's final message as +a JSON verdict. + +When Codex is logged in with an OpenAI API key, it sends review requests to +`gpt-5.6-luna` instead. For that login, give the reviewer route +`id = "gpt-5.6-luna"`. That route then receives every Codex request for the +`gpt-5.6-luna` model ID, not only review requests. + +Codex can send reviews to another model ID only through a full replacement model +catalog: set the `model_catalog_json` config key, and set +`auto_review_model_override` on the session model's entry. That catalog must +list every model Codex uses and changes with Codex releases, so a Switchyard +route is simpler. + +To run every action without a review, use Codex settings instead of a +Switchyard route. With `approval_policy = "never"` (`-a never`), Codex never +asks and sends no review requests. Commands still run inside the sandbox, and +anything the sandbox blocks fails back to the model: + +```toml +approval_policy = "never" +sandbox_mode = "workspace-write" + +[sandbox_workspace_write] +network_access = true # only if commands need the network +``` + +`--dangerously-bypass-approvals-and-sandbox` (`--yolo`, the Full Access preset) +runs everything with no sandbox and no approvals. To skip review only for some +commands, add an experimental `prefix_rule(pattern = [...], decision = "allow")` +rule to `~/.codex/rules/default.rules`. See Codex's +[Agent approvals & security](https://learn.chatgpt.com/docs/agent-approvals-security) +page. + +The installer, `scripts/linux/install.sh`, does not overwrite an existing +`~/.switchyard/composite.toml`. +If you installed Switchyard before this route was added, add the +`[targets.reviewer]` and `[routes.codex_auto_review]` blocks to that file by +hand, then restart the server. + ## Endpoints | Method | Path | Purpose | diff --git a/dev-server/README b/dev-server/README index d6a0f0922..a6a501c59 100644 --- a/dev-server/README +++ b/dev-server/README @@ -8,6 +8,7 @@ It has the following endpoints connected to Inference Hub: * `switchyard/random` -> oss-20b (above) and nvidia_dynamo/deepseek-ai/deepseek-v4-flash-nvfp4-elb * `switchyard/stage` -> nvidia/zai-org/glm-5.2 and aws/anthropic/bedrock-claude-opus-4-8 * `switchyard/classifier` -> same models as stage +* `codex-auto-review` -> same model as `switchyard/passthrough`, for Codex "Approve for me" reviews If you have an NVIDIA UNIX account and Silverfort I think you can ssh to it, ping for details. There's a `/home/README` with the infos once you're in. diff --git a/dev-server/config.toml b/dev-server/config.toml index 7ff4f10ed..2ce5a41c5 100644 --- a/dev-server/config.toml +++ b/dev-server/config.toml @@ -74,3 +74,16 @@ max_reviews = 3 gate_stall_turns = 30 gate_min_tool_results = 3 +# Codex "Approve for me" reviews + +# With "Approve for me" on, Codex sends each review request to the model +# codex-auto-review, unless Codex is logged in with an OpenAI API key. Reviews +# go to the efficient model because Codex waits for each one before it acts. +# Codex asks for low reasoning effort in every review, and extra_body only adds +# an effort when the request has none, so reviews run at low effort. + +[routes.codex_auto_review] +id = "codex-auto-review" +type = "passthrough" +target = "efficient" + diff --git a/examples/run_codex.sh b/examples/run_codex.sh index d019301a2..feae68450 100644 --- a/examples/run_codex.sh +++ b/examples/run_codex.sh @@ -46,6 +46,20 @@ classify_trigger = "user_turn" capable_target = "capable" efficient_target = "efficient" confidence_threshold = 0.5 + +# With "Approve for me" on, Codex sends each review request to the model +# codex-auto-review, unless Codex is logged in with an OpenAI API key. This +# route forwards those requests unchanged to codex-auto-review on the ChatGPT +# backend. To review with another model, point `target` at a small, fast +# model; Codex waits for each review and asks for low reasoning effort. +[targets.reviewer] +id = "codex-auto-review" +llm_client = "chatgpt_backend" + +[routes.codex_auto_review] +id = "codex-auto-review" +type = "passthrough" +target = "reviewer" EOF ./target/release/switchyard-server --config /tmp/composite.toml --dry-run diff --git a/scripts/config/composite.toml b/scripts/config/composite.toml index 2e0ba560d..4abef0e9b 100644 --- a/scripts/config/composite.toml +++ b/scripts/config/composite.toml @@ -38,3 +38,17 @@ classify_trigger = "user_turn" capable_target = "capable" efficient_target = "efficient" confidence_threshold = 0.5 + +# With "Approve for me" on, Codex sends each review request to the model +# codex-auto-review, unless Codex is logged in with an OpenAI API key. This +# route forwards those requests unchanged to codex-auto-review on the ChatGPT +# backend. To review with another model, point `target` at a small, fast +# model; Codex waits for each review and asks for low reasoning effort. +[targets.reviewer] +id = "codex-auto-review" +llm_client = "chatgpt_backend" + +[routes.codex_auto_review] +id = "codex-auto-review" +type = "passthrough" +target = "reviewer"