Skip to content
146 changes: 146 additions & 0 deletions docs/LOCAL-HARNESS-BACKENDS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,146 @@
# Local harness backends, for every product at once

How Codex and Claude Code become an option in *all* Redrob products — office,
cowork, browser, design, extension, canvas, query, recall, cad, reblend — from one
implementation in this engine, rather than ten adapters in ten repositories.

Companion documents: `docs/PROVIDER-AUTH.md` (what each vendor permits, and why we
never hold a token) and `docs/LOCAL-ENGINE-API.md` (the `/v1/chat/completions`
contract this rides on). redrob-cowork `docs/CODEX-RUNTIME.md` holds the working
prototype that proved the subprocess mechanics.

## The mistake this corrects

The first adapter was written inside redrob-cowork, behind that app's engine-spawn
hook. It works, and it found real bugs, but its home is wrong: it gives Cowork a
Codex option and gives the other nine products nothing. Ten products would mean
ten adapters, ten settings screens, ten credential stories, and ten places for the
subprocess environment bug described below to be re-introduced.

The adapter belongs here instead, because every product already reaches a model
through this engine — and the ones that still call the Console directly are exactly
the ones `/v1/chat/completions` was built to bring in.

## The mechanism already exists

`packages/core/src/config/plugin/local-provider.ts` admits a provider when two
conditions hold: the package is exactly `@ai-sdk/openai-compatible` (the trusted,
pinned one the Console provider itself uses), and the URL is on this machine or
network. That door was opened for Ollama, LM Studio, llama.cpp and vLLM, whose only
shared trait is that they **speak OpenAI-compatible chat-completions over a local
address**.

A harness can meet the same bar. Put a small local server in front of `codex exec`
or `claude -p` that speaks chat-completions, and it is admissible through a path
this engine already trusts — no new provider-loading machinery, no widening of the
arbitrary-npm-package refusal, no second credential store.

```
codex exec / claude -p subprocess; holds its OWN credential
▲
harness shim local, OpenAI-compatible, 127.0.0.1
▲
redrob-code engine registered via the existing local-provider path
▲
POST /v1/chat/completions one receiving route
▲
office · cowork · browser · design · extension · canvas · query · recall · cad · reblend
```

Models then select a backend by id — `codex/<model>`, `claude-code/<model>` — which
is the same way a local model is already selected, and needs no new request field.

### Why a shim here, having argued against one in Cowork

`docs/CODEX-RUNTIME.md` recommends *against* a protocol shim in Cowork and this
document recommends one. That is not an inconsistency, it is the surface being
different, and the difference is the whole argument.

A Cowork shim would have to emulate the **OpenCode session API** — sessions, the SSE
event shape, permissions — which is large, ours, and still changing. A shim that
falls subtly behind produces bugs that look like model bugs. A shim here emulates
**chat-completions**, which is small, published, frozen, and not ours to change. The
first is a maintenance liability; the second is an adapter against a stable
contract.

## Two tiers, and most products only need the first

This is the part that cannot be papered over: **Codex and Claude Code are agent
harnesses, not completion endpoints.** They run their own loop with their own
tools. So there are two levels of support, not one.

**Chat tier — all ten products.** Prompt in, text out. The harness's loop runs but
its tools are constrained to nothing the caller did not ask for. This covers every
product's ordinary AI use: rewrite this paragraph, summarise this sheet, answer
this question. It maps cleanly onto `/v1/chat/completions` and needs no per-product
code.

**Agent tier — cowork, code, cad.** The harness runs a real agent loop against a
workspace. chat-completions cannot express this, because that contract's rule is
"tools present → the caller owns them", and here the *runtime* owns the loop. This
needs ACP or a session surface, and it is separate work. Do not let it block the
chat tier.

A safety point that belongs in the chat tier and is easy to miss: a harness given a
writable sandbox will happily read and edit the user's filesystem. A user asking
Office to reword a sentence has not consented to an agent walking their disk. So
the shim pins `--sandbox read-only` with approvals set to never, and network access
off, unless a caller is on the agent tier and asked for more. The prototype already
defaults this way; the shim must not relax it for convenience.

The tool-ownership split surveyed earlier decides which products need more than the
chat tier:

- **No tool protocol** — query, recall, reblend. Chat tier is the whole story.
- **Host-owned tools** — office, browser, canvas. Chat tier works today. Giving the
harness their tools needs MCP servers (`~/.codex/config.toml`, or per-invocation
`--config mcp_servers.…`) or Codex app-server's `dynamicTools`, which lets the
tool stay in the host process. `dynamicTools` is the better fit and is labelled
experimental by OpenAI, so it is not a foundation to build on yet.
- **Engine-owned tools** — cowork, cad, extension, design. Under a harness backend
the engine's own tools are not in the loop at all. That is the agent tier's
problem to solve.

## What the products have to do

Almost nothing, which is the point.

Nothing at all to *work*: a product that names a model and calls the engine gets
the new backends when the engine gets them.

One thing to be *usable*: somewhere to turn it on, and three states rather than
one — runtime not installed, installed but not signed in, ready. Collapsing those
into "unavailable" strands the user, because the remedy differs and we are not
allowed to offer the vendor's login ourselves. The remedy we may show is "run
`codex login`" or "run `claude`".

Credentials need no work anywhere. The harness holds its own; the engine holds none
for it. Cowork already demonstrates the pattern for the BYOK case — it stores no
provider key and treats the engine's `auth.json` as the single source of truth.

## Sequencing

1. **Land `/v1/chat/completions`.** It is the receiving route for everything above
and it is not merged: the handler exists on `feature/v1-chat-completions`, and
`test:httpapi` fails without an `httpapi-exercise` scenario. Also needs the
`chatCompletions` capability flag, the SSE OpenAPI patch, and a route test.
Nothing here can ship before it.
2. **Build the shim with both backends.** One local chat-completions server; two
normalizers behind it. The prototype's split — event folding separated from the
subprocess — is what makes the second runtime a second normalizer rather than a
second architecture. Reuse it rather than re-deriving it.
3. **Register through the local-provider path**, and surface the backends in
`/v1/providers` with their three states so a product can render settings
without hardcoding a list.
4. **Verify against real binaries.** Neither runtime is installed on the build
host, so the live path — a real subscription actually paying for a turn — is
unproven until someone runs it on a machine with `codex` and `claude` signed in.
5. **Then the agent tier**, for cowork, code and cad, over ACP.

Two things to carry forward rather than discover later. `claude -p`'s
`stream-json` event schema has not been checked against the real binary — only the
flags are confirmed — so step 2 starts by reading it, not by assuming it mirrors
Codex's JSONL. And Anthropic has announced, then paused, a change that moves
third-party subscription usage onto a capped monthly credit; it currently still
draws from the subscription, but the trajectory is known, so the Claude backend
should surface usage state rather than assume it is free.
189 changes: 189 additions & 0 deletions docs/PROVIDER-AUTH.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,189 @@
# Provider authentication

How a user connects their own model access to this engine, and therefore to every
product built on it. One credential store, read by all of them, so a user sets up
once rather than once per app.

The goal this answers is "let people use the Claude or ChatGPT access they already
pay for". Part of that is available and part of it is not, and the split is not
technical — it is what each vendor's terms permit a third-party application to do.
So the shapes are enumerated first, then the design.

Not legal advice. Every claim below links the document it came from, checked
2026-09-22; re-check before shipping, because three of these pages changed in the
first half of this year.

## What each vendor actually allows a third-party app

| vendor | "sign in" for a third-party app | user's own API key | user's consumer subscription |
| --- | --- | --- | --- |
| Anthropic | **No** — expressly forbidden | Yes, expressly | **No** — prohibited, enforced with account bans |
| OpenAI | No self-serve registration; granted case by case | Yes | Only by embedding OpenAI's own Codex runtime |
| Google Gemini | OAuth exists but bills *your* Cloud project | Yes — the named supported path | **No** — prohibited, enforced with bans |
| GitHub Copilot | **Yes** — the Copilot SDK, officially | Yes | Yes, billed to the user's own subscription |
| OpenRouter | **Yes** — OAuth PKCE | Yes | n/a (it is BYOK by design) |
| Azure OpenAI | **Yes** — Microsoft Entra ID | Yes | n/a (user's own Azure resource) |
| Amazon Bedrock | No (SigV4 / Bedrock API keys) | Yes — the user's own AWS account | n/a |
| AWS Kiro | **No** | Kiro API key, subscriber-only | Only by driving the user's own installed CLI |

Sources, and the sentences that decide it:

- Anthropic, [Claude Code legal and compliance](https://code.claude.com/docs/en/legal-and-compliance):
"Anthropic does not permit third-party developers to offer Claude.ai login into
their own applications, or to route requests through Free, Pro, or Max plan
credentials on behalf of their users." And developers "should use API key
authentication through Claude Console or a supported cloud provider." Enforced:
accounts were banned for spoofing the Claude Code harness, and our own upstream
(OpenCode) removed Claude subscription AND Claude API key support on 2026-02-19
citing Anthropic legal requests ([The Register,
2026-02-20](https://www.theregister.com/software/2026/02/20/anthropic-clarifies-ban-on-third-party-tool-access-to-claude/5014546)).
The only sanctioned subscription route is shipping Claude Code itself,
unmodified, with its auth methods intact.
- OpenAI, [Codex authentication](https://developers.openai.com/codex/auth/):
"Sign in with ChatGPT" is documented for OpenAI's own surfaces only. Third
parties that have it — Zed, OpenClaw — get there by wrapping OpenAI's own Codex
runtime ([Zed](https://zed.dev/blog/chatgpt-subscription-in-zed)). Treat that as
a revocable product decision, not an entitlement: no terms clause grants it.
- Google, [Gemini CLI FAQ](https://github.com/google-gemini/gemini-cli/blob/main/docs/resources/faq.md):
"the supported and secure method is to use a Vertex AI or Google AI Studio API
key", and piggybacking Gemini CLI's OAuth "may be grounds for immediate
suspension or termination". Enforced in
[this thread](https://github.com/google-gemini/gemini-cli/discussions/20632).
Note the Gemini API terms also say the API is "not for consumer use" and
require paid services for EEA/UK/CH users.
- GitHub, [Copilot SDK OAuth setup](https://docs.github.com/en/copilot/how-tos/copilot-sdk/setup/github-oauth):
"Copilot requests are made on behalf of each authenticated user, using their
Copilot subscription… Your app never handles model API keys." A device flow is
documented for exactly our case, "Desktop applications where users interact
directly".
- OpenRouter, [OAuth PKCE](https://openrouter.ai/docs/guides/overview/auth/oauth):
send the user to `/auth` with a `code_challenge`, exchange the code for a
**user-controlled API key**. Loopback callback on any port, plus a headless
paste mode.
- Azure, [Entra ID auth](https://learn.microsoft.com/azure/ai-services/openai/how-to/managed-identity)
with the [device authorization grant](https://learn.microsoft.com/en-us/entra/identity-platform/v2-oauth2-device-code).
- Kiro, [Authentication](https://kiro.dev/docs/getting-started/authentication/):
subscriber `ksk_` API keys exist; the only documented embed path is driving the
user's own CLI over [ACP](https://kiro.dev/docs/cli/acp/).

### What this means for the product ask

"Sign in with Claude" and "Sign in with ChatGPT", as buttons in our own apps
spending the user's consumer subscription, are not available. Building them means
impersonating a first-party client, and the vendors ban accounts for it — the cost
lands on our users, not on us.

What IS available, and covers most of the intent:

1. **BYOK for every provider.** Anthropic and Google both name this as the
supported path for third-party tools. A user with a Claude API key or a Gemini
key connects in one step.
2. **Three real sign-in buttons**: GitHub Copilot, OpenRouter, Azure OpenAI. The
first two are the interesting ones — Copilot spends the user's own Copilot
subscription with GitHub's blessing, and OpenRouter fronts Claude and GPT
models behind an account login, which is the closest legitimate thing to what
was asked for.
3. **Optionally, unmodified first-party binaries.** Shipping Claude Code or the
Codex CLI as-is and letting it authenticate itself is sanctioned by both
vendors. It is a different product shape — their harness, not ours — so it is
recorded here as available rather than recommended.

## Design

### One store, in the engine

Credentials live where they already live: `Global.Path.data/auth.json`, through
`packages/redrob/src/auth`, whose `Info` union is already `Oauth | Api |
WellKnown`. `packages/core/src/console-key.ts` reads env → credential store →
that file, and `Integration` resolves a connection into a `Credential.Value` for
the provider layer.

Nothing new is invented, because the point is that products stop having stores of
their own. Today Office keeps keys in `userData/ai-settings.json`, Design in the
macOS keychain, the extension in `chrome.storage.local`, Query and Recall in
separate keyring services — five stores that never read each other, which is why
a user logs in again in every app.

A product reads the engine's store instead. It keeps its own as a cache if it
wants, but the engine's file is the source of truth, and the engine is what makes
the call.

### Products never hold a key

`POST /v1/chat/completions` (see `docs/LOCAL-ENGINE-API.md`) already carries the
engine's credential outward and takes none from the caller. That is the whole
mechanism: a product names a `model`, the engine resolves the provider and its
credential. Adding a provider is then a change in one place, and every downstream
app gets it without shipping a release.

This also removes a class of bug rather than moving it: a product that holds no
key cannot leak one, log one, or sync one.

### Connecting a provider

`redrob providers login` already implements both shapes — a generic OAuth flow
with `authorize()` plus `auto` and `code` callbacks, and an API-key path
(`packages/redrob/src/cli/cmd/providers.ts`). What is missing is not the flow but
its exposure: a product cannot drive it today.

Add, on the v2 surface beside the completions route:

- `GET /v1/providers` — what this engine can use, each entry declaring which auth
methods it accepts (`oauth`, `api-key`) and whether a credential is present.
This is what lets an app render a settings page without hardcoding a vendor
list that goes stale.
- `POST /v1/providers/:id/login` — starts a flow. For OAuth it returns the
authorization URL and an opaque attempt id; for an API key it accepts the key.
- `POST /v1/providers/:id/login/:attempt` — completes an OAuth attempt with the
authorization code, or reports that the loopback callback already completed it.
- `DELETE /v1/providers/:id/credential` — disconnect.

Three rules on those routes:

1. **The key never comes back out.** A response says a credential is present and
names it; it never returns the secret. A product that cannot read the key
cannot mishandle it, and this is also what keeps the loopback API from becoming
a credential-exfiltration endpoint if something else on the machine reaches it.
2. **The OAuth code stays in the engine.** The product gets the URL to open and an
opaque id, nothing else — the same split `apps/shell/src/main/redrob-connect.ts`
already uses between Electron's main process and its renderer.
3. **No caller-supplied base URL.** Keep this ban. It is not only policy: the
config layer admits a new openai-compatible provider only at a local address
(`isLocalEndpoint`), and honouring an arbitrary URL from a request would
forward the engine's own credential to whatever host the caller named. A local
model is selected by its `provider/model` id.

### Which providers get a login button

Ship BYOK for all of them. Add OAuth only where the vendor documents it for third
parties: GitHub Copilot, OpenRouter, Azure OpenAI. Anthropic, Google, OpenAI and
Bedrock get an API-key field and no button.

The UI should say which is which. A user who expects "sign in with Claude" and
finds a key field deserves the reason in one line — that Anthropic requires an API
key for third-party tools — rather than being left to assume the feature is
missing.

### Order of work

1. `GET /v1/providers` and the API-key path. This alone gives Office, Cowork and
Design one shared BYOK setup, and needs no vendor negotiation.
2. OpenRouter OAuth. Smallest real sign-in button and the one that reaches Claude
and GPT models legitimately.
3. GitHub Copilot via the Copilot SDK. A user's existing Copilot subscription,
sanctioned, with a documented device flow for desktop.
4. Azure OpenAI via Entra device code.

## What must not be built

- Reusing Claude Code's or the Codex CLI's OAuth client id to spend a consumer
subscription from our own harness. Anthropic and Google both ban accounts for
it; our upstream removed the code under legal pressure.
- Collecting, storing or proxying Claude.ai session tokens. Named explicitly in
Anthropic's compliance page.
- Paying for or intermediating another vendor's usage on a user's behalf. Also
named there, and it is what "just put our key in it" would amount to.

A compliance check worth keeping: this repository currently contains no Anthropic
OAuth client id and no `claude.ai` endpoint. The only hardcoded vendor OAuth is
xAI's (`packages/redrob/src/plugin/xai.ts`). Keep it that way.
Loading
Loading