diff --git a/.abcd/development/brief/04-surfaces/34-build.md b/.abcd/development/brief/04-surfaces/34-build.md index 2c513715b..9734c7adf 100644 --- a/.abcd/development/brief/04-surfaces/34-build.md +++ b/.abcd/development/brief/04-surfaces/34-build.md @@ -206,15 +206,19 @@ repository abcd manages has one, so a run is managed-only by construction. Each run directory is created one level at a time and proved real, the state file is replaced atomically inside an `os.Root`, and the reader decodes strictly, refusing an unknown field, a schema version it does not know, or a file stored -under a run id it does not name. The state is schema version 7. Version 7 +under a run id it does not name. The state is schema version 8. Version 8 +added the runner's record (itd-2609201916056194): the run's `fallbacks`, one +receipt per role a routed runner did not run, and the `route` a verified receipt +or a validator's recorded return names when a runner ran its agent. Version 7 added the landing (a lane's `landing`, the implementers' `receipts` it verified with the model each runner reported, and the captures its receipts declared fixed, `resolves`) and the run's captured `transcripts`. Version 6 added the fix-round cap (ruling DR1): the pace's `fix_rounds` and a lane's `hand_back`. Version 5 added the validate stage's record (a lane's `validation`). Each earlier version is the next one's strict subset, read as a -run that predates the addition (a version-5 run runs on the bundled cap) and -written back at version 7 by its next mutation; an earlier version carrying what +run that predates the addition (a version-5 run runs on the bundled cap, a +version-7 run is one the host ran every agent of) and +written back at version 8 by its next mutation; an earlier version carrying what only a later one writes is refused. Version 4 renamed the lane's stage (BU1, iss-2609291313276243): a lane's and a record line's `step` became `stage`, so "step" names only the spec's steps (`spec_step`, @@ -289,9 +293,41 @@ reported complete and closes no window. A stage whose body this build does not carry is refused naming the stage, the lane and the spec piece that delivers it, and the run is unchanged, ready to resume in a build that carries it. This build carries every stage of the -sequence. The process driver (piece 3) is the same loop -called by a process instead of a host, starting the named agent through the -runner and handing its receipt back. +sequence. + +**The runner** (piece 3, the process driver's loop half, and itd-2609201916056194). +A role's route is `roles..runner` in the layered configuration, the +repository's or the machine's: `host`, the default, or a runner the machine +enables under `runner.` (`claude`, `opencode`), with an optional model +route admitted against its provider's allowlist. The build verb reads the +configuration before it creates a run, and the step verb before each stage; a +fault, a model route off the allowlist included, is refused at the `runner` +stage before anything is created or launched, and its diagnostics (a role no +agent answers to) go to stderr. When a stage hands the lane to a role that is +routed to a runner, the step verb starts the runner itself, outside the run's +lock, in the lane's worktree, with the brief and the receipt path the host would +be handed, the claude runner with the role's tools granted without a prompt and +opencode under the permissions its own configuration sets; the runner's +transcript lands in abcd's history store, keyed on the repository's root +commit, and its receipt is handed back through the same receipt verb and +verified by the stage's own verifier, so a verified one completes the stage in +the same call. The verified receipt, or the validator's recorded return, names +the route that ran it (asked, ran, the model the runner reported), which is the +only field a runner-run review's record differs in from a host-run one's; the +record's receipt or verdict line names the runner that ran it. A runner that is absent, refuses, fails, runs past +its time, answers unparsably or writes a receipt the verifier refuses leaves the +lane awaiting: the call records one fallback receipt (the role, the runner asked +for, the reason and the route that runs it) in the state and the record, and +hands the host the await as the host-driven step does, naming the fallback. A +role left unset is the host's, and the call is the host-driven step byte for +byte. A step that re-tells an await starts nothing. The claude runner runs in +print mode with the bare flag, so the repository's hooks, plugins and +configured servers do not run; opencode runs in run mode with `--pure`. An +interrupt or a termination kills the runner's process group. The status and +record verbs count the fallbacks per runner and per role. A host session +always drives this: the no-host path, where a host-routed role goes to the +machine's `runner.fallback_host`, is the process driver's reversal of the +host-delegated boundary, and waits on the ADR decision 6 of the intent owes. ## The lane @@ -488,8 +524,10 @@ until the last step. verb reads a run's state back as its record: every lane with its spec step, branch and heads, the implementers' receipts the loop verified with the model each runner reported (as reported; the binary cannot verify it), every verdict -the loop recorded, round by round, the captures the lane fixed, its pull request -and what its landing did, the run's pending steps, the transcripts captured and +the loop recorded, round by round, the route that ran a receipt's or a return's +agent when a runner ran it, the captures the lane fixed, its pull request +and what its landing did, the run's pending steps, every fallback with the +count per runner and per role, the transcripts captured and the record's lines, in text and JSON. On a complete run, the record verb captures each transcript it is named into the history store as the history verb's capture of one path does, one capture per path, and records it in the @@ -519,6 +557,8 @@ the remedy as fields. - The shared run state and the claim the peers check reads: [`27-implement.md`](27-implement.md). - The pick: itd-2609211116005482 and its design record, spc-2609212015048113. +- The runner a routed role goes through: itd-2609201916056194 and its design + record, spc-2609221533057881 (`internal/core/runner`). - The plugin surface: `commands/build.md`. diff --git a/.abcd/development/intents/planned/itd-2609201916056194-abcd-runs-a-delegated-agent-through-a-command-line-model-run.md b/.abcd/development/intents/shipped/itd-2609201916056194-abcd-runs-a-delegated-agent-through-a-command-line-model-run.md similarity index 90% rename from .abcd/development/intents/planned/itd-2609201916056194-abcd-runs-a-delegated-agent-through-a-command-line-model-run.md rename to .abcd/development/intents/shipped/itd-2609201916056194-abcd-runs-a-delegated-agent-through-a-command-line-model-run.md index f298a5d9b..284838b48 100644 --- a/.abcd/development/intents/planned/itd-2609201916056194-abcd-runs-a-delegated-agent-through-a-command-line-model-run.md +++ b/.abcd/development/intents/shipped/itd-2609201916056194-abcd-runs-a-delegated-agent-through-a-command-line-model-run.md @@ -80,7 +80,11 @@ _None open._ ## Audit Notes -_Empty. Populated by intent-auditor when intent moves to shipped/._ +- 2026-09-30 — Closed on the fake-harness tests under the person's ruling RN1 (2026-09-30: "(a) SAME RULE, FINISH ON FAKES: the runner may close on the fake-harness tests with the three live checks (iss-2609301519558538) recorded as owed."). The acceptance criteria are proven against fake harnesses on a pinned PATH (the opencode review-route criterion structurally, and the no-host fallback path in the core only, since its surface waits on the ADR decision 6 owes); no real claude or opencode binary has run a role. The three live checks are owed and not yet run: iss-2609301519558538 tracks them and stays open. + + +Fidelity review OWED (receipt rcp-e5eee57c3e2a). + ## Grounds diff --git a/.abcd/development/specs/open/spc-2609221533057881-abcd-runs-a-delegated-agent-through-a-command-line-model-run.md b/.abcd/development/specs/closed/spc-2609221533057881-abcd-runs-a-delegated-agent-through-a-command-line-model-run.md similarity index 53% rename from .abcd/development/specs/open/spc-2609221533057881-abcd-runs-a-delegated-agent-through-a-command-line-model-run.md rename to .abcd/development/specs/closed/spc-2609221533057881-abcd-runs-a-delegated-agent-through-a-command-line-model-run.md index 10aab799e..9573cacb8 100644 --- a/.abcd/development/specs/open/spc-2609221533057881-abcd-runs-a-delegated-agent-through-a-command-line-model-run.md +++ b/.abcd/development/specs/closed/spc-2609221533057881-abcd-runs-a-delegated-agent-through-a-command-line-model-run.md @@ -45,3 +45,36 @@ The runner interface is the same shape the validator stage already defines for r | 6 bare flag | scope 2 | | 7 receipts differ only in route | scope 1, 2 | | 8 security review | scope 6 | + +## Progress + +The spec stays open: the live proof is owed (iss-2609301519558538) and no +ruling yet lets the intent close on fake harnesses. + +- **Landed: scopes 1 to 5 in the core** (`internal/core/runner`): the runner + interface and the one dispatcher, the claude adapter (print mode, `--bare`, + stream-json, `--permission-mode dontAsk`, the role's tools allowed) and the + opencode adapter (run mode, JSON events, `--pure`), the route through the + layered resolver, the fallback with one receipt writer and `Tally`, and the + allowlist admitted at the read and again before a launch; a model a + provider lists is admitted by its allowlist alone and refused only by a + configured `oracle.denylist` entry (adr-2609300107513982). +- **Landed: the loop wiring** (spc-2609202134338445 piece 3): `implement step` + drives through `loop.Drive`, which starts a routed role through the + dispatcher with the brief and receipt path the host would get, validates + its answer with the stage's own receipt verifier, stores its transcript in + abcd's history store, and stamps the verified receipt or recorded return + with the route that ran it (criteria 1 and 7, structurally); an unset role + leaves the step and the state byte-identical (criterion 2); a fallback is + recorded in the state's `fallbacks` and the record (criterion 3) and counted + per runner and per role by `implement status` and `implement record` + (criterion 4); `build` and `step` refuse a runner configuration fault, a + model off its allowlist included, before anything is created or launched + (criterion 5). The bare flag is asserted in the launch (criterion 6). +- **Not built: the no-host path at the surface.** `runner.fallback_host` + works in the core, but a surface that runs without a host session is the + process driver's reversal of the host-delegated boundary, which waits on + itd-2609201916151817's decision-6 ADR. +- **Owed: the live proof** of criterion 7 and the phase-1 unverified points + (iss-2609301519558538), and **criterion 8**, the security review, which + the lane's report lists point by point for the reviewer. diff --git a/.abcd/development/specs/open/spc-2609202134338445-one-verb-takes-a-single-intent-from-ready-to-delivered-witho.md b/.abcd/development/specs/open/spc-2609202134338445-one-verb-takes-a-single-intent-from-ready-to-delivered-witho.md index 77cb2f5c8..eb58ad93d 100644 --- a/.abcd/development/specs/open/spc-2609202134338445-one-verb-takes-a-single-intent-from-ready-to-delivered-witho.md +++ b/.abcd/development/specs/open/spc-2609202134338445-one-verb-takes-a-single-intent-from-ready-to-delivered-witho.md @@ -165,8 +165,15 @@ spec stays open until the last lane closes it. every commit on the lane's branch past its base, a passing definition of done's output and the report, each inside the lane's directory, or a refusal naming every gap. -- **Seam left, not built: piece 3**, the process driver. It waits on the runner - intent (itd-2609201916056194) and calls the same `Advance` and `Receipt`. +- **Landed in part (lane runner2): piece 3**, the process driver. `implement + step` drives through `loop.Drive`: when a stage hands the lane to a role + that `roles..runner` routes to a command-line runner + (itd-2609201916056194), the loop starts the agent itself through the runner + and hands its receipt back through `Receipt`, and the run record names the + runner that ran it (criterion 9); a role left on the host is handed to the + host exactly as before. Not built: the loop driving itself with no host + session, which is decision 6's reversal of the host-delegated boundary and + waits on the ADR that decision owes. - **Landed (lane fidelityOnce): piece 8**, the validators with the itd-58 verdict invariant. The validate stage hands the lane's head to a fresh ruthless reviewer and a fresh security reviewer, one at a time, and on the @@ -214,5 +221,5 @@ spec stays open until the last lane closes it. receipt carries `handback: {kind, reason, home}`, which the loop reads at the receipt, before the validators, discarding the lane's worktree and branch and ending the lane handed back. -- **Remaining: 11** (`--auto-plan` with its ADR), and piece 3 (the process - driver, on the runner). `--auto-plan` is not a flag yet. +- **Remaining: 11** (`--auto-plan` with its ADR), and piece 3's no-host + driving, behind the same ADR. `--auto-plan` is not a flag yet. diff --git a/.abcd/work/DECISIONS.md b/.abcd/work/DECISIONS.md index 86bd090e7..45e598384 100644 --- a/.abcd/work/DECISIONS.md +++ b/.abcd/work/DECISIONS.md @@ -2623,3 +2623,5 @@ together (the script's header says why there is no escape hatch). - 2026-09-30 — Correcting three points of the entry above after its review (lane fix-drainOwnRule of autonomous run A). The drain-rule offer of `ahoy install` is asked only of a person at a terminal, the itd-131 precedent the git identity question set, rather than behind a named opt-in flag: off a terminal neither its category question nor the offer is asked, so a piped answer stream keeps the order it had before the offer existed and a scripted yes never writes the record, and the run reports `drain_rule.offered` under `optional_skipped` naming the terminal as the way to be asked. The terminal gate was chosen over a `--drain-rule` flag because the record decides what an unattended agent may do, which a scripted answer is not a person's yes to, and a flag would hide the offer from the person at a terminal it is for. A checkout holding no release tag (a shallow clone fetches none) marks the anchor unknown rather than reading every deferral as lapsed: every record carrying a deferral is handed back as `deferred`, naming the missing tags and `git fetch --tags`, which keeps the rest of the dry run readable where refusing the whole plan would not. The rule's reader refuses, as malformed, a record that states any frontmatter key twice (not only a `drain_` key) and one whose frontmatter `id` disagrees with its file name, and reads each record through the capped trust-boundary reader, so a record that is a symlink or past the size cap refuses; every refusal of the rule exits 2 on the dry run as on the bare verb. - 2026-09-30 — Six entries above appear twice, verbatim: the five dated 2026-09-29 from "Two itd-111 follow-ups from its fidelity audit" to "Ruling J13", and the 2026-09-30 entry beginning "The 2026-09-29 itd111Follow entry above". Two histories carried them in opposite order relative to the 2026-09-30 BU1/BT1 entry (main below them, the implement-loop lanes above them), so joining them in integration 24b-3 kept main's order and repeated the six after BU1 in the lanes' order, the one merge result the append-only gate admits (every parent's lines kept in their order, DA002; no line beyond what the merge base held plus what each side added, DA003). Each pair is one decision recorded once: the first copy is the record, and the second repeats it (recorded by the integration lane of autonomous run A). - 2026-09-30 — An older site interface-string file keeps building: `abcd site setup` and `abcd site build` add to `site-src/ui.json` each label the allowlist declares and the file does not carry, with abcd's default words, name each on stderr, and change nothing else in it (the product thinker's ruling TG1 of 2026-09-30, relayed verbatim: "(b) ABCD ADDS THE MISSING LABELS: on the next site setup or site build, abcd adds only the missing required labels (with the default words); the project's own wording elsewhere in ui.json is never changed. No failure, no manual step; both intents stay impact: additive."). It is the one exception to "a file the repository owns once it exists is kept", recorded as adr-2609301720596683, which refines adr-47 and leaves decision 2's closed allowlist untouched: a blank declared label and an unknown key are still refused, and the site gate's own render never completes the file. itd-2609212103568351 and itd-2609212103572513 keep `impact: additive` (lane tgLabels of autonomous run A). +- 2026-09-30 — A command-line runner admits a harness reached through a directory, or a binary, that is group-writable only when the group is the system administrator group (gid 0 anywhere, gid 80 `admin` on darwin) and other cannot write it; other-writable stays refused whatever the group, and every other group stays refused (lane runner2Land of autonomous run A, `internal/core/runner/proc.go` `adminGroupWritableOnly`). Reason: the runner2 re-verification (reverify-runner2) found that on a Homebrew Mac `/opt/homebrew/bin` is `drwxrwsr-x` group admin, so a harness installed there was refused with the `chmod go-w` message; members of the administrator group can already act as root, so that write grants them nothing new, and asking a person to strip Homebrew's own directory mode would break Homebrew. +- 2026-09-30 — Narrowing the administrator-group exception in the entry above (lane fix-runnerAdmin of autonomous run A, `internal/core/runner/proc.go` `adminGroupWritableOnly`): a command-line runner admits a harness binary, or a directory it is reached through, that is group-writable (never other-writable) only on darwin and only when the group is gid 80 `admin`; gid 0 is refused on every OS, and gid 80 is refused off darwin. Reason: the entry above granted gid 0 on the premise that members of the administrator group can already act as root, which holds for darwin's admin group (its members may sudo by default) but not for Linux's gid 0 root group nor darwin's gid 0 wheel, whose membership does not by itself let someone act as root (the runner2Land review note). The Homebrew `/opt/homebrew/bin` case the exception exists for is group admin, so it stays admitted. diff --git a/.abcd/work/issues/open/iss-2609301519558538-deferred-owed-to-a-person-the-live-proof-of-the-command-line.md b/.abcd/work/issues/open/iss-2609301519558538-deferred-owed-to-a-person-the-live-proof-of-the-command-line.md new file mode 100644 index 000000000..e728e8870 --- /dev/null +++ b/.abcd/work/issues/open/iss-2609301519558538-deferred-owed-to-a-person-the-live-proof-of-the-command-line.md @@ -0,0 +1,15 @@ +--- +schema_version: 1 +id: "iss-2609301519558538" +slug: "deferred-owed-to-a-person-the-live-proof-of-the-command-line" +severity: "minor" +category: "future-work-seed" +source: "agent-finding" +found_during: "autonomous run A resumed 2026-09-25" +origin: researcher-authored +production_mode: hand-written +found_at: "internal/core/runner" +remedy: "Waits on a ruling: either the product thinker rules the runner may close on fakes with the live proof recorded as owed (as ruling L does for the three paid adapters), and spec close runs with Delivers, or a person runs checks (1) to (3) with their own credential and records each outcome on this record before the close." +--- + +Deferred, owed to a person: the live proof of the command-line runner (itd-2609201916056194). Phase 1 and 2 prove every criterion on fake harness binaries only; no ruling admits a live check of the runner (ruling L names the site setup, the API adapter and the decision adapter; H8 lists the three paid-service checks and itd-6's RepoPrompt run), so the intent cannot close. What a person must run: (1) the intent's first proof, criterion 7 live: a ruthless review of a real lane routed roles.ruthless-reviewer.runner=opencode, from a lane the host drives, its return recorded beside a host-run review's and differing only in the route; (2) the claude runner under --bare with --permission-mode dontAsk and --allowedTools, confirming a nested claude does not refuse under an inherited CLAUDECODE, and noting that the headless page (read 2026-09-30) says bare mode never reads OAuth or the keychain, so the claude runner needs ANTHROPIC_API_KEY, a paid API key rather than the person's subscription; (3) opencode run --format json --pure: the event shape (the CLI page does not document it; the adapter assumes step_start/text/step_finish with a sessionID and a part) and whether run mode prompts for a permission, which the timeout bounds and records as a fallback. diff --git a/.abcd/work/issues/resolved/iss-2609301557141123-the-claude-runner-takes-the-model-its-harness-s-init-event.md b/.abcd/work/issues/resolved/iss-2609301557141123-the-claude-runner-takes-the-model-its-harness-s-init-event.md new file mode 100644 index 000000000..4fe1700e1 --- /dev/null +++ b/.abcd/work/issues/resolved/iss-2609301557141123-the-claude-runner-takes-the-model-its-harness-s-init-event.md @@ -0,0 +1,23 @@ +--- +schema_version: 1 +id: "iss-2609301557141123" +slug: "the-claude-runner-takes-the-model-its-harness-s-init-event" +severity: "minor" +category: "bug" +source: "impl-review" +found_during: "autonomous run 2026-09-23" +origin: researcher-authored +production_mode: hand-written +found_at: "internal/core/runner/claude.go" +remedy: "Bound and shape the reported model as sessionRe bounds a session id (a plain id of at most 128 characters, a bracketed context suffix admitted), and refuse any other as an unparsable answer, so the route falls back and the fallback is recorded rather than the model written into the state." +resolution: "The init event's model is held to modelRe (bounded, plain, a bracketed suffix admitted); any other is an unparsable answer, so the route falls back and is recorded and the model never reaches state.json." +impact: fix +resolved_by: + commit: "70d49494f" +--- + +The claude runner takes the model its harness's init event reports unbounded and unshaped into the answer (internal/core/runner/claude.go parseClaude), and the loop writes it into state.json through the route record; a model string over 4 MiB pushes state.json past its read bound, so every later read refuses and the run is bricked, and control bytes travel into the record verbatim. + +## Grounds + +- pursued: a harness reporting a 5 MiB model or one carrying control bytes falls back with an unparsable reason and the run's state stays readable with no receipt carrying it; a real id such as claude-opus-4-6[1m] refused, or a state.json holding the odd model, would show it wrong diff --git a/.abcd/work/issues/resolved/iss-2609301557191251-the-runner-s-admit-internal-core-runner-proc-go-starts-a.md b/.abcd/work/issues/resolved/iss-2609301557191251-the-runner-s-admit-internal-core-runner-proc-go-starts-a.md new file mode 100644 index 000000000..0addb0a74 --- /dev/null +++ b/.abcd/work/issues/resolved/iss-2609301557191251-the-runner-s-admit-internal-core-runner-proc-go-starts-a.md @@ -0,0 +1,23 @@ +--- +schema_version: 1 +id: "iss-2609301557191251" +slug: "the-runner-s-admit-internal-core-runner-proc-go-starts-a" +severity: "minor" +category: "security" +source: "impl-review" +found_during: "autonomous run 2026-09-23" +origin: researcher-authored +production_mode: hand-written +found_at: "internal/core/runner/proc.go" +remedy: "After EvalSymlinks, stat the resolved binary, its directory and the PATH entry's directory, and refuse (absent, so the route falls back) any whose mode carries a group or other write bit, through the one fsutil test CallersAlone applies, exported as fsutil.WritableByOthers; ownership is not required, since a root-owned system binary is legitimate." +resolution: "admit stats the resolved binary, its directory and the PATH entry's directory and refuses any group or other can write, through fsutil.WritableByOthers, the one test CallersAlone applies." +impact: fix +resolved_by: + commit: "70d49494f" +--- + +The runner's admit (internal/core/runner/proc.go) starts a harness binary resolved on PATH even when the binary, or the directory PATH reaches it through, is writable by group or other: a chmod 777 directory holding claude was admitted, so any account that can write there chooses what the loop runs with the person's credentials. + +## Grounds + +- pursued: a chmod 777 PATH directory holding the fake, or a mode 0772 fake binary, is refused as absent naming group or other and never launched; a launch from either, or a 0755 system harness refused, would show it wrong diff --git a/ACKNOWLEDGEMENTS.md b/ACKNOWLEDGEMENTS.md index bf89b59f8..360192e2c 100644 --- a/ACKNOWLEDGEMENTS.md +++ b/ACKNOWLEDGEMENTS.md @@ -122,6 +122,13 @@ Ideas and methodologies that shaped the design — not code abcd depends on. outright instead of shipped, and the enforcing control sits at the execution layer. +- **Claude Code's print mode (Anthropic)** — the first command-line runner a + delegated role can be routed through (itd-2609201916056194, + `internal/core/runner`): print mode with `--bare`, so a target repository's + hooks, plugins and configured servers do not run, the stream-json event stream + as the transcript, and `--permission-mode dontAsk` with the role's tools + allowed, since no one is there to answer a prompt. + - **Cloudflare Workers and its v4 API (Cloudflare)** — the host the one provider behind `abcd site setup`'s hosting seam targets: an assets-only Worker deployed by the pinned `wrangler-action`, created, routed and @@ -240,6 +247,10 @@ Ideas and methodologies that shaped the design — not code abcd depends on. audience-by-placement ratification (adr-53) and the guide's self-contained-sections rule. +- **opencode (SST)** — the second command-line runner a delegated role can be + routed through (itd-2609201916056194, `internal/core/runner`): run mode with + JSON events and `--pure`, so no external plugin the repository configures + runs. - **OpenAI's Chat Completions API** — the protocol the OpenAI-compatible API adapter speaks (`internal/adapter/openaiapi`): the system and user messages a host's brief is rendered into, the sampling fields a row may set, and the diff --git a/commands/build.md b/commands/build.md index a27075702..bf9261d5b 100644 --- a/commands/build.md +++ b/commands/build.md @@ -196,7 +196,27 @@ returns hand the receipt back: ``` The lane advances only on a receipt that verifies. Running `implement step` while the -lane awaits a receipt re-tells what it awaits and moves nothing. When a lane is +lane awaits a receipt re-tells what it awaits and moves nothing. + +A role can run through a command-line runner instead of an agent you start. +`roles..runner` in the repository's or the machine's `.abcd/config.json` +names `host` (the default) or a runner, `claude` or `opencode`, that the machine +enables under `runner.` in `~/.abcd/config.json` (with an optional +`model` route, `/`, admitted against that provider's +allowlist). `build` reads this configuration before it creates the run, and a +fault, a model route off the allowlist included, is refused at the `runner` +stage with nothing created. When a stage hands the lane to a routed role, +`implement step` starts the runner itself with the brief and the receipt path +you would be handed, in the lane's worktree; its transcript goes to abcd's +history store and its receipt is verified exactly as yours would be, so a +verified one completes the stage in the same call, and the payload's `route` +names the runner that ran it (`asked`, `ran`, `model`). Do not start an agent +for it. When the runner is absent, refuses, fails, runs past its time or writes +a receipt that does not verify, the payload still names `awaiting` as usual and +adds `fallback` (the `role`, the runner `asked` for, the `reason` and the route +that runs it): start the agent yourself as above. Every fallback is recorded; +`implement status` and `implement record` count them per runner and per role. +Tell the user each fallback's reason. When a lane is done, the spec's next pending step opens the next lane, and the run record gets a line naming it, as the start line names the first. diff --git a/commands/implement.md b/commands/implement.md index d1e48e011..350cfb4ef 100644 --- a/commands/implement.md +++ b/commands/implement.md @@ -207,7 +207,8 @@ last, and a fourth reads its record at the end: `status` renders every run (or the one `--run` names): its pace and the layer each number came from, whether it is paused and until when, its lanes, each lane's spec step and next stage, what an awaiting lane waits on, the pending spec -steps and the run record. It writes nothing. +steps, the run's fallbacks from a routed runner to the host (`fallbacks`, and in +the text a count per runner and per role) and the run record. It writes nothing. A spec's **steps** and a lane's **stages** are two things: each spec step lands as one lane, and the loop takes the lane through its stages. `step` performs the @@ -221,6 +222,20 @@ while the lane awaits re-tells the await and moves nothing; a complete run says next lane and the run record names it. A stage that fails leaves the state as it was, so the next call performs it again, and a completed stage is never repeated. +When the stage hands the lane to a role that `roles..runner` routes to a +command-line runner (`claude` or `opencode`, enabled under `runner.` in +`~/.abcd/config.json`), `step` starts it itself, in the lane's worktree, with +the same brief and receipt path; the claude runner runs in print mode with +`--bare`, so the repository's hooks, plugins and configured servers do not run. +The runner's receipt is verified by the stage's own verifier: a verified one +completes the stage in the same call, and the result's `route` names the runner +that ran it. A runner that is absent, refuses, fails, runs past its time or +writes a receipt that does not verify leaves the lane awaiting, and the result +names `awaiting` as with no runner plus `fallback` (the role, the runner asked +for, the reason, the route that runs it); start the agent as for any await. A +`step` that re-tells an await starts no runner. The runner configuration is read +on every `step`; a fault is refused at the `runner` stage before anything runs. + `step` keeps the run's window clock, on the pace the run started with (`/abcd:build`). Once the window's working minutes have elapsed, `step` starts nothing, writes `next_eligible_at` (now plus the run's pause), records the @@ -311,8 +326,11 @@ Every landing step is recorded as it completes, so a killed `step` repeats the move that did not complete and finds what it made rather than making it twice. `record` renders a run's record: each lane with its receipts and the model each -runner reported, every verdict the loop recorded, the captures it fixed, its -pull request and landing, the transcripts captured, and the record's lines. +runner reported, every verdict the loop recorded, the route that ran a receipt's +or a return's agent when a runner ran it, the captures it fixed, its pull +request and landing, every fallback with its count per runner and per role +(`fallbacks`, `fallback_counts`), the transcripts captured, and the record's +lines. Without `--run` it reads the one run in progress, or else the latest run. With `--transcript ` (repeatable) on a complete run it captures each transcript into the history store as `history capture ` does, one capture per path, diff --git a/docs/reference/cli/commands.md b/docs/reference/cli/commands.md index 9d7627c5d..fdb2ec535 100644 --- a/docs/reference/cli/commands.md +++ b/docs/reference/cli/commands.md @@ -273,6 +273,12 @@ further for it, and `abcd implement step` refuses naming the hand-back. The run then moves one step per `abcd implement step`, driven by the host session. +The runner configuration is read before the run is created: roles..runner (host, +the default, or a runner) and the runners this machine enables under runner. in +~/.abcd/config.json, each model route admitted against its provider's allowlist. A fault, +a model route the allowlist does not admit included, is refused at the runner stage and +nothing is created or launched. + An issue id (iss-N, validated by shape) is built as one lane. Its checks are the repository's own drain rule, read as `abcd drain` reads it (the issue is open, nothing open blocks it, its category and severity are ones the rule takes, it carries a remedy a @@ -1777,7 +1783,9 @@ Render a loop run's record and capture its transcripts: Writes only with --trans Render a run's record: every lane with its spec step, branch and head, the implementers' receipts the loop verified with the model each runner reported, every verdict the loop recorded from a validator's return, the captures each lane fixed, its pull request and -what its landing did, the transcripts captured into the history store, and the record's +what its landing did, the route that ran each receipt's or return's agent when a runner +ran it, every fallback from a routed runner to the host with its count per runner and +per role, the transcripts captured into the history store, and the record's lines. Read-only unless --transcript is given. --transcript , repeatable, captures each transcript into the history store as @@ -1855,7 +1863,8 @@ Render the implement loop's runs in this checkout, lane by lane: Writes nothing; Render the runs `abcd build` started in this checkout, or the one --run names: the intent and spec, each lane with its spec step and next stage, what an awaiting lane -waits on, the pending spec steps, and the run record. Read-only: it writes nothing +waits on, the pending spec steps, the fallbacks from a routed runner to the host counted +per runner and per role, and the run record. Read-only: it writes nothing and creates nothing. Exit 2 when --run names no run. **Flags:** @@ -1918,6 +1927,22 @@ A stage whose body this abcd does not carry is refused naming the spec piece tha delivers it, and the run is unchanged. A stage that fails leaves the state as it was, so the next invocation performs it again; a completed stage is never repeated. +A role routed to a command-line runner (roles..runner: claude or opencode, enabled +under runner. in ~/.abcd/config.json) is started by the step itself when the stage +hands the lane out: the runner gets the brief and the receipt path the host would get, +runs in the lane's worktree (claude with the role's tools granted and nothing else asked, +opencode under its own permission configuration), its +transcript is stored in abcd's history store, and its receipt is verified by the stage's +own verifier, so a verified one completes the stage in the same call and the result and +the run record name the route that ran it. The claude runner runs in print mode with +--bare, so the repository's hooks, plugins and configured servers do not run; opencode +runs in run mode with --pure. A runner that is absent, refuses, fails, runs past its time, +or writes a receipt the verifier refuses leaves the lane awaiting and the host is handed +the role as with no runner, and the call records one fallback naming the role, the runner +asked for, the reason and the route that runs it. A role left unset is the host's, and +the call is exactly the host-driven step. A step that re-tells an await starts nothing. +An interrupt kills the runner's process group. + The run's window clock: once the run's working window has elapsed, the call starts nothing, writes next_eligible_at (now plus the run's pause) and exits 0 naming it; an agent already started may still hand back its receipt. Before next_eligible_at the call diff --git a/internal/core/implement/loop/brief_steps_test.go b/internal/core/implement/loop/brief_steps_test.go index be1d2147c..5aebab4f6 100644 --- a/internal/core/implement/loop/brief_steps_test.go +++ b/internal/core/implement/loop/brief_steps_test.go @@ -153,7 +153,7 @@ func TestTheBriefNamesWhatAnEarlierLaneOfTheRunBuilt(t *testing.T) { } continue } - if _, err := Advance(repo.Root(), id, steps, Options{}); err != nil { + if _, err := advance(repo.Root(), id, steps, Options{}); err != nil { t.Fatal(err) } } @@ -190,7 +190,7 @@ func TestABriefWhoseStepTheBaseListsOtherwiseIsRefused(t *testing.T) { repo.Write(specRel, specWithSteps(tc.steps)) repo.Commit("the steps change on the default branch") advanceTo(t, repo, start.RunID, StageBrief) - _, err = Advance(repo.Root(), start.RunID, DefaultStages(), Options{}) + _, err = advance(repo.Root(), start.RunID, DefaultStages(), Options{}) if r := mustRefusal(t, err); r.Stage != string(StageBrief) || !strings.Contains(r.Reason, tc.want) { t.Fatalf("want the brief refused naming %q: %+v", tc.want, r) } diff --git a/internal/core/implement/loop/brief_test.go b/internal/core/implement/loop/brief_test.go index daded5c30..5ccb47080 100644 --- a/internal/core/implement/loop/brief_test.go +++ b/internal/core/implement/loop/brief_test.go @@ -48,7 +48,7 @@ func advanceTo(t *testing.T, repo *gittest.Repo, runID string, want Stage) StepR if i := st.current(); i >= 0 && st.Lanes[i].Stage == want { return res } - if res, err = Advance(repo.Root(), runID, DefaultStages(), Options{}); err != nil { + if res, err = advance(repo.Root(), runID, DefaultStages(), Options{}); err != nil { t.Fatalf("advancing to %s: %v", want, err) } } @@ -68,7 +68,7 @@ func TestTheBriefNamesWhatItWasRenderedFrom(t *testing.T) { t.Fatal(err) } advanceTo(t, repo, start.RunID, StageBrief) - res, err := Advance(repo.Root(), start.RunID, DefaultStages(), Options{}) + res, err := advance(repo.Root(), start.RunID, DefaultStages(), Options{}) if err != nil { t.Fatal(err) } @@ -199,7 +199,7 @@ func TestTheBriefIsRenderedFromTheLaneBase(t *testing.T) { t.Fatal(err) } advanceTo(t, repo, start.RunID, StageBrief) - _, err = Advance(repo.Root(), start.RunID, DefaultStages(), Options{}) + _, err = advance(repo.Root(), start.RunID, DefaultStages(), Options{}) if r := mustRefusal(t, err); r.Stage != string(StageBrief) || !strings.Contains(r.Reason, "AGENTS.md") { t.Fatalf("want the missing conventions named: %+v", r) } @@ -220,7 +220,7 @@ func TestTheBriefIsRenderedFromTheLaneBase(t *testing.T) { t.Fatal(err) } advanceTo(t, repo, start.RunID, StageBrief) - _, err = Advance(repo.Root(), start.RunID, DefaultStages(), Options{}) + _, err = advance(repo.Root(), start.RunID, DefaultStages(), Options{}) if r := mustRefusal(t, err); r.Stage != string(StageBrief) || !strings.Contains(r.Reason, "not planned") || !strings.Contains(r.Reason, "drafts/") { t.Fatalf("want the base's bucket named: %+v", r) } @@ -315,7 +315,7 @@ func TestABriefSourceTheBaseHoldsAsASymlinkIsRefused(t *testing.T) { t.Fatal(err) } advanceTo(t, repo, start.RunID, StageBrief) - _, err = Advance(repo.Root(), start.RunID, DefaultStages(), Options{}) + _, err = advance(repo.Root(), start.RunID, DefaultStages(), Options{}) if r := mustRefusal(t, err); r.Stage != string(StageBrief) || !strings.Contains(r.Reason, "AGENTS.md") { t.Fatalf("want the linked AGENTS.md refused: %+v", r) } diff --git a/internal/core/implement/loop/drain_test.go b/internal/core/implement/loop/drain_test.go index 8dc9d43f1..e3410bd97 100644 --- a/internal/core/implement/loop/drain_test.go +++ b/internal/core/implement/loop/drain_test.go @@ -43,7 +43,7 @@ func handBackLaneOf(t *testing.T, repo *gittest.Repo, runID string, o Options, h if st.Lanes[0].Awaiting != nil { break } - if _, err := Advance(repo.Root(), runID, DefaultStages(), o); err != nil { + if _, err := advance(repo.Root(), runID, DefaultStages(), o); err != nil { t.Fatalf("advancing the drain's lane: %v", err) } } diff --git a/internal/core/implement/loop/drive.go b/internal/core/implement/loop/drive.go new file mode 100644 index 000000000..4099121f2 --- /dev/null +++ b/internal/core/implement/loop/drive.go @@ -0,0 +1,233 @@ +package loop + +// drive.go is the process driver (spc-2609202134338445 piece 3) and the +// runner's loop half (itd-2609201916056194, spc-2609221533057881): the same +// loop the host calls, with the agent a stage awaits started by the loop +// itself through the command-line runner its role is routed to +// (roles..runner), and its receipt handed back through the same +// Receipt a host calls. Nothing else changes: +// +// - a role left on the host (the default) is handed to the host exactly as +// advance hands it, with the same result and the same state, byte for byte +// (criterion 2); +// - a routed role is run through internal/core/runner's Dispatcher, whose +// validator IS the stage's own receipt verifier, run through Receipt under +// the run's lock, so a runner's answer is judged exactly as a host +// sub-agent's is and a verified one advances the lane there and then +// (criterion 1); the verified receipt, or the validator's recorded return, +// names the route that ran it, and nothing else about it differs +// (criterion 7); +// - a runner that is absent, refuses, fails, answers unparsably or writes a +// receipt the verifier refuses leaves the lane awaiting, writes one +// fallback receipt into the run's state (Fallbacks) and the record, and +// hands the role to the host with the reason (criterion 3), which the run +// record counts per runner and per role (criterion 4). +// +// The runner is started outside the run's lock, which is held only for the +// advance before it and the Receipt or the fallback write after it, so a +// status read or another checkout's step is never kept waiting on a model. +// +// A host session always drives this call: the no-host path, where abcd runs +// alone and a host-routed role goes to runner.fallback_host, is the process +// driver's reversal of the host-delegated boundary, which the ADR decision 6 +// of itd-2609201916151817 owes records before either path ships. + +import ( + "context" + "fmt" + "os" + "path/filepath" + "time" + + "github.com/intentdriven/abcd/internal/core/layered" + "github.com/intentdriven/abcd/internal/core/runner" +) + +// StageRunner is the record's stage for a role a runner ran, and StageFallback +// for one a runner did not. +const ( + StageRunner = "runner" + StageFallback = "fallback" +) + +// Runners is what the process driver starts a routed role through. +type Runners struct { + // Config is the runner configuration read when the call began + // (runner.Load); nil drives nothing and leaves every role on the host. + Config *runner.Config + // Transcripts is where each runner's transcript lands: abcd's own store + // (runner.HistoryStore) in production. + Transcripts runner.TranscriptStore + // Timeout bounds one runner's run; runner.DefaultTimeout when zero. + Timeout time.Duration +} + +// roleTools are the tools each role the loop starts is granted without a +// prompt when a runner runs it. A reviewer's are its agent definition's +// (agents/.md, which TestRoleToolsFollowTheAgentDefinitions holds this +// table to) plus Write, since its brief tells it to write its return to a +// path; the implementer's are what a lane's work takes; the auditor's are a +// reviewer's, its verdict written the same way. +var roleTools = map[string][]string{ + RoleImplementer: {"Read", "Edit", "Write", "Bash", "Grep", "Glob"}, + RoleRuthless: {"Read", "Grep", "Glob", "Bash", "Write"}, + RoleSecurity: {"Read", "Grep", "Glob", "Bash", "Write"}, + RoleAuditor: {"Read", "Grep", "Glob", "Bash", "Write"}, +} + +// toolsFor returns the tools a runner grants role; nil for a role the loop +// does not start. +func toolsFor(role string) []string { + t, ok := roleTools[role] + if !ok { + return nil + } + return append([]string(nil), t...) +} + +// Drive performs the next stage as advance does and, when the stage hands the +// lane to an agent whose role is routed to a runner, starts that agent through +// the runner and hands its receipt back through Receipt. A step that re-tells +// an await an earlier call began starts nothing. A role on the host, or a runner +// that did not run it, returns the await for the host to act on, as advance +// returns it; the latter also names the fallback it recorded. +func Drive(ctx context.Context, repoRoot, runID string, steps Stages, o Options, rs Runners) (StepResult, error) { + res, err := advance(repoRoot, runID, steps, o) + if err != nil || res.Awaiting == nil || !res.handed || rs.Config == nil { + // Nothing awaits, or the await is one an earlier call began: the + // host, or the runner that call started, is already on it. + return res, err + } + aw := *res.Awaiting + route := rs.Config.RouteFor(aw.Role) + if route.Runner == runner.Host { + return res, nil + } + st, err := ReadState(repoRoot, runID) + if err != nil { + return res, err + } + var worktree string + for _, l := range st.Lanes { + if l.ID == res.Lane { + worktree = l.Worktree + } + } + req := runner.Request{ + Role: aw.Role, + Brief: absIn(repoRoot, aw.Brief), + Receipt: absIn(repoRoot, aw.Receipt), + Dir: worktree, + Checkout: filepath.Clean(repoRoot), + Tools: toolsFor(aw.Role), + SessionID: fmt.Sprintf("%s-%s-%s-%d", runID, res.Lane, aw.Role, o.now().Unix()), + Timeout: rs.Timeout, + } + var done *StepResult + d := &runner.Dispatcher{ + Config: rs.Config, + HostSession: true, + Transcripts: rs.Transcripts, + Validate: func(ran string, _ runner.Request, ans runner.Answer) error { + r, err := receipt(repoRoot, runID, aw.Receipt, steps, o, + &runner.RouteRecord{Asked: route.Runner, Ran: ran, Model: ans.Model}) + if err != nil { + return err + } + done = &r + return nil + }, + Record: func(fb runner.FallbackReceipt) error { return recordFallback(repoRoot, runID, res.Lane, fb, o) }, + Now: o.Now, + } + out, err := d.Dispatch(ctx, req) + if err != nil { + return res, refuse(StageRunner, "", res.Lane, err.Error(), + "the lane still awaits its receipt: correct what the reason names and step again, or start the agent the step names by hand") + } + if out.Handoff { + res.Fallback = out.Fallback + if fb := out.Fallback; fb != nil { + res.Next = fmt.Sprintf("the %s runner did not run the %s (%s; recorded as a fallback), so the host runs it: %s", + fb.Asked, fb.Role, fb.Reason, res.Next) + } + return res, nil + } + if done == nil { + return res, refuse(StageRunner, "", res.Lane, "the runner ran the "+aw.Role+" but no receipt was handed back", + "the lane still awaits its receipt; step again") + } + r := *done + r.Route = &out.Receipt.Route + return r, nil +} + +// LoadRunners reads the runner configuration (runner.Load) at the roots a lane +// starts from, and refuses in the loop's refusal shape on any fault, a model +// route its provider's allowlist does not admit included, before a run is +// created or a runner launched (itd-2609201916056194 criterion 5). The +// configuration's diagnostics are the caller's to print. +func LoadRunners(r layered.Roots) (*runner.Config, error) { + c, err := runner.Load(r) + if err != nil { + return nil, refuse(StageRunner, "", "", err.Error(), + "correct the runner configuration the reason names; no runner is launched and nothing is written") + } + return c, nil +} + +// absIn is p read against the checkout root when it is relative. +func absIn(repoRoot, p string) string { + if filepath.IsAbs(p) { + return filepath.Clean(p) + } + return filepath.Join(repoRoot, filepath.FromSlash(p)) +} + +// recordFallback appends one fallback receipt to the run's state and a line to +// its record, under the lock. +func recordFallback(repoRoot, runID, laneID string, fb runner.FallbackReceipt, o Options) error { + return mutate(repoRoot, runID, func(_ *os.Root, st *State) (bool, error) { + now := o.now() + fb.At = now + st.Fallbacks = append(st.Fallbacks, fb) + st.Record = append(st.Record, Entry{At: now, Lane: laneID, Stage: StageFallback, + Note: fmt.Sprintf("the %s was routed to %s, which was %s (%s); %s runs it", fb.Role, fb.Asked, fb.Reason, fb.Detail, fb.Ran)}) + st.UpdatedAt = now + return true, nil + }) +} + +// stampRoute names the route that ran the agent on what the verifier recorded +// from its receipt: the implementer's verified receipt, or the validator's +// run whose return it was. +func stampRoute(lane *Lane, verified string, route *runner.RouteRecord) { + for i := len(lane.Receipts) - 1; i >= 0; i-- { + if lane.Receipts[i].Receipt == verified && lane.Receipts[i].Route == nil { + r := *route + lane.Receipts[i].Route = &r + break + } + } + if n := len(lane.Validation); n > 0 { + vs := lane.Validation[n-1].Validators + for i := range vs { + if vs[i].Return == verified && vs[i].Route == nil { + r := *route + vs[i].Route = &r + } + } + } +} + +// routeNote is the record's line for a role a runner ran. +func routeNote(role string, r runner.RouteRecord) string { + model := r.Model + if model == "" { + model = "none reported" + } + if r.Asked == r.Ran { + return fmt.Sprintf("the %s ran on the %s runner (model %s)", role, r.Ran, model) + } + return fmt.Sprintf("the %s, routed to %s, ran on the %s runner (model %s)", role, r.Asked, r.Ran, model) +} diff --git a/internal/core/implement/loop/drive_test.go b/internal/core/implement/loop/drive_test.go new file mode 100644 index 000000000..65b6e2660 --- /dev/null +++ b/internal/core/implement/loop/drive_test.go @@ -0,0 +1,526 @@ +package loop + +// drive_test.go proves the process driver (spc-2609202134338445 piece 3) and +// the runner's loop half (itd-2609201916056194 phase 2) without a model or a +// real harness. The test binary re-executes itself under the name of the +// harness a test puts on PATH (claude, opencode), and TestMain, seeing +// ABCD_LOOP_FAKE_HARNESS, plays that harness. PATH holds the fake's directory +// alone, so no real harness on the machine can be reached. +// +// The events each fake prints follow what the harnesses document, read on +// 2026-09-30. The claude CLI (code.claude.com/docs/en/headless): print mode's +// stream-json output is one JSON object per line, the system/init event +// carries the session metadata including the model, and the last line is a +// result message with the final text and the session; --bare skips hooks, +// plugins, MCP servers and CLAUDE.md, and never reads OAuth credentials or the +// keychain; dontAsk denies every call that would otherwise prompt while the +// --allowedTools entries run. opencode (opencode.ai/docs/cli): `run --format +// json` prints "raw JSON events", --pure (a global flag) runs "without +// external plugins", --dir, --file and --model provider/model; the page does +// NOT document the events' shape, so the step_start/text/step_finish lines +// with a sessionID and a part are the adapter's assumption, and the live +// shape is owed to a person's check. + +import ( + "bufio" + "bytes" + "context" + "encoding/json" + "fmt" + "os" + "path/filepath" + "reflect" + "strings" + "testing" + "time" + + "github.com/intentdriven/abcd/internal/core/layered" + "github.com/intentdriven/abcd/internal/core/runner" +) + +const ( + loopFakeEnv = "ABCD_LOOP_FAKE_HARNESS" + loopFakeLogEnv = "ABCD_LOOP_FAKE_LOG" +) + +func TestMain(m *testing.M) { + if mode := os.Getenv(loopFakeEnv); mode != "" { + os.Exit(loopFakeHarness(mode)) + } + os.Exit(m.Run()) +} + +// loopFakeHarness plays the harness its launch names: opencode's argv starts +// with "run". It writes what it was launched with under the log directory, +// then, by mode, writes the receipt its prompt names (a reviewer's return for +// a reviewer, an implementer's receipt otherwise) or does not. +func loopFakeHarness(mode string) int { + name := "claude" + if len(os.Args) > 1 && os.Args[1] == "run" { + name = "opencode" + } + if dir := os.Getenv(loopFakeLogEnv); dir != "" { + argv, _ := json.Marshal(os.Args[1:]) + _ = os.WriteFile(filepath.Join(dir, name+".argv.json"), argv, 0o600) + } + prompt := os.Args[len(os.Args)-1] + field := func(prefix string) string { + sc := bufio.NewScanner(strings.NewReader(prompt)) + for sc.Scan() { + if v, ok := strings.CutPrefix(sc.Text(), prefix); ok { + return strings.TrimSpace(v) + } + } + return "" + } + if mode == "ok" || strings.HasPrefix(mode, "model-") { + body := "{}\n" + if role := field("You are the "); strings.HasPrefix(role, RoleRuthless) { + body = reviewReturn + } + _ = os.WriteFile(field("Receipt: "), []byte(body), 0o600) + } + if name == "opencode" { + fmt.Println(`{"type":"step_start","sessionID":"ses_fake2","part":{"type":"step-start"}}`) + fmt.Println(`{"type":"text","sessionID":"ses_fake2","part":{"type":"text","text":"done"}}`) + fmt.Println(`{"type":"step_finish","sessionID":"ses_fake2","part":{"type":"step-finish"}}`) + return 0 + } + model := "fake-model" + switch mode { + case "model-huge": + model = strings.Repeat("m", 5<<20) + case "model-ctrl": + model = "fake\x1b[2Jmodel\r\n" + } + init, _ := json.Marshal(map[string]string{"type": "system", "subtype": "init", "session_id": "fake-session-2", "model": model}) + fmt.Println(string(init)) + fmt.Println(`{"type":"result","subtype":"success","is_error":false,"result":"done","session_id":"fake-session-2"}`) + return 0 +} + +// reviewReturn is a ruthless reviewer's return, the same bytes whichever route +// wrote it. +const reviewReturn = "# Review\n\nNothing to fix.\n\n### Verdict\n\nSHIP\n" + +// driveEnv is one test's harness set-up: the fakes on a PATH of their own and +// the log they write. +type driveEnv struct{ bin, log string } + +func newDriveEnv(t *testing.T, mode string, harnesses ...string) driveEnv { + t.Helper() + self, err := os.Executable() + if err != nil { + t.Fatal(err) + } + root := t.TempDir() + e := driveEnv{bin: filepath.Join(root, "bin"), log: filepath.Join(root, "log")} + for _, d := range []string{e.bin, e.log} { + if err := os.Mkdir(d, 0o700); err != nil { + t.Fatal(err) + } + } + for _, h := range harnesses { + if err := os.Symlink(self, filepath.Join(e.bin, h)); err != nil { + t.Fatal(err) + } + } + t.Setenv("PATH", e.bin) + t.Setenv(loopFakeEnv, mode) + t.Setenv(loopFakeLogEnv, e.log) + return e +} + +func (e driveEnv) argv(t *testing.T, harness string) []string { + t.Helper() + raw, err := os.ReadFile(filepath.Join(e.log, harness+".argv.json")) + if err != nil { + t.Fatalf("the %s fake was not launched: %v", harness, err) + } + var out []string + if err := json.Unmarshal(raw, &out); err != nil { + t.Fatal(err) + } + return out +} + +// runnerConfig reads a runner configuration from a machine and a repository +// config file written for the test. +func runnerConfig(t *testing.T, machine, repo string) *runner.Config { + t.Helper() + r := layered.Roots{Repo: t.TempDir(), Home: t.TempDir()} + for path, body := range map[string]string{ + filepath.Join(r.Home, ".abcd", "config.json"): machine, + filepath.Join(r.Repo, ".abcd", "config.json"): repo, + } { + if body == "" { + continue + } + if err := os.MkdirAll(filepath.Dir(path), 0o700); err != nil { + t.Fatal(err) + } + if err := os.WriteFile(path, []byte(body), 0o600); err != nil { + t.Fatal(err) + } + } + c, err := runner.Load(r) + if err != nil { + t.Fatalf("runner config: %v", err) + } + return c +} + +// memTranscripts is a transcript store that keeps what it is handed. +type memTranscripts struct{ stored []string } + +func (m *memTranscripts) Store(name string, req runner.Request, _ runner.Answer, raw []byte) error { + m.stored = append(m.stored, name+" "+req.Role+" "+fmt.Sprint(len(raw) > 0)) + return nil +} + +// driveSteps is fakeSteps with a worktree the runner can run in and an +// implement stage whose verifier reads the receipt the agent wrote. +func driveSteps(worktree string) Stages { + return Stages{ + {Name: StageWorktree, Piece: 6, Run: func(c Context, l *Lane) (Outcome, error) { + l.Branch, l.Worktree = "build/"+l.ID, worktree + return Outcome{Note: "worktree done"}, nil + }}, + {Name: StageBrief, Piece: 5, Run: func(c Context, l *Lane) (Outcome, error) { + l.Brief = RunRelDir + "/brief.md" + return Outcome{Note: "brief done"}, nil + }}, + {Name: StageImplement, Piece: 7, + Run: func(c Context, l *Lane) (Outcome, error) { + return Outcome{Await: &Await{Role: RoleImplementer, Brief: filepath.Join(c.RepoRoot, l.Brief), + Receipt: filepath.Join(c.RepoRoot, "receipt-"+l.ID+".json")}}, nil + }, + Verify: func(c Context, l *Lane, receipt string) error { + raw, err := os.ReadFile(receipt) + if err != nil || strings.TrimSpace(string(raw)) != "{}" { + return refuse("receipt", "", l.ID, "no receipt the contract accepts at "+receipt, "write it, then hand it back") + } + l.Receipts = append(l.Receipts, ReceiptRecord{Role: RoleImplementer, Receipt: receipt}) + return nil + }}, + {Name: StageValidate, Piece: 8, Run: func(Context, *Lane) (Outcome, error) { return Outcome{Note: "validate done"}, nil }}, + {Name: StageLand, Piece: 9, Run: func(Context, *Lane) (Outcome, error) { return Outcome{Note: "land done"}, nil }}, + } +} + +// startToImplement starts a run and takes its lane to the implement stage. +func startToImplement(t *testing.T) (root, id string, steps Stages, o Options) { + t.Helper() + repo := loopRepo(t, readyIntent("", settledQuestions), specWithSteps("")) + start, err := Start(repo.Root(), "itd-10", Options{}) + if err != nil { + t.Fatal(err) + } + wt := t.TempDir() + if err := os.WriteFile(filepath.Join(repo.Root(), filepath.FromSlash(RunRelDir), "brief.md"), []byte("# brief\n"), 0o600); err != nil { + t.Fatal(err) + } + steps = driveSteps(wt) + clock := time.Date(2026, 9, 30, 12, 0, 0, 0, time.UTC) + o = Options{Now: func() time.Time { return clock }} + for range 2 { + if _, err := advance(repo.Root(), start.RunID, steps, o); err != nil { + t.Fatal(err) + } + } + return repo.Root(), start.RunID, steps, o +} + +// TestAnUnsetRouteLeavesTheAwaitAndTheStateByteIdentical is criterion 2: with +// no runner configured, the driven step hands the host exactly what a plain +// step hands it, and writes exactly the same state. +func TestAnUnsetRouteLeavesTheAwaitAndTheStateByteIdentical(t *testing.T) { + root, id, steps, o := startToImplement(t) + newDriveEnv(t, "ok", "claude", "opencode") + path := filepath.Join(root, filepath.FromSlash(StateRelPath(id))) + before := stateBytes(t, root, id) + + plain, err := advance(root, id, steps, o) + if err != nil { + t.Fatal(err) + } + plainState := stateBytes(t, root, id) + if err := os.WriteFile(path, before, 0o600); err != nil { + t.Fatal(err) + } + store := &memTranscripts{} + driven, err := Drive(context.Background(), root, id, steps, o, Runners{Config: runnerConfig(t, "", ""), Transcripts: store}) + if err != nil { + t.Fatal(err) + } + if !reflect.DeepEqual(plain, driven) { + t.Fatalf("an unset route changed the step's result:\nplain %+v\ndriven %+v", plain, driven) + } + if !bytes.Equal(plainState, stateBytes(t, root, id)) { + t.Fatalf("an unset route changed the state:\nplain %s\ndriven %s", plainState, stateBytes(t, root, id)) + } + if len(store.stored) != 0 { + t.Fatalf("an unset route launched a runner: %v", store.stored) + } +} + +// TestARoutedRoleRunsThroughItsRunner is criterion 1 and the loop's criterion +// 9: a role routed to a runner is started by the loop itself with the brief +// and the receipt path the host would get, its receipt is verified by the +// stage's own verifier, its transcript is stored, and the record names the +// runner that ran it. +func TestARoutedRoleRunsThroughItsRunner(t *testing.T) { + root, id, steps, o := startToImplement(t) + env := newDriveEnv(t, "ok", "claude") + store := &memTranscripts{} + cfg := runnerConfig(t, `{"runner":{"claude":{}}}`, `{"roles":{"implementer":{"runner":"claude"}}}`) + res, err := Drive(context.Background(), root, id, steps, o, Runners{Config: cfg, Transcripts: store}) + if err != nil { + t.Fatal(err) + } + if res.PerformedStage != StageImplement || res.Stage != StageValidate || res.Awaiting != nil { + t.Fatalf("the runner's verified receipt completes the stage: %+v", res) + } + if res.Route == nil || res.Route.Asked != runner.Claude || res.Route.Ran != runner.Claude || res.Route.Model != "fake-model" { + t.Fatalf("the result names the route that ran: %+v", res.Route) + } + argv := strings.Join(env.argv(t, "claude"), "\n") + for _, want := range []string{"--bare", "--allowedTools=" + strings.Join(toolsFor(RoleImplementer), ","), + "Brief: " + filepath.Join(root, filepath.FromSlash(RunRelDir), "brief.md")} { + if !strings.Contains(argv, want) { + t.Fatalf("the launch carries %q:\n%s", want, argv) + } + } + if len(store.stored) != 1 || store.stored[0] != "claude implementer true" { + t.Fatalf("the transcript lands in abcd's store: %v", store.stored) + } + st, err := ReadState(root, id) + if err != nil { + t.Fatal(err) + } + rs := st.Lanes[0].Receipts + if len(rs) != 1 || rs[0].Route == nil || rs[0].Route.Ran != runner.Claude { + t.Fatalf("the verified receipt names the route that ran it: %+v", rs) + } + named := false + for _, e := range st.Record { + if e.Stage == StageRunner && strings.Contains(e.Note, "claude") { + named = true + } + } + if !named || len(st.Fallbacks) != 0 { + t.Fatalf("the run record names the runner and records no fallback: %+v %+v", st.Record, st.Fallbacks) + } +} + +// TestARunnerThatCannotRunTheRoleFallsBackAndIsRecorded is criterion 3: an +// absent runner, and a runner whose answer the stage's verifier refuses, each +// hand the role to the host, which is told what to start as before, and each +// leaves one fallback receipt naming the role, the runner, the reason and the +// route that ran. +func TestARunnerThatCannotRunTheRoleFallsBackAndIsRecorded(t *testing.T) { + for _, tc := range []struct { + name, mode, route string + harnesses []string + reason runner.Reason + }{ + {"absent", "ok", "opencode", nil, runner.ReasonAbsent}, + {"invalid", "noreceipt", "claude", []string{"claude"}, runner.ReasonInvalid}, + // A model past its bound or carrying control bytes is a refusal of + // the route, never written into the state: a 5 MiB one there would + // push state.json past its read bound and brick the run. + {"huge-model", "model-huge", "claude", []string{"claude"}, runner.ReasonUnparsable}, + {"control-model", "model-ctrl", "claude", []string{"claude"}, runner.ReasonUnparsable}, + } { + t.Run(tc.name, func(t *testing.T) { + root, id, steps, o := startToImplement(t) + newDriveEnv(t, tc.mode, tc.harnesses...) + cfg := runnerConfig(t, `{"runner":{"claude":{},"opencode":{}}}`, `{"roles":{"implementer":{"runner":"`+tc.route+`"}}}`) + res, err := Drive(context.Background(), root, id, steps, o, Runners{Config: cfg, Transcripts: &memTranscripts{}}) + if err != nil { + t.Fatal(err) + } + if res.Awaiting == nil || res.Awaiting.Role != RoleImplementer || res.PerformedStage != "" { + t.Fatalf("the host is handed the role: %+v", res) + } + fb := res.Fallback + if fb == nil || fb.Role != RoleImplementer || fb.Asked != tc.route || fb.Reason != tc.reason || fb.Ran != runner.Host { + t.Fatalf("the result names the fallback: %+v", fb) + } + st, err := ReadState(root, id) + if err != nil { + t.Fatal(err) + } + if len(st.Fallbacks) != 1 || st.Fallbacks[0] != *fb { + t.Fatalf("the state carries the one fallback receipt: %+v", st.Fallbacks) + } + if st.Lanes[0].Awaiting == nil { + t.Fatal("the lane still awaits the host's receipt") + } + if len(st.Lanes[0].Receipts) != 0 { + t.Fatalf("a refused route verified no receipt: %+v", st.Lanes[0].Receipts) + } + rec, err := ReadRecord(root, id) + if err != nil { + t.Fatal(err) + } + c := rec.FallbackCounts + if c.Total != 1 || c.ByRunner[tc.route] != 1 || c.ByRole[RoleImplementer] != 1 { + t.Fatalf("the run record counts the fallback per runner and per role (criterion 4): %+v", c) + } + }) + } +} + +// TestAReviewThroughARunnerDiffersFromAHostReviewOnlyInItsRoute is criterion +// 7's structural half: one review, returned with the same bytes once by the +// host and once by the opencode runner, is recorded by the stage's own +// verifier with records that differ in the route alone. +func TestAReviewThroughARunnerDiffersFromAHostReviewOnlyInItsRoute(t *testing.T) { + review := func(t *testing.T, routed bool) ValidatorRun { + repo := loopRepo(t, readyIntent("", settledQuestions), specWithSteps("")) + start, err := Start(repo.Root(), "itd-10", Options{}) + if err != nil { + t.Fatal(err) + } + wt := t.TempDir() + steps := driveSteps(wt) + const briefRel, returnRel = "review/brief.md", "review/return.md" + if err := os.MkdirAll(filepath.Join(repo.Root(), "review"), 0o700); err != nil { + t.Fatal(err) + } + if err := os.WriteFile(filepath.Join(repo.Root(), "review", "brief.md"), []byte("# review\n"), 0o600); err != nil { + t.Fatal(err) + } + steps[2] = StageDef{Name: StageImplement, Piece: 7, + Run: func(c Context, l *Lane) (Outcome, error) { + if len(l.Validation) == 0 { + l.Validation = []ValidationRound{{Round: 1, HeadSHA: strings.Repeat("a", 40), + Validators: []ValidatorRun{{Role: RoleRuthless, Brief: briefRel, Return: returnRel}}}} + } + return Outcome{Await: &Await{Role: RoleRuthless, Brief: briefRel, Receipt: returnRel}}, nil + }, + Verify: func(c Context, l *Lane, receipt string) error { + l.Awaiting.Role = RoleRuthless + return verifyValidation(c, l, receipt) + }} + o := Options{Now: func() time.Time { return time.Date(2026, 9, 30, 12, 0, 0, 0, time.UTC) }} + for range 2 { + if _, err := advance(repo.Root(), start.RunID, steps, o); err != nil { + t.Fatal(err) + } + } + if routed { + newDriveEnv(t, "ok", "opencode") + cfg := runnerConfig(t, `{"runner":{"opencode":{}}}`, `{"roles":{"ruthless-reviewer":{"runner":"opencode"}}}`) + res, err := Drive(context.Background(), repo.Root(), start.RunID, steps, o, Runners{Config: cfg, Transcripts: &memTranscripts{}}) + if err != nil || res.PerformedStage != StageImplement { + t.Fatalf("the opencode review is verified: %+v %v", res, err) + } + } else { + res, err := advance(repo.Root(), start.RunID, steps, o) + if err != nil || res.Awaiting == nil { + t.Fatalf("the host is handed the review: %+v %v", res, err) + } + if err := os.WriteFile(filepath.Join(repo.Root(), returnRel), []byte(reviewReturn), 0o600); err != nil { + t.Fatal(err) + } + if _, err := Receipt(repo.Root(), start.RunID, returnRel, steps, o); err != nil { + t.Fatal(err) + } + } + st, err := ReadState(repo.Root(), start.RunID) + if err != nil { + t.Fatal(err) + } + return st.Lanes[0].Validation[0].Validators[0] + } + host, routed := review(t, false), review(t, true) + if host.Route != nil { + t.Fatalf("a host-run review names no runner route: %+v", host.Route) + } + if routed.Route == nil || routed.Route.Asked != runner.OpenCode || routed.Route.Ran != runner.OpenCode { + t.Fatalf("the opencode review names its route: %+v", routed.Route) + } + if host.Verdict != "SHIP" { + t.Fatalf("the host review is recorded: %+v", host) + } + routed.Route = nil + if !reflect.DeepEqual(host, routed) { + t.Fatalf("the two reviews differ beyond the route:\nhost %+v\nrouted %+v", host, routed) + } +} + +// TestAVersion7StateCarryingARunnersRecordIsRefused: version 7 never wrote a +// fallback or a route, so a file of that version carrying one is refused, and +// one without either is read and written back at the current version. +func TestAVersion7StateCarryingARunnersRecordIsRefused(t *testing.T) { + root, id, _, _ := startToImplement(t) + path := filepath.Join(root, filepath.FromSlash(StateRelPath(id))) + current := stateBytes(t, root, id) + cur := fmt.Sprintf(`"schema_version": %d,`, SchemaVersion) + old := strings.Replace(string(current), cur, `"schema_version": 7,`, 1) + if err := os.WriteFile(path, []byte(old), 0o600); err != nil { + t.Fatal(err) + } + if st, err := ReadState(root, id); err != nil || st.SchemaVersion != SchemaVersion { + t.Fatalf("a version-7 file is read as the current version: %v", err) + } + carrying := strings.Replace(old, `"record": [`, `"fallbacks": [{"at":"2026-09-30T12:00:00Z","role":"implementer","asked":"claude","reason":"absent","detail":"x","ran":"host"}], + "record": [`, 1) + if err := os.WriteFile(path, []byte(carrying), 0o600); err != nil { + t.Fatal(err) + } + if _, err := ReadState(root, id); err == nil { + t.Fatal("a version-7 file carrying a fallback must be refused") + } +} + +// TestRoleToolsFollowTheAgentDefinitions: a reviewer run through a runner is +// granted the tools its agent definition names, plus Write for the return its +// brief tells it to write, and nothing its definition does not name. +func TestRoleToolsFollowTheAgentDefinitions(t *testing.T) { + for _, role := range []string{RoleRuthless, RoleSecurity} { + raw, err := os.ReadFile(filepath.Join("..", "..", "..", "..", "agents", role+".md")) + if err != nil { + t.Fatal(err) + } + var def []string + for _, ln := range strings.Split(string(raw), "\n") { + if v, ok := strings.CutPrefix(ln, "tools:"); ok { + for _, tool := range strings.Split(v, ",") { + def = append(def, strings.TrimSpace(tool)) + } + break + } + } + want := append(def, "Write") + if got := toolsFor(role); !reflect.DeepEqual(got, want) { + t.Fatalf("%s: tools = %v, want its definition's %v plus Write", role, got, def) + } + } + if got := toolsFor("scribe"); got != nil { + t.Fatalf("a role the loop does not start is granted nothing: %v", got) + } +} + +// TestAStepThatReTellsAnAwaitStartsNoRunner: once a runner's fallback has +// handed the host the role, stepping again re-tells the await; it neither +// starts the runner again nor records a second fallback. +func TestAStepThatReTellsAnAwaitStartsNoRunner(t *testing.T) { + root, id, steps, o := startToImplement(t) + newDriveEnv(t, "ok") + cfg := runnerConfig(t, `{"runner":{"opencode":{}}}`, `{"roles":{"implementer":{"runner":"opencode"}}}`) + for range 2 { + if _, err := Drive(context.Background(), root, id, steps, o, Runners{Config: cfg, Transcripts: &memTranscripts{}}); err != nil { + t.Fatal(err) + } + } + st, err := ReadState(root, id) + if err != nil { + t.Fatal(err) + } + if len(st.Fallbacks) != 1 { + t.Fatalf("one fallback for one hand-out, got %d", len(st.Fallbacks)) + } +} diff --git a/internal/core/implement/loop/fixrounds_test.go b/internal/core/implement/loop/fixrounds_test.go index 318205d7c..605e18df0 100644 --- a/internal/core/implement/loop/fixrounds_test.go +++ b/internal/core/implement/loop/fixrounds_test.go @@ -35,7 +35,7 @@ func TestAnUndecidedCriterionReopensTheWorkLikeANotMet(t *testing.T) { if a.Verdict != "INCONCLUSIVE" || a.Pass { t.Fatalf("an undecided audit does not pass the round: %+v", a) } - res, err := Advance(repo.Root(), id, stages, Options{}) + res, err := advance(repo.Root(), id, stages, Options{}) if err != nil || res.Awaiting == nil || res.Awaiting.Role != RoleImplementer || res.Stage != StageValidate { t.Fatalf("an undecided audit hands the lane to a fresh implementer, never to its landing: %+v %v", res, err) } @@ -140,7 +140,7 @@ func TestALaneThatExhaustsItsFixRoundsIsHandedBack(t *testing.T) { passRound(t, repo, id, stages, RoleRuthless, RoleSecurity) handBack(t, repo, id, stages, RoleAuditor, "NOT_MET") - res, err := Advance(repo.Root(), id, stages, Options{}) + res, err := advance(repo.Root(), id, stages, Options{}) if err != nil { t.Fatal(err) } @@ -173,7 +173,7 @@ func TestALaneThatExhaustsItsFixRoundsIsHandedBack(t *testing.T) { } before := stateBytes(t, repo.Root(), id) - _, err = Advance(repo.Root(), id, stages, Options{}) + _, err = advance(repo.Root(), id, stages, Options{}) if r := mustRefusal(t, err); r.Stage != string(StageHandedBack) || !strings.Contains(r.Reason, "unachievable") || !strings.Contains(r.Remedy, "itd-10") { t.Fatalf("a handed-back lane starts nothing further, and says why: %+v", r) } @@ -235,7 +235,7 @@ func TestAHandedBackPickNamesThePickFalsified(t *testing.T) { } handBack(t, repo, id, stages, RoleRuthless, reviewerReturn("FIX FIRST")) passRound(t, repo, id, stages, RoleSecurity, RoleAuditor) - res, err := Advance(repo.Root(), id, stages, Options{}) + res, err := advance(repo.Root(), id, stages, Options{}) if err != nil || res.HandBack == nil { t.Fatalf("with no fix round allowed, the first failing round hands the lane back: %+v %v", res, err) } diff --git a/internal/core/implement/loop/issue_test.go b/internal/core/implement/loop/issue_test.go index b2184d07d..b6e12f675 100644 --- a/internal/core/implement/loop/issue_test.go +++ b/internal/core/implement/loop/issue_test.go @@ -132,7 +132,7 @@ func issueAwaiting(t *testing.T) (*gittest.Repo, string, Lane, string) { t.Fatal(err) } advanceTo(t, repo, start.RunID, StageImplement) - if _, err := Advance(repo.Root(), start.RunID, DefaultStages(), Options{}); err != nil { + if _, err := advance(repo.Root(), start.RunID, DefaultStages(), Options{}); err != nil { t.Fatal(err) } st, err := ReadState(repo.Root(), start.RunID) @@ -197,7 +197,7 @@ func TestALaneReportHandBackStopsTheLaneAndDiscardsItsWork(t *testing.T) { if out := repo.Git("branch", "--list", l.Branch); strings.TrimSpace(out) != "" { t.Fatalf("the lane's branch is discarded: %q", out) } - if _, err := Advance(repo.Root(), runID, DefaultStages(), Options{}); err == nil { + if _, err := advance(repo.Root(), runID, DefaultStages(), Options{}); err == nil { t.Fatal("a handed-back lane refuses every later step") } @@ -260,7 +260,7 @@ func TestAnIssueLaneLandsOnePullRequestThatResolvesItsIssue(t *testing.T) { f := &landFixture{repo: repo, bare: bare, gh: gh, issue: eligibleIssue, runID: start.RunID, stages: DefaultStages()} stepTo(t, repo, f.runID, f.stages, StageImplement) - if res, err := Advance(repo.Root(), f.runID, f.stages, Options{}); err != nil || res.Awaiting == nil { + if res, err := advance(repo.Root(), f.runID, f.stages, Options{}); err != nil || res.Awaiting == nil { t.Fatalf("the implement stage awaits an implementer: %+v %v", res, err) } l := currentLane(t, repo, f.runID) diff --git a/internal/core/implement/loop/land_attribution_test.go b/internal/core/implement/loop/land_attribution_test.go index a7d89d7b5..38c0f71c3 100644 --- a/internal/core/implement/loop/land_attribution_test.go +++ b/internal/core/implement/loop/land_attribution_test.go @@ -38,7 +38,7 @@ func TestALandingWhoseReceiptReportsNoModelIsRefused(t *testing.T) { l := f.validated(t) head := l.HeadSHA f.step(t) // prepare - _, err := Advance(f.repo.Root(), f.runID, f.stages, Options{}) + _, err := advance(f.repo.Root(), f.runID, f.stages, Options{}) r := mustRefusal(t, err) if r.Stage != string(StageLand) || !strings.Contains(r.Reason, "model") || !strings.Contains(r.Remedy, "model") { t.Fatalf("a landing with no reported model is refused naming the model: %+v", r) @@ -110,7 +110,7 @@ func TestARefusingCommitMsgHookStopsTheLanding(t *testing.T) { seen := filepath.Join(t.TempDir(), "seen") hook := installHook(t, f, "commit-msg", "#!/bin/sh\ncat \"$1\" > '"+seen+"'\necho 'commit-msg: refused by the test' >&2\nexit 1\n") f.step(t) // prepare - _, err := Advance(f.repo.Root(), f.runID, f.stages, Options{}) + _, err := advance(f.repo.Root(), f.runID, f.stages, Options{}) r := mustRefusal(t, err) if r.Stage != string(StageLand) || !strings.Contains(r.Reason, "refused by the test") { t.Fatalf("a refusing commit-msg hook stops the landing, naming what it said: %+v", r) diff --git a/internal/core/implement/loop/land_test.go b/internal/core/implement/loop/land_test.go index 43e04fbf9..c50870ba1 100644 --- a/internal/core/implement/loop/land_test.go +++ b/internal/core/implement/loop/land_test.go @@ -118,7 +118,7 @@ func (f *landFixture) validated(t *testing.T) Lane { t.Helper() repo, id := f.repo, f.runID stepTo(t, repo, id, f.stages, StageImplement) - res, err := Advance(repo.Root(), id, f.stages, Options{}) + res, err := advance(repo.Root(), id, f.stages, Options{}) if err != nil || res.Awaiting == nil { t.Fatalf("the implement stage awaits an implementer: %+v %v", res, err) } @@ -142,7 +142,7 @@ func (f *landFixture) validated(t *testing.T) Lane { // step advances the run once and fails the test on an error. func (f *landFixture) step(t *testing.T) StepResult { t.Helper() - res, err := Advance(f.repo.Root(), f.runID, f.stages, Options{}) + res, err := advance(f.repo.Root(), f.runID, f.stages, Options{}) if err != nil { t.Fatalf("landing step: %v", err) } @@ -246,7 +246,7 @@ func TestTheLandingClosesTheSpecResolvesTheCapturesAndArmsTheMerge(t *testing.T) } // No preflight receipt: refused, nothing pushed. - _, err := Advance(f.repo.Root(), f.runID, f.stages, Options{}) + _, err := advance(f.repo.Root(), f.runID, f.stages, Options{}) r := mustRefusal(t, err) if r.Stage != string(StageLand) || !strings.Contains(r.Reason, "preflight receipt") || !strings.Contains(r.Remedy, "preflight") { t.Fatalf("a landing without the preflight receipt is refused naming it: %+v", r) @@ -287,7 +287,7 @@ func TestTheLandingClosesTheSpecResolvesTheCapturesAndArmsTheMerge(t *testing.T) } // Not merged yet: the loop waits, and cleans nothing up. - _, err = Advance(f.repo.Root(), f.runID, f.stages, Options{}) + _, err = advance(f.repo.Root(), f.runID, f.stages, Options{}) if r := mustRefusal(t, err); !r.Contention || !strings.Contains(r.Reason, "not on") { t.Fatalf("an unmerged lane waits for its merge: %+v", r) } @@ -348,7 +348,7 @@ func TestNothingIsPushedAfterArming(t *testing.T) { late := laneCommit(t, f.repo, l, "late.txt") preflighted(t, l, late) for range 3 { - _, _ = Advance(f.repo.Root(), f.runID, f.stages, Options{}) + _, _ = advance(f.repo.Root(), f.runID, f.stages, Options{}) } if got := f.remoteBranch(t, l.Branch); got != pushed { t.Fatalf("nothing is pushed after arming: the remote moved from %s to %s", pushed, got) @@ -410,7 +410,7 @@ func TestAClosedPullRequestIsRefusedAndNothingIsCleanedUp(t *testing.T) { if err := os.WriteFile(filepath.Join(f.gh, "state"), []byte("CLOSED\n"), 0o644); err != nil { t.Fatal(err) } - _, err := Advance(f.repo.Root(), f.runID, f.stages, Options{}) + _, err := advance(f.repo.Root(), f.runID, f.stages, Options{}) if r := mustRefusal(t, err); r.Contention || !strings.Contains(r.Reason, "closed") { t.Fatalf("a pull request closed without merging is refused, not waited on: %+v", r) } @@ -447,7 +447,7 @@ func TestAKilledLandingResumesAtTheStepThatDidNotComplete(t *testing.T) { f.validated(t) f.stages = stages f.step(t) - if _, err := Advance(f.repo.Root(), f.runID, stages, Options{}); err == nil { + if _, err := advance(f.repo.Root(), f.runID, stages, Options{}); err == nil { t.Fatal("the kill surfaces") } f.step(t) @@ -458,7 +458,7 @@ func TestAKilledLandingResumesAtTheStepThatDidNotComplete(t *testing.T) { } preflighted(t, l, l.HeadSHA) f.step(t) - if _, err := Advance(f.repo.Root(), f.runID, stages, Options{}); err == nil { + if _, err := advance(f.repo.Root(), f.runID, stages, Options{}); err == nil { t.Fatal("the kill surfaces") } f.step(t) diff --git a/internal/core/implement/loop/lane_test.go b/internal/core/implement/loop/lane_test.go index 99ad997c4..b375109b0 100644 --- a/internal/core/implement/loop/lane_test.go +++ b/internal/core/implement/loop/lane_test.go @@ -50,7 +50,7 @@ func TestTheWorktreeStepMakesTheLaneInTheStore(t *testing.T) { if err != nil { t.Fatal(err) } - res, err := Advance(repo.Root(), start.RunID, DefaultStages(), Options{}) + res, err := advance(repo.Root(), start.RunID, DefaultStages(), Options{}) if err != nil { t.Fatal(err) } @@ -143,7 +143,7 @@ func TestTheWorktreeStepRefusesAPathThatCouldEscapeTheStore(t *testing.T) { if err := writeState(root, st); err != nil { t.Fatal(err) } - _, err = Advance(repo.Root(), start.RunID, DefaultStages(), Options{}) + _, err = advance(repo.Root(), start.RunID, DefaultStages(), Options{}) r := mustRefusal(t, err) if r.Stage != string(StageWorktree) { t.Fatalf("want the worktree stage to refuse: %+v", r) @@ -175,7 +175,7 @@ func TestTheWorktreeStepRefusesASymlinkedStore(t *testing.T) { if err := os.Symlink(elsewhere, filepath.Join(os.Getenv("HOME"), ".abcd", "worktrees")); err != nil { t.Fatal(err) } - _, err = Advance(repo.Root(), start.RunID, DefaultStages(), Options{}) + _, err = advance(repo.Root(), start.RunID, DefaultStages(), Options{}) if r := mustRefusal(t, err); r.Stage != string(StageWorktree) || !strings.Contains(r.Reason, "real directories") { t.Fatalf("want the symlinked store refused: %+v", r) } @@ -200,7 +200,7 @@ func TestTheWorktreeStepNeverAdoptsWhatItDidNotMake(t *testing.T) { if err := os.WriteFile(filepath.Join(squat, "mine.txt"), []byte("the user's\n"), 0o600); err != nil { t.Fatal(err) } - _, err = Advance(repo.Root(), start.RunID, DefaultStages(), Options{}) + _, err = advance(repo.Root(), start.RunID, DefaultStages(), Options{}) if r := mustRefusal(t, err); r.Stage != string(StageWorktree) || !strings.Contains(r.Reason, "occupied") { t.Fatalf("want the occupied path refused: %+v", r) } @@ -241,7 +241,7 @@ func TestTheWorktreeStepRefusesAStoreLevelAnyoneElseCanWrite(t *testing.T) { if err := os.Chmod(level, tc.mode); err != nil { t.Fatal(err) } - _, err = Advance(repo.Root(), start.RunID, DefaultStages(), Options{}) + _, err = advance(repo.Root(), start.RunID, DefaultStages(), Options{}) r := mustRefusal(t, err) if r.Stage != string(StageWorktree) || !strings.Contains(r.Reason, "writable by its group or by every user") { t.Fatalf("want the writable store level refused: %+v", r) @@ -267,7 +267,7 @@ func TestTheWorktreeStepMakesTheStoreTheCallersAlone(t *testing.T) { if err != nil { t.Fatal(err) } - if _, err := Advance(repo.Root(), start.RunID, DefaultStages(), Options{}); err != nil { + if _, err := advance(repo.Root(), start.RunID, DefaultStages(), Options{}); err != nil { t.Fatal(err) } home := os.Getenv("HOME") diff --git a/internal/core/implement/loop/loop.go b/internal/core/implement/loop/loop.go index 126525284..74023f29f 100644 --- a/internal/core/implement/loop/loop.go +++ b/internal/core/implement/loop/loop.go @@ -1,7 +1,7 @@ package loop // loop.go is the step interface (spec piece 2; decision 5's default driver): -// Start creates a run after the checks, Advance performs the lane's next stage +// Start creates a run after the checks, advance performs the lane's next stage // and exits, Receipt verifies what an agent stage waited on and advances, and // Status reads. Each takes the run tier's lock, reads the state first and // writes it last; a stage that fails leaves the state exactly as it was. @@ -18,6 +18,7 @@ import ( "github.com/intentdriven/abcd/internal/core/intent" "github.com/intentdriven/abcd/internal/core/layered" "github.com/intentdriven/abcd/internal/core/recordid" + "github.com/intentdriven/abcd/internal/core/runner" "github.com/intentdriven/abcd/internal/core/statusblock" "github.com/intentdriven/abcd/internal/fsutil" "github.com/intentdriven/abcd/internal/gitutil" @@ -187,7 +188,7 @@ type StartResult struct { Next string `json:"next"` } -// StepResult is what Advance and Receipt return. +// StepResult is what advance and Receipt return. type StepResult struct { RunID string `json:"run_id"` Lane string `json:"lane,omitempty"` @@ -205,7 +206,17 @@ type StepResult struct { // HandBack is set when the lane was handed back to the person: this call // stopped it, or it stood stopped when the run was started again. HandBack *HandBack `json:"hand_back,omitempty"` - Next string `json:"next"` + // Route is the route that ran the agent when the process driver started + // it through a runner (Drive); absent when the host is to run it. + Route *runner.RouteRecord `json:"route,omitempty"` + // Fallback is the fallback receipt this call recorded when the runner a + // role is routed to did not run it and the host is handed the role. + Fallback *runner.FallbackReceipt `json:"fallback,omitempty"` + Next string `json:"next"` + // handed is true when this call handed the lane to an agent, false when + // it re-told an await an earlier call began: only the call that hands the + // work out may start a runner for it. + handed bool } // Start resumes the live run for key, or runs the checks and, when every one @@ -546,13 +557,13 @@ func openNextLaneRecorded(st *State, now time.Time) { Note: fmt.Sprintf("%s opened for step %d of %s (%s)", l.ID, l.SpecStep, st.Spec, l.StepTitle)}) } -// Advance performs the next stage of the run's current lane and returns. A lane +// advance performs the next stage of the run's current lane and returns. A lane // that awaits a receipt performs nothing and re-tells what it awaits; a run // that is complete says so; a run paused by its window clock is refused until // next_eligible_at. A stage whose body this build does not carry is refused // naming the piece that delivers it. The state is written only after a body // succeeds, and then once. -func Advance(repoRoot, runID string, steps Stages, o Options) (StepResult, error) { +func advance(repoRoot, runID string, steps Stages, o Options) (StepResult, error) { var res StepResult err := mutate(repoRoot, runID, func(root *os.Root, st *State) (bool, error) { now := o.now() @@ -640,6 +651,7 @@ func Advance(repoRoot, runID string, steps Stages, o Options) (StepResult, error } st.UpdatedAt = now res = laneResult(*st, lane, performed) + res.handed = out.Await != nil return true, nil }) return res, err @@ -696,7 +708,14 @@ func pausedMove(lane Lane, until time.Time) string { // lane awaits one, when the path is not the one the stage named, when this build // carries no verifier for the stage, and when the verifier refuses it; in every // refusal the lane stays where it was. A verified receipt completes the stage. -func Receipt(repoRoot, runID, receipt string, steps Stages, o Options) (StepResult, error) { +func Receipt(repoRoot, runID, receiptPath string, steps Stages, o Options) (StepResult, error) { + return receipt(repoRoot, runID, receiptPath, steps, o, nil) +} + +// receipt is Receipt. A route that is not nil is the runner that ran the +// agent (Drive): it is stamped on what the verifier recorded from the receipt, +// and the record names it. The host's receipt carries none. +func receipt(repoRoot, runID, receipt string, steps Stages, o Options, route *runner.RouteRecord) (StepResult, error) { var res StepResult err := mutate(repoRoot, runID, func(root *os.Root, st *State) (bool, error) { now := o.now() @@ -722,6 +741,10 @@ func Receipt(repoRoot, runID, receipt string, steps Stages, o Options) (StepResu } return false, refuse("receipt", "", lane.ID, err.Error(), "correct what the reason names, then hand the receipt back") } + if route != nil { + stampRoute(&lane, lane.Awaiting.Receipt, route) + st.Record = append(st.Record, Entry{At: now, Lane: lane.ID, Stage: StageRunner, Note: routeNote(lane.Awaiting.Role, *route)}) + } if lane.HandBack != nil { // The lane's own receipt handed the work back: the verifier has // discarded it, and the lane ends here, before the validators. diff --git a/internal/core/implement/loop/loop_test.go b/internal/core/implement/loop/loop_test.go index 617a6e88f..cf306c9c0 100644 --- a/internal/core/implement/loop/loop_test.go +++ b/internal/core/implement/loop/loop_test.go @@ -515,7 +515,7 @@ func TestTheHostDrivesTheLoopEndToEnd(t *testing.T) { for lane := 1; lane <= 2; lane++ { for _, want := range []Stage{StageWorktree, StageBrief} { - res, err := Advance(repo.Root(), id, steps, Options{}) + res, err := advance(repo.Root(), id, steps, Options{}) if err != nil { t.Fatal(err) } @@ -523,7 +523,7 @@ func TestTheHostDrivesTheLoopEndToEnd(t *testing.T) { t.Fatalf("lane %d: performed %q, want %q (%+v)", lane, res.PerformedStage, want, res) } } - res, err := Advance(repo.Root(), id, steps, Options{}) + res, err := advance(repo.Root(), id, steps, Options{}) if err != nil { t.Fatal(err) } @@ -537,7 +537,7 @@ func TestTheHostDrivesTheLoopEndToEnd(t *testing.T) { // Asking again tells the same thing and moves nothing. before := stateBytes(t, repo.Root(), id) - again, err := Advance(repo.Root(), id, steps, Options{}) + again, err := advance(repo.Root(), id, steps, Options{}) if err != nil || again.Awaiting == nil || again.Awaiting.Receipt != receipt || again.PerformedStage != "" { t.Fatalf("a step while awaiting re-tells the await: %+v %v", again, err) } @@ -567,7 +567,7 @@ func TestTheHostDrivesTheLoopEndToEnd(t *testing.T) { t.Fatalf("a verified receipt completes the agent step: %+v", got) } for _, want := range []Stage{StageValidate, StageLand} { - res, err := Advance(repo.Root(), id, steps, Options{}) + res, err := advance(repo.Root(), id, steps, Options{}) if err != nil || res.PerformedStage != want { t.Fatalf("lane %d: performed %q, want %q (%v)", lane, res.PerformedStage, want, err) } @@ -580,7 +580,7 @@ func TestTheHostDrivesTheLoopEndToEnd(t *testing.T) { if !st.Complete() || len(st.Lanes) != 2 || st.Lanes[1].SpecStep != 2 || st.Lanes[1].Branch != "build/lane-2" { t.Fatalf("both spec steps landed through their own lanes: %+v", st) } - fin, err := Advance(repo.Root(), id, steps, Options{}) + fin, err := advance(repo.Root(), id, steps, Options{}) if err != nil || !fin.Complete { t.Fatalf("a complete run says so: %+v %v", fin, err) } @@ -601,18 +601,18 @@ func TestAKilledStepRepeatsAndACompletedStepDoesNot(t *testing.T) { t.Fatal(err) } f := &fakeSteps{calls: map[Stage]int{}, failing: StageBrief} - if _, err := Advance(repo.Root(), start.RunID, f.steps(), Options{}); err != nil { + if _, err := advance(repo.Root(), start.RunID, f.steps(), Options{}); err != nil { t.Fatal(err) } before := stateBytes(t, repo.Root(), start.RunID) - if _, err := Advance(repo.Root(), start.RunID, f.steps(), Options{}); err == nil { + if _, err := advance(repo.Root(), start.RunID, f.steps(), Options{}); err == nil { t.Fatal("a step that fails must report it") } if !bytes.Equal(before, stateBytes(t, repo.Root(), start.RunID)) { t.Fatal("a step that did not complete must leave the state as it was") } f.failing = "" - res, err := Advance(repo.Root(), start.RunID, f.steps(), Options{}) + res, err := advance(repo.Root(), start.RunID, f.steps(), Options{}) if err != nil || res.PerformedStage != StageBrief { t.Fatalf("the next invocation performs the step that did not complete: %+v %v", res, err) } @@ -635,7 +635,7 @@ func TestAStepThisBuildDoesNotCarryIsRefusedByName(t *testing.T) { for i := range bare { bare[i].Run, bare[i].Verify = nil, nil } - _, err = Advance(repo.Root(), start.RunID, bare, Options{}) + _, err = advance(repo.Root(), start.RunID, bare, Options{}) r := mustRefusal(t, err) if r.Stage != string(StageWorktree) || r.Lane != "lane-1" || !strings.Contains(r.Reason, "piece 6") { t.Fatalf("want the unbuilt step and its piece named: %+v", r) @@ -667,7 +667,7 @@ func TestAPauseRefusesUntilNextEligibleAt(t *testing.T) { t.Fatal(err) } f := &fakeSteps{calls: map[Stage]int{}} - _, err = Advance(repo.Root(), start.RunID, f.steps(), Options{Now: func() time.Time { return now }}) + _, err = advance(repo.Root(), start.RunID, f.steps(), Options{Now: func() time.Time { return now }}) r := mustRefusal(t, err) if r.Stage != "pause" || !r.Contention || !strings.Contains(r.Reason, "2026-09-25T13:00:00Z") { t.Fatalf("want the pause named: %+v", r) @@ -675,7 +675,7 @@ func TestAPauseRefusesUntilNextEligibleAt(t *testing.T) { if f.calls[StageWorktree] != 0 { t.Fatal("a paused run performs nothing") } - res, err := Advance(repo.Root(), start.RunID, f.steps(), Options{Now: func() time.Time { return later }}) + res, err := advance(repo.Root(), start.RunID, f.steps(), Options{Now: func() time.Time { return later }}) if err != nil || res.PerformedStage != StageWorktree { t.Fatalf("at next_eligible_at the loop moves again: %+v %v", res, err) } @@ -827,7 +827,7 @@ func TestAReceiptNamedThroughASymlinkedPathIsTheReceiptAwaited(t *testing.T) { f := &fakeSteps{calls: map[Stage]int{}} var res StepResult for range 3 { - if res, err = Advance(repo.Root(), start.RunID, f.steps(), Options{}); err != nil { + if res, err = advance(repo.Root(), start.RunID, f.steps(), Options{}); err != nil { t.Fatal(err) } } @@ -883,7 +883,7 @@ func TestTheRecordNamesEachLaneAsItOpens(t *testing.T) { } continue } - if _, err := Advance(repo.Root(), id, steps, Options{}); err != nil { + if _, err := advance(repo.Root(), id, steps, Options{}); err != nil { t.Fatal(err) } } diff --git a/internal/core/implement/loop/next_test.go b/internal/core/implement/loop/next_test.go index 19e404085..44a1e89e0 100644 --- a/internal/core/implement/loop/next_test.go +++ b/internal/core/implement/loop/next_test.go @@ -171,7 +171,7 @@ func TestNextWritesTheReasonAsTheLaneFirstCommit(t *testing.T) { t.Fatalf("the pick starts the build's own run: %+v", st) } - if _, err := Advance(repo.Root(), res.Start.RunID, DefaultStages(), Options{}); err != nil { + if _, err := advance(repo.Root(), res.Start.RunID, DefaultStages(), Options{}); err != nil { t.Fatal(err) } st, err = ReadState(repo.Root(), res.Start.RunID) @@ -213,7 +213,7 @@ func TestNextWritesTheReasonAsTheLaneFirstCommit(t *testing.T) { // The same lane from here: the brief, then the implementer's receipt, which // may not count the pick's commit. advanceTo(t, repo, st.RunID, StageImplement) - if _, err := Advance(repo.Root(), st.RunID, DefaultStages(), Options{}); err != nil { + if _, err := advance(repo.Root(), st.RunID, DefaultStages(), Options{}); err != nil { t.Fatal(err) } st, _ = ReadState(repo.Root(), st.RunID) @@ -365,7 +365,7 @@ func TestAReceiptRefusesABranchThatDroppedThePick(t *testing.T) { } runID := res.Start.RunID advanceTo(t, repo, runID, StageImplement) - if _, err := Advance(repo.Root(), runID, DefaultStages(), Options{}); err != nil { + if _, err := advance(repo.Root(), runID, DefaultStages(), Options{}); err != nil { t.Fatal(err) } st, err := ReadState(repo.Root(), runID) diff --git a/internal/core/implement/loop/pace_test.go b/internal/core/implement/loop/pace_test.go index 9ec3de233..507ded93d 100644 --- a/internal/core/implement/loop/pace_test.go +++ b/internal/core/implement/loop/pace_test.go @@ -253,12 +253,12 @@ func TestAnElapsedWindowStartsNothingAndWritesNextEligibleAt(t *testing.T) { f := &fakeSteps{calls: map[Stage]int{}} steps := f.steps() for i, want := range []Stage{StageWorktree, StageBrief} { - res, err := Advance(repo.Root(), id, steps, at(time.Duration(i+1)*time.Minute)) + res, err := advance(repo.Root(), id, steps, at(time.Duration(i+1)*time.Minute)) if err != nil || res.PerformedStage != want { t.Fatalf("inside the window the loop moves: %+v %v", res, err) } } - res, err := Advance(repo.Root(), id, steps, at(5*time.Minute)) + res, err := advance(repo.Root(), id, steps, at(5*time.Minute)) if err != nil || res.Awaiting == nil { t.Fatalf("the implementer is started inside the window: %+v %v", res, err) } @@ -266,7 +266,7 @@ func TestAnElapsedWindowStartsNothingAndWritesNextEligibleAt(t *testing.T) { // The window elapses while the implementer works. closed := t0.Add(61 * time.Minute) - res, err = Advance(repo.Root(), id, steps, at(61*time.Minute)) + res, err = advance(repo.Root(), id, steps, at(61*time.Minute)) if err != nil { t.Fatalf("closing the window is not a failure: %v", err) } @@ -294,7 +294,7 @@ func TestAnElapsedWindowStartsNothingAndWritesNextEligibleAt(t *testing.T) { // Before next_eligible_at nothing moves and the state is unchanged. before := stateBytes(t, repo.Root(), id) - _, err = Advance(repo.Root(), id, steps, at(80*time.Minute)) + _, err = advance(repo.Root(), id, steps, at(80*time.Minute)) r := mustRefusal(t, err) if r.Stage != "pause" || !r.Contention || !strings.Contains(r.Reason, wantNext.Format(time.RFC3339)) { t.Fatalf("a step inside the pause is refused naming the time: %+v", r) @@ -304,7 +304,7 @@ func TestAnElapsedWindowStartsNothingAndWritesNextEligibleAt(t *testing.T) { } // At next_eligible_at a new window opens and the loop moves again. - res, err = Advance(repo.Root(), id, steps, at(91*time.Minute)) + res, err = advance(repo.Root(), id, steps, at(91*time.Minute)) if err != nil || res.PerformedStage != StageValidate { t.Fatalf("after the pause the loop moves: %+v %v", res, err) } @@ -350,7 +350,7 @@ func TestAVersionOneStateIsReadAsAnUnpacedRun(t *testing.T) { t.Fatalf("a version-1 run is unpaced: %+v", st) } f := &fakeSteps{calls: map[Stage]int{}} - res, err := Advance(repo.Root(), start.RunID, f.steps(), Options{}) + res, err := advance(repo.Root(), start.RunID, f.steps(), Options{}) if err != nil || res.PerformedStage != StageWorktree { t.Fatalf("a version-1 run steps on, days after it started: %+v %v", res, err) } diff --git a/internal/core/implement/loop/receipt_test.go b/internal/core/implement/loop/receipt_test.go index c657a524e..10a91ab39 100644 --- a/internal/core/implement/loop/receipt_test.go +++ b/internal/core/implement/loop/receipt_test.go @@ -21,7 +21,7 @@ func awaitingLane(t *testing.T) (*gittest.Repo, string, Lane, string) { t.Fatal(err) } advanceTo(t, repo, start.RunID, StageImplement) - res, err := Advance(repo.Root(), start.RunID, DefaultStages(), Options{}) + res, err := advance(repo.Root(), start.RunID, DefaultStages(), Options{}) if err != nil { t.Fatal(err) } diff --git a/internal/core/implement/loop/record.go b/internal/core/implement/loop/record.go index e8e034cba..98d7adbf4 100644 --- a/internal/core/implement/loop/record.go +++ b/internal/core/implement/loop/record.go @@ -15,6 +15,8 @@ import ( "os" "slices" "time" + + "github.com/intentdriven/abcd/internal/core/runner" ) // StageTranscript is the run record's stage for a captured transcript, and @@ -40,7 +42,12 @@ type RunRecord struct { Lanes []RecordLane `json:"lanes"` Pending []PendingStep `json:"pending"` Transcripts []Transcript `json:"transcripts"` - Record []Entry `json:"record"` + // Fallbacks are the run's fallback receipts, and FallbackCounts their + // count per runner asked for and per role (itd-2609201916056194 + // criterion 4). + Fallbacks []runner.FallbackReceipt `json:"fallbacks"` + FallbackCounts runner.Counts `json:"fallback_counts"` + Record []Entry `json:"record"` } // RecordLane is one lane of the run record. @@ -75,13 +82,20 @@ type RecordVerdict struct { Pass bool `json:"pass"` // Return is the validator's return the verdict was parsed from. Return string `json:"return"` + // Route is the runner route that ran the validator; absent when the host + // ran it. + Route *runner.RouteRecord `json:"route,omitempty"` } // recordOf renders a run's state as its record. func recordOf(st State) RunRecord { rec := RunRecord{RunID: st.RunID, Key: st.Key, Intent: st.Intent, Spec: st.Spec, Driver: st.Driver, Complete: st.Complete(), CreatedAt: st.CreatedAt, UpdatedAt: st.UpdatedAt, Pace: st.Pace, - Lanes: []RecordLane{}, Pending: st.Pending, Transcripts: st.Transcripts, Record: st.Record} + Lanes: []RecordLane{}, Pending: st.Pending, Transcripts: st.Transcripts, Record: st.Record, + Fallbacks: st.Fallbacks, FallbackCounts: runner.Tally(st.Fallbacks)} + if rec.Fallbacks == nil { + rec.Fallbacks = []runner.FallbackReceipt{} + } if rec.Pending == nil { rec.Pending = []PendingStep{} } @@ -104,7 +118,7 @@ func recordOf(st State) RunRecord { continue } rl.Verdicts = append(rl.Verdicts, RecordVerdict{Round: r.Round, HeadSHA: r.HeadSHA, Role: v.Role, - Verdict: v.Verdict, Pass: v.Pass, Return: v.Return}) + Verdict: v.Verdict, Pass: v.Pass, Return: v.Return, Route: v.Route}) } } for _, r := range l.Resolves { diff --git a/internal/core/implement/loop/stage_test.go b/internal/core/implement/loop/stage_test.go index 89ee5bfe2..e877fe4c2 100644 --- a/internal/core/implement/loop/stage_test.go +++ b/internal/core/implement/loop/stage_test.go @@ -27,7 +27,7 @@ func TestTheStateAndTheResultNameTheLaneStage(t *testing.T) { t.Fatal(err) } f := &fakeSteps{calls: map[Stage]int{}} - res, err := Advance(repo.Root(), start.RunID, f.steps(), Options{}) + res, err := advance(repo.Root(), start.RunID, f.steps(), Options{}) if err != nil { t.Fatal(err) } @@ -135,7 +135,7 @@ func TestAVersionThreeStateIsMigratedOnReadAndNeverRewrittenByTheRead(t *testing } f := &fakeSteps{calls: map[Stage]int{}} - res, err := Advance(repo.Root(), start.RunID, f.steps(), Options{}) + res, err := advance(repo.Root(), start.RunID, f.steps(), Options{}) if err != nil || res.PerformedStage != StageBrief { t.Fatalf("a migrated run steps on from where it stood: %+v %v", res, err) } diff --git a/internal/core/implement/loop/state.go b/internal/core/implement/loop/state.go index 89d79a9e2..4669f36ae 100644 --- a/internal/core/implement/loop/state.go +++ b/internal/core/implement/loop/state.go @@ -33,10 +33,10 @@ // // and its worktree in the machine-scoped store, // ~/.abcd/worktrees//-. The -// process driver (piece 3, waiting on the runner intent itd-2609201916056194) -// is the same loop called by a process instead of a host: it starts the agent an -// Await names through the runner and then calls Receipt, so it needs no seam -// beyond the two this package exports. +// process driver (piece 3, drive.go) is the same loop: Drive performs the next +// stage and, when it hands the lane to a role routed to a command-line runner +// (itd-2609201916056194), starts that agent through the runner and hands its +// receipt back through Receipt. // // Core never writes to stdout; the CLI front door formats what these functions // return. @@ -53,6 +53,7 @@ import ( "time" "github.com/intentdriven/abcd/internal/core/jsonstrict" + "github.com/intentdriven/abcd/internal/core/runner" "github.com/intentdriven/abcd/internal/fsutil" ) @@ -113,7 +114,18 @@ const lockFileName = ".lock" // `transcripts`. Version 6 is its strict subset, read as a run nothing has // landed yet and written back at version 7; a version-6 file carrying any of // them is not one version 6 wrote, and is refused. -const SchemaVersion = 7 +// +// Version 8 added the runner's record (itd-2609201916056194, spc-2609202134338445 +// piece 3): the run's `fallbacks`, one receipt per role a runner did not run, +// and the `route` a verified receipt or a validator's return names when a +// runner, not the host, ran its agent. Version 7 is its strict subset, read as +// a run the host ran every agent of and written back at version 8; a version-7 +// file carrying either is not one version 7 wrote, and is refused. +const SchemaVersion = 8 + +// schemaVersionUnrouted is the version before the runner's record: read, never +// written. +const schemaVersionUnrouted = 7 // schemaVersionUnlanded is the version before the landing: read, never // written. @@ -233,6 +245,11 @@ type State struct { // Transcripts are the transcripts the run's record captured into the // history store once the run was complete, one capture per path (piece 10). Transcripts []Transcript `json:"transcripts,omitempty"` + // Fallbacks are the run's fallback receipts, one per role a runner it was + // routed to did not run (itd-2609201916056194 criterion 3): the role, the + // runner asked for, the reason and the route that ran instead. The run's + // summary counts them per runner and per role (runner.Tally). + Fallbacks []runner.FallbackReceipt `json:"fallbacks,omitempty"` } // Transcript is one transcript the run record captured into the history store. @@ -313,6 +330,10 @@ type ReceiptRecord struct { // Model is the model the runner reported, as reported; empty when it // reported none. The binary cannot verify it. Model string `json:"model,omitempty"` + // Route is the route that ran the agent when a runner ran it: the runner + // asked for, the one that ran and the model it reported. Absent when the + // host ran it, so a host-run receipt reads as it always has. + Route *runner.RouteRecord `json:"route,omitempty"` } // HandBack is a lane stopped and handed back to the person, with what the last @@ -371,6 +392,10 @@ type ValidatorRun struct { // Audit is the fidelity audit's request and reading, on the // intent-auditor's run. Audit *AuditRun `json:"audit,omitempty"` + // Route is the route that ran the validator when a runner ran it; absent + // when the host ran it. It is the one field a runner-run review's record + // differs in from a host-run one's (itd-2609201916056194 criterion 7). + Route *runner.RouteRecord `json:"route,omitempty"` } // AuditRun is the fidelity audit the lane that closes the spec takes, once, over @@ -471,6 +496,29 @@ func (s State) landed() bool { return false } +// routed reports whether the state carries anything only a version-8 run +// writes: a fallback receipt, or a runner's route on a receipt or a return. +func (s State) routed() bool { + if len(s.Fallbacks) > 0 { + return true + } + for _, l := range s.Lanes { + for _, r := range l.Receipts { + if r.Route != nil { + return true + } + } + for _, v := range l.Validation { + for _, vr := range v.Validators { + if vr.Route != nil { + return true + } + } + } + } + return false +} + // validated reports whether the state carries anything only the validate // stage writes. func (s State) validated() bool { @@ -572,7 +620,10 @@ func readStateIn(root *os.Root, runID string) (State, error) { case st.SchemaVersion <= schemaVersionUnlanded && st.landed(): return State{}, refuse("state", "", "", fmt.Sprintf("%s is schema version %d but carries a landing, a verified receipt or a captured transcript, which version %d never wrote", rel, st.SchemaVersion, st.SchemaVersion), "the loop is the file's only writer; restore it or remove the run directory "+runRel(runID)) - case st.SchemaVersion >= schemaVersionUnpaced && st.SchemaVersion <= schemaVersionUnlanded: + case st.SchemaVersion <= schemaVersionUnrouted && st.routed(): + return State{}, refuse("state", "", "", fmt.Sprintf("%s is schema version %d but carries a fallback or a runner's route, which version %d never wrote", rel, st.SchemaVersion, st.SchemaVersion), + "the loop is the file's only writer; restore it or remove the run directory "+runRel(runID)) + case st.SchemaVersion >= schemaVersionUnpaced && st.SchemaVersion <= schemaVersionUnrouted: // Read as the current version, its stages already carried over by // decodeState when it named them `step`; the next write carries it, and // this read writes nothing. diff --git a/internal/core/implement/loop/validate_test.go b/internal/core/implement/loop/validate_test.go index 1df6d8abc..96556301d 100644 --- a/internal/core/implement/loop/validate_test.go +++ b/internal/core/implement/loop/validate_test.go @@ -41,7 +41,7 @@ func stepTo(t *testing.T, repo *gittest.Repo, runID string, stages Stages, want if i := st.current(); i >= 0 && st.Lanes[i].Stage == want && st.Lanes[i].Awaiting == nil { return st.Lanes[i] } - if _, err := Advance(repo.Root(), runID, stages, Options{}); err != nil { + if _, err := advance(repo.Root(), runID, stages, Options{}); err != nil { t.Fatalf("advancing to %s: %v", want, err) } } @@ -68,7 +68,7 @@ func currentLane(t *testing.T, repo *gittest.Repo, runID string) Lane { func implemented(t *testing.T, repo *gittest.Repo, runID string, stages Stages, file string) Lane { t.Helper() stepTo(t, repo, runID, stages, StageImplement) - res, err := Advance(repo.Root(), runID, stages, Options{}) + res, err := advance(repo.Root(), runID, stages, Options{}) if err != nil || res.Awaiting == nil || res.Awaiting.Role != RoleImplementer { t.Fatalf("the implement stage awaits an implementer: %+v %v", res, err) } @@ -134,7 +134,7 @@ func auditorVerdict(t *testing.T, request, verdict string) string { // hands back body as that agent's return. It returns the receipt's result. func handBack(t *testing.T, repo *gittest.Repo, runID string, stages Stages, role, body string) StepResult { t.Helper() - res, err := Advance(repo.Root(), runID, stages, Options{}) + res, err := advance(repo.Root(), runID, stages, Options{}) if err != nil { t.Fatal(err) } @@ -193,7 +193,7 @@ func TestTheValidatorsAreFreshAgentsAndOnlyTheLoopRecordsAVerdict(t *testing.T) id, stages := start.RunID, DefaultStages() l := implemented(t, repo, id, stages, "one.txt") - res, err := Advance(repo.Root(), id, stages, Options{}) + res, err := advance(repo.Root(), id, stages, Options{}) if err != nil || res.Awaiting == nil || res.Awaiting.Role != RoleRuthless { t.Fatalf("the first validator is a fresh ruthless reviewer: %+v %v", res, err) } @@ -218,7 +218,7 @@ func TestTheValidatorsAreFreshAgentsAndOnlyTheLoopRecordsAVerdict(t *testing.T) handBack(t, repo, id, stages, RoleSecurity, reviewerReturn("APPROVE")) handBack(t, repo, id, stages, RoleAuditor, "MET") - done, err := Advance(repo.Root(), id, stages, Options{}) + done, err := advance(repo.Root(), id, stages, Options{}) if err != nil || done.PerformedStage != StageValidate || done.Stage != StageLand { t.Fatalf("a passing round completes the validate stage: %+v %v", done, err) } @@ -237,7 +237,7 @@ func TestTheValidatorsAreFreshAgentsAndOnlyTheLoopRecordsAVerdict(t *testing.T) t.Fatalf("the audit's receipt and range are the run's, for the close to consume: %+v", a) } // The landing (piece 9) takes the lane from here, one step at a time. - res, err = Advance(repo.Root(), id, stages, Options{}) + res, err = advance(repo.Root(), id, stages, Options{}) if err != nil || res.Stage != StageLand || res.PerformedStage != "" { t.Fatalf("the landing's first step leaves the lane at land: %+v %v", res, err) } @@ -257,7 +257,7 @@ func TestTheAuditRunsOnceOnTheClosingLaneOverTheWholeDelivery(t *testing.T) { first := implemented(t, repo, id, stages, "one.txt") passRound(t, repo, id, stages, RoleRuthless, RoleSecurity) - res, err := Advance(repo.Root(), id, stages, Options{}) + res, err := advance(repo.Root(), id, stages, Options{}) if err != nil || res.PerformedStage != StageValidate { t.Fatalf("a lane that does not close the spec completes its validation without an audit: %+v %v", res, err) } @@ -268,7 +268,7 @@ func TestTheAuditRunsOnceOnTheClosingLaneOverTheWholeDelivery(t *testing.T) { if v := st.Lanes[0].Validation[0].Validators; len(v) != 2 { t.Fatalf("the non-closing lane's validators are the two reviewers: %+v", v) } - if _, err := Advance(repo.Root(), id, stages, Options{}); err != nil { + if _, err := advance(repo.Root(), id, stages, Options{}); err != nil { t.Fatal(err) } @@ -294,7 +294,7 @@ func TestTheAuditRunsOnceOnTheClosingLaneOverTheWholeDelivery(t *testing.T) { if a == nil || a.BaseSHA != st.Lanes[0].BaseSHA || a.HeadSHA != closing.HeadSHA { t.Fatalf("the audit's range is the whole delivery: %+v", a) } - if res, err := Advance(repo.Root(), id, stages, Options{}); err != nil || res.PerformedStage != StageValidate { + if res, err := advance(repo.Root(), id, stages, Options{}); err != nil || res.PerformedStage != StageValidate { t.Fatalf("the closing lane's passing round completes its validation: %+v %v", res, err) } } @@ -303,7 +303,7 @@ func TestTheAuditRunsOnceOnTheClosingLaneOverTheWholeDelivery(t *testing.T) { // hands back its receipt, naming commits (the lane's own when it made none). func fixed(t *testing.T, repo *gittest.Repo, runID string, stages Stages, report string, commits ...string) { t.Helper() - res, err := Advance(repo.Root(), runID, stages, Options{}) + res, err := advance(repo.Root(), runID, stages, Options{}) if err != nil || res.Awaiting == nil || res.Awaiting.Role != RoleImplementer { t.Fatalf("a round that did not pass hands its findings to a fresh implementer: %+v %v", res, err) } @@ -372,7 +372,7 @@ func TestAFindingIsAppliedByAFreshImplementerOrRejectedInWriting(t *testing.T) { // Round 3 judges the same head again, and passes. passRound(t, repo, id, stages, RoleRuthless, RoleSecurity, RoleAuditor) - if res, err := Advance(repo.Root(), id, stages, Options{}); err != nil || res.PerformedStage != StageValidate { + if res, err := advance(repo.Root(), id, stages, Options{}); err != nil || res.PerformedStage != StageValidate { t.Fatalf("a passing round completes the stage: %+v %v", res, err) } st, _ := ReadState(repo.Root(), id) @@ -402,7 +402,7 @@ func TestAReportCarryingAVerdictIsRefusedAtTheAdvance(t *testing.T) { } passRound(t, repo, id, stages, RoleRuthless, RoleSecurity, RoleAuditor) before := stateBytes(t, repo.Root(), id) - _, err = Advance(repo.Root(), id, stages, Options{}) + _, err = advance(repo.Root(), id, stages, Options{}) r := mustRefusal(t, err) if r.Stage != string(StageValidate) || !strings.Contains(r.Reason, RunRelDir+"/"+id+"/lane-1/"+ReportFileName) || !strings.Contains(r.Reason, "verdict") { t.Fatalf("the refusal names the report carrying a verdict: %+v", r) @@ -413,7 +413,7 @@ func TestAReportCarryingAVerdictIsRefusedAtTheAdvance(t *testing.T) { if err := os.WriteFile(report, []byte("built it; the reviewers judge it\n"), 0o600); err != nil { t.Fatal(err) } - if res, err := Advance(repo.Root(), id, stages, Options{}); err != nil || res.PerformedStage != StageValidate { + if res, err := advance(repo.Root(), id, stages, Options{}); err != nil || res.PerformedStage != StageValidate { t.Fatalf("the loop's own recorded SHIP lets the advance proceed: %+v %v", res, err) } }) @@ -448,7 +448,7 @@ func TestAReturnTheLoopCannotReadAVerdictFromIsRefused(t *testing.T) { if tc.role == RoleAuditor { passRound(t, repo, id, stages, RoleSecurity) } - res, err := Advance(repo.Root(), id, stages, Options{}) + res, err := advance(repo.Root(), id, stages, Options{}) if err != nil || res.Awaiting == nil || res.Awaiting.Role != tc.role { t.Fatalf("want the %s handed out: %+v %v", tc.role, res, err) } diff --git a/internal/core/runner/adapter_test.go b/internal/core/runner/adapter_test.go new file mode 100644 index 000000000..c0739edb8 --- /dev/null +++ b/internal/core/runner/adapter_test.go @@ -0,0 +1,456 @@ +package runner + +import ( + "context" + "crypto/rand" + "encoding/hex" + "errors" + "os" + "path/filepath" + "runtime" + "slices" + "strconv" + "strings" + "syscall" + "testing" + "time" +) + +// TestClaudeLaunchPassesTheBareFlag is criterion 6: the claude CLI runner +// launches in print mode with --bare, so the target repository's hooks and +// configured servers do not run, and grants the role's contract's tools +// without a prompt. +func TestClaudeLaunchPassesTheBareFlag(t *testing.T) { + f := newFake(t, "ok", Claude) + ans, transcript, err := newClaude("").Run(context.Background(), f.request("ruthless-reviewer")) + if err != nil { + t.Fatalf("run: %v", err) + } + argv := f.argv(t, Claude) + for _, want := range []string{"--print", "--bare", "--output-format", "stream-json", "--no-session-persistence", + "--permission-mode", "dontAsk", "--allowedTools=Read,Grep"} { + if !slices.Contains(argv, want) { + t.Errorf("claude argv %q lacks %q", argv, want) + } + } + if argv[len(argv)-2] != "--" { + t.Errorf("the prompt is not behind the end-of-options marker: %q", argv) + } + if ans.SessionID != "fake-session-1" || ans.Model != "fake-model" || ans.Text != "done" { + t.Errorf("answer = %+v", ans) + } + if !strings.Contains(string(transcript), `"type":"result"`) { + t.Errorf("transcript does not carry the event stream: %q", transcript) + } +} + +// TestClaudeModelRouteReachesTheLaunch: a configured model route is passed as +// the model the harness asks for, as one argument. +func TestClaudeModelRouteReachesTheLaunch(t *testing.T) { + f := newFake(t, "ok", Claude) + if _, _, err := newClaude("local/qwen3-coder").Run(context.Background(), f.request("scribe")); err != nil { + t.Fatalf("run: %v", err) + } + if argv := f.argv(t, Claude); !slices.Contains(argv, "--model=qwen3-coder") { + t.Errorf("claude argv %q lacks the model", argv) + } +} + +// TestOpenCodeLaunch: the opencode runner uses run mode with raw JSON events, +// no external plugins, the repository as its directory, and the brief attached. +func TestOpenCodeLaunch(t *testing.T) { + f := newFake(t, "ok", OpenCode) + req := f.request("ruthless-reviewer") + ans, _, err := newOpenCode("openrouter/qwen/qwen3-coder").Run(context.Background(), req) + if err != nil { + t.Fatalf("run: %v", err) + } + argv := f.argv(t, OpenCode) + if argv[0] != "run" { + t.Errorf("opencode argv %q does not start with run", argv) + } + for _, want := range []string{"--format=json", "--pure", "--dir=" + req.Dir, "--file=" + req.Brief, + "--model=openrouter/qwen/qwen3-coder"} { + if !slices.Contains(argv, want) { + t.Errorf("opencode argv %q lacks %q", argv, want) + } + } + if argv[len(argv)-2] != "--" { + t.Errorf("the prompt is not behind the end-of-options marker: %q", argv) + } + if ans.SessionID != "ses_fake1" || ans.Text != "done" || ans.Model != "openrouter/qwen/qwen3-coder" { + t.Errorf("answer = %+v", ans) + } +} + +// TestSameBriefAndContractEveryRoute is criterion 1's input half: every +// runner is handed the one prompt prompt renders, naming the role, the brief +// and the receipt path the host sub-agent is handed. +func TestSameBriefAndContractEveryRoute(t *testing.T) { + f := newFake(t, "ok", Claude, OpenCode) + req := f.request("security-reviewer") + if _, _, err := newClaude("").Run(context.Background(), req); err != nil { + t.Fatal(err) + } + if _, _, err := newOpenCode("").Run(context.Background(), req); err != nil { + t.Fatal(err) + } + want := prompt(req) + for _, h := range []string{Claude, OpenCode} { + argv := f.argv(t, h) + if got := argv[len(argv)-1]; got != want { + t.Errorf("%s prompt = %q, want %q", h, got, want) + } + } + for _, part := range []string{"security-reviewer", "Brief: " + req.Brief, "Receipt: " + req.Receipt} { + if !strings.Contains(want, part) { + t.Errorf("prompt %q lacks %q", want, part) + } + } +} + +// TestFailureKinds: each way a harness can fail is a Failure naming its kind. +func TestFailureKinds(t *testing.T) { + for _, tc := range []struct { + mode string + harness string + want Reason + }{ + {"exit1", Claude, ReasonFailed}, + {"garbage", Claude, ReasonUnparsable}, + {"refuse", Claude, ReasonRefused}, + {"exit1", OpenCode, ReasonFailed}, + {"garbage", OpenCode, ReasonUnparsable}, + {"refuse", OpenCode, ReasonRefused}, + } { + t.Run(tc.harness+"-"+tc.mode, func(t *testing.T) { + f := newFake(t, tc.mode, tc.harness) + var r Runner = newClaude("") + if tc.harness == OpenCode { + r = newOpenCode("") + } + _, _, err := r.Run(context.Background(), f.request("scribe")) + var fl *Failure + if !errors.As(err, &fl) || fl.Reason != tc.want { + t.Fatalf("err = %v, want a %s failure", err, tc.want) + } + }) + } +} + +// TestAModelPastItsShapeIsARefusalOfTheRoute: the model a harness's init +// event reports is bounded and shaped as a session id is. One past the bound, +// or one carrying control bytes, fails the route (a fallback, recorded), +// never an answer that carries it on into the run's state; a real id with a +// bracketed suffix is an answer. +func TestAModelPastItsShapeIsARefusalOfTheRoute(t *testing.T) { + for _, mode := range []string{"hugemodel", "ctrlmodel"} { + t.Run(mode, func(t *testing.T) { + f := newFake(t, mode, Claude) + ans, _, err := newClaude("").Run(context.Background(), f.request("scribe")) + var fl *Failure + if !errors.As(err, &fl) || fl.Reason != ReasonUnparsable { + t.Fatalf("err = %v, want an unparsable failure", err) + } + if ans.Model != "" || strings.Contains(fl.Detail, "\x1b") || len(fl.Detail) > 200 { + t.Fatalf("the model travelled on: answer %d bytes, detail %q", len(ans.Model), fl.Detail) + } + }) + } + f := newFake(t, "suffixmodel", Claude) + ans, _, err := newClaude("").Run(context.Background(), f.request("scribe")) + if err != nil || ans.Model != "claude-opus-4-6[1m]" { + t.Fatalf("a real model id is an answer: %+v %v", ans, err) + } +} + +// TestAHarnessOthersCanWriteIsRefused: a harness binary, or the directory it +// is reached through, that group or other can write is refused before launch +// and named as such: anyone who can write it chooses what runs. +func TestAHarnessOthersCanWriteIsRefused(t *testing.T) { + t.Run("directory", func(t *testing.T) { + f := newFake(t, "ok", Claude) + if err := os.Chmod(f.bin, 0o777); err != nil { + t.Fatal(err) + } + assertWritableRefused(t, f) + }) + t.Run("binary", func(t *testing.T) { + f := newFake(t, "ok") + self, _ := os.Executable() + raw, err := os.ReadFile(self) + if err != nil { + t.Fatal(err) + } + bin := filepath.Join(f.bin, Claude) + if err := os.WriteFile(bin, raw, 0o700); err != nil { + t.Fatal(err) + } + if err := os.Chmod(bin, 0o772); err != nil { + t.Fatal(err) + } + assertWritableRefused(t, f) + }) +} + +// TestAHarnessOnlyTheAdministratorGroupCanWriteIsAdmitted: a directory PATH +// reaches a harness through that is group-writable, never other-writable, and +// whose group is darwin's admin group (gid 80: the Homebrew /opt/homebrew/bin +// shape) is admitted on darwin, since members of that group can sudo by +// default. gid 0 is refused on every OS (being in Linux's root group or +// darwin's wheel does not by itself let a member act as root), gid 80 is +// refused off darwin, and any other-writable directory is refused whatever +// its group. +func TestAHarnessOnlyTheAdministratorGroupCanWriteIsAdmitted(t *testing.T) { + var admin []uint32 + nonAdmin := []uint32{0, 20, 1000} + if runtime.GOOS == "darwin" { + admin = append(admin, 80) + } else { + nonAdmin = append(nonAdmin, 80) + } + withGroup := func(t *testing.T, gid uint32) { + t.Helper() + prev := fileGroup + fileGroup = func(os.FileInfo) (uint32, bool) { return gid, true } + t.Cleanup(func() { fileGroup = prev }) + } + for _, gid := range admin { + t.Run("0775 gid "+strconv.Itoa(int(gid))+" admits", func(t *testing.T) { + f := newFake(t, "ok", Claude) + if err := os.Chmod(f.bin, 0o775); err != nil { + t.Fatal(err) + } + withGroup(t, gid) + if _, _, err := newClaude("").Run(context.Background(), f.request("scribe")); err != nil { + t.Fatalf("run: %v, want a harness only the administrator group can write admitted", err) + } + if !f.launched(Claude) { + t.Fatal("the harness was not launched") + } + }) + } + for _, gid := range []uint32{0, 80} { + t.Run("0777 gid "+strconv.Itoa(int(gid))+" refuses", func(t *testing.T) { + f := newFake(t, "ok", Claude) + if err := os.Chmod(f.bin, 0o777); err != nil { + t.Fatal(err) + } + withGroup(t, gid) + assertWritableRefused(t, f) + }) + } + for _, gid := range nonAdmin { + t.Run("0775 gid "+strconv.Itoa(int(gid))+" refuses", func(t *testing.T) { + f := newFake(t, "ok", Claude) + if err := os.Chmod(f.bin, 0o775); err != nil { + t.Fatal(err) + } + withGroup(t, gid) + assertWritableRefused(t, f) + }) + } + t.Run("0775 unknown group refuses", func(t *testing.T) { + f := newFake(t, "ok", Claude) + if err := os.Chmod(f.bin, 0o775); err != nil { + t.Fatal(err) + } + prev := fileGroup + fileGroup = func(os.FileInfo) (uint32, bool) { return 0, false } + t.Cleanup(func() { fileGroup = prev }) + assertWritableRefused(t, f) + }) +} + +func assertWritableRefused(t *testing.T, f fakeEnv) { + t.Helper() + _, _, err := newClaude("").Run(context.Background(), f.request("scribe")) + var fl *Failure + if !errors.As(err, &fl) || fl.Reason != ReasonAbsent || !strings.Contains(fl.Detail, "group or other can write") { + t.Fatalf("err = %v, want the writable harness refused as absent, naming why", err) + } + if f.launched(Claude) { + t.Fatal("a harness others can write was launched") + } +} + +// TestAbsentBinaryIsAbsent: a harness that is not on PATH is absent, and +// nothing runs. +func TestAbsentBinaryIsAbsent(t *testing.T) { + f := newFake(t, "ok") + _, _, err := newClaude("").Run(context.Background(), f.request("scribe")) + var fl *Failure + if !errors.As(err, &fl) || fl.Reason != ReasonAbsent { + t.Fatalf("err = %v, want absent", err) + } +} + +// TestBinaryInsideTheRepositoryIsRefused: a PATH entry that resolves inside +// the repository the role runs in is repository content, never run. +func TestBinaryInsideTheRepositoryIsRefused(t *testing.T) { + f := newFake(t, "ok") + self, _ := os.Executable() + planted := filepath.Join(f.repo, "tools") + if err := os.Mkdir(planted, 0o700); err != nil { + t.Fatal(err) + } + raw, err := os.ReadFile(self) + if err != nil { + t.Fatal(err) + } + if err := os.WriteFile(filepath.Join(planted, Claude), raw, 0o700); err != nil { + t.Fatal(err) + } + t.Setenv("PATH", planted) + _, _, err = newClaude("").Run(context.Background(), f.request("scribe")) + var fl *Failure + if !errors.As(err, &fl) || fl.Reason != ReasonAbsent { + t.Fatalf("err = %v, want the planted binary refused as absent", err) + } + if f.launched(Claude) { + t.Fatal("the binary inside the repository was launched") + } +} + +// TestLaunchEnvironmentIsScrubbed: an inherited repository-selection or config +// injection variable never reaches the harness, and the launch runs in the +// repository directory with no stdin. +func TestLaunchEnvironmentIsScrubbed(t *testing.T) { + f := newFake(t, "ok", Claude) + t.Setenv("GIT_DIR", "/elsewhere/.git") + t.Setenv("GIT_CONFIG_PARAMETERS", "'core.hooksPath'='/elsewhere'") + if _, _, err := newClaude("").Run(context.Background(), f.request("scribe")); err != nil { + t.Fatal(err) + } + env, _ := os.ReadFile(filepath.Join(f.log, Claude+".env")) + for _, bad := range []string{"GIT_DIR=", "GIT_CONFIG_PARAMETERS="} { + if strings.Contains(string(env), "\n"+bad) || strings.HasPrefix(string(env), bad) { + t.Errorf("the harness inherited %s", bad) + } + } + cwd, _ := os.ReadFile(filepath.Join(f.log, Claude+".cwd")) + want, _ := filepath.EvalSymlinks(f.repo) + if got, _ := filepath.EvalSymlinks(string(cwd)); got != want { + t.Errorf("harness ran in %q, want %q", got, want) + } +} + +// TestNoCredentialInArgumentOrError: the harness's own credential stays in its +// environment; abcd never puts it on the command line, and a harness that +// echoes it while failing does not carry it into the error abcd returns. +func TestNoCredentialInArgumentOrError(t *testing.T) { + f := newFake(t, "exit1", Claude) + b := make([]byte, 12) + _, _ = rand.Read(b) + cred := "abcd-fake-credential-" + hex.EncodeToString(b) + t.Setenv(fakeCredEnv, cred) + _, _, err := newClaude("").Run(context.Background(), f.request("scribe")) + if err == nil { + t.Fatal("a failing harness returned no error") + } + if strings.Contains(err.Error(), cred) { + t.Errorf("the error carries the credential: %v", err) + } + for _, a := range f.argv(t, Claude) { + if strings.Contains(a, cred) { + t.Errorf("an argument carries the credential: %q", a) + } + } + env, _ := os.ReadFile(filepath.Join(f.log, Claude+".env")) + if !strings.Contains(string(env), fakeCredEnv+"="+cred) { + t.Error("the harness did not receive its own credential from the environment") + } +} + +// TestTimeoutKillsTheProcessGroup: a harness past its time is killed with +// every process in the group the runner started, and reported as failed. +func TestTimeoutKillsTheProcessGroup(t *testing.T) { + f := newFake(t, "hang", Claude) + req := f.request("scribe") + req.Timeout = 2 * time.Second + _, _, err := newClaude("").Run(context.Background(), req) + var fl *Failure + if !errors.As(err, &fl) || fl.Reason != ReasonFailed || !strings.Contains(fl.Detail, "time") { + t.Fatalf("err = %v, want a failure for running out of time", err) + } + raw, rerr := os.ReadFile(filepath.Join(f.log, "child.pid")) + if rerr != nil { + t.Fatalf("the fake never started its child: %v", rerr) + } + pid, _ := strconv.Atoi(string(raw)) + deadline := time.Now().Add(10 * time.Second) + for syscall.Kill(pid, 0) == nil { + if time.Now().After(deadline) { + _ = syscall.Kill(pid, syscall.SIGKILL) // our own fake's child, by its pid + t.Fatalf("the harness's child %d outlived the kill", pid) + } + time.Sleep(50 * time.Millisecond) + } +} + +// TestOutputIsBounded: a harness writing past the bound is cut off and +// reported, never read into memory whole. +func TestOutputIsBounded(t *testing.T) { + f := newFake(t, "flood", Claude) + c := newClaude("") + c.launch.maxStdout = 64 << 10 + _, transcript, err := c.Run(context.Background(), f.request("scribe")) + var fl *Failure + if !errors.As(err, &fl) || !strings.Contains(fl.Detail, "bound") { + t.Fatalf("err = %v, want a failure naming the output bound", err) + } + if len(transcript) > 64<<10+4096 { + t.Errorf("transcript is %d bytes, past the bound", len(transcript)) + } +} + +// TestRequestIsCheckedBeforeLaunch: a relative path, a tool list that would +// split the flag, or a role that is not a plain name refuses before anything +// runs. +func TestRequestIsCheckedBeforeLaunch(t *testing.T) { + f := newFake(t, "ok", Claude) + for name, mut := range map[string]func(*Request){ + "relative brief": func(r *Request) { r.Brief = "brief.md" }, + "relative dir": func(r *Request) { r.Dir = "repo" }, + "comma tool": func(r *Request) { r.Tools = []string{"Read,Bash"} }, + "role": func(r *Request) { r.Role = "../x" }, + "no session": func(r *Request) { r.SessionID = "" }, + } { + req := f.request("scribe") + mut(&req) + if _, _, err := newClaude("").Run(context.Background(), req); err == nil { + t.Errorf("%s: accepted", name) + } + } + if f.launched(Claude) { + t.Error("a refused request launched the harness") + } +} + +// TestBinaryInsideTheCheckoutIsRefused: a lane's worktree lives outside the +// checkout the run belongs to, so a program planted in that checkout and put +// on PATH is refused too, though it is not inside the directory the role +// runs in. +func TestBinaryInsideTheCheckoutIsRefused(t *testing.T) { + f := newFake(t, "ok") + self, _ := os.Executable() + checkout := t.TempDir() + planted := filepath.Join(checkout, "tools") + if err := os.Mkdir(planted, 0o700); err != nil { + t.Fatal(err) + } + if err := os.Symlink(self, filepath.Join(planted, OpenCode)); err != nil { + t.Fatal(err) + } + t.Setenv("PATH", planted) + req := f.request("scribe") + req.Checkout = checkout + _, _, err := newOpenCode("").Run(context.Background(), req) + var fl *Failure + if !errors.As(err, &fl) || fl.Reason != ReasonAbsent { + t.Fatalf("err = %v, want the binary planted in the checkout refused as absent", err) + } + if f.launched(OpenCode) { + t.Fatal("the binary inside the checkout was launched") + } +} diff --git a/internal/core/runner/claude.go b/internal/core/runner/claude.go new file mode 100644 index 000000000..daeceb13e --- /dev/null +++ b/internal/core/runner/claude.go @@ -0,0 +1,148 @@ +package runner + +// claude.go is the claude CLI runner (spc-2609221533057881 scope 2, criterion +// 6): print mode with the bare flag, so the target repository's hooks, plugins, +// CLAUDE.md auto-discovery and configured servers do not run untrusted +// (the intent's Decision 4); the stream-json event stream as the transcript; +// no session persisted by the harness, since abcd's store keeps the record; +// and the role's contract's tools granted without a prompt, everything else +// denied rather than asked (dontAsk), because no one is there to answer. +// +// Under --bare the harness reads its Anthropic credential from its own +// environment or its own settings only; abcd passes the environment through +// and never names a key. + +import ( + "bytes" + "context" + "encoding/json" + "path/filepath" + "strings" +) + +// ClaudeCLI is the claude CLI runner. +type ClaudeCLI struct { + // model is the model route, /, or "" for the harness's + // own default. + model string + launch launcher +} + +// newClaude returns the claude CLI runner asking for model, a +// / route the configuration admitted, or "" for the harness's +// default. +func newClaude(model string) *ClaudeCLI { return &ClaudeCLI{model: model, launch: defaultLauncher()} } + +// Name is the route's name. +func (*ClaudeCLI) Name() string { return Claude } + +// args is the launch's argv after the binary. +func (c *ClaudeCLI) args(req Request) []string { + a := []string{ + "--print", "--bare", + "--output-format", "stream-json", "--verbose", + "--no-session-persistence", + "--permission-mode", "dontAsk", + } + if len(req.Tools) > 0 { + a = append(a, "--allowedTools="+strings.Join(req.Tools, ",")) + } + // The brief and the receipt sit in the run's lane directory, outside the + // working tree the role runs in. + dirs := []string{filepath.Dir(req.Brief)} + if d := filepath.Dir(req.Receipt); d != dirs[0] { + dirs = append(dirs, d) + } + for _, d := range dirs { + a = append(a, "--add-dir="+d) + } + if _, m, ok := strings.Cut(c.model, "/"); ok { + a = append(a, "--model="+m) + } + return append(a, "--", prompt(req)) +} + +// Run runs the role through the claude CLI. +func (c *ClaudeCLI) Run(ctx context.Context, req Request) (Answer, []byte, error) { + if err := req.check(); err != nil { + return Answer{}, nil, err + } + bin, err := c.launch.admit(Claude, "claude", req.Dir, req.Checkout) + if err != nil { + return Answer{}, nil, err + } + res, err := c.launch.run(ctx, Claude, bin, c.args(req), req.Dir, req.timeout()) + transcript := res.transcript() + if err != nil { + return Answer{}, transcript, err + } + ans, err := parseClaude(res.stdout) + return ans, transcript, err +} + +// claudeEvent is the part of a stream-json event the runner reads. +type claudeEvent struct { + Type string `json:"type"` + Subtype string `json:"subtype"` + SessionID string `json:"session_id"` + Model string `json:"model"` + IsError bool `json:"is_error"` + Result string `json:"result"` +} + +// parseClaude reads the event stream: one JSON object per line, the init +// event's model and session, and the closing result event. A line that is not +// an event, or a stream with no result, is unparsable; a result that reports +// an error is a refusal. +func parseClaude(out []byte) (Answer, error) { + var ans Answer + var result *claudeEvent + for _, line := range bytes.Split(out, []byte("\n")) { + line = bytes.TrimSpace(line) + if len(line) == 0 { + continue + } + var ev claudeEvent + if json.Unmarshal(line, &ev) != nil || ev.Type == "" { + return Answer{}, fail(Claude, ReasonUnparsable, "a line of its output is not a stream-json event") + } + if ev.Type == "system" && ev.Subtype == "init" { + if ev.Model != "" && !modelRe.MatchString(ev.Model) { + // Refused rather than trimmed: an answer is never recorded + // under a model the harness did not report. + return Answer{}, fail(Claude, ReasonUnparsable, "its init event reports a model that is not a plain model id of at most 128 characters") + } + ans.Model = ev.Model + } + if ans.SessionID == "" && ev.SessionID != "" { + ans.SessionID = ev.SessionID + } + if ev.Type == "result" { + e := ev + result = &e + } + } + if result == nil { + return Answer{}, fail(Claude, ReasonUnparsable, "its output carries no result event") + } + if result.IsError || result.Subtype != "success" { + return Answer{}, fail(Claude, ReasonRefused, "its result reports %s", safeWord(result.Subtype)) + } + ans.Text = result.Result + return ans, nil +} + +// safeWord renders a harness-supplied status word for a detail: a short token +// of plain characters, else "an error". Nothing longer from the harness +// reaches a detail. +func safeWord(s string) string { + if len(s) == 0 || len(s) > 40 { + return "an error" + } + for _, r := range s { + if !(r == '_' || r == '-' || r >= 'a' && r <= 'z' || r >= 'A' && r <= 'Z' || r >= '0' && r <= '9') { + return "an error" + } + } + return s +} diff --git a/internal/core/runner/config.go b/internal/core/runner/config.go new file mode 100644 index 000000000..d14300199 --- /dev/null +++ b/internal/core/runner/config.go @@ -0,0 +1,271 @@ +package runner + +// config.go reads the route per role and the runners this machine enabled, +// through the layered configuration resolver (the 2026-09-25 ruling that +// placed roles..runner in layered.Config): +// +// - roles..runner names host (the default when unset) or a runner. +// The repository and the machine may set it; the higher layer wins per +// role. A role no agent answers to is a diagnostic and is skipped, as the +// oracle's routes are. +// - runner. enables a shipped runner (claude, opencode) on this +// machine, with an optional model route, /, admitted +// against that provider's allowlist when the configuration is read +// (adr-2609221009491186; its Decision 2 superseded by adr-2609300107513982, +// so abcd bundles no vendor denylist and the allowlist alone decides, and +// an oracle.denylist entry the configuration writes still refuses a model +// it matches), so a route off the list is refused before any runner can +// be launched. +// - runner.fallback_host names the enabled runner that runs a role when abcd +// runs with no host session (the intent's Decision 2). +// +// The runner namespace is the machine's alone: which harness abcd starts, and +// on whose credential it runs, is the person's own machine's to say, as a +// provider block is (the product thinker's ruling AA(b) of 2026-09-29, which +// keeps a repository from spending the person's paid key). A repository that +// declares it is refused, naming the machine's file. A role a repository +// routes to a runner the machine did not enable runs nowhere new: the dispatch +// records it as an absent runner and falls back. + +import ( + "fmt" + "sort" + "strings" + + "github.com/intentdriven/abcd/internal/core/layered" + "github.com/intentdriven/abcd/internal/core/oracle" +) + +// The configuration keys, as every refusal names them. +const ( + rolesKey = "roles" + runnerKey = "runner" + fallbackKey = "fallback_host" + modelKey = "model" +) + +// roleImplementer is the loop's implementer, the one role outside the agent +// roster (the loop's RoleImplementer; spelled here so the loop can import this +// package). +const roleImplementer = "implementer" + +// Route is where a role runs and the layer that said so. +type Route struct { + Runner string + Layer layered.Layer + Origin string +} + +// RunnerConfig is one runner this machine enabled. +type RunnerConfig struct { + Name string + // Model is the admitted / route, "" for the harness's own + // default. + Model string + Origin string +} + +// Config is the runner configuration one invocation read. +type Config struct { + roles map[string]Route + runners map[string]RunnerConfig + fallback string + api *oracle.APIConfig + // Diagnostics are the non-fatal reports the read produced, one line each, + // for a front door to print on stderr. + Diagnostics []string +} + +// Load reads and validates the runner configuration. A fault is an error +// naming the file and the key; none falls through to a default. +func Load(r layered.Roots) (*Config, error) { + s, err := layered.Load(layered.Config, r) + if err != nil { + return nil, fmt.Errorf("runner: %w", err) + } + if err := s.Claim(rolesKey+".*", runnerKey); err != nil { + return nil, fmt.Errorf("runner: %w", err) + } + if err := s.Claim(runnerKey, append([]string{fallbackKey}, runnerNames...)...); err != nil { + return nil, fmt.Errorf("runner: %w", err) + } + for _, n := range runnerNames { + if err := s.Claim(runnerKey+"."+n, modelKey); err != nil { + return nil, fmt.Errorf("runner: %w", err) + } + } + found, err := s.Lookup(runnerKey) + if err != nil { + return nil, fmt.Errorf("runner: %w", err) + } + for _, f := range found { + if f.Layer != layered.Machine { + return nil, fmt.Errorf("runner: %s (%s layer): %s is set here, but which harness abcd starts, and on whose "+ + "credential, is this machine's alone to say; set it in %s and remove it from %s", + f.Origin, f.Layer, runnerKey, layered.Config.MachineOrigin(), f.Origin) + } + } + api, err := oracle.LoadAPI(r) + if err != nil { + return nil, fmt.Errorf("runner: the model routes are admitted against the provider configuration: %w", err) + } + c := &Config{roles: map[string]Route{}, runners: map[string]RunnerConfig{}, api: api} + if err := c.readRunners(s); err != nil { + return nil, err + } + if err := c.readFallback(s); err != nil { + return nil, err + } + if err := c.readRoles(s); err != nil { + return nil, err + } + return c, nil +} + +func (c *Config) readRunners(s *layered.Stack) error { + for _, n := range runnerNames { + found, err := s.Lookup(runnerKey + "." + n) + if err != nil { + return fmt.Errorf("runner: %w", err) + } + if len(found) == 0 { + continue + } + rc := RunnerConfig{Name: n, Origin: found[0].Origin} + key := runnerKey + "." + n + "." + modelKey + mf, err := s.Lookup(key) + if err != nil { + return fmt.Errorf("runner: %w", err) + } + if len(mf) > 0 { + where := fmt.Sprintf("%s (%s layer): %s", mf[0].Origin, mf[0].Layer, key) + text, err := layered.Decode[string](mf[0].Raw) + if err != nil { + return fmt.Errorf("runner: %s: %w", where, err) + } + if err := c.admitModel(text); err != nil { + return fmt.Errorf("runner: %s is %q, %w", where, layered.BoundKey(text), err) + } + rc.Model = text + } + c.runners[n] = rc + } + return nil +} + +// admitModel is the allowlist check a model route passes when it is read and +// again before a runner is launched. +func (c *Config) admitModel(route string) error { + provider, model, ok := strings.Cut(route, "/") + if !ok || provider == "" || model == "" { + return fmt.Errorf("which is not /") + } + if err := c.api.Admit(provider, model); err != nil { + return fmt.Errorf("which is refused before any runner is launched: %w", err) + } + return nil +} + +func (c *Config) readFallback(s *layered.Stack) error { + key := runnerKey + "." + fallbackKey + v, err := layered.Get(s, key, "", func(name string) error { + if _, ok := c.runners[name]; !ok { + return fmt.Errorf("the fallback host is a runner this machine enables under %s. (%s)", + runnerKey, strings.Join(runnerNames, ", ")) + } + return nil + }) + if err != nil { + return fmt.Errorf("runner: %w", err) + } + c.fallback = v.V + return nil +} + +func (c *Config) readRoles(s *layered.Stack) error { + names := map[string]bool{} + for _, l := range []layered.Layer{layered.Flag, layered.Repo, layered.Machine} { + ns, err := s.Members(l, rolesKey) + if err != nil { + return fmt.Errorf("runner: %w", err) + } + for _, n := range ns { + names[n] = true + } + } + sorted := make([]string, 0, len(names)) + for n := range names { + sorted = append(sorted, n) + } + sort.Strings(sorted) + known := map[string]bool{roleImplementer: true} + for _, a := range oracle.Roster() { + known[a] = true + } + for _, role := range sorted { + if !roleRe.MatchString(role) { + return fmt.Errorf("runner: %s.%s: the role is not a plain lower-case name", rolesKey, layered.BoundKey(role)) + } + key := rolesKey + "." + role + "." + runnerKey + found, err := s.Lookup(key) + if err != nil { + return fmt.Errorf("runner: %w", err) + } + if len(found) == 0 { + continue + } + win := found[0] + where := fmt.Sprintf("%s (%s layer): %s", win.Origin, win.Layer, key) + if !known[role] { + c.Diagnostics = append(c.Diagnostics, fmt.Sprintf("runner: %s names %q, which is neither an agent in the roster "+ + "nor the implementer; the route is skipped and the remaining routes apply", where, role)) + continue + } + name, err := layered.Decode[string](win.Raw) + if err != nil { + return fmt.Errorf("runner: %s: %w", where, err) + } + if name != Host && !isRunnerName(name) { + return fmt.Errorf("runner: %s is %q; a role runs on %s or one of the runners %s", + where, layered.BoundKey(name), Host, strings.Join(runnerNames, ", ")) + } + c.roles[role] = Route{Runner: name, Layer: win.Layer, Origin: win.Origin} + } + return nil +} + +func isRunnerName(n string) bool { + for _, r := range runnerNames { + if n == r { + return true + } + } + return false +} + +// RouteFor returns where role runs: its configured route, or the host from +// the bundled default. +func (c *Config) RouteFor(role string) Route { + if r, ok := c.roles[role]; ok { + return r + } + return Route{Runner: Host, Layer: layered.Bundled, Origin: "bundled"} +} + +// Runner returns the runner this machine enabled under name. +func (c *Config) Runner(name string) (RunnerConfig, bool) { + rc, ok := c.runners[name] + return rc, ok +} + +// FallbackHost is the runner that runs a role when there is no host session, +// "" when none is configured. +func (c *Config) FallbackHost() string { return c.fallback } + +// adapter builds the enabled runner's adapter. +func (c *Config) adapter(rc RunnerConfig) Runner { + if rc.Name == OpenCode { + return newOpenCode(rc.Model) + } + return newClaude(rc.Model) +} diff --git a/internal/core/runner/config_test.go b/internal/core/runner/config_test.go new file mode 100644 index 000000000..1b1aeabdb --- /dev/null +++ b/internal/core/runner/config_test.go @@ -0,0 +1,149 @@ +package runner + +import ( + "os" + "path/filepath" + "strings" + "testing" + + "github.com/intentdriven/abcd/internal/core/layered" +) + +// localProvider is a machine provider block that holds no key: the allowlist a +// runner's model route is admitted against. +const localProvider = `"oracle":{"api":{"local":{"base_url":"http://localhost:11434/v1","models":["qwen3-coder"]}}}` + +// roots writes the machine and repository config files (either may be "") and +// returns the roots that read them. +func roots(t *testing.T, machine, repo string) layered.Roots { + t.Helper() + r := layered.Roots{Repo: t.TempDir(), Home: t.TempDir()} + put := func(path, body string) { + if body == "" { + return + } + if err := os.MkdirAll(filepath.Dir(path), 0o700); err != nil { + t.Fatal(err) + } + if err := os.WriteFile(path, []byte(body), 0o600); err != nil { + t.Fatal(err) + } + } + put(filepath.Join(r.Home, ".abcd", "config.json"), machine) + put(filepath.Join(r.Repo, ".abcd", "config.json"), repo) + return r +} + +func mustLoad(t *testing.T, machine, repo string) *Config { + t.Helper() + c, err := Load(roots(t, machine, repo)) + if err != nil { + t.Fatalf("load: %v", err) + } + return c +} + +// TestUnsetRoleIsHost is criterion 2's configuration half: with nothing +// configured every role routes to the host, from the bundled default. +func TestUnsetRoleIsHost(t *testing.T) { + c := mustLoad(t, "", "") + r := c.RouteFor("ruthless-reviewer") + if r.Runner != Host || r.Layer != layered.Bundled { + t.Fatalf("route = %+v, want the bundled host", r) + } +} + +// TestRoleRoutedByRepoToMachineRunner: the repository names the runner for a +// role; the machine enables the runner and its model route. +func TestRoleRoutedByRepoToMachineRunner(t *testing.T) { + c := mustLoad(t, + `{`+localProvider+`,"runner":{"fallback_host":"claude","claude":{},"opencode":{"model":"local/qwen3-coder"}}}`, + `{"roles":{"ruthless-reviewer":{"runner":"opencode"}}}`) + r := c.RouteFor("ruthless-reviewer") + if r.Runner != OpenCode || r.Layer != layered.Repo || r.Origin != ".abcd/config.json" { + t.Fatalf("route = %+v", r) + } + rc, ok := c.Runner(OpenCode) + if !ok || rc.Model != "local/qwen3-coder" { + t.Fatalf("runner = %+v, %v", rc, ok) + } + if c.FallbackHost() != Claude { + t.Fatalf("fallback host = %q", c.FallbackHost()) + } +} + +// TestModelOffTheAllowlistIsRefused is criterion 5: a runner whose model +// route its provider does not list is refused when the configuration is read, +// before any runner can be launched, naming the list. +func TestModelOffTheAllowlistIsRefused(t *testing.T) { + for name, model := range map[string]string{ + "not listed": "local/llama-3", + "provider unknown": "elsewhere/qwen3-coder", + "not provider/model": "qwen3-coder", + } { + _, err := Load(roots(t, `{`+localProvider+`,"runner":{"opencode":{"model":"`+model+`"}}}`, "")) + if err == nil { + t.Errorf("%s: %q admitted", name, model) + continue + } + if !strings.Contains(err.Error(), "runner.opencode.model") { + t.Errorf("%s: refusal does not name the key: %v", name, err) + } + } +} + +// TestAllowlistAloneDecides follows adr-2609300107513982: abcd bundles no +// vendor denylist, so a vendor-prefixed model a provider lists is admitted, +// and only an oracle.denylist entry the configuration writes refuses it. +func TestAllowlistAloneDecides(t *testing.T) { + const listed = `"oracle":{"api":{"local":{"base_url":"http://localhost:11434/v1","models":["anthropic/claude-opus"]}}` + c := mustLoad(t, `{`+listed+`},"runner":{"opencode":{"model":"local/anthropic/claude-opus"}}}`, "") + if rc, ok := c.Runner(OpenCode); !ok || rc.Model != "local/anthropic/claude-opus" { + t.Fatalf("runner = %+v, %v; a listed model is admitted by the allowlist alone", rc, ok) + } + _, err := Load(roots(t, + `{`+listed+`,"denylist":["anthropic/*"]},"runner":{"opencode":{"model":"local/anthropic/claude-opus"}}}`, "")) + if err == nil || !strings.Contains(err.Error(), "oracle.denylist") { + t.Fatalf("err = %v, want the configured denylist entry to refuse the route", err) + } +} + +// TestRunnerBlocksAreTheMachines: a repository cannot enable a runner or set +// its model: which harness abcd launches, and on whose credential, is the +// person's own machine's to say. +func TestRunnerBlocksAreTheMachines(t *testing.T) { + _, err := Load(roots(t, "", `{"runner":{"claude":{}}}`)) + if err == nil || !strings.Contains(err.Error(), "~/.abcd/config.json") { + t.Fatalf("err = %v, want a refusal naming the machine's file", err) + } +} + +// TestConfigRefusals: every other fault is loud. +func TestConfigRefusals(t *testing.T) { + for name, tc := range map[string][2]string{ + "unknown runner": {`{"runner":{"claude":{}}}`, `{"roles":{"scribe":{"runner":"gemini"}}}`}, + "unknown role key": {"", `{"roles":{"scribe":{"runnr":"host"}}}`}, + "unknown runner key": {`{"runner":{"claude":{"binary":"/tmp/x"}}}`, ""}, + "fallback not enabled": {`{"runner":{"fallback_host":"opencode","claude":{}}}`, ""}, + "fallback host is host": {`{"runner":{"fallback_host":"host","claude":{}}}`, ""}, + "runner not a string": {"", `{"roles":{"scribe":{"runner":1}}}`}, + "unknown runner namespace": {`{"runner":{"aider":{}}}`, ""}, + } { + if _, err := Load(roots(t, tc[0], tc[1])); err == nil { + t.Errorf("%s: loaded", name) + } + } +} + +// TestRoleOutsideTheRosterIsADiagnostic: a role no agent answers to is named +// and skipped, as the oracle's routes are; the rest apply. +func TestRoleOutsideTheRosterIsADiagnostic(t *testing.T) { + c := mustLoad(t, `{"runner":{"claude":{}}}`, + `{"roles":{"not-an-agent":{"runner":"claude"},"implementer":{"runner":"claude"}}}`) + if len(c.Diagnostics) != 1 || !strings.Contains(c.Diagnostics[0], "not-an-agent") { + t.Fatalf("diagnostics = %q", c.Diagnostics) + } + if r := c.RouteFor("implementer"); r.Runner != Claude { + t.Fatalf("implementer route = %+v", r) + } +} diff --git a/internal/core/runner/dispatch.go b/internal/core/runner/dispatch.go new file mode 100644 index 000000000..63e3f890e --- /dev/null +++ b/internal/core/runner/dispatch.go @@ -0,0 +1,173 @@ +package runner + +// dispatch.go is the one place the fallback is decided (spc-2609221533057881 +// scope 4). A role routed to a runner runs through it; when the runner is +// absent, refuses, fails, answers unparsably, or its answer fails the +// contract's validator, the role goes to the host session, or with no host +// session to the host the operator configured, and exactly one receipt is +// written for the event. Because the branch and the receipt writer are one, +// the count the run's summary reports is every fallback there was. + +import ( + "context" + "errors" + "fmt" + "strings" + "time" + + "github.com/intentdriven/abcd/internal/termsafe" +) + +// maxDetail bounds a validator's refusal as a receipt carries it. +const maxDetail = 300 + +// Dispatcher runs roles by their routes. +type Dispatcher struct { + // Config is the runner configuration read at the lane's start. + Config *Config + // HostSession is true when a host session drives the loop and can take a + // role handed back to it. + HostSession bool + // Validate is the contract's validator: the same check the host sub-agent's + // answer passes. It is handed the runner that ran the role, so a caller + // whose validator also records the answer (the loop's receipt verifier) + // can name the route that ran. + Validate func(runner string, req Request, ans Answer) error + // Transcripts is where every runner's transcript lands. + Transcripts TranscriptStore + // Record appends one fallback receipt to the run's state. + Record func(FallbackReceipt) error + // Now is the clock the receipts are stamped with; time.Now when nil. + Now func() time.Time +} + +// Outcome is one dispatch's result. +type Outcome struct { + // Handoff is true when the host session runs the role, exactly as it does + // with no runner configured: the caller hands the host the brief and + // awaits its receipt. + Handoff bool + // Answer is the runner's parsed answer when a runner ran the role. + Answer *Answer + // Fallback is the receipt this dispatch recorded, nil when none. + Fallback *FallbackReceipt + // Receipt is the role's receipt block. + Receipt RoleReceipt +} + +// Dispatch runs req by its role's route. +func (d *Dispatcher) Dispatch(ctx context.Context, req Request) (Outcome, error) { + if d.Config == nil || d.Validate == nil || d.Transcripts == nil || d.Record == nil { + return Outcome{}, errors.New("runner: a dispatch needs the configuration, the contract's validator, " + + "the transcript store and the run's receipt writer") + } + if err := req.check(); err != nil { + return Outcome{}, err + } + route := d.Config.RouteFor(req.Role) + landing := Host + if !d.HostSession { + landing = d.Config.FallbackHost() + if landing == "" { + return Outcome{}, fmt.Errorf("runner: there is no host session and no %s.%s is configured, so a role "+ + "has nowhere to land; set it in the machine's config to the runner that stands in for the host", + runnerKey, fallbackKey) + } + } + out := Outcome{Receipt: newRoleReceipt(req, route.Runner)} + + if route.Runner == Host { + if d.HostSession { + out.Handoff = true + out.Receipt.Route.Ran = Host + return out, nil + } + ans, err := d.runOn(ctx, landing, req) + if err != nil { + return Outcome{}, fmt.Errorf("runner: the configured host %s did not run %s: %w", landing, req.Role, err) + } + return out.ran(landing, ans), nil + } + + ans, err := d.runOn(ctx, route.Runner, req) + if err == nil { + return out.ran(route.Runner, ans), nil + } + var fl *Failure + if !errors.As(err, &fl) { + return Outcome{}, err + } + fb := FallbackReceipt{At: d.now(), Role: req.Role, Asked: route.Runner, Reason: fl.Reason, Detail: fl.Detail, Ran: landing} + if landing == route.Runner { + fb.Ran = none + } + if rerr := d.Record(fb); rerr != nil { + return Outcome{}, fmt.Errorf("runner: the fallback receipt for %s was not recorded: %w", req.Role, rerr) + } + out.Fallback = &fb + if fb.Ran == none { + return Outcome{}, fmt.Errorf("runner: %s %s for %s and it is the configured host, so nothing else can run it: %s", + route.Runner, fl.Reason, req.Role, fl.Detail) + } + if d.HostSession { + out.Handoff = true + out.Receipt.Route.Ran = Host + return out, nil + } + ans, err = d.runOn(ctx, landing, req) + if err != nil { + return Outcome{}, fmt.Errorf("runner: %s %s for %s, and the configured host %s did not run it either: %w", + route.Runner, fl.Reason, req.Role, landing, err) + } + return out.ran(landing, ans), nil +} + +// none is the route a receipt names when nothing could run the role. +const none = "none" + +func (o Outcome) ran(route string, ans Answer) Outcome { + o.Answer = &ans + o.Receipt.Route.Ran = route + o.Receipt.Route.Model = ans.Model + return o +} + +func (d *Dispatcher) now() time.Time { + if d.Now != nil { + return d.Now() + } + return time.Now().UTC() +} + +// runOn runs req on the named runner, stores its transcript and validates its +// answer. A *Failure is a reason to fall back; any other error is not. +func (d *Dispatcher) runOn(ctx context.Context, name string, req Request) (Answer, error) { + rc, ok := d.Config.Runner(name) + if !ok { + return Answer{}, fail(name, ReasonAbsent, "it is not enabled on this machine (%s.%s in the machine's config)", runnerKey, name) + } + if rc.Model != "" { + // Admitted when the configuration was read; consulted again so a + // runner never launches on a stale answer. + if err := d.Config.admitModel(rc.Model); err != nil { + return Answer{}, fmt.Errorf("runner: %s.%s.%s is %q, %w", runnerKey, name, modelKey, rc.Model, err) + } + } + ans, transcript, err := d.Config.adapter(rc).Run(ctx, req) + if len(transcript) > 0 { + if serr := d.Transcripts.Store(name, req, ans, transcript); serr != nil { + return Answer{}, fmt.Errorf("runner: the %s transcript for %s was not stored: %w", name, req.Role, serr) + } + } + if err != nil { + return Answer{}, err + } + if verr := d.Validate(name, req, ans); verr != nil { + detail := termsafe.Sanitize(verr.Error()) + if len(detail) > maxDetail { + detail = strings.ToValidUTF8(detail[:maxDetail], "") + "..." + } + return Answer{}, fail(name, ReasonInvalid, "the contract's validator refused its answer: %s", detail) + } + return ans, nil +} diff --git a/internal/core/runner/dispatch_test.go b/internal/core/runner/dispatch_test.go new file mode 100644 index 000000000..fbf6dac46 --- /dev/null +++ b/internal/core/runner/dispatch_test.go @@ -0,0 +1,316 @@ +package runner + +import ( + "context" + "os" + "os/exec" + "path/filepath" + "reflect" + "strings" + "testing" + "time" + + "github.com/intentdriven/abcd/internal/core/history" + "github.com/intentdriven/abcd/internal/gittest" +) + +// memStore is a transcript store the dispatch tests read back. +type memStore struct{ got []stored } + +type stored struct { + runner, role, session string + raw []byte +} + +func (m *memStore) Store(runner string, req Request, ans Answer, raw []byte) error { + m.got = append(m.got, stored{runner, req.Role, ans.SessionID, raw}) + return nil +} + +// harness is the dispatcher every test builds: fixed clock, the contract's +// validator, an in-memory store, and the fallback receipts collected. +type harness struct { + d *Dispatcher + store *memStore + receipts []FallbackReceipt +} + +var fixed = time.Date(2026, 9, 30, 4, 0, 0, 0, time.UTC) + +func newHarness(t *testing.T, c *Config, hostSession bool) *harness { + t.Helper() + h := &harness{store: &memStore{}} + h.d = &Dispatcher{ + Config: c, + HostSession: hostSession, + Validate: validReceipt, + Transcripts: h.store, + Record: func(r FallbackReceipt) error { + h.receipts = append(h.receipts, r) + return nil + }, + Now: func() time.Time { return fixed }, + } + return h +} + +const routedMachine = `{` + localProvider + `,"runner":{"fallback_host":"claude","claude":{},"opencode":{"model":"local/qwen3-coder"}}}` +const routedRepo = `{"roles":{"ruthless-reviewer":{"runner":"opencode"}}}` + +// TestUnsetRoleHandsToHostUnchanged is criterion 2: a role left unset goes to +// the host session exactly as today: no runner launched, no fallback receipt. +func TestUnsetRoleHandsToHostUnchanged(t *testing.T) { + f := newFake(t, "ok", Claude, OpenCode) + h := newHarness(t, mustLoad(t, routedMachine, routedRepo), true) + out, err := h.d.Dispatch(context.Background(), f.request("security-reviewer")) + if err != nil { + t.Fatal(err) + } + if !out.Handoff || out.Answer != nil || out.Fallback != nil || len(h.receipts) != 0 { + t.Fatalf("outcome = %+v, receipts %v", out, h.receipts) + } + if out.Receipt.Route.Asked != Host || out.Receipt.Route.Ran != Host { + t.Fatalf("route = %+v", out.Receipt.Route) + } + if f.launched(Claude) || f.launched(OpenCode) { + t.Fatal("a runner was launched for an unset role") + } +} + +// TestRoutedRoleRunsThroughItsRunner is criterion 1: the routed role runs +// through the runner, its answer is validated by the contract's validator, and +// its transcript lands in the store. +func TestRoutedRoleRunsThroughItsRunner(t *testing.T) { + f := newFake(t, "ok", Claude, OpenCode) + h := newHarness(t, mustLoad(t, routedMachine, routedRepo), true) + req := f.request("ruthless-reviewer") + out, err := h.d.Dispatch(context.Background(), req) + if err != nil { + t.Fatal(err) + } + if out.Handoff || out.Answer == nil || out.Fallback != nil { + t.Fatalf("outcome = %+v", out) + } + if out.Receipt.Route.Asked != OpenCode || out.Receipt.Route.Ran != OpenCode { + t.Fatalf("route = %+v", out.Receipt.Route) + } + if len(h.store.got) != 1 || h.store.got[0].runner != OpenCode || h.store.got[0].role != req.Role || + !strings.Contains(string(h.store.got[0].raw), `"type":"text"`) { + t.Fatalf("transcripts = %+v", h.store.got) + } + if f.launched(Claude) { + t.Fatal("the fallback host ran for a runner that succeeded") + } +} + +// TestFallbackOnEveryFailureKind is criterion 3 with a host session: an +// absent, refusing, failing, unparsable or invalid runner hands the role to the +// host, and one receipt names the role, the runner asked for, the reason and +// the route that ran. +func TestFallbackOnEveryFailureKind(t *testing.T) { + for _, tc := range []struct { + name, mode string + onPath []string + want Reason + }{ + {"absent", "ok", nil, ReasonAbsent}, + {"refuses", "refuse", []string{OpenCode}, ReasonRefused}, + {"fails", "exit1", []string{OpenCode}, ReasonFailed}, + {"unparsable", "garbage", []string{OpenCode}, ReasonUnparsable}, + {"invalid answer", "noreceipt", []string{OpenCode}, ReasonInvalid}, + } { + t.Run(tc.name, func(t *testing.T) { + f := newFake(t, tc.mode, tc.onPath...) + h := newHarness(t, mustLoad(t, routedMachine, routedRepo), true) + out, err := h.d.Dispatch(context.Background(), f.request("ruthless-reviewer")) + if err != nil { + t.Fatal(err) + } + if !out.Handoff || out.Fallback == nil || len(h.receipts) != 1 { + t.Fatalf("outcome = %+v, receipts %v", out, h.receipts) + } + r := h.receipts[0] + if r.Role != "ruthless-reviewer" || r.Asked != OpenCode || r.Reason != tc.want || r.Ran != Host || + !r.At.Equal(fixed) || r.Detail == "" { + t.Fatalf("receipt = %+v", r) + } + if out.Receipt.Route.Asked != OpenCode || out.Receipt.Route.Ran != Host { + t.Fatalf("role receipt route = %+v", out.Receipt.Route) + } + }) + } +} + +// TestRoutedToADisabledRunnerFallsBack: a role routed to a runner this +// machine has not enabled is an absent runner, recorded, not a silent host run. +func TestRoutedToADisabledRunnerFallsBack(t *testing.T) { + f := newFake(t, "ok", Claude, OpenCode) + h := newHarness(t, mustLoad(t, `{"runner":{"claude":{}}}`, routedRepo), true) + out, err := h.d.Dispatch(context.Background(), f.request("ruthless-reviewer")) + if err != nil { + t.Fatal(err) + } + if !out.Handoff || len(h.receipts) != 1 || h.receipts[0].Reason != ReasonAbsent { + t.Fatalf("outcome = %+v, receipts %v", out, h.receipts) + } + if f.launched(OpenCode) { + t.Fatal("a runner the machine did not enable was launched") + } +} + +// TestNoHostSessionFallsBackToTheConfiguredHost is criterion 3 with no host +// session: the operator's configured fallback host runs the role, and the +// receipt names it as the route that ran. +func TestNoHostSessionFallsBackToTheConfiguredHost(t *testing.T) { + f := newFake(t, "ok", Claude) // opencode absent, claude present + h := newHarness(t, mustLoad(t, routedMachine, routedRepo), false) + out, err := h.d.Dispatch(context.Background(), f.request("ruthless-reviewer")) + if err != nil { + t.Fatal(err) + } + if out.Handoff || out.Answer == nil || len(h.receipts) != 1 { + t.Fatalf("outcome = %+v, receipts %v", out, h.receipts) + } + if r := h.receipts[0]; r.Asked != OpenCode || r.Ran != Claude || r.Reason != ReasonAbsent { + t.Fatalf("receipt = %+v", r) + } + if len(h.store.got) != 1 || h.store.got[0].runner != Claude { + t.Fatalf("transcripts = %+v", h.store.got) + } +} + +// TestNoHostSessionUnsetRoleRunsOnTheConfiguredHost: with no host session the +// configured host is the host, so an unset role runs there with no fallback. +func TestNoHostSessionUnsetRoleRunsOnTheConfiguredHost(t *testing.T) { + f := newFake(t, "ok", Claude) + h := newHarness(t, mustLoad(t, routedMachine, ""), false) + out, err := h.d.Dispatch(context.Background(), f.request("scribe")) + if err != nil { + t.Fatal(err) + } + if out.Answer == nil || out.Fallback != nil || out.Receipt.Route.Ran != Claude || out.Receipt.Route.Asked != Host { + t.Fatalf("outcome = %+v", out) + } +} + +// TestNoHostAndNoFallbackHostIsRefusedBeforeLaunch: decision 2 says there is +// always a landing, so a dispatch with neither is refused before any runner +// starts. +func TestNoHostAndNoFallbackHostIsRefusedBeforeLaunch(t *testing.T) { + f := newFake(t, "ok", Claude, OpenCode) + h := newHarness(t, mustLoad(t, `{`+localProvider+`,"runner":{"opencode":{"model":"local/qwen3-coder"}}}`, routedRepo), false) + _, err := h.d.Dispatch(context.Background(), f.request("ruthless-reviewer")) + if err == nil || !strings.Contains(err.Error(), "runner.fallback_host") { + t.Fatalf("err = %v, want a refusal naming runner.fallback_host", err) + } + if f.launched(OpenCode) || f.launched(Claude) { + t.Fatal("a runner was launched with no landing configured") + } +} + +// TestFallbackHostFailingToo is an error naming both: nothing lands silently. +func TestFallbackHostFailingToo(t *testing.T) { + f := newFake(t, "exit1", Claude, OpenCode) + h := newHarness(t, mustLoad(t, routedMachine, routedRepo), false) + _, err := h.d.Dispatch(context.Background(), f.request("ruthless-reviewer")) + if err == nil || !strings.Contains(err.Error(), OpenCode) || !strings.Contains(err.Error(), Claude) { + t.Fatalf("err = %v, want both failures named", err) + } + if len(h.receipts) != 1 { + t.Fatalf("receipts = %v; the fallback that was tried is still recorded", h.receipts) + } +} + +// TestTallyCountsPerRunnerAndPerRole is criterion 4's core: the fallback count +// per runner and per role, from the receipts a run recorded. +func TestTallyCountsPerRunnerAndPerRole(t *testing.T) { + rs := []FallbackReceipt{ + {Role: "ruthless-reviewer", Asked: OpenCode, Reason: ReasonAbsent, Ran: Host}, + {Role: "ruthless-reviewer", Asked: OpenCode, Reason: ReasonFailed, Ran: Host}, + {Role: "security-reviewer", Asked: OpenCode, Reason: ReasonInvalid, Ran: Host}, + {Role: "implementer", Asked: Claude, Reason: ReasonRefused, Ran: Host}, + } + got := Tally(rs) + want := Counts{ + Total: 4, + ByRunner: map[string]int{OpenCode: 3, Claude: 1}, + ByRole: map[string]int{"ruthless-reviewer": 2, "security-reviewer": 1, "implementer": 1}, + } + if !reflect.DeepEqual(got, want) { + t.Fatalf("tally = %+v, want %+v", got, want) + } + if z := Tally(nil); z.Total != 0 || z.ByRunner == nil || z.ByRole == nil { + t.Fatalf("empty tally = %+v; a run with no fallback still reports zero, not nothing", z) + } +} + +// TestReceiptsDifferOnlyInRoute is criterion 7's core: the role receipt a +// runner's run produces and the one a host run produces carry the same role, +// brief and contract, and differ only in the route block. +func TestReceiptsDifferOnlyInRoute(t *testing.T) { + f := newFake(t, "ok", Claude, OpenCode) + req := f.request("ruthless-reviewer") + viaRunner, err := newHarness(t, mustLoad(t, routedMachine, routedRepo), true).d.Dispatch(context.Background(), req) + if err != nil { + t.Fatal(err) + } + viaHost, err := newHarness(t, mustLoad(t, routedMachine, ""), true).d.Dispatch(context.Background(), req) + if err != nil { + t.Fatal(err) + } + a, b := viaRunner.Receipt, viaHost.Receipt + if a.Route == b.Route { + t.Fatal("the two receipts do not name different routes") + } + a.Route, b.Route = RouteRecord{}, RouteRecord{} + if !reflect.DeepEqual(a, b) { + t.Fatalf("receipts differ beyond the route:\n runner %+v\n host %+v", a, b) + } +} + +// TestTranscriptLandsInTheHistoryStore is criterion 1's store half: the +// production store writes the runner's transcript into abcd's own +// transcript store, redacted, under the runner as its tool. +func TestTranscriptLandsInTheHistoryStore(t *testing.T) { + home := t.TempDir() + t.Setenv("HOME", home) + repo := t.TempDir() + git := exec.Command("git", "init", "-q", repo) + git.Env = gittest.Env(t) + if out, err := git.CombinedOutput(); err != nil { + t.Fatalf("git init: %v %s", err, out) + } + rootSHA := strings.Repeat("ab", 20) + s := HistoryStore{RepoRoot: repo, RootSHA: rootSHA} + req := Request{Role: "ruthless-reviewer", SessionID: "run-2609300400000000-lane-1-ruthless-reviewer"} + raw := []byte(`{"type":"text","sessionID":"ses_fake1","part":{"type":"text","text":"done"}}` + "\n") + if err := s.Store(OpenCode, req, Answer{SessionID: "ses_fake1"}, raw); err != nil { + t.Fatalf("store: %v", err) + } + recs, err := history.List(repo, rootSHA) + if err != nil { + t.Fatal(err) + } + if len(recs) != 1 || recs[0].SourceTool != OpenCode || recs[0].SessionID != "ses_fake1" || + recs[0].AgentType != "ruthless-reviewer" { + t.Fatalf("records = %+v", recs) + } + if _, err := os.Stat(filepath.Join(home, ".abcd", "transcripts", rootSHA)); err != nil { + t.Fatalf("the store is not the user-level one: %v", err) + } +} + +// TestDispatchNeedsAValidator: the answer is validated the way the host's is, +// so a dispatcher without the contract's validator refuses rather than taking +// any answer. +func TestDispatchNeedsAValidator(t *testing.T) { + f := newFake(t, "ok", Claude, OpenCode) + h := newHarness(t, mustLoad(t, routedMachine, routedRepo), true) + h.d.Validate = nil + if _, err := h.d.Dispatch(context.Background(), f.request("ruthless-reviewer")); err == nil { + t.Fatal("dispatched with no validator") + } + if f.launched(OpenCode) { + t.Fatal("launched before the dispatcher was found incomplete") + } +} diff --git a/internal/core/runner/main_test.go b/internal/core/runner/main_test.go new file mode 100644 index 000000000..f2fde9b31 --- /dev/null +++ b/internal/core/runner/main_test.go @@ -0,0 +1,227 @@ +package runner + +// main_test.go is the fake command-line harness every test drives. No test +// reaches a real harness or a paid model: the test binary re-executes itself +// under the name of the harness a test puts on PATH (claude, opencode), and +// TestMain, seeing ABCD_RUNNER_FAKE, plays that harness instead of running the +// tests. What it does is chosen by the mode in ABCD_RUNNER_FAKE; what it saw +// (its argv and its environment) it writes under ABCD_RUNNER_FAKE_LOG, so a +// test asserts on the launch the runner really made. + +import ( + "bufio" + "encoding/json" + "fmt" + "os" + "os/exec" + "path/filepath" + "strconv" + "strings" + "testing" + "time" +) + +const ( + fakeModeEnv = "ABCD_RUNNER_FAKE" + fakeLogEnv = "ABCD_RUNNER_FAKE_LOG" + // fakeCredEnv is the credential variable a harness reads its key from; + // the test sets it to a value built at run time and asserts it never + // reaches an argument, an error or a receipt. + fakeCredEnv = "ANTHROPIC_API_KEY" +) + +func TestMain(m *testing.M) { + if mode := os.Getenv(fakeModeEnv); mode != "" { + os.Exit(fakeHarness(mode)) + } + os.Exit(m.Run()) +} + +// fakeHarness plays the harness its own name says it is. +func fakeHarness(mode string) int { + // The runner executes the resolved binary, so argv[0] names the test + // binary; which harness this is shows in the launch itself. + name := "claude" + if len(os.Args) > 1 && os.Args[1] == "run" { + name = "opencode" + } + logDir := os.Getenv(fakeLogEnv) + if logDir != "" && mode != "sleeper" { + argv, _ := json.Marshal(os.Args[1:]) + _ = os.WriteFile(filepath.Join(logDir, name+".argv.json"), argv, 0o600) + _ = os.WriteFile(filepath.Join(logDir, name+".env"), []byte(strings.Join(os.Environ(), "\n")), 0o600) + wd, _ := os.Getwd() + _ = os.WriteFile(filepath.Join(logDir, name+".cwd"), []byte(wd), 0o600) + } + prompt := "" + if len(os.Args) > 1 { + prompt = os.Args[len(os.Args)-1] + } + switch mode { + case "sleeper": + time.Sleep(60 * time.Second) + return 0 + case "hang": + child := exec.Command(os.Args[0]) + child.Env = append(os.Environ(), fakeModeEnv+"=sleeper") + if err := child.Start(); err != nil { + return 3 + } + _ = os.WriteFile(filepath.Join(logDir, "child.pid"), []byte(strconv.Itoa(child.Process.Pid)), 0o600) + time.Sleep(60 * time.Second) + return 0 + case "flood": + chunk := strings.Repeat("x", 4096) + for i := 0; i < 1024; i++ { + fmt.Fprint(os.Stdout, chunk) + } + return 0 + case "exit1": + cred := os.Getenv(fakeCredEnv) + fmt.Fprintf(os.Stdout, "auth failed for key %s\n", cred) + fmt.Fprintf(os.Stderr, "error: invalid key %s\n", cred) + return 1 + case "garbage": + fmt.Fprintln(os.Stdout, "this is not a structured event stream") + return 0 + case "refuse": + if name == "opencode" { + fmt.Fprintln(os.Stdout, `{"type":"error","sessionID":"ses_fake1","error":{"name":"APIError","data":{"message":"model refused"}}}`) + } else { + fmt.Fprintln(os.Stdout, `{"type":"system","subtype":"init","session_id":"fake-session-1","model":"fake-model"}`) + fmt.Fprintln(os.Stdout, `{"type":"result","subtype":"error_during_execution","is_error":true,"result":"refused","session_id":"fake-session-1"}`) + } + return 0 + case "hugemodel", "ctrlmodel", "suffixmodel": + // A harness whose init event reports a model: one past any bound, + // one carrying control bytes, and a real id's bracketed suffix. + model := map[string]string{ + "hugemodel": strings.Repeat("m", 5<<20), + "ctrlmodel": "fake\x1b[2Jmodel\r\n", + "suffixmodel": "claude-opus-4-6[1m]", + }[mode] + init, _ := json.Marshal(map[string]string{"type": "system", "subtype": "init", "session_id": "fake-session-1", "model": model}) + fmt.Fprintln(os.Stdout, string(init)) + fmt.Fprintln(os.Stdout, `{"type":"result","subtype":"success","is_error":false,"result":"done","session_id":"fake-session-1"}`) + return 0 + case "ok", "noreceipt": + if mode == "ok" { + if rec := promptField(prompt, "Receipt: "); rec != "" { + _ = os.WriteFile(rec, []byte(`{"ok":true}`), 0o600) + } + } + if name == "opencode" { + fmt.Fprintln(os.Stdout, `{"type":"step_start","sessionID":"ses_fake1","part":{"type":"step-start"}}`) + fmt.Fprintln(os.Stdout, `{"type":"text","sessionID":"ses_fake1","part":{"type":"text","text":"done"}}`) + fmt.Fprintln(os.Stdout, `{"type":"step_finish","sessionID":"ses_fake1","part":{"type":"step-finish"}}`) + } else { + fmt.Fprintln(os.Stdout, `{"type":"system","subtype":"init","session_id":"fake-session-1","model":"fake-model"}`) + fmt.Fprintln(os.Stdout, `{"type":"assistant","message":{"content":[{"type":"text","text":"done"}]},"session_id":"fake-session-1"}`) + fmt.Fprintln(os.Stdout, `{"type":"result","subtype":"success","is_error":false,"result":"done","session_id":"fake-session-1"}`) + } + return 0 + } + fmt.Fprintf(os.Stderr, "fake harness: unknown mode %q\n", mode) + return 2 +} + +// promptField returns the value of the prompt line that starts with prefix. +func promptField(prompt, prefix string) string { + sc := bufio.NewScanner(strings.NewReader(prompt)) + for sc.Scan() { + if v, ok := strings.CutPrefix(sc.Text(), prefix); ok { + return strings.TrimSpace(v) + } + } + return "" +} + +// fakeEnv is one test's fake harness set-up: a PATH directory holding the +// harnesses as links to the test binary, a log directory, and a repository +// directory with a brief in it. +type fakeEnv struct { + bin, log, repo, lane string +} + +// newFake puts the named harnesses on PATH (and nothing else of the test's), +// sets the mode, and returns the directories. +func newFake(t *testing.T, mode string, harnesses ...string) fakeEnv { + t.Helper() + self, err := os.Executable() + if err != nil { + t.Fatal(err) + } + root := t.TempDir() + f := fakeEnv{ + bin: filepath.Join(root, "bin"), + log: filepath.Join(root, "log"), + repo: filepath.Join(root, "repo"), + lane: filepath.Join(root, "lane"), + } + for _, d := range []string{f.bin, f.log, f.repo, f.lane} { + if err := os.Mkdir(d, 0o700); err != nil { + t.Fatal(err) + } + } + for _, h := range harnesses { + if err := os.Symlink(self, filepath.Join(f.bin, h)); err != nil { + t.Fatal(err) + } + } + if err := os.WriteFile(filepath.Join(f.lane, "brief.md"), []byte("# brief\n"), 0o600); err != nil { + t.Fatal(err) + } + t.Setenv("PATH", f.bin) + t.Setenv(fakeModeEnv, mode) + t.Setenv(fakeLogEnv, f.log) + return f +} + +// request is the request every test hands a runner. +func (f fakeEnv) request(role string) Request { + return Request{ + Role: role, + Brief: filepath.Join(f.lane, "brief.md"), + Receipt: filepath.Join(f.lane, "receipt.json"), + Dir: f.repo, + Tools: []string{"Read", "Grep"}, + SessionID: "run-2609300400000000-lane-1-" + role, + Timeout: 20 * time.Second, + } +} + +// argv returns what the named fake harness was launched with. +func (f fakeEnv) argv(t *testing.T, harness string) []string { + t.Helper() + raw, err := os.ReadFile(filepath.Join(f.log, harness+".argv.json")) + if err != nil { + t.Fatalf("the %s fake was not launched: %v", harness, err) + } + var out []string + if err := json.Unmarshal(raw, &out); err != nil { + t.Fatal(err) + } + return out +} + +// launched reports whether the named fake harness ran at all. +func (f fakeEnv) launched(harness string) bool { + _, err := os.Stat(filepath.Join(f.log, harness+".argv.json")) + return err == nil +} + +// validReceipt is the contract's validator the tests hand the dispatcher: the +// receipt exists and is the JSON the fake writes. +func validReceipt(_ string, req Request, _ Answer) error { + raw, err := os.ReadFile(req.Receipt) + if err != nil { + return fmt.Errorf("no receipt at the contract's path") + } + var v struct { + OK bool `json:"ok"` + } + if json.Unmarshal(raw, &v) != nil || !v.OK { + return fmt.Errorf("the receipt is not the contract's shape") + } + return nil +} diff --git a/internal/core/runner/opencode.go b/internal/core/runner/opencode.go new file mode 100644 index 000000000..7c5559b78 --- /dev/null +++ b/internal/core/runner/opencode.go @@ -0,0 +1,109 @@ +package runner + +// opencode.go is the opencode runner (spc-2609221533057881 scope 2): run mode +// with raw JSON events, the repository as its directory, the brief attached, +// and no external plugins (--pure), the nearest the harness offers to the +// claude CLI's bare mode, so a plugin the target repository configures does +// not run. Permissions are left at the harness's own configuration: abcd does +// not pass the flag that approves every request unasked. +// +// The runner reaches opencode through run mode only; its server's session +// endpoints, where a server is already up, are a later adapter. + +import ( + "bytes" + "context" + "encoding/json" +) + +// OpenCodeCLI is the opencode runner. +type OpenCodeCLI struct { + // model is the model route, / as opencode spells it too, + // or "" for the harness's own default. + model string + launch launcher +} + +// newOpenCode returns the opencode runner asking for model, a +// / route the configuration admitted, or "". +func newOpenCode(model string) *OpenCodeCLI { + return &OpenCodeCLI{model: model, launch: defaultLauncher()} +} + +// Name is the route's name. +func (*OpenCodeCLI) Name() string { return OpenCode } + +func (o *OpenCodeCLI) args(req Request) []string { + a := []string{"run", "--format=json", "--pure", "--dir=" + req.Dir, "--file=" + req.Brief} + if o.model != "" { + a = append(a, "--model="+o.model) + } + return append(a, "--", prompt(req)) +} + +// Run runs the role through opencode. +func (o *OpenCodeCLI) Run(ctx context.Context, req Request) (Answer, []byte, error) { + if err := req.check(); err != nil { + return Answer{}, nil, err + } + bin, err := o.launch.admit(OpenCode, "opencode", req.Dir, req.Checkout) + if err != nil { + return Answer{}, nil, err + } + res, err := o.launch.run(ctx, OpenCode, bin, o.args(req), req.Dir, req.timeout()) + transcript := res.transcript() + if err != nil { + return Answer{}, transcript, err + } + ans, err := parseOpenCode(res.stdout) + if err != nil { + return Answer{}, transcript, err + } + ans.Model = o.model + return ans, transcript, nil +} + +// openCodeEvent is the part of a run-mode JSON event the runner reads. +type openCodeEvent struct { + Type string `json:"type"` + SessionID string `json:"sessionID"` + Part struct { + Type string `json:"type"` + Text string `json:"text"` + } `json:"part"` +} + +// parseOpenCode reads the event stream: one JSON object per line. An error +// event is a refusal; a line that is not an event, or a stream with no text +// and no finished step, is unparsable. The last text part is the final +// message. +func parseOpenCode(out []byte) (Answer, error) { + var ans Answer + seen := false + for _, line := range bytes.Split(out, []byte("\n")) { + line = bytes.TrimSpace(line) + if len(line) == 0 { + continue + } + var ev openCodeEvent + if json.Unmarshal(line, &ev) != nil || ev.Type == "" { + return Answer{}, fail(OpenCode, ReasonUnparsable, "a line of its output is not a JSON event") + } + if ans.SessionID == "" && ev.SessionID != "" { + ans.SessionID = ev.SessionID + } + switch ev.Type { + case "error": + return Answer{}, fail(OpenCode, ReasonRefused, "it reported an error event") + case "text": + ans.Text = ev.Part.Text + seen = true + case "step_finish": + seen = true + } + } + if !seen { + return Answer{}, fail(OpenCode, ReasonUnparsable, "its output carries no text and no finished step") + } + return ans, nil +} diff --git a/internal/core/runner/proc.go b/internal/core/runner/proc.go new file mode 100644 index 000000000..ee47dcf60 --- /dev/null +++ b/internal/core/runner/proc.go @@ -0,0 +1,264 @@ +package runner + +// proc.go is the one place this package starts a process. The trust boundary: +// a runner starts a harness with a prompt and a repository path, so +// +// - the argv is a vector handed to exec, never a shell line, and the prompt +// travels as one argument behind the end-of-options marker; +// - the binary is resolved on PATH by its fixed name and refused when it is +// not absolute, when other can write it or its directory, when a group +// other than darwin's admin group (gid 80, on darwin only) can write +// either, or when it resolves inside the repository the role runs in, or +// the checkout a lane's worktree belongs to, lexically or through a +// symlink (a PATH entry into either is repository content, never run); +// - the environment is the parent's with every git repository-selection and +// config-injection variable scrubbed (gitutil.ScrubbedEnv), so an inherited +// GIT_DIR cannot aim the role's git at another repository, while the +// person's own git identity and the harness's own credential variables +// pass through untouched: abcd never logs a harness in; +// - stdin is the null device, the working directory is the repository; +// - stdout and stderr are each bounded; a harness writing past the bound is +// cut off and the run fails naming the bound; +// - the child leads its own process group, and a run past its time, or one +// whose context ends, is killed with that group through the pid this +// handle holds: never a pattern, never another process; +// - an error is written by abcd and never carries the harness's output. + +import ( + "bytes" + "context" + "errors" + "fmt" + "os" + "os/exec" + "path/filepath" + "runtime" + "syscall" + "time" + + "github.com/intentdriven/abcd/internal/fsutil" + "github.com/intentdriven/abcd/internal/gitutil" +) + +// Output bounds. A transcript is the whole event stream of one role's run; +// stderr is diagnostics only. +const ( + defaultMaxStdout = 32 << 20 + defaultMaxStderr = 64 << 10 +) + +// pipeGrace is how long a run keeps reading output once the harness has exited +// or been killed: a descendant that left the group can hold the pipe open. +const pipeGrace = 5 * time.Second + +// launcher starts one harness binary. Its fields are the seams the tests +// narrow; production uses defaultLauncher. +type launcher struct { + lookPath func(string) (string, error) + maxStdout int + maxStderr int +} + +func defaultLauncher() launcher { + return launcher{lookPath: exec.LookPath, maxStdout: defaultMaxStdout, maxStderr: defaultMaxStderr} +} + +// procResult is what one run produced. +type procResult struct { + stdout, stderr []byte + // overflow names the stream that passed its bound, "" when neither did. + overflow string +} + +// admit resolves name on PATH and refuses a result that is not absolute or +// that lies inside any of repos (the directory the role runs in, and the +// checkout it belongs to), lexically or after symlink resolution. +func (l launcher) admit(runner, name string, repos ...string) (string, error) { + p, err := l.lookPath(name) + if err != nil { + return "", fail(runner, ReasonAbsent, "%s is not on PATH", name) + } + if !filepath.IsAbs(p) { + return "", fail(runner, ReasonAbsent, "%s resolves to a relative path, which is never run", name) + } + resolved, err := filepath.EvalSymlinks(p) + if err != nil { + return "", fail(runner, ReasonAbsent, "%s does not resolve to a file", name) + } + // Whoever can write the binary, the directory PATH reaches it through, or + // the directory it resolves into chooses what runs. + for _, c := range []string{resolved, filepath.Dir(resolved), filepath.Dir(filepath.Clean(p))} { + fi, err := os.Stat(c) + if err != nil { + return "", fail(runner, ReasonAbsent, "%s could not be examined before it is run", name) + } + if fsutil.WritableByOthers(fi) && !adminGroupWritableOnly(fi) { + return "", fail(runner, ReasonAbsent, "%s, or a directory it is reached through, is one group or other can write, "+ + "so it is never run; chmod go-w it", name) + } + } + var guards []string + for _, repo := range repos { + if repo == "" { + continue + } + guards = append(guards, filepath.Clean(repo)) + if g, err := filepath.EvalSymlinks(repo); err == nil { + guards = append(guards, g) + } + } + fold := fsutil.CaseFoldingFS() + for _, g := range guards { + for _, c := range []string{filepath.Clean(p), resolved} { + if fsutil.PathWithin(c, g, fold) { + return "", fail(runner, ReasonAbsent, "%s resolves inside the repository the role runs in or the checkout it belongs to; "+ + "a program there is repository content and is never run", name) + } + } + } + return resolved, nil +} + +// adminGroupWritableOnly reports whether fi's only write bit beyond its +// owner's is the group's, and that group is darwin's admin group (gid 80), +// on darwin only. Homebrew installs /opt/homebrew/bin as root- or +// user-owned, group admin, mode 0775, so a harness installed through it is +// reached through a group-writable directory. Members of darwin's admin +// group can sudo by default, so that write grants them nothing they do not +// hold, and the directory is admitted. gid 0 is never admitted: being in +// Linux's root group, or darwin's wheel, does not by itself let a member act +// as root. Other-writable is never admitted, whatever the group, and neither +// is any other group, gid 80 off darwin, nor a group that cannot be read. +func adminGroupWritableOnly(fi os.FileInfo) bool { + if fi.Mode().Perm()&0o002 != 0 { + return false + } + gid, ok := fileGroup(fi) + if !ok { + return false + } + return runtime.GOOS == "darwin" && gid == 80 +} + +// fileGroup reads the owning group of a stat result; ok is false when the +// platform does not report one. A variable so a test can name a group it +// cannot chown to. +var fileGroup = func(fi os.FileInfo) (uint32, bool) { + st, ok := fi.Sys().(*syscall.Stat_t) + if !ok { + return 0, false + } + return st.Gid, true +} + +// run starts bin with args in dir and waits for it, at most timeout. +func (l launcher) run(ctx context.Context, runner, bin string, args []string, dir string, timeout time.Duration) (procResult, error) { + // #nosec G204 -- bin is a fixed harness name resolved by admit (absolute, + // outside the repository); args are a vector built by the adapter, never a + // shell line. + cmd := exec.Command(bin, args...) + cmd.Dir = dir + cmd.Env = gitutil.ScrubbedEnv() + cmd.Stdin = nil + var out, errb bytes.Buffer + ow := &bounded{w: &out, remaining: l.maxStdout} + ew := &bounded{w: &errb, remaining: l.maxStderr} + cmd.Stdout = ow + cmd.Stderr = ew + cmd.SysProcAttr = &syscall.SysProcAttr{Setpgid: true} + cmd.WaitDelay = pipeGrace + if err := cmd.Start(); err != nil { + return procResult{}, fail(runner, ReasonFailed, "%s could not be started", filepath.Base(bin)) + } + timer := time.NewTimer(timeout) + defer timer.Stop() + done := make(chan error, 1) + go func() { done <- cmd.Wait() }() + var werr error + killed := "" + select { + case werr = <-done: + case <-timer.C: + killed = fmt.Sprintf("did not finish within the time allowed (%s)", timeout) + case <-ctx.Done(): + killed = "was stopped: the run it belongs to ended" + } + if killed != "" { + // The group this child leads (Setpgid), through the pid this handle + // holds; the leader is not yet reaped, so the group id is still its. + _ = syscall.Kill(-cmd.Process.Pid, syscall.SIGKILL) + <-done + } + res := procResult{stdout: out.Bytes(), stderr: errb.Bytes()} + switch { + case ow.overflowed: + res.overflow = "stdout" + case ew.overflowed: + res.overflow = "stderr" + } + name := filepath.Base(bin) + switch { + case killed != "": + return res, fail(runner, ReasonFailed, "%s %s; its process group was killed", name, killed) + case res.overflow != "": + return res, fail(runner, ReasonFailed, "%s wrote past the %s bound (%d bytes); the rest was discarded", + name, res.overflow, l.bound(res.overflow)) + case errors.Is(werr, exec.ErrWaitDelay): + return res, fail(runner, ReasonFailed, "%s exited, but a process it started left its group and held its output", name) + case werr != nil: + var ee *exec.ExitError + if errors.As(werr, &ee) { + return res, fail(runner, ReasonFailed, "%s exited with status %d", name, ee.ExitCode()) + } + return res, fail(runner, ReasonFailed, "%s did not complete", name) + } + return res, nil +} + +func (l launcher) bound(stream string) int { + if stream == "stderr" { + return l.maxStderr + } + return l.maxStdout +} + +// bounded keeps at most remaining bytes and drops the rest without failing the +// writer, recording that it dropped some. +type bounded struct { + w *bytes.Buffer + remaining int + overflowed bool +} + +func (b *bounded) Write(p []byte) (int, error) { + n := len(p) + k := n + if k > b.remaining { + k = b.remaining + b.overflowed = true + } + if k > 0 { + b.w.Write(p[:k]) + b.remaining -= k + } + return n, nil +} + +// transcript is what a run's record keeps: the event stream, then the +// harness's stderr under a marker line when it wrote any. The store redacts it +// on write. +func (r procResult) transcript() []byte { + if len(r.stderr) == 0 { + return r.stdout + } + var b bytes.Buffer + b.Write(r.stdout) + if len(r.stdout) > 0 && r.stdout[len(r.stdout)-1] != '\n' { + b.WriteByte('\n') + } + // The marker opens with "==", never a "---" run, so no frontmatter reader + // takes it for a block delimiter. + b.WriteString("== abcd runner: the harness's stderr ==\n") + b.Write(r.stderr) + return b.Bytes() +} diff --git a/internal/core/runner/receipt.go b/internal/core/runner/receipt.go new file mode 100644 index 000000000..7f869b620 --- /dev/null +++ b/internal/core/runner/receipt.go @@ -0,0 +1,65 @@ +package runner + +// receipt.go is what a dispatch leaves in the record: the role's receipt +// block, the same for every route but the route it names (criterion 7), and +// the fallback receipt, one per fallback, which Tally counts per runner and +// per role for the run's summary (criteria 3 and 4). + +import "time" + +// RouteRecord is the route a role asked for and the one that ran it. +type RouteRecord struct { + // Asked is the route the configuration named: host or a runner. + Asked string `json:"asked"` + // Ran is the route that ran the role: host, a runner, or none. + Ran string `json:"ran"` + // Model is the model the runner that ran it reported; "" on the host, + // whose own receipt reports it. + Model string `json:"model"` +} + +// RoleReceipt is a role run's receipt block. Everything but Route is the +// request's, so two runs of one request differ in their route alone. +type RoleReceipt struct { + Role string `json:"role"` + Brief string `json:"brief"` + // Contract is the receipt path the contract names. + Contract string `json:"contract"` + Route RouteRecord `json:"route"` +} + +func newRoleReceipt(req Request, asked string) RoleReceipt { + return RoleReceipt{Role: req.Role, Brief: req.Brief, Contract: req.Receipt, Route: RouteRecord{Asked: asked}} +} + +// FallbackReceipt is one fallback: the role, the runner asked for, the reason, +// and the route that ran instead. +type FallbackReceipt struct { + At time.Time `json:"at"` + Role string `json:"role"` + Asked string `json:"asked"` + Reason Reason `json:"reason"` + // Detail is abcd's own account of the reason; it never carries the + // harness's output. + Detail string `json:"detail"` + Ran string `json:"ran"` +} + +// Counts is the fallback intel a run's summary reports. +type Counts struct { + Total int `json:"total"` + ByRunner map[string]int `json:"by_runner"` + ByRole map[string]int `json:"by_role"` +} + +// Tally counts fallback receipts per runner asked for and per role. A run with +// none reports zero, with empty maps rather than none. +func Tally(rs []FallbackReceipt) Counts { + c := Counts{ByRunner: map[string]int{}, ByRole: map[string]int{}} + for _, r := range rs { + c.Total++ + c.ByRunner[r.Asked]++ + c.ByRole[r.Role]++ + } + return c +} diff --git a/internal/core/runner/runner.go b/internal/core/runner/runner.go new file mode 100644 index 000000000..8dc18a027 --- /dev/null +++ b/internal/core/runner/runner.go @@ -0,0 +1,196 @@ +// Package runner runs a delegated role through a command-line harness the +// operator named, with the same brief, inputs and output contract the host's +// own sub-agent gets (itd-2609201916056194, spc-2609221533057881). +// +// A role's route is roles..runner in the layered configuration +// (config.go): host, the default, or a runner this machine enabled under +// runner.. The shipped runners are the claude CLI in print mode with the +// bare flag (claude.go) and opencode's run mode (opencode.go). Each is started +// as a process (proc.go): the argv is a vector and never a shell line, the +// binary is resolved on PATH and refused inside the repository, the +// environment is the parent's scrubbed of every git repository-selection and +// config-injection variable, stdin is the null device, output is bounded, and +// a run past its time is killed with the process group it leads. +// +// The fallback is decided in one place (dispatch.go): an absent, refusing, +// failing or unparsable runner, or an answer the contract's validator refuses, +// hands the role to the host session, or with no host session to the host the +// operator configured (runner.fallback_host), and writes one receipt naming +// the role, the runner asked for, the reason and the route that ran. Tally +// counts those receipts per runner and per role for the run's summary. +// +// The package never prints. It starts processes only through Dispatch and the +// adapters' Run, and never logs a harness in: a harness's credential is its +// own, read from its own environment or configuration, and abcd never puts one +// on a command line, in an error or in a receipt. +package runner + +import ( + "context" + "errors" + "fmt" + "path/filepath" + "regexp" + "strings" + "time" + + "github.com/intentdriven/abcd/internal/termsafe" +) + +// The routes a role may name. Host is the default. +const ( + Host = "host" + Claude = "claude" + OpenCode = "opencode" +) + +// runnerNames are the runners this build ships, in the order a refusal names +// them. +var runnerNames = []string{Claude, OpenCode} + +// DefaultTimeout bounds one run when the request names no timeout. +const DefaultTimeout = 60 * time.Minute + +// Request is what a role is handed: the same the host sub-agent is handed. +type Request struct { + // Role is the agent the run plays: an agent in the roster, or the + // implementer. + Role string + // Brief is the absolute path of the brief the role follows; it states the + // task and the output contract. + Brief string + // Receipt is the absolute path the contract says the answer is written to. + Receipt string + // Dir is the absolute path of the working tree the role runs in. + Dir string + // Checkout is the absolute path of the checkout the run belongs to when + // Dir is a worktree of it (a lane's worktree lives outside it); a program + // inside it is repository content and is never run, as one inside Dir is + // not. Empty when Dir is the checkout itself. + Checkout string + // Tools are the tools the role's contract grants, granted without a + // prompt; none granted when empty. + Tools []string + // SessionID names the run's transcript record when the harness reports no + // session of its own. + SessionID string + // Timeout bounds the run; DefaultTimeout when zero. + Timeout time.Duration +} + +// Answer is the one answer shape every adapter parses its harness's events +// into. The contract's own answer is the receipt file; this is what the +// harness said about the run. +type Answer struct { + // Text is the harness's final message. + Text string + // Model is the model the harness reports, or the one it was asked for when + // it reports none. + Model string + // SessionID is the harness's own session id, "" when it reports none. + SessionID string +} + +// Runner is one route a role can run through. +type Runner interface { + // Name is the route's name, as roles..runner spells it. + Name() string + // Run runs the role and returns the parsed answer and the transcript, the + // harness's raw event stream. A failure is a *Failure; the transcript is + // returned whenever the harness produced one, failure or not. + Run(ctx context.Context, req Request) (Answer, []byte, error) +} + +// Reason is why a runner did not answer: the kinds the fallback records. +type Reason string + +const ( + // ReasonAbsent: the runner is not enabled on this machine, or its binary + // is not on PATH or is refused. + ReasonAbsent Reason = "absent" + // ReasonRefused: the harness ran and reported that it refused or erred. + ReasonRefused Reason = "refused" + // ReasonFailed: the harness could not start, exited non-zero, ran out of + // time or overran its output bound. + ReasonFailed Reason = "failed" + // ReasonUnparsable: the harness's output is not its structured events. + ReasonUnparsable Reason = "unparsable" + // ReasonInvalid: the answer failed the contract's validator. + ReasonInvalid Reason = "invalid" +) + +// Failure is a runner's failure. Detail is written by abcd, never copied from +// the harness's output, so it cannot carry anything the harness printed, a +// credential it echoed included. +type Failure struct { + Runner string + Reason Reason + Detail string +} + +func (f *Failure) Error() string { + return fmt.Sprintf("runner %s %s: %s", f.Runner, f.Reason, f.Detail) +} + +func fail(runner string, reason Reason, format string, args ...any) *Failure { + return &Failure{Runner: runner, Reason: reason, Detail: fmt.Sprintf(format, args...)} +} + +var ( + roleRe = regexp.MustCompile(`^[a-z0-9][a-z0-9-]{0,63}$`) + sessionRe = regexp.MustCompile(`^[A-Za-z0-9._-]{1,128}$`) + // modelRe is a model id as a harness reports it (claude-opus-4-6, + // provider/model, a bracketed context suffix): bounded and plain, as a + // session id is, since it is written into the run's state and records. + modelRe = regexp.MustCompile(`^[A-Za-z0-9][A-Za-z0-9._:/-]{0,127}(\[[A-Za-z0-9._-]{1,16}\])?$`) + // toolRe is one tool as a contract names it (Read, Bash(git log:*)). A + // comma or a line break would split the one flag the tools travel in. + toolRe = regexp.MustCompile(`^[A-Za-z][A-Za-z0-9_]*(\([^,\r\n()]{1,120}\))?$`) +) + +// check refuses a request no runner may be started with, before any launch. +func (r Request) check() error { + if !roleRe.MatchString(r.Role) { + return fmt.Errorf("runner: role %q is not a plain lower-case name", termsafe.Sanitize(r.Role)) + } + if r.Checkout != "" && (!filepath.IsAbs(r.Checkout) || filepath.Clean(r.Checkout) != r.Checkout) { + return fmt.Errorf("runner: the checkout %q is not a clean absolute path", termsafe.Sanitize(r.Checkout)) + } + for _, p := range []struct{ name, path string }{{"brief", r.Brief}, {"receipt", r.Receipt}, {"directory", r.Dir}} { + if !filepath.IsAbs(p.path) || filepath.Clean(p.path) != p.path { + return fmt.Errorf("runner: the %s %q is not a clean absolute path", p.name, termsafe.Sanitize(p.path)) + } + } + for _, t := range r.Tools { + if !toolRe.MatchString(t) { + return fmt.Errorf("runner: tool %q is not a tool name as a contract spells it", termsafe.Sanitize(t)) + } + } + if !sessionRe.MatchString(r.SessionID) { + return errors.New("runner: the request names no session id for its transcript") + } + if r.Timeout < 0 { + return errors.New("runner: a negative timeout") + } + return nil +} + +func (r Request) timeout() time.Duration { + if r.Timeout == 0 { + return DefaultTimeout + } + return r.Timeout +} + +// prompt is the one prompt every route hands a role: the role, the brief and +// the receipt path, which are what the host sub-agent is handed. Everything +// else the role needs is in the brief. +func prompt(r Request) string { + var b strings.Builder + fmt.Fprintf(&b, "You are the %s agent, started by abcd.\n", r.Role) + fmt.Fprintf(&b, "Brief: %s\n", r.Brief) + fmt.Fprintf(&b, "Receipt: %s\n", r.Receipt) + b.WriteString("Read the brief and follow it exactly: it states your task and the output contract. " + + "Write your receipt to the Receipt path, in the shape the brief names.\n") + return b.String() +} diff --git a/internal/core/runner/transcript.go b/internal/core/runner/transcript.go new file mode 100644 index 000000000..0a73c0e9c --- /dev/null +++ b/internal/core/runner/transcript.go @@ -0,0 +1,37 @@ +package runner + +// transcript.go is where a runner's transcript lands: abcd's own transcript +// store (internal/core/history), whatever the harness keeps of its own. The +// store redacts on write and refuses a transcript whose redaction leaves a +// blocking span, so what a harness echoed never reaches the record unredacted. + +import "github.com/intentdriven/abcd/internal/core/history" + +// TranscriptStore keeps one runner's transcript of one role's run. +type TranscriptStore interface { + Store(runner string, req Request, ans Answer, raw []byte) error +} + +// HistoryStore is the production store: the repository's lane of the +// user-level transcript store, keyed on its root commit. +type HistoryStore struct { + RepoRoot string + RootSHA string +} + +// Store captures raw as a native transcript produced by the runner, recorded +// under the harness's own session id when it reported one and the request's +// otherwise, with the role as its agent type. +func (h HistoryStore) Store(runner string, req Request, ans Answer, raw []byte) error { + sid := ans.SessionID + if !sessionRe.MatchString(sid) { + sid = req.SessionID + } + _, err := history.Capture(h.RepoRoot, h.RootSHA, raw, history.CaptureMeta{ + SessionID: sid, + Kind: history.RouteNative, + Tool: runner, + AgentType: req.Role, + }) + return err +} diff --git a/internal/fsutil/fsutil.go b/internal/fsutil/fsutil.go index 411d9a4ca..2ac520319 100644 --- a/internal/fsutil/fsutil.go +++ b/internal/fsutil/fsutil.go @@ -303,7 +303,7 @@ func ReadDeclaration(path string, limit int64) ([]byte, DeclarationRefusal, erro // that must judge a link as itself Lstat's and refuses it before calling, as // ReadDeclaration does. func CallersAlone(path string, fi os.FileInfo) error { - if fi.Mode().Perm()&0o022 != 0 { + if WritableByOthers(fi) { return ErrDeclarationWritable } // An unreadable owner is refused too: "I could not learn who owns this" and @@ -315,6 +315,14 @@ func CallersAlone(path string, fi os.FileInfo) error { return nil } +// WritableByOthers reports whether fi's mode carries a group or other write +// bit: the one test of "someone else could write this" that CallersAlone +// applies, exported for a check that judges the mode alone, such as a harness +// binary a runner is about to start, which root may own. +func WritableByOthers(fi os.FileInfo) bool { + return fi.Mode().Perm()&0o022 != 0 +} + // ReadGuardedInRoot is ReadGuarded resolved inside an os.Root containment // scope. rel is a slash-separated path relative to root; every component is // resolved by the OS within root, so a symlinked ANCESTOR directory — the shape diff --git a/internal/surface/cli/build.go b/internal/surface/cli/build.go index 4e1aa6845..6294c89da 100644 --- a/internal/surface/cli/build.go +++ b/internal/surface/cli/build.go @@ -1,17 +1,22 @@ package cli import ( + "context" "encoding/json" "errors" "fmt" "io" "os" + "os/signal" "path/filepath" + "sort" "strings" + "syscall" "time" "github.com/intentdriven/abcd/internal/core/implement/loop" "github.com/intentdriven/abcd/internal/core/layered" + "github.com/intentdriven/abcd/internal/core/runner" "github.com/intentdriven/abcd/internal/fsutil" "github.com/intentdriven/abcd/internal/gitutil" "github.com/intentdriven/abcd/internal/termsafe" @@ -125,6 +130,11 @@ func newBuildCommand(asJSON *bool) *cobra.Command { "handed back: it stops as unachievable with the last round's findings, the run starts nothing\n" + "further for it, and `abcd implement step` refuses naming the hand-back.\n\n" + "The run then moves one step per `abcd implement step`, driven by the host session.\n\n" + + "The runner configuration is read before the run is created: roles..runner (host,\n" + + "the default, or a runner) and the runners this machine enables under runner. in\n" + + "~/.abcd/config.json, each model route admitted against its provider's allowlist. A fault,\n" + + "a model route the allowlist does not admit included, is refused at the runner stage and\n" + + "nothing is created or launched.\n\n" + "An issue id (iss-N, validated by shape) is built as one lane. Its checks are the\n" + "repository's own drain rule, read as `abcd drain` reads it (the issue is open, nothing\n" + "open blocks it, its category and severity are ones the rule takes, it carries a remedy a\n" + @@ -148,6 +158,9 @@ func newBuildCommand(asJSON *bool) *cobra.Command { for _, n := range notes { fmt.Fprintln(cmd.ErrOrStderr(), termsafe.Sanitize(n)) } + if _, err := loadRunners(cmd, roots); err != nil { + return loopFail(cmd.OutOrStdout(), *asJSON, prefix, err) + } o := loop.Options{Session: session, Roots: &roots} if cmd.Flags().Changed("pace") { o.Pace = &pace @@ -234,6 +247,9 @@ func newBuildNextCommand(asJSON *bool) *cobra.Command { for _, n := range notes { fmt.Fprintln(cmd.ErrOrStderr(), termsafe.Sanitize(n)) } + if _, err := loadRunners(cmd, roots); err != nil { + return loopFail(cmd.OutOrStdout(), *asJSON, prefix, err) + } o := loop.Options{Session: session, Roots: &roots} if cmd.Flags().Changed("pace") { o.Pace = &pace @@ -376,7 +392,8 @@ func newImplementStatusCommand(asJSON *bool) *cobra.Command { Use: "status [--run ]", Long: "Render the runs `abcd build` started in this checkout, or the one --run names: the\n" + "intent and spec, each lane with its spec step and next stage, what an awaiting lane\n" + - "waits on, the pending spec steps, and the run record. Read-only: it writes nothing\n" + + "waits on, the pending spec steps, the fallbacks from a routed runner to the host counted\n" + + "per runner and per role, and the run record. Read-only: it writes nothing\n" + "and creates nothing. Exit 2 when --run names no run.", Args: cobra.NoArgs, RunE: func(cmd *cobra.Command, _ []string) error { @@ -426,6 +443,7 @@ func newImplementStatusCommand(asJSON *bool) *cobra.Command { renderLaneLine(w, l) } renderPending(w, st.Pending) + renderFallbacks(w, runner.Tally(st.Fallbacks)) fmt.Fprintf(w, " record: %d line(s)\n", len(st.Record)) for _, e := range st.Record { fmt.Fprintf(w, " %s %-9s %s %s\n", e.At.Format("2006-01-02T15:04:05Z"), termsafe.Sanitize(e.Stage), @@ -455,6 +473,13 @@ func renderStepResult(w io.Writer, verb string, res loop.StepResult) { case res.Complete: fmt.Fprintf(w, "%s: %s is complete\n", verb, res.RunID) } + if r := res.Route; r != nil { + fmt.Fprintf(w, " ran on: the %s runner (asked %s; model %s)\n", termsafe.Sanitize(r.Ran), termsafe.Sanitize(r.Asked), termsafe.Sanitize(modelOrNone(r.Model))) + } + if fb := res.Fallback; fb != nil { + fmt.Fprintf(w, " fallback: %s was routed to %s, which was %s (%s); %s runs it\n", termsafe.Sanitize(fb.Role), termsafe.Sanitize(fb.Asked), + fb.Reason, termsafe.Sanitize(fsutil.RedactHome(fb.Detail)), termsafe.Sanitize(fb.Ran)) + } fmt.Fprintf(w, "next: %s\n", termsafe.Sanitize(fsutil.RedactHome(res.Next))) } @@ -507,6 +532,21 @@ func newImplementStepCommand(asJSON *bool) *cobra.Command { "A stage whose body this abcd does not carry is refused naming the spec piece that\n" + "delivers it, and the run is unchanged. A stage that fails leaves the state as it was,\n" + "so the next invocation performs it again; a completed stage is never repeated.\n\n" + + "A role routed to a command-line runner (roles..runner: claude or opencode, enabled\n" + + "under runner. in ~/.abcd/config.json) is started by the step itself when the stage\n" + + "hands the lane out: the runner gets the brief and the receipt path the host would get,\n" + + "runs in the lane's worktree (claude with the role's tools granted and nothing else asked,\n" + + "opencode under its own permission configuration), its\n" + + "transcript is stored in abcd's history store, and its receipt is verified by the stage's\n" + + "own verifier, so a verified one completes the stage in the same call and the result and\n" + + "the run record name the route that ran it. The claude runner runs in print mode with\n" + + "--bare, so the repository's hooks, plugins and configured servers do not run; opencode\n" + + "runs in run mode with --pure. A runner that is absent, refuses, fails, runs past its time,\n" + + "or writes a receipt the verifier refuses leaves the lane awaiting and the host is handed\n" + + "the role as with no runner, and the call records one fallback naming the role, the runner\n" + + "asked for, the reason and the route that runs it. A role left unset is the host's, and\n" + + "the call is exactly the host-driven step. A step that re-tells an await starts nothing.\n" + + "An interrupt kills the runner's process group.\n\n" + "The run's window clock: once the run's working window has elapsed, the call starts\n" + "nothing, writes next_eligible_at (now plus the run's pause) and exits 0 naming it; an\n" + "agent already started may still hand back its receipt. Before next_eligible_at the call\n" + @@ -524,11 +564,30 @@ func newImplementStepCommand(asJSON *bool) *cobra.Command { if err != nil { return loopFail(cmd.OutOrStdout(), *asJSON, prefix, err) } - res, err := loop.Advance(root, id, loop.DefaultStages(), loop.Options{}) + roots, notes := layered.RootsFor(root) + for _, n := range notes { + fmt.Fprintln(cmd.ErrOrStderr(), termsafe.Sanitize(n)) + } + cfg, err := loadRunners(cmd, roots) + if err != nil { + return loopFail(cmd.OutOrStdout(), *asJSON, prefix, err) + } + // An interrupt or a termination ends the context, and the runner + // kills the process group it started through its own handle: a + // harness never outlives the step that started it. + ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM) + defer stop() + res, err := loop.Drive(ctx, root, id, loop.DefaultStages(), loop.Options{}, + loop.Runners{Config: cfg, Transcripts: &lazyHistoryStore{cmd: cmd}}) if err != nil { return loopFail(cmd.OutOrStdout(), *asJSON, prefix, err) } res.Awaiting = redactAwait(res.Awaiting) + if res.Fallback != nil { + fb := *res.Fallback + fb.Detail = fsutil.RedactHome(fb.Detail) + res.Fallback = &fb + } return render(cmd.OutOrStdout(), *asJSON, res, func(w io.Writer) { renderStepResult(w, "step", res) }) }, } @@ -599,7 +658,9 @@ func newImplementRecordCommand(asJSON *bool) *cobra.Command { Long: "Render a run's record: every lane with its spec step, branch and head, the implementers'\n" + "receipts the loop verified with the model each runner reported, every verdict the loop\n" + "recorded from a validator's return, the captures each lane fixed, its pull request and\n" + - "what its landing did, the transcripts captured into the history store, and the record's\n" + + "what its landing did, the route that ran each receipt's or return's agent when a runner\n" + + "ran it, every fallback from a routed runner to the host with its count per runner and\n" + + "per role, the transcripts captured into the history store, and the record's\n" + "lines. Read-only unless --transcript is given.\n\n" + "--transcript , repeatable, captures each transcript into the history store as\n" + "`abcd history capture ` does, one capture per path, and records it in the run's\n" + @@ -668,10 +729,10 @@ func renderRunRecord(w io.Writer, rec loop.RunRecord) { if model == "" { model = "none reported" } - fmt.Fprintf(w, " receipt %s (%s; model %s)\n", termsafe.Sanitize(r.Receipt), termsafe.Sanitize(r.Role), termsafe.Sanitize(model)) + fmt.Fprintf(w, " receipt %s (%s; model %s)%s\n", termsafe.Sanitize(r.Receipt), termsafe.Sanitize(r.Role), termsafe.Sanitize(model), routeSuffix(r.Route)) } for _, v := range l.Verdicts { - fmt.Fprintf(w, " round %d at %s: %s %s\n", v.Round, shortSHA(v.HeadSHA), termsafe.Sanitize(v.Role), termsafe.Sanitize(v.Verdict)) + fmt.Fprintf(w, " round %d at %s: %s %s%s\n", v.Round, shortSHA(v.HeadSHA), termsafe.Sanitize(v.Role), termsafe.Sanitize(v.Verdict), routeSuffix(v.Route)) } if len(l.Resolves) > 0 { fmt.Fprintf(w, " resolves %s\n", termsafe.Sanitize(strings.Join(l.Resolves, ", "))) @@ -682,6 +743,11 @@ func renderRunRecord(w io.Writer, rec loop.RunRecord) { } } renderPending(w, rec.Pending) + renderFallbacks(w, rec.FallbackCounts) + for _, fb := range rec.Fallbacks { + fmt.Fprintf(w, " %s %s asked %s: %s (%s); %s ran it\n", fb.At.Format("2006-01-02T15:04:05Z"), termsafe.Sanitize(fb.Role), + termsafe.Sanitize(fb.Asked), fb.Reason, termsafe.Sanitize(fsutil.RedactHome(fb.Detail)), termsafe.Sanitize(fb.Ran)) + } fmt.Fprintf(w, " transcripts: %d captured into the history store\n", len(rec.Transcripts)) for _, t := range rec.Transcripts { how := "stored" @@ -700,6 +766,78 @@ func renderRunRecord(w io.Writer, rec loop.RunRecord) { } } +// renderFallbacks renders a run's fallback counts per runner and per role +// (itd-2609201916056194 criterion 4); a run with none renders nothing. +func renderFallbacks(w io.Writer, c runner.Counts) { + if c.Total == 0 { + return + } + fmt.Fprintf(w, " fallbacks: %d (by runner: %s; by role: %s)\n", c.Total, countList(c.ByRunner), countList(c.ByRole)) +} + +// countList renders a count map as "name n, name n", in name order. +func countList(m map[string]int) string { + names := make([]string, 0, len(m)) + for n := range m { + names = append(names, n) + } + sort.Strings(names) + parts := make([]string, 0, len(names)) + for _, n := range names { + parts = append(parts, fmt.Sprintf("%s %d", termsafe.Sanitize(n), m[n])) + } + return strings.Join(parts, ", ") +} + +// routeSuffix names the runner that ran a receipt's or a return's agent; the +// host's carry none. +func routeSuffix(r *runner.RouteRecord) string { + if r == nil { + return "" + } + return fmt.Sprintf(" — ran on the %s runner (asked %s; model %s)", termsafe.Sanitize(r.Ran), termsafe.Sanitize(r.Asked), termsafe.Sanitize(modelOrNone(r.Model))) +} + +func modelOrNone(m string) string { + if m == "" { + return "none reported" + } + return m +} + +// loadRunners reads the runner configuration a lane starts from and prints its +// diagnostics on stderr; a fault is the loop's refusal, before anything is +// created or launched. +func loadRunners(cmd *cobra.Command, roots layered.Roots) (*runner.Config, error) { + cfg, err := loop.LoadRunners(roots) + if err != nil { + return nil, err + } + for _, d := range cfg.Diagnostics { + fmt.Fprintln(cmd.ErrOrStderr(), termsafe.Sanitize(fsutil.RedactHome(d))) + } + return cfg, nil +} + +// lazyHistoryStore is the runner's transcript store: the repository's lane of +// abcd's own history store, keyed on its root commit, resolved on the first +// transcript so a step that starts no runner resolves nothing. +type lazyHistoryStore struct { + cmd *cobra.Command + store runner.TranscriptStore +} + +func (l *lazyHistoryStore) Store(name string, req runner.Request, ans runner.Answer, raw []byte) error { + if l.store == nil { + repoRoot, rootSHA, err := historyStore(l.cmd) + if err != nil { + return err + } + l.store = runner.HistoryStore{RepoRoot: repoRoot, RootSHA: rootSHA} + } + return l.store.Store(name, req, ans, raw) +} + // renderLanding renders what a lane's landing has done so far. func renderLanding(w io.Writer, pr int, ld *loop.Landing) { if ld == nil { diff --git a/internal/surface/cli/build_runner_surface_test.go b/internal/surface/cli/build_runner_surface_test.go new file mode 100644 index 000000000..b2aeac25b --- /dev/null +++ b/internal/surface/cli/build_runner_surface_test.go @@ -0,0 +1,115 @@ +package cli + +import ( + "encoding/json" + "os" + "os/exec" + "path/filepath" + "strings" + "testing" +) + +// writeMachineConfig writes the machine layer's config under the test's HOME. +func writeMachineConfig(t *testing.T, body string) { + t.Helper() + dir := filepath.Join(os.Getenv("HOME"), ".abcd") + if err := os.MkdirAll(dir, 0o700); err != nil { + t.Fatal(err) + } + if err := os.WriteFile(filepath.Join(dir, "config.json"), []byte(body), 0o600); err != nil { + t.Fatal(err) + } +} + +// TestBuildRefusesARunnerModelOffItsAllowlistBeforeTheRun is +// itd-2609201916056194 criterion 5 at the front door: a runner whose model +// route its provider does not list is refused when the lane is about to +// start, naming the key, and no run is created. +func TestBuildRefusesARunnerModelOffItsAllowlistBeforeTheRun(t *testing.T) { + repo := buildRepo(t) + writeMachineConfig(t, `{"oracle":{"api":{"local":{"base_url":"http://localhost:11434/v1","models":["qwen3-coder"]}}},`+ + `"runner":{"opencode":{"model":"local/llama-3"}}}`) + ref := refusalDocs(t, 2, "build", "itd-10", "--json") + if ref["stage"] != "runner" || !strings.Contains(ref["reason"].(string), "runner.opencode.model") { + t.Fatalf("the refusal names the runner's model key: %+v", ref) + } + runDirAbsent(t, repo.Root()) +} + +// TestAStepWhoseRunnerIsAbsentHandsTheHostTheRoleAndCountsIt is criteria 3 +// and 4 at the front door: `implement step` at a stage whose role is routed to +// a runner that is not on PATH hands the host the role as before, names the +// fallback, and `implement status` and `implement record` count it per runner +// and per role. PATH holds git's own directory alone, so no harness on the +// machine can be reached. +func TestAStepWhoseRunnerIsAbsentHandsTheHostTheRoleAndCountsIt(t *testing.T) { + repo := buildRepo(t) + repo.Write("AGENTS.md", "# AGENTS.md\n\n- Run make check.\n") + repo.Commit("conventions") + mustImplement(t, "build", "itd-10", "--json") + for _, want := range []string{"worktree", "brief"} { + if res := mustStep(t, "implement", "step", "--json"); res.PerformedStage != want { + t.Fatalf("step = %+v, want %s", res, want) + } + } + writeMachineConfig(t, `{"runner":{"opencode":{}}}`) + repo.Write(".abcd/config.json", `{"roles":{"implementer":{"runner":"opencode"}}}`) + git, err := exec.LookPath("git") + if err != nil { + t.Fatal(err) + } + if git, err = filepath.EvalSymlinks(git); err != nil { + t.Fatal(err) + } + t.Setenv("PATH", filepath.Dir(git)) + for _, h := range []string{"opencode", "claude"} { + if p, err := exec.LookPath(h); err == nil { + t.Fatalf("a real %s is reachable at %s; the test runs no real harness", h, p) + } + } + + var res struct { + Awaiting *struct { + Role string `json:"role"` + } `json:"awaiting"` + Fallback *struct { + Role string `json:"role"` + Asked string `json:"asked"` + Reason string `json:"reason"` + Ran string `json:"ran"` + } `json:"fallback"` + } + if err := json.Unmarshal([]byte(mustImplement(t, "implement", "step", "--json")), &res); err != nil { + t.Fatal(err) + } + if res.Awaiting == nil || res.Awaiting.Role != "implementer" { + t.Fatalf("the host is handed the implementer: %+v", res) + } + if fb := res.Fallback; fb == nil || fb.Asked != "opencode" || fb.Reason != "absent" || fb.Ran != "host" { + t.Fatalf("the step names the fallback: %+v", res.Fallback) + } + text := mustImplement(t, "implement", "step") + if !strings.Contains(text, "awaits the implementer's receipt") { + t.Fatalf("a step while awaiting re-tells the await: %s", text) + } + status := mustImplement(t, "implement", "status") + if !strings.Contains(status, "fallbacks: 1 (by runner: opencode 1; by role: implementer 1)") { + t.Fatalf("status counts the fallback per runner and per role:\n%s", status) + } + var rec struct { + FallbackCounts struct { + Total int `json:"total"` + ByRunner map[string]int `json:"by_runner"` + ByRole map[string]int `json:"by_role"` + } `json:"fallback_counts"` + } + if err := json.Unmarshal([]byte(mustImplement(t, "implement", "record", "--json")), &rec); err != nil { + t.Fatal(err) + } + if c := rec.FallbackCounts; c.Total != 1 || c.ByRunner["opencode"] != 1 || c.ByRole["implementer"] != 1 { + t.Fatalf("the run record counts the fallback: %+v", c) + } + if text := mustImplement(t, "implement", "record"); !strings.Contains(text, "fallbacks: 1 (by runner: opencode 1; by role: implementer 1)") { + t.Fatalf("the record's text counts the fallback:\n%s", text) + } +}