Skip to content

feat(bench): add product-level agent benchmark harness - #428

Open
e54-bot wants to merge 1 commit into
mainfrom
wright-414-agent-benchmark
Open

e54-bot wants to merge 1 commit into
mainfrom
wright-414-agent-benchmark

Conversation

@e54-bot

@e54-bot e54-bot commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • Add benchmarks/agent/agent_bench.py (stdlib only): runs any agent command against a scenario workspace in baseline (no wright on PATH) or wright (logging shim) conditions with an identical task, then grades deterministically.
  • Add four raw Workshop scenarios: greenfield, understanding, modification, diagnosis/repair.
  • Define wright-agent-bench/v1 in docs/agent-benchmark.md: allowed context, scenario format, checks with failure layers, unsafe-edit detection, runtime-only claims, Wright-use trace (correction rounds, exit 3/4 gaps).
  • validate requires each reference to pass and each seed to fail; wired into the existing non-blocking CI benchmark job.

Refs #414. Not closing: only raw Workshop is exercised; OPY and DEL/OSTW scenarios wait on owner-declared capability support, and no real agent run has been recorded.

Test plan

  • python3 benchmarks/agent/agent_bench.py validate
  • python3 -m unittest discover -s benchmarks/agent (includes a fake-agent run proving baseline has 0 Wright invocations and the wright condition traces them)

Add a stdlib harness, four raw Workshop scenarios (greenfield, understanding, modification, diagnosis), a baseline-vs-Wright condition, and a wright-agent-bench/v1 result contract. Scenario validity is checked in the CI benchmark job.

Refs #414

@Teakowa Teakowa left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified locally on db33668: agent_bench.py validate passes all four scenarios against a real target/debug/wright, the unittest suite passes, and the harness's CLI assumptions (check/lint JSON envelopes, while-without-wait, serve symbols op, exit >= 3 as owner/environment gap) match docs/cli/machine-contract.md. CI is green and the issue scoping (Refs #414, raw Workshop only) is honest. A few actionable defects:

1. CI never validates the content this PR adds. validate runs only inside the benchmark job, gated on needs.paths.outputs.opy == 'true', but benchmarks/** is not in the rust_core path filter in .github/workflows/ci.yml. A later PR touching only scenario files skips the benchmark job; a prompt.md-only PR skips CI entirely via paths-ignore: "**/*.md". The only automated guard on scenario validity never fires for the changes it guards. Add benchmarks/** to the rust_core filter.

2. declares-winner accepts a wrong outcome. benchmarks/agent/scenarios/greenfield-elimination-race/scenario.json accepts "Declare Match Draw" as satisfying "the first player to reach 7 points wins". A draw does not meet the requirement; the check certifies failure. Drop the alternative (keep Declare Player Victory).

3. excludes-self-kills requires one exact spelling. The check requires literal Attacker != Victim, but the equivalent canonical form Compare(Attacker, !=, Victim) parses cleanly (wright check exits 0, no diagnostics) while failing this check — a false agent-layer failure on a correct solution. The scenario's own reference already uses Compare(...) for the winner condition, so this spelling is plausible agent output. Add it to the text alternatives. Related: modify-target-score's rule-count-preserved uses min: 6 only, so an agent that adds an extra rule still passes; consider an exact/max bound on contains if "change nothing else" is meant to be enforced.

4. Timed-out trials are graded while the agent may still be running. agent_bench.py run_trial uses subprocess.run(shell=True, timeout=...); on timeout only the shell is killed, so a surviving agent child can keep mutating workspace/ while grade() runs — nondeterministic results exactly when trials time out. Use start_new_session=True and kill the process group on TimeoutExpired.

5. Test crashes instead of skipping when target/ is absent. test_agent_bench.py setUp: mkdtemp(dir=ROOT/"target") raises FileNotFoundError when WRIGHT_BIN points to an external binary and the repo has no target/ (reproduced in a clean worktree). Create the dir first or use the default tempdir.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

2 participants