Conversation
Add a stdlib harness, four raw Workshop scenarios (greenfield, understanding, modification, diagnosis), a baseline-vs-Wright condition, and a wright-agent-bench/v1 result contract. Scenario validity is checked in the CI benchmark job. Refs #414
Teakowa
left a comment
There was a problem hiding this comment.
Verified locally on db33668: agent_bench.py validate passes all four scenarios against a real target/debug/wright, the unittest suite passes, and the harness's CLI assumptions (check/lint JSON envelopes, while-without-wait, serve symbols op, exit >= 3 as owner/environment gap) match docs/cli/machine-contract.md. CI is green and the issue scoping (Refs #414, raw Workshop only) is honest. A few actionable defects:
1. CI never validates the content this PR adds. validate runs only inside the benchmark job, gated on needs.paths.outputs.opy == 'true', but benchmarks/** is not in the rust_core path filter in .github/workflows/ci.yml. A later PR touching only scenario files skips the benchmark job; a prompt.md-only PR skips CI entirely via paths-ignore: "**/*.md". The only automated guard on scenario validity never fires for the changes it guards. Add benchmarks/** to the rust_core filter.
2. declares-winner accepts a wrong outcome. benchmarks/agent/scenarios/greenfield-elimination-race/scenario.json accepts "Declare Match Draw" as satisfying "the first player to reach 7 points wins". A draw does not meet the requirement; the check certifies failure. Drop the alternative (keep Declare Player Victory).
3. excludes-self-kills requires one exact spelling. The check requires literal Attacker != Victim, but the equivalent canonical form Compare(Attacker, !=, Victim) parses cleanly (wright check exits 0, no diagnostics) while failing this check — a false agent-layer failure on a correct solution. The scenario's own reference already uses Compare(...) for the winner condition, so this spelling is plausible agent output. Add it to the text alternatives. Related: modify-target-score's rule-count-preserved uses min: 6 only, so an agent that adds an extra rule still passes; consider an exact/max bound on contains if "change nothing else" is meant to be enforced.
4. Timed-out trials are graded while the agent may still be running. agent_bench.py run_trial uses subprocess.run(shell=True, timeout=...); on timeout only the shell is killed, so a surviving agent child can keep mutating workspace/ while grade() runs — nondeterministic results exactly when trials time out. Use start_new_session=True and kill the process group on TimeoutExpired.
5. Test crashes instead of skipping when target/ is absent. test_agent_bench.py setUp: mkdtemp(dir=ROOT/"target") raises FileNotFoundError when WRIGHT_BIN points to an external binary and the repo has no target/ (reproduced in a clean worktree). Create the dir first or use the default tempdir.
Summary
benchmarks/agent/agent_bench.py(stdlib only): runs any agent command against a scenario workspace inbaseline(nowrighton PATH) orwright(logging shim) conditions with an identical task, then grades deterministically.wright-agent-bench/v1indocs/agent-benchmark.md: allowed context, scenario format, checks with failure layers, unsafe-edit detection, runtime-only claims, Wright-use trace (correction rounds, exit 3/4 gaps).validaterequires each reference to pass and each seed to fail; wired into the existing non-blocking CI benchmark job.Refs #414. Not closing: only raw Workshop is exercised; OPY and DEL/OSTW scenarios wait on owner-declared capability support, and no real agent run has been recorded.
Test plan
python3 benchmarks/agent/agent_bench.py validatepython3 -m unittest discover -s benchmarks/agent(includes a fake-agent run proving baseline has 0 Wright invocations and the wright condition traces them)