Evaluation of this test
' + for key,label in [('question','Question'),('test_quality','Checks'),('limitation','Limits'),('decision','Conclusion'),('next_action','Next step')]: + content+=f'{label}
{esc(result["assessment"][key])}
' + content+='BBEH MINI SUBSET · CALIBRATION ONLY
'+esc(result['assessment']['finding'])+'
| Workflow | Correct / received | Missing / 12 | Score / 100 |
|---|---|---|---|
| {labels[arm]} | {s["correct"]}/{s["evaluated"]} | {s["missing"]} | {score} |
Missing-answer bounds show the range if every missing answer were wrong or correct. They are not confidence intervals. Sonnet’s one correct answer out of five received is not a complete 12-question accuracy score.
Four questions each from Multistep Arithmetic, Web of Lies and Hyperbaton formed the calibration. Another eight per family were reserved for the main comparison. The predefined gate required all 36 calibration responses and 3–9 correct answers per baseline, plus sufficient remaining runtime. Seven Sonnet timeouts prevented that gate from passing.
The planned main comparison would have tested Terra review against Sonnet self-review, plain revision and Terra alone, using shared original answers and an official deterministic scorer. Those review arms were never launched. Both participants used medium effort: claude-sonnet-5 and gpt-5.6-terra.
Eight answers match the normalized reference except for surrounding angle brackets, which the official scorer does not accept: three Sonnet originals, three revisions and two Terra answers. The official scores above are unchanged. This diagnostic means the low score cannot be interpreted purely as weak reasoning.
{spend["calls"]}/220 calls used: 17 Sonnet and 12 Terra. The controller ran for about six minutes. No conditional main calls were launched. Known list-equivalent cost was ${spend["known_list_equivalent_usd"]:.4f}, with seven calls unpriced; this is an incomplete cost estimate, not a total bill.
{esc(result["assessment"][key])}
' + content+='This is a selected subset of BBEH Mini, not an official leaderboard submission. Calibration inputs, answers, references, hashes and charged attempts are available in the data download. Unused main answer keys are excluded.
' + content+=f'Official benchmark commit: {esc(result["benchmark"]["commit"])}.
Official BBEH source ↗ · Download results ↓ · All test assessments → · Earlier Sonnet/Terra planning results →
' + return shell('Co-Evolution · BBEH calibration and assessment',content) diff --git a/benchmarks/site/build-evaluations.py b/benchmarks/site/build-evaluations.py index 7244d7e..96abad9 100644 --- a/benchmarks/site/build-evaluations.py +++ b/benchmarks/site/build-evaluations.py @@ -78,6 +78,9 @@ def build(): result=json.loads((PUBLIC/e['data']).read_text(encoding='utf-8')) if result.get('schema')=='planbench-results/1.0': (PUBLIC/e['page']).write_text(render_planbench(result,e['data'],e['id']),encoding='utf-8',newline='\n') + elif result.get('schema')=='bbeh-results/1.0': + spec=importlib.util.spec_from_file_location('bbeh_page',SITE/'bbeh-page.py');module=importlib.util.module_from_spec(spec);spec.loader.exec_module(module) + (PUBLIC/e['page']).write_text(module.render(result,shell),encoding='utf-8',newline='\n') print('Rendered test assessments and PlanBench outcome.') if __name__=='__main__':build() diff --git a/benchmarks/site/observatory.html b/benchmarks/site/observatory.html index c9dfe28..7f95e35 100644 --- a/benchmarks/site/observatory.html +++ b/benchmarks/site/observatory.html @@ -29,7 +29,7 @@One model writes the code. Another reviews it.
See what changes in accuracy, cost, and reliability.
Real repository issues. Official evaluation. Open evidence. Read the test assessments → · Sonnet/Terra planning results → · Astra/Fable →
+Real repository issues. Official evaluation. Open evidence. Read the test assessments → · BBEH reasoning screen → · Sonnet/Terra planning results → · Astra/Fable →
diff --git a/benchmarks/site/public/bbeh-results.json b/benchmarks/site/public/bbeh-results.json new file mode 100644 index 0000000..581eceb --- /dev/null +++ b/benchmarks/site/public/bbeh-results.json @@ -0,0 +1,2380 @@ +{ + "schema": "bbeh-results/1.0", + "title": "BBEH compact reasoning screen", + "generated_at": "2026-09-14T19:55:09.708499+00:00", + "completion": "gate-rejected", + "benchmark": { + "name": "BIG-Bench Extra Hard Mini subset", + "commit": "80d12ca916b7158f22293fcf3144f4d3d854d4be", + "families": [ + "multistep_arithmetic", + "web_of_lies", + "hyperbaton" + ], + "full_benchmark_score": false, + "official_scorer_canonical_sha256": "4b4f06e5babb015de2ba639bae995a5526182ffb5b5890af54dfcf038580eb34" + }, + "models": { + "sonnet": "claude-sonnet-5", + "codex": "gpt-5.6-terra" + }, + "effort": "medium", + "gate": { + "passed": false, + "scoreable": false, + "correct": { + "A": 1, + "B": 1, + "E": 3 + }, + "denominator": 12, + "required_range": [ + 3, + 9 + ], + "estimated_main_seconds": 1522.5327, + "remaining_seconds": 5877.079237937927 + }, + "calibration": { + "scores": { + "A": { + "workflow": "Sonnet original", + "correct": 1, + "incorrect": 4, + "missing": 7, + "evaluated": 5, + "planned": 12, + "score": null, + "score_bounds": [ + 8.333333333333334, + 66.66666666666667 + ] + }, + "B": { + "workflow": "Sonnet plain revision", + "correct": 1, + "incorrect": 4, + "missing": 7, + "evaluated": 5, + "planned": 12, + "score": null, + "score_bounds": [ + 8.333333333333334, + 66.66666666666667 + ] + }, + "E": { + "workflow": "Terra alone", + "correct": 3, + "incorrect": 9, + "missing": 0, + "evaluated": 12, + "planned": 12, + "score": 25.0, + "score_bounds": [ + 25.0, + 25.0 + ] + } + }, + "outcomes": [ + { + "question": "web_of_lies-e5dbc0fee8cb0bf8", + "family": "web_of_lies", + "arm": "A", + "response": "**Reasoning summary:**\n\n- Sal=true ⇒ Gwenn=false. Octavia/Ocean form an opposite pair (exactly one true), so among {Gwenn,Octavia,Ocean} exactly one is true → Kandi's \"exactly one\" claim is true → Kandi=true.\n- Marcus/Mariana form a mirror pair (exactly one true). Shenna's claim \"exactly two of Kandi(true), Marcus, Mariana\" is satisfied regardless of which of Marcus/Mariana is true → Shenna=true, but Marcus's individual value stays undetermined (no other clue pins it) → **Marcus = unknown**.\n- Chain: Jamey's, Antwan's, Jaymie's, Alexis's claims all fail regardless of the underlying opposite-pairs → all false → Teressa=true (Alexis lies) → Sima's claim fails regardless of Maynard/Osborn → Sima=false → Ryan's claim holds regardless of Virginia/Odell → Ryan=true → Dallas's \"only one of Ryan,Mitchell,Oprah lies\" holds regardless (Mitchell/Oprah opposite) → Dallas=true → Bernita (\"Dallas lies\") = false.\n- Mandy/Mason opposite pair + Bernita false gives 2 liars among {Bernita,Mandy,Mason} → Willian's \"only one lies\" is false → **Willian = no**.\n- Melvin/Osian opposite pair → Andree's conjunction (all three true) is always false → Andree=false → Fletcher (says Andree true) = false. With Fletcher=false, Shalonda's \"exactly one or all three\" of {Fletcher,Opal,Ozzie} fails for every possible Opal/Ozzie value (they're a matching pair) → **Shalonda = no**.\n\nThe final answer is:BBEH MINI SUBSET · CALIBRATION ONLY
Terra completed 12/12 calibration answers and scored 3/12 (25%) under the official scorer. Sonnet completed five originals, with one correct; seven originals timed out at 120 seconds. Its five available plain revisions also had one correct, with seven dependent revisions blocked. The controller stopped after 29 calls; the conditional 168-call main comparison did not run.
| Workflow | Correct / received | Missing / 12 | Score / 100 |
|---|---|---|---|
| Sonnet original | 1/5 | 7 | Unavailable Bounds 8.3–66.7% |
| Sonnet plain revision | 1/5 | 7 | Unavailable Bounds 8.3–66.7% |
| Terra alone | 3/12 | 0 | 25.0 |
Missing-answer bounds show the range if every missing answer were wrong or correct. They are not confidence intervals. Sonnet’s one correct answer out of five received is not a complete 12-question accuracy score.
Four questions each from Multistep Arithmetic, Web of Lies and Hyperbaton formed the calibration. Another eight per family were reserved for the main comparison. The predefined gate required all 36 calibration responses and 3–9 correct answers per baseline, plus sufficient remaining runtime. Seven Sonnet timeouts prevented that gate from passing.
The planned main comparison would have tested Terra review against Sonnet self-review, plain revision and Terra alone, using shared original answers and an official deterministic scorer. Those review arms were never launched. Both participants used medium effort: claude-sonnet-5 and gpt-5.6-terra.
Eight answers match the normalized reference except for surrounding angle brackets, which the official scorer does not accept: three Sonnet originals, three revisions and two Terra answers. The official scores above are unchanged. This diagnostic means the low score cannot be interpreted purely as weak reasoning.
29/220 calls used: 17 Sonnet and 12 Terra. The controller ran for about six minutes. No conditional main calls were launched. Known list-equivalent cost was $0.8945, with seven calls unpriced; this is an incomplete cost estimate, not a total bill.
Does this compact BBEH subset provide usable headroom for a low-compute comparison of Co-Evolution with cheaper workflows?
The 12 calibration and 24 disjoint main questions were frozen before generation. Participant prompts excluded answer keys. All 22 received answers could be extracted, and their unchanged text was graded with the pinned official deterministic scorer after generation froze. Publication replays that scorer and verifies input and response hashes. Missing answers remain unavailable, not incorrect.
This is a 12-question, three-family suitability screen, not a full BBEH score or a Co-Evolution effectiveness test. Sonnet full-set accuracy is unknown, with bounds of 8.3–66.7%. Eight officially incorrect responses (three A, three B, two E) match the normalized reference except for surrounding angle brackets. This diagnostic does not change the scores, but shows that formatting affects the apparent difficulty. Seven timed-out calls have no recorded price, so the cost total is incomplete.
Reject this bundle under the configured limits and stop at the predeclared gate. The observed scores are below the earlier ceiling, but Sonnet timeouts and answer-format sensitivity prevent a clean reasoning comparison. No positive or negative Co-Evolution effect has been measured. The gate avoided launching 168 conditional main calls.
Publish this calibration and its assessment, then close the run. A separate plan should first address exact answer-format instructions and runtime suitability using fresh calibration questions. Web of Lies completed across all baselines here, but four questions and substantial formatting effects do not establish it as a suitably difficult replacement. Do not change settings or continue this run.
This is a selected subset of BBEH Mini, not an official leaderboard submission. Calibration inputs, answers, references, hashes and charged attempts are available in the data download. Unused main answer keys are excluded.
Official benchmark commit: 80d12ca916b7158f22293fcf3144f4d3d854d4be.
Official BBEH source ↗ · Download results ↓ · All test assessments → · Earlier Sonnet/Terra planning results →
EVALUATING THE TESTS
Every current result publication needs an assessment of what was tested, what the results establish, their limitations, and the next decision. Scores alone do not establish practical benefit.
complete · 200/200 scored plans; smoke tasks excluded
Does Terra review improve Sonnet plans beyond plain revision or independent Sonnet self-review on the same fixed PlanBench subset?
Sonnet originals scored 48/50 (96%). Plain revision, Sonnet self-review, and Terra review followed by Sonnet revision each scored 50/50 (100%). The two original plans executed legally but failed to satisfy the complete goal; all three revision approaches fixed them. Terra review added 0 points over either revision control and 4 points over the original draft.
All 200 scored plans were produced and evaluated after generation froze. The same pinned 50 tasks, original drafts, role prompts, upstream PDDL extractor and official VAL were used across arms. The two failed originals were goal failures, not parsing errors. Model routing, budget allocation, retry accounting, and exact exported plan hashes were checked; no validator feedback reached participants.
The 96% original score is still ceiling-limited, and both controls reached 100%, leaving no room for a positive D-C effect. Fifty tasks and one response per arm cannot establish general equivalence. The [0,0] paired bootstrap interval resamples ties, not population certainty. Before any scored task, a readiness amendment changed both models from high to medium effort and gave Claude 8,192 combined reasoning/response tokens; six initial calls remain charged and excluded. Comparisons with the earlier high-effort Astra run are not model-only. Validity does not measure optimality, human rework, or software delivery.
Use plain Sonnet revision for this benchmark: it achieved the same 100% score at an estimated $0.0552 per task versus $0.0655 with Terra review. Terra review cost about 19% more than plain revision. It was about 20% cheaper than Sonnet self-review ($0.0821), so it was the more economical reviewer here, but it did not outperform the cheaper plain-revision workflow. The predeclared score-improvement threshold was not met.
Do not spend another full run on this saturated subset. If continuing to measure accuracy gains, specify a harder recognized benchmark or task stratum before running, keep matched revision controls, and retain cost and coverage reporting. Treat the observed reviewer cost advantage separately from claims of better planning quality or productivity.
partial · 195/200 scored plans; smoke tasks excluded
Does Fable critique improve Astra valid-plan rate beyond self-review and plain revision?
Original Astra, plain revision and self-review each produced 50/50 valid plans (100%). Fable review delivered 45/50 plans: 45 valid, 0 invalid and 5 missing. Cross-model review gained zero points on the jointly evaluated tasks. With the missing outcomes unresolved, its full-cohort score can only be 90–100.
The official 110-task corpus, fixed 50-task manifest, upstream PDDL extractor and VAL were pinned. Offline valid/invalid/malformed fixtures and duplicate-free resume passed. Candidates were frozen before scoring. All arms share the same original per task; no validator feedback reached participants.
This is a continued, one-generation-per-arm 50-task subset, not the complete leaderboard. Two documented continuations preserved earlier successes and charged attempts; known request-specific refusals could receive one identical retry within the original reserve. Observed workflow latency includes queueing and recovery suspension, so it does not isolate intrinsic review latency. Static public tasks may have training exposure. Symbolic validity does not measure human rework or software delivery. Astra token caps are prompt targets, not an enforced CLI limit; provider effort labels do not establish equal compute. The [0,0] empirical bootstrap interval resamples only ties; it does not prove population-level equivalence. VAL validity also does not assess shortest-plan optimality.
Partial measurement: paired observed outcomes are descriptive, and missing tasks prevent a complete fixed-50 benchmark claim. Even the most favorable missing outcomes cannot meet the predeclared score-improvement threshold. The completed original-draft arm is also ceiling-limited under the predeclared rule. For this benchmark, prefer the original Astra workflow: it already reached every goal. Additional review produced no measured validity gain and introduced delivery failures and extra compute.
Do not repeat the same matrix on this saturated subset. Keep the simpler workflow for tasks at this level. Any further study should use a separately specified benchmark with more headroom, preserve a fresh scored cohort, and measure application outcomes before claiming productivity benefits.
partial · 194/198 plans; Astra 194 and Fable 187 judgments
Does cross-model critique improve written plans beyond original drafts, plain revision and matched self-review?
Across 18 cross-model workflows, post-hoc equal-brief mean gains over drafts were +7.48 Astra and +5.58 Fable points; over matched self-review, +4.08 and +1.94. The predeclared Sonnet-with-Terra contrast was +9.00 Astra but -1.00 Fable over five paired briefs. These are rubric points, not productivity percentages.
The two judges were isolated from authorship and each other; supporting citations were validated. Successful outputs and charged attempts were preserved through recovery. Exported coverage and row means were checked against saved data, and original study hashes matched.
Only six synthetic briefs underlie the overlapping comparisons. Generation resumed after partial grading and beyond the original deadline. Astra flagged critical violations in 110/194 judgments, Fable in 4/187. No task execution, human editing time or downstream rework was measured. Post-hoc aggregate gains are exploratory.
There is a positive average plan-score signal, but no established general productivity benefit or reliable superiority over plain revision. Do not require multi-model review for every task on this evidence alone.
Read the separately reported PlanBench follow-up for objective symbolic-plan validity. Keep those scores separate from judge-rated writing quality; neither study alone measures human productivity.
mixed · 50-task subset; current and interim rows have separate denominators
Do the recorded coding workflows solve more repository issues under the official SWE-bench evaluator?
The completed base50-light rows report Sonnet solo at 39/50 (78%), Sonnet followed by Terra at 42/50 (84%), and Terra solo at 33/50 (66%). The first two differ by three solved tasks, or six percentage points. Other rows include partial samples and different run cohorts.
The site consumes saved official evaluator exports with task-level outcomes and provenance. Source-embedding and archive-preservation checks protect the published snapshot. This assessment does not rerun the coding experiments or treat site regression tests as benchmark performance.
These are workflow-level outcomes on one public 50-task subset, not full SWE-bench leaderboard results. Different executors and incomplete self-review controls limit isolation of reviewer benefit. Partial cohorts and single-shot models cannot be pooled with complete coding-agent rows; some resource accounting is estimated or incomplete.
The observed six-point gain is worth investigating, but is not by itself proof that cross-vendor review caused it or that the benefit generalizes. Use the paired task evidence and matched controls rather than ranking partial percentages.
Preserve this coding snapshot and its distinct cohorts. Future claims should compare fixed matched workflows on the same task set, disclose missingness and resources, and refresh this assessment when the underlying results change.