Skip to content

fix: alfworld benchmark graph - #554

Open
dalongbao (dalongbao) wants to merge 2 commits into
microsoft:mainfrom
dalongbao:main
Open

fix: alfworld benchmark graph#554
dalongbao (dalongbao) wants to merge 2 commits into
microsoft:mainfrom
dalongbao:main

Conversation

@dalongbao

Copy link
Copy Markdown
Contributor

No description provided.

Copilot AI lite review requested due to automatic review settings August 21, 2026 17:59

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Updates the skills benchmark documentation to align with a regenerated ALFWorld “success vs finale cost” chart, reflecting the baseline presentation used across the benchmark plots.

Changes:

  • Adjust the “Performance versus finale cost” narrative in skills/README.md to describe a single pristine-baseline reference point for each chart.
  • Regenerate agent-lightning-alfworld-success-finale-cost.svg to reflect the corrected ALFWorld finale-cost visualization (axes/ticks/layout and baseline depiction).

Reviewed changes

Copilot reviewed 1 out of 2 changed files in this pull request and generated 1 comment.

File Description
skills/README.md Updates the explanatory text for the finale-cost plots to match the charts’ baseline/reference-point presentation.
skills/assets/agent-lightning-alfworld-success-finale-cost.svg Updates the ALFWorld finale-cost graph output to the corrected version.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread skills/README.md
#### Performance versus finale cost

The selected-budget views use the groups with the strongest aggregate skill-over-control lift: \$5 for SpreadsheetBench and \$10 for OfficeQA and ALFWorld. Every harness/treatment point is one of three runs; the x-axis is that run's finale cost, and the y-axis is held-out SpreadsheetBench accuracy, OfficeQA correctness, or ALFWorld success. Finale cost measures LLM gateway spend, so an ALFWorld deterministic controller can have exactly \$0 finale cost while still executing and scoring real environment steps; coincident zero-cost ALFWorld results are offset slightly along the x-axis so each replicate remains visible. SpreadsheetBench and OfficeQA show their aggregate pristine-baseline results as single reference points. The dotted ALFWorld baseline is a score-only reference: the corrected records do not include baseline deployment cost, so assigning it an x-coordinate would invent data.
The selected-budget views use the groups with the strongest aggregate skill-over-control lift: \$5 for SpreadsheetBench and \$10 for OfficeQA and ALFWorld. Every harness/treatment point is one of three runs; the x-axis is that run's finale cost, and the y-axis is held-out SpreadsheetBench accuracy, OfficeQA correctness, or ALFWorld success. Finale cost measures LLM gateway spend, so an ALFWorld deterministic controller can have exactly \$0 finale cost while still executing and scoring real environment steps; coincident zero-cost ALFWorld results are offset slightly along the x-axis so each replicate remains visible. Each chart shows its aggregate pristine-baseline result as a single reference point.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants