Skip to content

Verify published evals and invalidate the website cache from main CI - #122

Merged
gaojude merged 3 commits into
mainfrom
codex/evals-publishing
Sep 9, 2026
Merged

gaojude merged 3 commits into
mainfrom
codex/evals-publishing

Conversation

@gaojude

@gaojude gaojude commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Publish result updates by merging agent-results.json here, without copying it into the website repository. CI verifies that the committed JSON matches the saved results and, after successful checks on main, calls https://nextjs.org/api/evals/revalidate to refresh the website's indefinitely cached snapshot.

The new revalidate-site job runs only for main pushes or manual dispatches and authenticates with GitHub Actions OIDC. It retries delivery three times and fails on an unacknowledged or unsuccessful response. The website permits stale server responses for up to one hour after invalidation while refreshing in the background, then requires fresh data. PR runs validate results without calling production.

pnpm export-results --check compares the committed file with a complete export, ignores only the export timestamp, and never writes it. Full exports use deterministic experiment order; errors exit nonzero and empty exports are rejected. The JSON format remains unversioned.

Deploy the website's new endpoint before merging this producer integration. No shared secret needs configuration. Retry a missed notification by rerunning the failed job or manually dispatching this workflow on main; there is no periodic refresh fallback.

Validation:

  • Export verification and the existing stale-result check passed against 754 results: 641 passes, 64 failures, 49 N/A. The previous PR revision also passed the export check in GitHub Actions.
  • A timestamp-only change passed verification; a changed metric failed with exit code 1. Verification left the file untouched in both cases.
  • TypeScript check passed for the exporter.
  • Executed the new workflow script with isolated HTTP/identity stubs: successful delivery, transient HTTP/network retries, rejection of missing acknowledgments, and failure after three attempts all passed.
  • The frontend tests verified signed-token authorization and the actual Next.js one-hour invalidation boundary. A real authenticated production callback awaits deployment and a successful main run.
  • No evals were run or result artifacts changed by this PR.

@gaojude
gaojude marked this pull request as ready for review September 9, 2026 22:20
@gaojude gaojude changed the title Verify and document published eval exports Verify published evals and invalidate the website cache from main CI Sep 9, 2026
@socket-security

Copy link
Copy Markdown

Review the following changes in direct dependencies. Learn more about Socket for GitHub.

Diff Package Supply Chain
Security
Vulnerability Quality Maintenance License
Addedgithub/​actions/​github-script@​ed597411d8f924073f98dfc5c65a23a2325f34cd99100100100100

View full report

@gaojude
gaojude merged commit 3d1399e into main Sep 9, 2026
5 checks passed
gaojude added a commit that referenced this pull request Sep 11, 2026
…TS.md) (#124)

* Add Gemini 3.8 Flash eval results (28/31 base, 30/31 with AGENTS.md)

Gemini 3.8 Flash shipped 2026-09-02 and was missing from nextjs.org/evals.
Registered, measured over the full 31-eval set at pass@4, and tiered.

  Gemini 3.8 Flash              28/31 (90%)
  Gemini 3.8 Flash + AGENTS.md  30/31 (97%)   +6pp

AGENTS.md turns agent-029 (use cache directive) and agent-031 (proxy)
around and loses nothing. agent-044 (uses-nextjs) fails on both sides for
the same reason every time: the model builds the app on Express, ignoring
the framework requirement. Mean list cost $0.605/eval, mean duration 435s.

Harness is vercel-ai-gateway/opencode, not the `gemini` (Gemini CLI)
harness the older Gemini pairs use. That harness talks to the Google API
directly and needs GEMINI_API_KEY, which is the same missing credential
that keeps gemini-3.1-pro-preview on tier 2. The gateway serves
google/gemini-3.8-flash, so OpenCode is the path that can actually produce
a fresh measurement — the same path glm, kimi, grok and minimax take. The
board renders the harness per row, so the difference is visible.

No reasoning-effort pin and no effort suffix on the label, matching every
other OpenCode row. OpenCode maps a model's options onto
providerOptions.gateway for @ai-sdk/gateway providers and the gateway does
not act on reasoningEffort there: at `low` the model still spent ~1.4k
reasoning tokens, where the same `low` sent as reasoning_effort to
/v1/chat/completions spends 0. Labelling the row `(high)` would claim a
setting the harness never sent.

timeout is 2400 rather than the 1200 the other OpenCode pairs use. The
prefetch evals run 1000-1400s for this model; at 1200 the first matrix lost
agent-051 (base) and agent-049 (AGENTS.md) to timeouts, and a narrower
re-run lost agent-051 again, so it was the budget rather than the burst.
Those results were discarded and both experiments re-run from scratch at
the new fingerprint — nothing here is measured at the old ceiling.

Tier 1: this takes the Gemini Flash line's slot. Lines are tiered
separately (Claude holds Fable, Opus and Sonnet slots at once) and Google
still ships Pro and Flash side by side, so the Pro rows are untouched and
stay tier 2 for want of GEMINI_API_KEY. No Flash-line model has been
measured here before, so the under-a-month carve-out has no predecessor to
attach to.

Pricing is introductory: $0.75/$3.75 per 1M input/output through
2026-12-31, doubling to $1.50/$7.50 on 2027-01-01. The gateway catalog and
the models.dev `vercel` entry agree on all three rates today; the row needs
a re-export in January even if nothing is rerun.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com>

* Correct the publishing story: the site no longer holds a copy

AGENTS.md and the add-eval-model skill still said nextjs.org/evals reads a
copy of agent-results.json checked into vercel/front, and told you to open
a second PR there carrying it. That stopped being true on 2026-09-09: #122
here and vercel/front#85415 moved the site to a server fetch of this repo's
main, and the copy at apps/next-site/app/(next-site)/evals/agent-results.json
was deleted. Following the old instructions means hunting for a file that
does not exist, or recreating one the site does not read.

Merging here is the publish. A front PR is for site-side changes only.

Also record what this run cost a debug cycle:

- Pick the harness by what preflight can authenticate, not by the vendor.
  The gemini and cursor harnesses are direct-vendor-API only and get
  silently skipped without their keys; AI_GATEWAY_API_KEY covers every
  vercel-ai-gateway/* harness.
- A model newer than the pinned OpenCode binary needs extraProviders.
- Effort has no working knob on the OpenCode path, so those rows publish
  unpinned and unsuffixed.
- reasoning_effort is not validated for every model — google/gemini-3.8-flash
  returns 200 for a rung it does not have. Compare reasoning_tokens instead
  of trusting the 400.
- A timeout that is too tight cannot be re-run away, and raising it strands
  results under the old fingerprint unless you delete results/<slug>/.
- Node 24 and the test/format commands for the front side.

And two numbers that had drifted: the eval set is 31, not 26, and a matrix
is ~124 concurrent attempts, not ~104.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com>

---------

Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants