Verify published evals and invalidate the website cache from main CI - #122
Merged
Merged
Conversation
gaojude
marked this pull request as ready for review
September 9, 2026 22:20
aurorascharff
approved these changes
Sep 9, 2026
|
Review the following changes in direct dependencies. Learn more about Socket for GitHub.
|
gaojude
added a commit
that referenced
this pull request
Sep 11, 2026
…TS.md) (#124) * Add Gemini 3.8 Flash eval results (28/31 base, 30/31 with AGENTS.md) Gemini 3.8 Flash shipped 2026-09-02 and was missing from nextjs.org/evals. Registered, measured over the full 31-eval set at pass@4, and tiered. Gemini 3.8 Flash 28/31 (90%) Gemini 3.8 Flash + AGENTS.md 30/31 (97%) +6pp AGENTS.md turns agent-029 (use cache directive) and agent-031 (proxy) around and loses nothing. agent-044 (uses-nextjs) fails on both sides for the same reason every time: the model builds the app on Express, ignoring the framework requirement. Mean list cost $0.605/eval, mean duration 435s. Harness is vercel-ai-gateway/opencode, not the `gemini` (Gemini CLI) harness the older Gemini pairs use. That harness talks to the Google API directly and needs GEMINI_API_KEY, which is the same missing credential that keeps gemini-3.1-pro-preview on tier 2. The gateway serves google/gemini-3.8-flash, so OpenCode is the path that can actually produce a fresh measurement — the same path glm, kimi, grok and minimax take. The board renders the harness per row, so the difference is visible. No reasoning-effort pin and no effort suffix on the label, matching every other OpenCode row. OpenCode maps a model's options onto providerOptions.gateway for @ai-sdk/gateway providers and the gateway does not act on reasoningEffort there: at `low` the model still spent ~1.4k reasoning tokens, where the same `low` sent as reasoning_effort to /v1/chat/completions spends 0. Labelling the row `(high)` would claim a setting the harness never sent. timeout is 2400 rather than the 1200 the other OpenCode pairs use. The prefetch evals run 1000-1400s for this model; at 1200 the first matrix lost agent-051 (base) and agent-049 (AGENTS.md) to timeouts, and a narrower re-run lost agent-051 again, so it was the budget rather than the burst. Those results were discarded and both experiments re-run from scratch at the new fingerprint — nothing here is measured at the old ceiling. Tier 1: this takes the Gemini Flash line's slot. Lines are tiered separately (Claude holds Fable, Opus and Sonnet slots at once) and Google still ships Pro and Flash side by side, so the Pro rows are untouched and stay tier 2 for want of GEMINI_API_KEY. No Flash-line model has been measured here before, so the under-a-month carve-out has no predecessor to attach to. Pricing is introductory: $0.75/$3.75 per 1M input/output through 2026-12-31, doubling to $1.50/$7.50 on 2027-01-01. The gateway catalog and the models.dev `vercel` entry agree on all three rates today; the row needs a re-export in January even if nothing is rerun. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com> * Correct the publishing story: the site no longer holds a copy AGENTS.md and the add-eval-model skill still said nextjs.org/evals reads a copy of agent-results.json checked into vercel/front, and told you to open a second PR there carrying it. That stopped being true on 2026-09-09: #122 here and vercel/front#85415 moved the site to a server fetch of this repo's main, and the copy at apps/next-site/app/(next-site)/evals/agent-results.json was deleted. Following the old instructions means hunting for a file that does not exist, or recreating one the site does not read. Merging here is the publish. A front PR is for site-side changes only. Also record what this run cost a debug cycle: - Pick the harness by what preflight can authenticate, not by the vendor. The gemini and cursor harnesses are direct-vendor-API only and get silently skipped without their keys; AI_GATEWAY_API_KEY covers every vercel-ai-gateway/* harness. - A model newer than the pinned OpenCode binary needs extraProviders. - Effort has no working knob on the OpenCode path, so those rows publish unpinned and unsuffixed. - reasoning_effort is not validated for every model — google/gemini-3.8-flash returns 200 for a rung it does not have. Compare reasoning_tokens instead of trusting the 400. - A timeout that is too tight cannot be re-run away, and raising it strands results under the old fingerprint unless you delete results/<slug>/. - Node 24 and the test/format commands for the front side. And two numbers that had drifted: the eval set is 31, not 26, and a matrix is ~124 concurrent attempts, not ~104. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com> --------- Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Publish result updates by merging
agent-results.jsonhere, without copying it into the website repository. CI verifies that the committed JSON matches the saved results and, after successful checks onmain, callshttps://nextjs.org/api/evals/revalidateto refresh the website's indefinitely cached snapshot.The new
revalidate-sitejob runs only formainpushes or manual dispatches and authenticates with GitHub Actions OIDC. It retries delivery three times and fails on an unacknowledged or unsuccessful response. The website permits stale server responses for up to one hour after invalidation while refreshing in the background, then requires fresh data. PR runs validate results without calling production.pnpm export-results --checkcompares the committed file with a complete export, ignores only the export timestamp, and never writes it. Full exports use deterministic experiment order; errors exit nonzero and empty exports are rejected. The JSON format remains unversioned.Deploy the website's new endpoint before merging this producer integration. No shared secret needs configuration. Retry a missed notification by rerunning the failed job or manually dispatching this workflow on
main; there is no periodic refresh fallback.Validation:
mainrun.