Add Gemini 3.8 Flash to nextjs.org/evals (28/31 base, 30/31 with AGENTS.md) - #124
Merged
Merged
Conversation
Gemini 3.8 Flash shipped 2026-09-02 and was missing from nextjs.org/evals. Registered, measured over the full 31-eval set at pass@4, and tiered. Gemini 3.8 Flash 28/31 (90%) Gemini 3.8 Flash + AGENTS.md 30/31 (97%) +6pp AGENTS.md turns agent-029 (use cache directive) and agent-031 (proxy) around and loses nothing. agent-044 (uses-nextjs) fails on both sides for the same reason every time: the model builds the app on Express, ignoring the framework requirement. Mean list cost $0.605/eval, mean duration 435s. Harness is vercel-ai-gateway/opencode, not the `gemini` (Gemini CLI) harness the older Gemini pairs use. That harness talks to the Google API directly and needs GEMINI_API_KEY, which is the same missing credential that keeps gemini-3.1-pro-preview on tier 2. The gateway serves google/gemini-3.8-flash, so OpenCode is the path that can actually produce a fresh measurement — the same path glm, kimi, grok and minimax take. The board renders the harness per row, so the difference is visible. No reasoning-effort pin and no effort suffix on the label, matching every other OpenCode row. OpenCode maps a model's options onto providerOptions.gateway for @ai-sdk/gateway providers and the gateway does not act on reasoningEffort there: at `low` the model still spent ~1.4k reasoning tokens, where the same `low` sent as reasoning_effort to /v1/chat/completions spends 0. Labelling the row `(high)` would claim a setting the harness never sent. timeout is 2400 rather than the 1200 the other OpenCode pairs use. The prefetch evals run 1000-1400s for this model; at 1200 the first matrix lost agent-051 (base) and agent-049 (AGENTS.md) to timeouts, and a narrower re-run lost agent-051 again, so it was the budget rather than the burst. Those results were discarded and both experiments re-run from scratch at the new fingerprint — nothing here is measured at the old ceiling. Tier 1: this takes the Gemini Flash line's slot. Lines are tiered separately (Claude holds Fable, Opus and Sonnet slots at once) and Google still ships Pro and Flash side by side, so the Pro rows are untouched and stay tier 2 for want of GEMINI_API_KEY. No Flash-line model has been measured here before, so the under-a-month carve-out has no predecessor to attach to. Pricing is introductory: $0.75/$3.75 per 1M input/output through 2026-12-31, doubling to $1.50/$7.50 on 2027-01-01. The gateway catalog and the models.dev `vercel` entry agree on all three rates today; the row needs a re-export in January even if nothing is rerun. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com>
AGENTS.md and the add-eval-model skill still said nextjs.org/evals reads a copy of agent-results.json checked into vercel/front, and told you to open a second PR there carrying it. That stopped being true on 2026-09-09: #122 here and vercel/front#85415 moved the site to a server fetch of this repo's main, and the copy at apps/next-site/app/(next-site)/evals/agent-results.json was deleted. Following the old instructions means hunting for a file that does not exist, or recreating one the site does not read. Merging here is the publish. A front PR is for site-side changes only. Also record what this run cost a debug cycle: - Pick the harness by what preflight can authenticate, not by the vendor. The gemini and cursor harnesses are direct-vendor-API only and get silently skipped without their keys; AI_GATEWAY_API_KEY covers every vercel-ai-gateway/* harness. - A model newer than the pinned OpenCode binary needs extraProviders. - Effort has no working knob on the OpenCode path, so those rows publish unpinned and unsuffixed. - reasoning_effort is not validated for every model — google/gemini-3.8-flash returns 200 for a rung it does not have. Compare reasoning_tokens instead of trusting the 400. - A timeout that is too tight cannot be re-run away, and raising it strands results under the old fingerprint unless you delete results/<slug>/. - Node 24 and the test/format commands for the front side. And two numbers that had drifted: the eval set is 31, not 26, and a matrix is ~124 concurrent attempts, not ~104. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com>
Contributor
Author
|
@gaojude ready for review. GitHub refuses a formal review request here — the only GitHub identity on the devbox that produced this is The judgement calls worth a second opinion, in order:
Also please look at the second commit on its own — it corrects |
franklbh
approved these changes
Sep 11, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Gemini 3.8 Flash shipped on 2026-09-02 and was missing from nextjs.org/evals, whose last run was 2026-09-10. This registers it, measures it over the full 31-eval set at pass@4, and tiers it. Everything published here is measured — no hand-edited numbers.
Results
agent-029-use-cache-directiveandagent-031-proxy-middlewareboth flip fail→pass. Without the docs the model reaches forunstable_cache()/revalidateTag()instead of'use cache', and exports bothmiddlewareandproxyfromproxy.tswhen the eval requires onlyproxy— exactly the "this is not the Next.js you know" failures the pointer exists to fix.agent-044-uses-nextjsfails on both sides, the same way in all eight attempts: the model builds a working book-tracking app on Express and never reaches for Next.js. That is the one genuine model failure in the pair.Harness: OpenCode over the AI Gateway, not the Gemini CLI
The older Gemini pairs use the
geminiharness, which talks to the Google API directly and readsGEMINI_API_KEY— the same missing credential that keepsgemini-3.1-pro-previewon tier 2 despite qualifying for tier 1. The gateway does servegoogle/gemini-3.8-flash, sovercel-ai-gateway/opencodeis the path that can actually produce a fresh measurement, and it is the path glm, kimi, grok and minimax already take. The board rendersagentHarnessper row, so the row says OpenCode and the difference is visible rather than implied.No effort pin, and why the label says so
The gateway catalog advertises a
low/medium/highladder for this id, andhighwould be the rung to publish at per "Reasoning effort" in the README. OpenCode has no working knob for it on this path: it maps a model'soptionsontoproviderOptions.gatewayfor@ai-sdk/gatewayproviders, and the gateway does not act onreasoningEffortthere. Atlowthe model still burned ~1.4k reasoning tokens on a probe prompt, where the samelowsent asreasoning_effortto/v1/chat/completionsburns 0. So these runs are at the provider default, like every other OpenCode row on the board (kimi-k3, glm-5.2 and grok-4.6 all expose ladders and all publish unpinned). Labelling the row(high)would assert a setting the harness never sent.Worth noting for the next model:
google/gemini-3.8-flashreturns 200 forreasoning_effort: "xhigh", a rung it does not have. The 400-enumerates-the-set trick the README leans on does not hold for every provider.timeout: 2400, and a discarded matrixThe other current OpenCode pairs use 1200. "Flash" is about per-token speed, not about how long an agentic loop runs: the prefetch evals land at 1000–1400s for this model. At 1200 the first full matrix lost
agent-051(base) andagent-049(AGENTS.md) to timeouts, and a narrower re-run at 4-way concurrency lostagent-051again — the budget, not the burst. Timeouts are deleted rather than counted, so a ceiling that tight does not publish a wrong number, it just makes the matrix impossible to finish.timeoutis part of the fingerprint, so I raised it once, deletedresults/gemini-3.8-flash*and re-ran both experiments from scratch. Nothing in this PR is measured at the old ceiling, and no eval isnotAvailable.Tiering
gemini-3.8-flashtakes the Gemini Flash line's tier-1 slot. Lines are tiered separately here — Claude holds Fable, Opus and Sonnet slots at once — and Google still ships Pro and Flash side by side, so this does not displace the Pro line. The under-a-month carve-out has no predecessor to attach to: no Flash-line model has ever been measured on this board. Gemini 3.0 / 3.1 Pro Preview stay tier 2 for the existing reason — the Gemini CLI harness they use cannot be rerun withoutGEMINI_API_KEY.Pricing
$0.75 / $3.75 / $0.075per 1M input / output / cache-read. The gateway catalog and the models.devvercelentry agree on all three as of 2026-09-11. This is introductory pricing through 2026-12-31 — input and output double to$1.50 / $7.50on 2027-01-01, so this row needs a re-export in January even if nothing is rerun. Noted inMODEL_PRICINGnext to the rates.Second commit: the publishing story was out of date
AGENTS.mdand theadd-eval-modelskill still said the site reads a copy ofagent-results.jsonchecked intovercel/front, and told you to open a second PR there carrying it. That stopped being true on 2026-09-09 (#122 here, vercel/front#85415): the site server-fetches this repo'smainand the copy was deleted. Following the old instructions means recreating a file the site does not read. Corrected, along with the detours this run cost a debug cycle — harness selection under missing vendor keys,extraProvidersfor a model newer than the pinned binary, the unvalidatedreasoning_effort, the fingerprint/stale-results trap when raising a timeout — and two numbers that had drifted (31 evals, not 26; ~124 concurrent attempts, not ~104).Cost
$50.30 at list for the 86 retained runs. The discarded 1200s matrix was about the same size, so total spend across both was roughly $100.
Verification
The
agent-results.jsondiff is purely additive — the new experiment plusexportedAt. No other model's numbers moved, and noresults/*/summary.jsonchurn fromsync-evals.Site side
Merging this to
mainis the publish; CI invalidates the page's cache. The paired PR invercel/frontcarries no results — see vercel/front#85735.🤖 Generated with Claude Code