Skip to content

Add Gemini 3.8 Flash to nextjs.org/evals (28/31 base, 30/31 with AGENTS.md) - #124

Merged
gaojude merged 2 commits into
mainfrom
gemini-3-8-flash-evals
Sep 11, 2026
Merged

gaojude merged 2 commits into
mainfrom
gemini-3-8-flash-evals

Conversation

@gaojude

@gaojude gaojude commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

Gemini 3.8 Flash shipped on 2026-09-02 and was missing from nextjs.org/evals, whose last run was 2026-09-10. This registers it, measures it over the full 31-eval set at pass@4, and tiers it. Everything published here is measured — no hand-edited numbers.

Results

Success rate
Gemini 3.8 Flash 28/31 (90%)
Gemini 3.8 Flash + AGENTS.md 30/31 (97%) +6pp
  • AGENTS.md wins two, loses none. agent-029-use-cache-directive and agent-031-proxy-middleware both flip fail→pass. Without the docs the model reaches for unstable_cache()/revalidateTag() instead of 'use cache', and exports both middleware and proxy from proxy.ts when the eval requires only proxy — exactly the "this is not the Next.js you know" failures the pointer exists to fix.
  • agent-044-uses-nextjs fails on both sides, the same way in all eight attempts: the model builds a working book-tracking app on Express and never reaches for Next.js. That is the one genuine model failure in the pair.
  • Mean list cost $0.605/eval, mean duration 435s.

Harness: OpenCode over the AI Gateway, not the Gemini CLI

The older Gemini pairs use the gemini harness, which talks to the Google API directly and reads GEMINI_API_KEY — the same missing credential that keeps gemini-3.1-pro-preview on tier 2 despite qualifying for tier 1. The gateway does serve google/gemini-3.8-flash, so vercel-ai-gateway/opencode is the path that can actually produce a fresh measurement, and it is the path glm, kimi, grok and minimax already take. The board renders agentHarness per row, so the row says OpenCode and the difference is visible rather than implied.

No effort pin, and why the label says so

The gateway catalog advertises a low/medium/high ladder for this id, and high would be the rung to publish at per "Reasoning effort" in the README. OpenCode has no working knob for it on this path: it maps a model's options onto providerOptions.gateway for @ai-sdk/gateway providers, and the gateway does not act on reasoningEffort there. At low the model still burned ~1.4k reasoning tokens on a probe prompt, where the same low sent as reasoning_effort to /v1/chat/completions burns 0. So these runs are at the provider default, like every other OpenCode row on the board (kimi-k3, glm-5.2 and grok-4.6 all expose ladders and all publish unpinned). Labelling the row (high) would assert a setting the harness never sent.

Worth noting for the next model: google/gemini-3.8-flash returns 200 for reasoning_effort: "xhigh", a rung it does not have. The 400-enumerates-the-set trick the README leans on does not hold for every provider.

timeout: 2400, and a discarded matrix

The other current OpenCode pairs use 1200. "Flash" is about per-token speed, not about how long an agentic loop runs: the prefetch evals land at 1000–1400s for this model. At 1200 the first full matrix lost agent-051 (base) and agent-049 (AGENTS.md) to timeouts, and a narrower re-run at 4-way concurrency lost agent-051 again — the budget, not the burst. Timeouts are deleted rather than counted, so a ceiling that tight does not publish a wrong number, it just makes the matrix impossible to finish.

timeout is part of the fingerprint, so I raised it once, deleted results/gemini-3.8-flash* and re-ran both experiments from scratch. Nothing in this PR is measured at the old ceiling, and no eval is notAvailable.

Tiering

gemini-3.8-flash takes the Gemini Flash line's tier-1 slot. Lines are tiered separately here — Claude holds Fable, Opus and Sonnet slots at once — and Google still ships Pro and Flash side by side, so this does not displace the Pro line. The under-a-month carve-out has no predecessor to attach to: no Flash-line model has ever been measured on this board. Gemini 3.0 / 3.1 Pro Preview stay tier 2 for the existing reason — the Gemini CLI harness they use cannot be rerun without GEMINI_API_KEY.

Pricing

$0.75 / $3.75 / $0.075 per 1M input / output / cache-read. The gateway catalog and the models.dev vercel entry agree on all three as of 2026-09-11. This is introductory pricing through 2026-12-31 — input and output double to $1.50 / $7.50 on 2027-01-01, so this row needs a re-export in January even if nothing is rerun. Noted in MODEL_PRICING next to the rates.

Second commit: the publishing story was out of date

AGENTS.md and the add-eval-model skill still said the site reads a copy of agent-results.json checked into vercel/front, and told you to open a second PR there carrying it. That stopped being true on 2026-09-09 (#122 here, vercel/front#85415): the site server-fetches this repo's main and the copy was deleted. Following the old instructions means recreating a file the site does not read. Corrected, along with the detours this run cost a debug cycle — harness selection under missing vendor keys, extraProviders for a model newer than the pinned binary, the unvalidated reasoning_effort, the fingerprint/stale-results trap when raising a timeout — and two numbers that had drifted (31 evals, not 26; ~124 concurrent attempts, not ~104).

Cost

$50.30 at list for the 86 retained runs. The discarded 1200s matrix was about the same size, so total spend across both was roughly $100.

Verification

pnpm typecheck                 ok
pnpm test:cost                 9/9
node scripts/check-stale.mjs   Eval cache OK
pnpm export-results --check    Verified (930 total / 722 pass / 59 fail / 149 N/A)
pnpm status 'gemini-3.8-flash*'  Everything up to date

The agent-results.json diff is purely additive — the new experiment plus exportedAt. No other model's numbers moved, and no results/*/summary.json churn from sync-evals.

Site side

Merging this to main is the publish; CI invalidates the page's cache. The paired PR in vercel/front carries no results — see vercel/front#85735.

🤖 Generated with Claude Code

vercel Bot and others added 2 commits September 11, 2026 18:38
Gemini 3.8 Flash shipped 2026-09-02 and was missing from nextjs.org/evals.
Registered, measured over the full 31-eval set at pass@4, and tiered.

  Gemini 3.8 Flash              28/31 (90%)
  Gemini 3.8 Flash + AGENTS.md  30/31 (97%)   +6pp

AGENTS.md turns agent-029 (use cache directive) and agent-031 (proxy)
around and loses nothing. agent-044 (uses-nextjs) fails on both sides for
the same reason every time: the model builds the app on Express, ignoring
the framework requirement. Mean list cost $0.605/eval, mean duration 435s.

Harness is vercel-ai-gateway/opencode, not the `gemini` (Gemini CLI)
harness the older Gemini pairs use. That harness talks to the Google API
directly and needs GEMINI_API_KEY, which is the same missing credential
that keeps gemini-3.1-pro-preview on tier 2. The gateway serves
google/gemini-3.8-flash, so OpenCode is the path that can actually produce
a fresh measurement — the same path glm, kimi, grok and minimax take. The
board renders the harness per row, so the difference is visible.

No reasoning-effort pin and no effort suffix on the label, matching every
other OpenCode row. OpenCode maps a model's options onto
providerOptions.gateway for @ai-sdk/gateway providers and the gateway does
not act on reasoningEffort there: at `low` the model still spent ~1.4k
reasoning tokens, where the same `low` sent as reasoning_effort to
/v1/chat/completions spends 0. Labelling the row `(high)` would claim a
setting the harness never sent.

timeout is 2400 rather than the 1200 the other OpenCode pairs use. The
prefetch evals run 1000-1400s for this model; at 1200 the first matrix lost
agent-051 (base) and agent-049 (AGENTS.md) to timeouts, and a narrower
re-run lost agent-051 again, so it was the budget rather than the burst.
Those results were discarded and both experiments re-run from scratch at
the new fingerprint — nothing here is measured at the old ceiling.

Tier 1: this takes the Gemini Flash line's slot. Lines are tiered
separately (Claude holds Fable, Opus and Sonnet slots at once) and Google
still ships Pro and Flash side by side, so the Pro rows are untouched and
stay tier 2 for want of GEMINI_API_KEY. No Flash-line model has been
measured here before, so the under-a-month carve-out has no predecessor to
attach to.

Pricing is introductory: $0.75/$3.75 per 1M input/output through
2026-12-31, doubling to $1.50/$7.50 on 2027-01-01. The gateway catalog and
the models.dev `vercel` entry agree on all three rates today; the row needs
a re-export in January even if nothing is rerun.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com>
AGENTS.md and the add-eval-model skill still said nextjs.org/evals reads a
copy of agent-results.json checked into vercel/front, and told you to open
a second PR there carrying it. That stopped being true on 2026-09-09: #122
here and vercel/front#85415 moved the site to a server fetch of this repo's
main, and the copy at apps/next-site/app/(next-site)/evals/agent-results.json
was deleted. Following the old instructions means hunting for a file that
does not exist, or recreating one the site does not read.

Merging here is the publish. A front PR is for site-side changes only.

Also record what this run cost a debug cycle:

- Pick the harness by what preflight can authenticate, not by the vendor.
  The gemini and cursor harnesses are direct-vendor-API only and get
  silently skipped without their keys; AI_GATEWAY_API_KEY covers every
  vercel-ai-gateway/* harness.
- A model newer than the pinned OpenCode binary needs extraProviders.
- Effort has no working knob on the OpenCode path, so those rows publish
  unpinned and unsuffixed.
- reasoning_effort is not validated for every model — google/gemini-3.8-flash
  returns 200 for a rung it does not have. Compare reasoning_tokens instead
  of trusting the 400.
- A timeout that is too tight cannot be re-run away, and raising it strands
  results under the old fingerprint unless you delete results/<slug>/.
- Node 24 and the test/format commands for the front side.

And two numbers that had drifted: the eval set is 31, not 26, and a matrix
is ~124 concurrent attempts, not ~104.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com>

gaojude commented Sep 11, 2026

Copy link
Copy Markdown
Contributor Author

@gaojude ready for review. GitHub refuses a formal review request here — the only GitHub identity on the devbox that produced this is gaojude, so the PR is authored by you and POST /pulls/124/requested_reviewers answers 422 Review cannot be requested from pull request author. Flagging it this way instead; please reassign to a second reviewer if you want one on the record.

The judgement calls worth a second opinion, in order:

  1. Harness. Published as OpenCode, not Gemini CLI, because the gemini harness needs GEMINI_API_KEY and this box has none — the same gap that keeps gemini-3.1-pro-preview on tier 2. That makes the Flash row not strictly comparable to the two Pro rows above it. Reasonable, or should Gemini rows stay on one harness even at the cost of not measuring this model at all?
  2. No (high) suffix. The gateway advertises a low/medium/high ladder, but OpenCode's providerOptions.gateway.reasoningEffort demonstrably does not reach the model (evidence in the PR body), so the runs are at the provider default — same as every other OpenCode row. I chose a bare label over a claim the harness cannot back.
  3. timeout: 2400 and the discarded first matrix. Both experiments were re-run from scratch after the raise; nothing published here was measured at 1200.
  4. Tier 1 as a separate Flash line, leaving the Pro rows untouched.

Also please look at the second commit on its own — it corrects AGENTS.md and the add-eval-model skill, which still described the deleted agent-results.json copy in front and told the next person to open a PR carrying it.

@gaojude
gaojude merged commit cbb7d9e into main Sep 11, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants