Add Claude Fable 5.1 (high) to nextjs.org/evals - #121
Merged
Merged
Conversation
Adds the experiment pair, display name, and list price for Claude Fable 5.1, released 2026-09-01, so the model can be run and published on nextjs.org/evals alongside the Claude Fable 5 (high) row already there. No results yet — this is the config half of the README's "Adding a new model", the same posture as #116's registration commit. Sourced, not inferred: - Model id. `claude-fable-5.1` is what the AI Gateway serves: catalog entry `anthropic/claude-fable-5.1`, reached through the gateway's Anthropic-compatible endpoint the way the existing `claude-fable-5` pair reaches Fable 5. Verified against the live gateway, which answers 404 `model_not_found` for an id it does not know rather than falling back — so a 200 for this string is real resolution, not the trap that sank the bogus gpt-5.6 ids. - Effort `high`, per the README's "Reasoning effort" rule, and confirmed before spending a matrix on it: the gateway's own 400 for this id enumerates none/minimal/low/medium/high/xhigh/max, and `reasoning_effort: high` returns 200. Encoded in the display name, not the slug — the slug carries the version, as with the claude-fable-5 pair whose label is already "Claude Fable 5 (high)". - List price 10/50/0.25/12.5 per 1M is the gateway catalog's entry for this id. Input, output, and cache write match Fable 5; cache reads are a quarter of it ($0.25 vs $1), which is not cosmetic given how heavily Claude Code caches. The pair uses `isNextApp` from lib/setup.js, which the claude-fable-5 configs predate. The eval set now includes the two framework-choice fixtures (agent-044, agent-045) that start empty, and installing next@canary or writing a Next.js-naming AGENTS.md into those answers the question they exist to ask. Deliberately left alone: - TIER_1 — an experiment with no results is not exported at all, so the tier belongs to the PR that lands the run. The comment there records the arithmetic that PR needs: Fable 5.1 takes the Fable line's tier-1 slot outright and claude-fable-5 drops to tier 2, since 5 was already being measured here on 2026-06-09, at least 84 days earlier, past the retention policy's under-a-month carve-out. - agent-results.json — regenerating it is `pnpm export-results` after a real run, not a hand edit. Confirmed the export is unchanged apart from its `exportedAt` timestamp: an unmeasured experiment does not reach the board. Both slugs go into ACCEPTED_STALE so `eval-cache-check` stays green while results are empty: `agent-eval status` reports all 26 evals as new for a never-run experiment, and check-stale.mjs fails on any unaccepted one. Drop the two entries in the PR that lands the run. Verified: `pnpm typecheck` clean, `pnpm test:cost` 9/9, and — after `pnpm sync-evals 071a2343c509751585cd9f77ae66e8c30daf2ea8`, the SHA CI pins — `node scripts/check-stale.mjs` reports "Eval cache OK", `pnpm eval:dry claude-fable-5.1` resolves to `vercel-ai-gateway/claude-code` / `claude-fable-5.1`, 26 evals x 4 runs, and `pnpm eval:smoke claude-fable-5.1` passed agent-000 on the first attempt in 185.5s against a real Vercel Sandbox. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com>
A model addition that stops at the config changes nothing anyone can see: `export-results` only exports experiments that have results, so the PR looks complete while the board is untouched. The README already rejects that state — "a staging post, not a destination" — but nothing put the rule where an agent reads it before starting, and the failure mode has repeated. Adds `.agents/skills/add-eval-model/SKILL.md`: the whole path from an empty checkout to a PR with numbers in it, ordered so the expensive step is the middle of the task rather than the end of it. It leads with the deliverable (a landed run, and the cost and wall clock that implies) and carries the gotchas that cost time here: - the sandbox token triple is all-or-nothing, and one or two of the three fails as a confusing OIDC error - the gateway is the source of truth for both the model id (an unknown id 404s, so a 200 is real resolution) and the effort rung (a bad value 400s with the real set enumerated) - `eval:smoke` exits 1 and persists nothing even when the eval passes - all `runs` start concurrently, so `earlyExit` trims the tail rather than saving 4x, and the two experiments should run sequentially - a subset run publishes `notAvailable` cells that read as model failure - `sync-evals` refingerprints unrelated cached results, and that churn should not reach the commit - copy the newest existing pair, not the oldest: the `isNextApp` guard matters now that the framework-choice fixtures are in the set Skills live in `.agents/skills/` with `.claude/skills` symlinked to it, so Claude Code and any harness that reads `.agents/` get the same copy. `.gitignore` now ignores `.claude/*` with `!.claude/skills` rather than all of `.claude`, since local Claude state should still stay out. Root `AGENTS.md` indexes the skills and states the finish-the-run expectation, with `CLAUDE.md` symlinked to it — the same pointer pattern the `--agents-md` experiment configs write into their fixtures. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com>
Lands the run for the pair registered earlier in this branch, so Fable 5.1 actually reaches the board instead of sitting registered and unexported. Results, 26 evals, pass@4 through Claude Code on the AI Gateway: claude-fable-5.1 24/26 (92%) claude-fable-5.1--agents-md 24/26 (92%) Same two failures on both sides, and both are genuine four-attempt model failures rather than flakes: - agent-040-instant — every run reached for Suspense boundaries instead of exporting the `instant` config with `prefetch: "static"` that the eval asks for. 0/4 on both experiments. - agent-044-uses-nextjs — the framework-choice fixture. The model built a working book tracker on a plain `node:http` server with no `next` dependency at all. 0/4 on both, and the AGENTS.md pointer did not save it, which is notable given that naming the framework is the one thing that file does. AGENTS.md delta is 0: it flips no eval either way. Its only visible effect is on agent-031-proxy-middleware, which the base pair passes on the third attempt and the AGENTS.md pair passes on the first — a pass@4 pass in both columns, so it does not move the score. Against the Fable 5 row it replaces, at the same 92%: $0.62 vs $1.78 mean list cost per eval, and 198s vs 268s mean duration. Most of the cost gap is the cache-read rate this branch already recorded ($0.25 vs $1.00 per 1M) — Claude Code's traffic is mostly cache reads. Tiering: claude-fable-5.1 takes the Fable line's only tier-1 slot and claude-fable-5 drops to tier 2. 5.1 shipped 2026-09-01 and 5 was already being measured here on 2026-06-09, at least 84 days earlier — past the retention policy's under-a-month carve-out, so the two do not share the slot. Fable 5's results are fresh as of #120, so it needs no ACCEPTED_STALE entries on the way down. Also drops the two ACCEPTED_STALE entries added for these slugs, which existed only to keep eval-cache-check green while results were empty. Verified: pnpm typecheck clean, pnpm test:cost 9/9, `agent-eval status 'claude-fable-5.1*'` reports "Everything up to date", and node scripts/check-stale.mjs reports "Eval cache OK" against the CI-pinned next.js SHA. agent-results.json is `pnpm export-results` output, not a hand edit, and the committed transcripts were scanned for gateway/Vercel/GitHub credentials before commit (cf. #108). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com>
Merging in this repo does not change nextjs.org/evals — the page reads a copy of agent-results.json checked into vercel/front, so a task that stops at the export is still invisible. The skill now ends where the site does. Adds the copy path, notes that the page is fully data-driven off that file so no component changes are needed, and records the two things the README's bare `cp` glosses over: the export writes no trailing newline and front runs Prettier, and the copy is normally a release or two behind, so the diff carries earlier refreshes as well as yours. Also notes that front rejects API-authored commits (409, verified signatures required) — clone sparse and push, rather than burning a turn on the contents API. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com>
devjiwonchoi
approved these changes
Sep 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds Claude Fable 5.1 (released 2026-09-01) to the board: experiment pair, list price, a full pass@4 run, the export, and the tiering change. Companion PR that puts it on the site: vercel/front#85397.
Results
26 evals, pass@4, Claude Code on the AI Gateway at effort
high:Same score, roughly a third of the cost and 70s less per eval. Most of the cost gap is the cache-read rate — $0.25 vs $1.00 per 1M — which dominates because Claude Code's traffic is mostly cache reads.
Two failures, identical on both sides and genuine four-attempt model failures rather than flakes:
agent-040-instant(0/4 both) — every run reached for Suspense boundaries instead of exporting theinstantconfig withprefetch: "static"the eval asks for.agent-044-uses-nextjs(0/4 both) — the framework-choice fixture. The model built a working book tracker on a plainnode:httpserver with nonextdependency. The AGENTS.md pointer did not save it, which is notable given that naming the framework is the one thing that file does.AGENTS.md delta is 0 — it flips no eval. Its only visible effect is
agent-031-proxy-middleware, which the base pair passes on the third attempt and the AGENTS.md pair passes on the first; a pass@4 pass in both columns, so it does not move the score.Tiering
claude-fable-5.1takes the Fable line's only tier-1 slot;claude-fable-5drops to tier 2. 5.1 shipped 2026-09-01 and 5 was already being measured here on 2026-06-09 — at least 84 days earlier, past the retention policy's under-a-month carve-out, so the two do not share the slot. Fable 5's results are fresh as of #120, so it needs noACCEPTED_STALEentries on the way down.Sourced, not inferred
claude-fable-5.1anthropic/claude-fable-5.1, through the gateway's Anthropic-compatible endpoint the wayclaude-fable-5goes. The gateway 404smodel_not_foundfor ids it does not serve, so a 200 is real resolution — the check that would have caught the bogus gpt-5.6 ids.highnone/minimal/low/medium/high/xhigh/max, andhighreturns 200.10 / 50 / 0.25 / 12.5per 1MThe pair guards its
setupwithisNextAppfromlib/setup.js, which theclaude-fable-5configs predate — the framework-choice fixtures (agent-044,agent-045) start empty, and installingnext@canaryor writing a Next.js-namingAGENTS.mdinto them answers the question they exist to ask.Also in this PR: an
add-eval-modelskillThe first pass at this change stopped at the config and handed back a "ready to run" PR. That state is exactly what the README rejects — "a staging post, not a destination" — and because
export-resultsonly exports experiments that have results, such a PR looks complete while the board is untouched..agents/skills/add-eval-model/SKILL.mdwrites the whole path down where an agent reads it before starting: the deliverable is a landed run, with the cost and wall clock that implies, plus the gotchas that cost time here (all-or-nothing sandbox token triple; the gateway as source of truth for id and effort rung;eval:smokeexiting 1 and persisting nothing even on a pass; allrunsstarting concurrently soearlyExittrims the tail rather than saving 4x; a subset run publishingnotAvailablecells that read as model failure;sync-evalsrefingerprinting unrelated results).Skills live in
.agents/skills/with.claude/skillssymlinked to it, so Claude Code and any harness reading.agents/get one copy..gitignorenow ignores.claude/*with!.claude/skillsinstead of all of.claude, so local Claude state still stays out. RootAGENTS.mdindexes the skills and states the finish-the-run expectation, withCLAUDE.mdsymlinked to it — the same pointer pattern the--agents-mdconfigs write into their fixtures.Verification
pnpm typecheckclean;pnpm test:cost9/9agent-eval status 'claude-fable-5.1*'—Everything up to date — nothing to runnode scripts/check-stale.mjsat the CI-pinned next.js SHA —Eval cache OKagent-results.jsonis verbatimpnpm export-resultsoutput, not a hand editresults/*/summary.jsonchurn thatsync-evalsrefingerprints into unrelated experiments was reverted rather than committed🤖 Generated with Claude Code