Skip to content

Add Claude Fable 5.1 (high) to nextjs.org/evals - #121

Merged
gaojude merged 4 commits into
mainfrom
add-fable-5-1-to-evals
Sep 9, 2026
Merged

gaojude merged 4 commits into
mainfrom
add-fable-5-1-to-evals

Conversation

@gaojude

@gaojude gaojude commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Adds Claude Fable 5.1 (released 2026-09-01) to the board: experiment pair, list price, a full pass@4 run, the export, and the tiering change. Companion PR that puts it on the site: vercel/front#85397.

Results

26 evals, pass@4, Claude Code on the AI Gateway at effort high:

Fable 5.1 (high) Fable 5 (high) — the row it replaces
base 24/26 (92%) 24/26 (92%)
+ AGENTS.md 24/26 (92%) 24/26 (92%)
mean list cost / eval $0.62 $1.78
mean duration 198s 268s

Same score, roughly a third of the cost and 70s less per eval. Most of the cost gap is the cache-read rate — $0.25 vs $1.00 per 1M — which dominates because Claude Code's traffic is mostly cache reads.

Two failures, identical on both sides and genuine four-attempt model failures rather than flakes:

  • agent-040-instant (0/4 both) — every run reached for Suspense boundaries instead of exporting the instant config with prefetch: "static" the eval asks for.
  • agent-044-uses-nextjs (0/4 both) — the framework-choice fixture. The model built a working book tracker on a plain node:http server with no next dependency. The AGENTS.md pointer did not save it, which is notable given that naming the framework is the one thing that file does.

AGENTS.md delta is 0 — it flips no eval. Its only visible effect is agent-031-proxy-middleware, which the base pair passes on the third attempt and the AGENTS.md pair passes on the first; a pass@4 pass in both columns, so it does not move the score.

Tiering

claude-fable-5.1 takes the Fable line's only tier-1 slot; claude-fable-5 drops to tier 2. 5.1 shipped 2026-09-01 and 5 was already being measured here on 2026-06-09 — at least 84 days earlier, past the retention policy's under-a-month carve-out, so the two do not share the slot. Fable 5's results are fresh as of #120, so it needs no ACCEPTED_STALE entries on the way down.

Sourced, not inferred

Source
Model id claude-fable-5.1 AI Gateway catalog entry anthropic/claude-fable-5.1, through the gateway's Anthropic-compatible endpoint the way claude-fable-5 goes. The gateway 404s model_not_found for ids it does not serve, so a 200 is real resolution — the check that would have caught the bogus gpt-5.6 ids.
Effort high The README's "Reasoning effort" rule, confirmed before spending the matrix: the gateway's own 400 enumerates none/minimal/low/medium/high/xhigh/max, and high returns 200.
List price 10 / 50 / 0.25 / 12.5 per 1M Same catalog entry. In, out and cache write match Fable 5; cache reads are a quarter of it, which is what the cost column above is mostly measuring.

The pair guards its setup with isNextApp from lib/setup.js, which the claude-fable-5 configs predate — the framework-choice fixtures (agent-044, agent-045) start empty, and installing next@canary or writing a Next.js-naming AGENTS.md into them answers the question they exist to ask.

Also in this PR: an add-eval-model skill

The first pass at this change stopped at the config and handed back a "ready to run" PR. That state is exactly what the README rejects — "a staging post, not a destination" — and because export-results only exports experiments that have results, such a PR looks complete while the board is untouched.

.agents/skills/add-eval-model/SKILL.md writes the whole path down where an agent reads it before starting: the deliverable is a landed run, with the cost and wall clock that implies, plus the gotchas that cost time here (all-or-nothing sandbox token triple; the gateway as source of truth for id and effort rung; eval:smoke exiting 1 and persisting nothing even on a pass; all runs starting concurrently so earlyExit trims the tail rather than saving 4x; a subset run publishing notAvailable cells that read as model failure; sync-evals refingerprinting unrelated results).

Skills live in .agents/skills/ with .claude/skills symlinked to it, so Claude Code and any harness reading .agents/ get one copy. .gitignore now ignores .claude/* with !.claude/skills instead of all of .claude, so local Claude state still stays out. Root AGENTS.md indexes the skills and states the finish-the-run expectation, with CLAUDE.md symlinked to it — the same pointer pattern the --agents-md configs write into their fixtures.

Verification

  • pnpm typecheck clean; pnpm test:cost 9/9
  • agent-eval status 'claude-fable-5.1*' — Everything up to date — nothing to run
  • node scripts/check-stale.mjs at the CI-pinned next.js SHA — Eval cache OK
  • agent-results.json is verbatim pnpm export-results output, not a hand edit
  • Committed transcripts scanned for gateway / Vercel / GitHub credentials before commit (cf. Scrub leaked AI Gateway credentials from committed transcripts #108)
  • results/*/summary.json churn that sync-evals refingerprints into unrelated experiments was reverted rather than committed

🤖 Generated with Claude Code

vercel Bot and others added 3 commits September 9, 2026 19:11
Adds the experiment pair, display name, and list price for Claude Fable
5.1, released 2026-09-01, so the model can be run and published on
nextjs.org/evals alongside the Claude Fable 5 (high) row already there.
No results yet — this is the config half of the README's "Adding a new
model", the same posture as #116's registration commit.

Sourced, not inferred:

- Model id. `claude-fable-5.1` is what the AI Gateway serves: catalog
  entry `anthropic/claude-fable-5.1`, reached through the gateway's
  Anthropic-compatible endpoint the way the existing `claude-fable-5`
  pair reaches Fable 5. Verified against the live gateway, which answers
  404 `model_not_found` for an id it does not know rather than falling
  back — so a 200 for this string is real resolution, not the trap that
  sank the bogus gpt-5.6 ids.
- Effort `high`, per the README's "Reasoning effort" rule, and confirmed
  before spending a matrix on it: the gateway's own 400 for this id
  enumerates none/minimal/low/medium/high/xhigh/max, and `reasoning_effort:
  high` returns 200. Encoded in the display name, not the slug — the slug
  carries the version, as with the claude-fable-5 pair whose label is
  already "Claude Fable 5 (high)".
- List price 10/50/0.25/12.5 per 1M is the gateway catalog's entry for
  this id. Input, output, and cache write match Fable 5; cache reads are a
  quarter of it ($0.25 vs $1), which is not cosmetic given how heavily
  Claude Code caches.

The pair uses `isNextApp` from lib/setup.js, which the claude-fable-5
configs predate. The eval set now includes the two framework-choice
fixtures (agent-044, agent-045) that start empty, and installing
next@canary or writing a Next.js-naming AGENTS.md into those answers the
question they exist to ask.

Deliberately left alone:

- TIER_1 — an experiment with no results is not exported at all, so the
  tier belongs to the PR that lands the run. The comment there records
  the arithmetic that PR needs: Fable 5.1 takes the Fable line's tier-1
  slot outright and claude-fable-5 drops to tier 2, since 5 was already
  being measured here on 2026-06-09, at least 84 days earlier, past the
  retention policy's under-a-month carve-out.
- agent-results.json — regenerating it is `pnpm export-results` after a
  real run, not a hand edit. Confirmed the export is unchanged apart from
  its `exportedAt` timestamp: an unmeasured experiment does not reach the
  board.

Both slugs go into ACCEPTED_STALE so `eval-cache-check` stays green while
results are empty: `agent-eval status` reports all 26 evals as new for a
never-run experiment, and check-stale.mjs fails on any unaccepted one.
Drop the two entries in the PR that lands the run.

Verified: `pnpm typecheck` clean, `pnpm test:cost` 9/9, and — after
`pnpm sync-evals 071a2343c509751585cd9f77ae66e8c30daf2ea8`, the SHA CI
pins — `node scripts/check-stale.mjs` reports "Eval cache OK",
`pnpm eval:dry claude-fable-5.1` resolves to
`vercel-ai-gateway/claude-code` / `claude-fable-5.1`, 26 evals x 4 runs,
and `pnpm eval:smoke claude-fable-5.1` passed agent-000 on the first
attempt in 185.5s against a real Vercel Sandbox.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com>
A model addition that stops at the config changes nothing anyone can
see: `export-results` only exports experiments that have results, so the
PR looks complete while the board is untouched. The README already
rejects that state — "a staging post, not a destination" — but nothing
put the rule where an agent reads it before starting, and the failure
mode has repeated.

Adds `.agents/skills/add-eval-model/SKILL.md`: the whole path from an
empty checkout to a PR with numbers in it, ordered so the expensive step
is the middle of the task rather than the end of it. It leads with the
deliverable (a landed run, and the cost and wall clock that implies) and
carries the gotchas that cost time here:

- the sandbox token triple is all-or-nothing, and one or two of the three
  fails as a confusing OIDC error
- the gateway is the source of truth for both the model id (an unknown id
  404s, so a 200 is real resolution) and the effort rung (a bad value 400s
  with the real set enumerated)
- `eval:smoke` exits 1 and persists nothing even when the eval passes
- all `runs` start concurrently, so `earlyExit` trims the tail rather than
  saving 4x, and the two experiments should run sequentially
- a subset run publishes `notAvailable` cells that read as model failure
- `sync-evals` refingerprints unrelated cached results, and that churn
  should not reach the commit
- copy the newest existing pair, not the oldest: the `isNextApp` guard
  matters now that the framework-choice fixtures are in the set

Skills live in `.agents/skills/` with `.claude/skills` symlinked to it, so
Claude Code and any harness that reads `.agents/` get the same copy.
`.gitignore` now ignores `.claude/*` with `!.claude/skills` rather than all
of `.claude`, since local Claude state should still stay out. Root
`AGENTS.md` indexes the skills and states the finish-the-run expectation,
with `CLAUDE.md` symlinked to it — the same pointer pattern the
`--agents-md` experiment configs write into their fixtures.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com>
Lands the run for the pair registered earlier in this branch, so Fable
5.1 actually reaches the board instead of sitting registered and
unexported.

Results, 26 evals, pass@4 through Claude Code on the AI Gateway:

  claude-fable-5.1              24/26 (92%)
  claude-fable-5.1--agents-md   24/26 (92%)

Same two failures on both sides, and both are genuine four-attempt model
failures rather than flakes:

- agent-040-instant — every run reached for Suspense boundaries instead of
  exporting the `instant` config with `prefetch: "static"` that the eval
  asks for. 0/4 on both experiments.
- agent-044-uses-nextjs — the framework-choice fixture. The model built a
  working book tracker on a plain `node:http` server with no `next`
  dependency at all. 0/4 on both, and the AGENTS.md pointer did not save
  it, which is notable given that naming the framework is the one thing
  that file does.

AGENTS.md delta is 0: it flips no eval either way. Its only visible effect
is on agent-031-proxy-middleware, which the base pair passes on the third
attempt and the AGENTS.md pair passes on the first — a pass@4 pass in both
columns, so it does not move the score.

Against the Fable 5 row it replaces, at the same 92%: $0.62 vs $1.78 mean
list cost per eval, and 198s vs 268s mean duration. Most of the cost gap is
the cache-read rate this branch already recorded ($0.25 vs $1.00 per 1M) —
Claude Code's traffic is mostly cache reads.

Tiering: claude-fable-5.1 takes the Fable line's only tier-1 slot and
claude-fable-5 drops to tier 2. 5.1 shipped 2026-09-01 and 5 was already
being measured here on 2026-06-09, at least 84 days earlier — past the
retention policy's under-a-month carve-out, so the two do not share the
slot. Fable 5's results are fresh as of #120, so it needs no ACCEPTED_STALE
entries on the way down.

Also drops the two ACCEPTED_STALE entries added for these slugs, which
existed only to keep eval-cache-check green while results were empty.

Verified: pnpm typecheck clean, pnpm test:cost 9/9, `agent-eval status
'claude-fable-5.1*'` reports "Everything up to date", and node
scripts/check-stale.mjs reports "Eval cache OK" against the CI-pinned
next.js SHA. agent-results.json is `pnpm export-results` output, not a hand
edit, and the committed transcripts were scanned for gateway/Vercel/GitHub
credentials before commit (cf. #108).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com>
@gaojude gaojude changed the title Register Claude Fable 5.1 (high) for nextjs.org/evals Add Claude Fable 5.1 (high) to nextjs.org/evals Sep 9, 2026
Merging in this repo does not change nextjs.org/evals — the page reads a
copy of agent-results.json checked into vercel/front, so a task that stops
at the export is still invisible. The skill now ends where the site does.

Adds the copy path, notes that the page is fully data-driven off that file
so no component changes are needed, and records the two things the README's
bare `cp` glosses over: the export writes no trailing newline and front runs
Prettier, and the copy is normally a release or two behind, so the diff
carries earlier refreshes as well as yours.

Also notes that front rejects API-authored commits (409, verified signatures
required) — clone sparse and push, rather than burning a turn on the
contents API.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Co-Authored-By: Jude Gao <32973745+gaojude@users.noreply.github.com>
@gaojude
gaojude merged commit 6288d12 into main Sep 9, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants