[world] Make the sealed log opt-in instead of default-on - #3735
Conversation
Stamping spec 7 on every new run has two unmet preconditions. The Python runtime validates `specVersion <= 6` and rejects a spec-7 `run_started` outright, so every run it serves is unrunnable. The docs for the flag already say to leave it off where a runtime pinning its own spec range has not caught up; nothing was enforcing that. Pre-assigned positions are also stranding runs. A spec-7 run stalls between a step outcome and the resume that should follow it, and only the queue's own redelivery moves it on ~860s later (~1086s for a step dispatch). Measured on the e2e team over 2026-08-21 20:00-23:00 UTC, against the same backend in the same window: spec 7 stalled 380/28500 runs (1.33%) versus spec 6 at 8/18819 (0.04%). Reading is unchanged: SPEC_VERSION_MAX_SUPPORTED stays 7 and the runtime still floors at slot identity, so runs created while the flag was on stay readable and can still be advanced. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
🦋 Changeset detectedLatest commit: 862f04e The changes in this PR will be included in the next version bump. This PR includes changesets to release 20 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
🧪 E2E Test Results❌ Some tests failed ❌ Failed E2E Tests▲ Vercel Production (1 failed)nextjs-webpack-quickjs (1 failed):
|
| Passed | Failed | Skipped | Total | |
|---|---|---|---|---|
| ❌ ▲ Vercel Production | 3577 | 1 | 742 | 4320 |
| ✅ 💻 Local Development | 3922 | 0 | 558 | 4480 |
| ✅ 📦 Local Production | 3922 | 0 | 558 | 4480 |
| ✅ 🐘 Local Postgres | 3922 | 0 | 558 | 4480 |
| ✅ 🪟 Windows | 320 | 0 | 0 | 320 |
| ✅ 🌐 Cross-language Conformance | 9 | 0 | 132 | 141 |
| ✅ vercel-multi-region | 27 | 0 | 0 | 27 |
| Total | 15699 | 1 | 2548 | 18248 |
Details by Category
❌ ▲ Vercel Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-node | 132 | 0 | 28 |
| ✅ astro-quickjs | 132 | 0 | 28 |
| ✅ example-node | 132 | 0 | 28 |
| ✅ example-quickjs | 132 | 0 | 28 |
| ✅ express-node | 132 | 0 | 28 |
| ✅ express-quickjs | 132 | 0 | 28 |
| ✅ fastify-node | 132 | 0 | 28 |
| ✅ fastify-quickjs | 132 | 0 | 28 |
| ✅ hono-node | 132 | 0 | 28 |
| ✅ hono-quickjs | 132 | 0 | 28 |
| ✅ nest-node | 132 | 0 | 28 |
| ✅ nest-quickjs | 132 | 0 | 28 |
| ✅ nextjs-turbopack-node | 157 | 0 | 3 |
| ✅ nextjs-turbopack-quickjs | 157 | 0 | 3 |
| ✅ nextjs-webpack-node | 157 | 0 | 3 |
| ❌ nextjs-webpack-quickjs | 156 | 1 | 3 |
| ✅ nitro-node | 132 | 0 | 28 |
| ✅ nitro-quickjs | 132 | 0 | 28 |
| ✅ nuxt-node | 132 | 0 | 28 |
| ✅ nuxt-quickjs | 132 | 0 | 28 |
| ✅ python-node | 8 | 0 | 152 |
| ✅ sveltekit-node | 151 | 0 | 9 |
| ✅ sveltekit-quickjs | 151 | 0 | 9 |
| ✅ tanstack-start-node | 132 | 0 | 28 |
| ✅ tanstack-start-quickjs | 132 | 0 | 28 |
| ✅ vite-node | 132 | 0 | 28 |
| ✅ vite-quickjs | 132 | 0 | 28 |
✅ 💻 Local Development
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 26 |
| ✅ astro-stable-quickjs | 134 | 0 | 26 |
| ✅ express-stable-node | 134 | 0 | 26 |
| ✅ express-stable-quickjs | 134 | 0 | 26 |
| ✅ fastify-stable-node | 134 | 0 | 26 |
| ✅ fastify-stable-quickjs | 134 | 0 | 26 |
| ✅ hono-stable-node | 134 | 0 | 26 |
| ✅ hono-stable-quickjs | 134 | 0 | 26 |
| ✅ nest-stable-node | 134 | 0 | 26 |
| ✅ nest-stable-quickjs | 134 | 0 | 26 |
| ✅ nextjs-turbopack-canary-node | 141 | 0 | 19 |
| ✅ nextjs-turbopack-canary-quickjs | 141 | 0 | 19 |
| ✅ nextjs-turbopack-stable-node | 160 | 0 | 0 |
| ✅ nextjs-turbopack-stable-quickjs | 160 | 0 | 0 |
| ✅ nextjs-webpack-canary-node | 141 | 0 | 19 |
| ✅ nextjs-webpack-canary-quickjs | 141 | 0 | 19 |
| ✅ nextjs-webpack-stable-node | 160 | 0 | 0 |
| ✅ nextjs-webpack-stable-quickjs | 160 | 0 | 0 |
| ✅ nitro-stable-node | 134 | 0 | 26 |
| ✅ nitro-stable-quickjs | 134 | 0 | 26 |
| ✅ nuxt-stable-node | 134 | 0 | 26 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 26 |
| ✅ sveltekit-stable-node | 153 | 0 | 7 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 7 |
| ✅ tanstack-start-node | 134 | 0 | 26 |
| ✅ tanstack-start-quickjs | 134 | 0 | 26 |
| ✅ vite-stable-node | 134 | 0 | 26 |
| ✅ vite-stable-quickjs | 134 | 0 | 26 |
✅ 📦 Local Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 26 |
| ✅ astro-stable-quickjs | 134 | 0 | 26 |
| ✅ express-stable-node | 134 | 0 | 26 |
| ✅ express-stable-quickjs | 134 | 0 | 26 |
| ✅ fastify-stable-node | 134 | 0 | 26 |
| ✅ fastify-stable-quickjs | 134 | 0 | 26 |
| ✅ hono-stable-node | 134 | 0 | 26 |
| ✅ hono-stable-quickjs | 134 | 0 | 26 |
| ✅ nest-stable-node | 134 | 0 | 26 |
| ✅ nest-stable-quickjs | 134 | 0 | 26 |
| ✅ nextjs-turbopack-canary-node | 141 | 0 | 19 |
| ✅ nextjs-turbopack-canary-quickjs | 141 | 0 | 19 |
| ✅ nextjs-turbopack-stable-node | 160 | 0 | 0 |
| ✅ nextjs-turbopack-stable-quickjs | 160 | 0 | 0 |
| ✅ nextjs-webpack-canary-node | 141 | 0 | 19 |
| ✅ nextjs-webpack-canary-quickjs | 141 | 0 | 19 |
| ✅ nextjs-webpack-stable-node | 160 | 0 | 0 |
| ✅ nextjs-webpack-stable-quickjs | 160 | 0 | 0 |
| ✅ nitro-stable-node | 134 | 0 | 26 |
| ✅ nitro-stable-quickjs | 134 | 0 | 26 |
| ✅ nuxt-stable-node | 134 | 0 | 26 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 26 |
| ✅ sveltekit-stable-node | 153 | 0 | 7 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 7 |
| ✅ tanstack-start-node | 134 | 0 | 26 |
| ✅ tanstack-start-quickjs | 134 | 0 | 26 |
| ✅ vite-stable-node | 134 | 0 | 26 |
| ✅ vite-stable-quickjs | 134 | 0 | 26 |
✅ 🐘 Local Postgres
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 134 | 0 | 26 |
| ✅ astro-stable-quickjs | 134 | 0 | 26 |
| ✅ express-stable-node | 134 | 0 | 26 |
| ✅ express-stable-quickjs | 134 | 0 | 26 |
| ✅ fastify-stable-node | 134 | 0 | 26 |
| ✅ fastify-stable-quickjs | 134 | 0 | 26 |
| ✅ hono-stable-node | 134 | 0 | 26 |
| ✅ hono-stable-quickjs | 134 | 0 | 26 |
| ✅ nest-stable-node | 134 | 0 | 26 |
| ✅ nest-stable-quickjs | 134 | 0 | 26 |
| ✅ nextjs-turbopack-canary-node | 141 | 0 | 19 |
| ✅ nextjs-turbopack-canary-quickjs | 141 | 0 | 19 |
| ✅ nextjs-turbopack-stable-node | 160 | 0 | 0 |
| ✅ nextjs-turbopack-stable-quickjs | 160 | 0 | 0 |
| ✅ nextjs-webpack-canary-node | 141 | 0 | 19 |
| ✅ nextjs-webpack-canary-quickjs | 141 | 0 | 19 |
| ✅ nextjs-webpack-stable-node | 160 | 0 | 0 |
| ✅ nextjs-webpack-stable-quickjs | 160 | 0 | 0 |
| ✅ nitro-stable-node | 134 | 0 | 26 |
| ✅ nitro-stable-quickjs | 134 | 0 | 26 |
| ✅ nuxt-stable-node | 134 | 0 | 26 |
| ✅ nuxt-stable-quickjs | 134 | 0 | 26 |
| ✅ sveltekit-stable-node | 153 | 0 | 7 |
| ✅ sveltekit-stable-quickjs | 153 | 0 | 7 |
| ✅ tanstack-start-node | 134 | 0 | 26 |
| ✅ tanstack-start-quickjs | 134 | 0 | 26 |
| ✅ vite-stable-node | 134 | 0 | 26 |
| ✅ vite-stable-quickjs | 134 | 0 | 26 |
✅ 🪟 Windows
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack-node | 160 | 0 | 0 |
| ✅ nextjs-turbopack-quickjs | 160 | 0 | 0 |
✅ 🌐 Cross-language Conformance
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ python | 9 | 0 | 132 |
✅ vercel-multi-region
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack | 27 | 0 | 0 |
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 188787ms → this run 211947ms (Δ +23160ms, +12%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): ℹ️ Metric definitions & methodologyStreams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 41 total
Full trace: |
The conformance suite pinned a run's stamped specVersion to SPEC_VERSION_CURRENT. What a World is told to stamp is mintedSpecVersion(), and the two differ whenever a version is readable before it is mintable — the normal mid-bump state, not a conformance defect. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
No backport to This commit flips the To override, re-run the Backport to stable workflow manually via |
* Revert "[world] Make the sealed log opt-in instead of default-on (#3735)" Reverts b2cac62. New runs are stamped at spec 7 again, now that a read which cannot see past an unfilled position waits for it instead of reporting a log that ends there (workflow-server: derive the in-request seal poll budget from the staleness bound). Two things are kept from #3735 rather than reverted: - the world-testing conformance floor at mintedSpecVersion(), which was wrong for any staged bump and not specific to this default - a note on mintedSpecVersion recording what default-on rests on: the events density requirement, and that a sealed log meets it by repair rather than by construction, so the READ has to wait Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * TEMPORARY: point world-vercel at workflow-server#839 preview Validating the seal-poll-budget fix end to end with spec 7 on. Reverted before merge; the override lint guard is expected to fail meanwhile. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Revert "TEMPORARY: point world-vercel at workflow-server#839 preview" This reverts commit 5e17cc9. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Summary
WORKFLOW_SEALED_LOGto off, so new runs are stamped at the slot-identity spec version rather than the sealed log (spec 7)WORKFLOW_SEALED_LOG=1opts inSPEC_VERSION_MAX_SUPPORTEDstays 7 and the runtime still floors at slot identitymintedSpecVersion()rather thanSPEC_VERSION_CURRENTWhy
Main's E2E has been red continuously since #3634 landed. Two unmet preconditions for stamping spec 7 on every new run:
1. The Python runtime cannot read a sealed log
It validates
specVersion <= 6and rejects a spec-7run_startedoutright:Every run it serves is unrunnable, which is why
E2E Python ConformanceandE2E Vercel Prod Tests (python - node)fail. The flag's own docs already say to leave it off where a runtime that pins its own spec range has not caught up; nothing was enforcing it.2. A truncated sealed-log read makes a completed step look in-flight, and inline ownership then waits out a full 860s lease
This one is a composition bug between two individually-sound features.
Spec 7 assigns a log position before the write commits, so a fan-out nearly always has a position whose writer is still in flight.
events.tssays so directly:The densifying read truncates below that hole, and a truncated page comes back empty with
hasMore: false. So a resume replay can read a log that stops below thestep_completedevents — the completed steps still look pending.backfillSealedLogHolesalready documents the resulting wedge:What makes this a timed stall rather than a permanent wedge is inline step ownership (#2780). Those pending steps still carry an
ownerMessageIdon their lateststep_startedwith nostep_retrying, soisStepOwnershipActiveis true and the lease has time left. The dispatch decision table then deliberately does not re-dispatch, because dispatching would double-execute a step that looks like it is still running:It arms a delayed run continuation at
lastStartedAt + INLINE_OWNERSHIP_LEASE_SECONDSinstead. That constant is 860. When the backstop fires the hole is long sealed, the log reads dense, and the run finishes in ~2s.The data matches the constant, measured from the anchoring
step_startedon the e2e team, 2026-08-21 20:00-23:30 UTC:step_completedstep_failedSub-second variance over 383 samples: this is the lease, not queue jitter. (A second, noisier family sits at ~1086s after a
step_started, stddev 56s — the owning message's redelivery re-stamping a fresh ownership epoch.)Spec 6 cannot produce it: positions are allocated by the write that occupies them, so there is no hole, no truncation, and the replay always sees the completions. Same backend, same hour, spec-6 runs coming from PR branches whose base predates #3634:
A representative run (
abortAnyInStepWorkflow, two parallel steps completing 168ms apart):At a 60s e2e test timeout this surfaces as a diffuse spread of unrelated failures across nearly every framework lane.
Scope
This is the kill switch the flag was built for, not a fix for the composition. It buys back a green main and unblocks the Python lanes.
The actual fix is upstream of the flag and worth doing separately. Two candidates, and the first looks right: a truncated read must not feed the ownership decision. Truncation means "the log might not end here", so concluding a step is still pending from a page that admits it cut itself short is unsound — that path should re-read or re-drive rather than arm an 860s backstop. Relatedly, sealing a hole currently repairs the log but nothing re-drives the runs whose readers were truncated by it; a seal probably owes a resume.
Note this is a separate cause from #3709. That addresses run-status long-poll pool starvation after #3570, which is why main went red on 2026-08-20; these stalls begin 2026-08-21 ~20:00 UTC when #3634 landed. Both are needed for a green main.
Compatibility
Runs created while the flag was on keep working. Verified across two Worlds over one storage directory, writer opted in and reader on the new default:
assertWorldSupportsRuntimeProtocolaccepts[6, 7], so the World this now selects is admitted; that floor was put there for exactly this rollback.Test plan
pnpm --filter @workflow/world test(160 passed)pnpm --filter @workflow/world-local test(557 passed)pnpm --filter @workflow/world-vercel test(548 passed)pnpm --filter @workflow/core test(104 files passed, 1 skipped)vitest run packages/world-testing/test/embedded.test.ts(13 passed, incl.numbers events by position)turbo run typecheck --filter=@workflow/world --filter=@workflow/corebiome checkon the changed filesmintedSpecVersionacross unset/""/0/false/1/true/malformed, and the cross-flag readback above🤖 Generated with Claude Code