A process.exec step in rote captures the first 65,536 bytes of stdout
and discards the rest. The step exits 0. The runner reports the run completed. The linter cannot see
it.
I found it because it was quietly corrupting my own play's output, then measured how much of the public play registry stands in the same place: 249 published plays consume a process step's stdout, and 221 of them — 89% — never read the truncation flag.
Rendered writeup, with every truncation drawn to scale: https://aryangorde6.github.io/rote-stdout-truncation/ Filed upstream: modiqo/rote-releases#10 The audit as a runnable play: https://play.modiqo.ai/aryangorde/truncation-audit
rote play run https://play.modiqo.ai/aryangorde/truncation-audit root=~/.rote
A process.exec step captures at most 65,536 bytes into body.stdout.text. It sets
body.stdout.truncated: true alongside bytes and preview_bytes, and writes the complete payload
to .rote/artifacts/processes/@N/stdout.json — a path a steps_with_presentation body cannot open.
The full data exists on disk and is unreachable from the code that needs it.
The cap is exact, not approximate. Across 4,825 process observations recorded on one machine, 14 were truncated, and every one kept precisely 65,536 bytes.
Reproduced on rote 0.80.0 with a three-line play:
steps:
big:
type: process.exec
argv: [python3, -c, "print('x'*100000)"]emitted by command : 100001 bytes
captured text len : 65536
truncated flag : true
step exit : { kind: "code", code: 0 }
run status : succeeded
It is undocumented. 65536 appears twice in rote guidance typescript play-creation — once as
browser.extract's unrelated truncated field, once as an author-supplied maxBytes argument.
Neither describes this cap.
The failure is invisible at every gate that exists to catch failures:
| Gate | What it reports |
|---|---|
| the child process | exits 0 |
| the step | completed |
| the runner | N/N completed, 0 failed, 0 blocked |
| the run | exits 0 |
rote play lint |
passes — fixtures are small, so it never reproduces |
It only appears on real input, at real scale, in production. There are two failure modes, and the second is far worse:
- The prefix breaks the parser.
JSON.parsethrows on the cut object. You get an error and blame the provider for a malformed response. Annoying, but loud. - The prefix is still valid input. If the step emits newline-delimited records — or the body
splits on
\nand filters — the truncated data parses cleanly.
Mode 2 gives you a confident, wrong answer with no error anywhere. Not a crash: a number that is simply too low. That is the one that ships.
Eight distinct steps across five published plays, recorded on one laptop over two days of ordinary use. A ninth cut belongs to the synthetic three-line repro I wrote to confirm the cap, and is excluded here:
| Play | Step | Produced | Carried | Never read |
|---|---|---|---|---|
| repo-leak-doctor | history_check |
1,346,888 | 65,536 | 95.1% |
| repo-leak-doctor | tracked_files |
587,958 | 65,536 | 88.9% |
| pr-review-verdict | fetch_inline |
448,889 | 65,536 | 85.4% |
| tech-debt | scan_suppressions |
155,641 | 65,536 | 57.9% |
| repo-onboarding-brief | read_runner_scripts |
146,366 | 65,536 | 55.2% |
| pr-review-verdict | fetch_reviews |
133,553 | 65,536 | 50.9% |
| repo-onboarding-brief | read_secondary_ci |
101,672 | 65,536 | 35.5% |
| review-gate | merge_checks |
82,822 | 65,536 | 20.9% |
The last one is mine. review-gate answers one question before you merge: are there unresolved
review threads or failing checks? Pointed at
DataDog/dd-trace-js#10068 — a pull request
carrying 917 checks — it returned a confident verdict computed from four fifths of the evidence.
A user running the audit reported that it only ever saw steps that were cut — his step had stopped being cut and started timing out instead, bypassing every degrade path he had written. He was right, and my first explanation of my own bug was wrong. I said a killed step was invisible by construction: the run fails, never reaches the presentation plane, no record is written.
That is not what happens. Forty-eight failed runs on this machine wrote presentation records, and the timed-out one was among them:
run_20260905_084310.629_1 status: failed
step evaluate -> status: "failed"
output.diagnostic.response_id: 2
output.diagnostic.exit: { kind: "timed_out", timeout_ms: 30000 }
The rule, measured across all 896 step outcomes on disk:
| Outcome | body |
diagnostic |
Steps |
|---|---|---|---|
| completed | yes | — | 787 |
| failed | — | yes | 55 |
| blocked | — | — | 50 |
| failed (no observation) | — | — | 4 |
A completed step files its observation under outcome.output.body; a failed step files it
under outcome.output.diagnostic. A killed step is a failed step. My audit read bodies, so it
reported zero timeouts on a machine that had one — and the record had been sitting there the whole
time, step name attached.
It was never structural. It was a place I had not looked, and I had called it a hole in the runtime.
Truncation and timeout are two outcomes of one condition: the payload exceeded what the step could carry. Only the first sets a flag anyone can read.
302 packages pulled and statically analysed — 298 published, plus local working copies of my own four plays, excluded from every published figure below so my own work is not double-counted.
| Class | n | % |
|---|---|---|
| consumes stdout, never reads the flag | 221 | 73.2 |
| consumes stdout, reads the flag | 32 | 10.6 |
| no stdout read | 22 | 7.3 |
| no process steps at all | 16 | 5.3 |
| legacy no-steps body | 11 | 3.6 |
Of the 249 published packages that consume process stdout, 221 — 89% — never read the truncation flag. The 28 that do belong to seven owners.
My first denominator was wrong, and the way it was wrong is instructive.
rote registry play list has no list-all — it requires an owner, org or community. So I enumerated
by sweeping search queries: 30 queries surfaced 292 plays, 45 surfaced 532. It was tempting to call
532 "the registry."
It is not. Per-owner listing is exhaustive per owner, so I listed all 122 owners directly and got 697. Every play search had found was in that set; 165 were not. Search had missed 24% of the registry — and no number of additional queries would have revealed that, because a search sweep cannot report what it failed to surface.
If you want to enumerate that registry: list owners, do not sweep queries.
The instinct is to ask for a bigger buffer. That is the wrong shape of fix — it moves the cliff without removing it. The rule that generalises: don't move the payload, move the answer.
Aggregate where the rows still exist, inside the step, before stdout is ever written. gh's --jq
and --template run in-process, so filtering there costs nothing. Carry out a bounded number of
detail rows plus an explicit retained-vs-total count, so any list you print is a labelled lower
bound rather than a confidently wrong number.
Applied to the step that started this, filtering passing checks server-side instead of client-side:
before 82,822 bytes truncated
after 173 bytes clean — check count still exact
A 99.8% reduction with no loss of the actual answer. The payload was never the point.
per_page=100on a REST call is safe.- An unpaginated GraphQL body, a
gh pr view --jsonstatus rollup, and rawgit log -porgit ls-filespiped into the body are not. - Detect it with
body.stdout?.truncated === true— it is on the typedProcessExecStreamthe presentation SDK hands you.
The finding changed five times, and the tool caused every correction. Each wrong answer was confident.
- "Four authors handle this." Wrong. All four were matching their own unrelated clamping logic
that happened to use the word
truncated. - "Only my play handles it." True across the 26 plays I had installed. False across 302.
- "45 are guarded." Wrong again. The proximity heuristic — looking for
truncatednearstdout— had a 29% false-positive rate, 12 of 41, matching things likefindings_truncatedbesidestdout_bytes. A false "this is handled" is the worst possible output for a tool like this, so the detector now requires member access on the stream itself, with frontmatter and comments stripped first. - "Search found 532, so listing finds 31% more." The share of the registry actually missed is 165/697 = 24%. 31% was 165/532 — a different question, silently answered.
- "The audit reads 14% of the available evidence." It reads about 84%. The script I wrote to verify the play walked only single-shape step outputs and skipped 3,328 fan-out items — so I trusted a throwaway checker over the audited tool without ever running one against the other. This one reached another person before I caught it.
Three of the five corrections came from other people running the play and reporting what it got wrong. Every one was a failure mode I had not personally hit — which is the argument for being wrong in public quickly rather than careful in private.
Using the tool at scale also found defects in the tool itself, including the one I like most: its own output reached 41 KB against the 64 KB cap at 302 plays. The truncation auditor was on course to truncate its own evidence. Compacting the index to positional rows brought it back to 29.5 KB.
| Path | What it is |
|---|---|
index.html |
the rendered writeup, served at the Pages URL above |
audit/scan_runs.py |
run-history scanner: reads the response log and the presentation plane, joins them on response_id, reports cuts and timeouts apart |
audit/scan_plays.py |
static scanner: reads installed play sources and separates plays that check the flag from those that consume stdout without it |
Both scanners aggregate inside the step and emit a bounded summary, for the reason the writeup describes — a naive version of this audit would truncate on its own evidence.
All four are public, read-only, and need nothing beyond gh:
| Play | What it answers |
|---|---|
| truncation-audit | which of your plays silently drop data, and which already have |
| flake-finder | which CI jobs are flaky, grouped by job across runs rather than by run |
| reviewer-finder | who should review this diff, by breadth of prior contact, flagging bus-factor-1 files |
| review-gate | unresolved threads and failing checks, before merge |
Built for the Rote Playoffs, September 2026.