Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Cut at 65,536

A process.exec step in rote captures the first 65,536 bytes of stdout and discards the rest. The step exits 0. The runner reports the run completed. The linter cannot see it.

I found it because it was quietly corrupting my own play's output, then measured how much of the public play registry stands in the same place: 249 published plays consume a process step's stdout, and 221 of them — 89% — never read the truncation flag.

Rendered writeup, with every truncation drawn to scale: https://aryangorde6.github.io/rote-stdout-truncation/ Filed upstream: modiqo/rote-releases#10 The audit as a runnable play: https://play.modiqo.ai/aryangorde/truncation-audit

rote play run https://play.modiqo.ai/aryangorde/truncation-audit root=~/.rote

What the cap actually is

A process.exec step captures at most 65,536 bytes into body.stdout.text. It sets body.stdout.truncated: true alongside bytes and preview_bytes, and writes the complete payload to .rote/artifacts/processes/@N/stdout.json — a path a steps_with_presentation body cannot open. The full data exists on disk and is unreachable from the code that needs it.

The cap is exact, not approximate. Across 4,825 process observations recorded on one machine, 14 were truncated, and every one kept precisely 65,536 bytes.

Reproduced on rote 0.80.0 with a three-line play:

steps:
  big:
    type: process.exec
    argv: [python3, -c, "print('x'*100000)"]
emitted by command : 100001 bytes
captured text len  : 65536
truncated flag     : true
step exit          : { kind: "code", code: 0 }
run status         : succeeded

It is undocumented. 65536 appears twice in rote guidance typescript play-creation — once as browser.extract's unrelated truncated field, once as an author-supplied maxBytes argument. Neither describes this cap.

Why nothing reports it

The failure is invisible at every gate that exists to catch failures:

Gate What it reports
the child process exits 0
the step completed
the runner N/N completed, 0 failed, 0 blocked
the run exits 0
rote play lint passes — fixtures are small, so it never reproduces

It only appears on real input, at real scale, in production. There are two failure modes, and the second is far worse:

  1. The prefix breaks the parser. JSON.parse throws on the cut object. You get an error and blame the provider for a malformed response. Annoying, but loud.
  2. The prefix is still valid input. If the step emits newline-delimited records — or the body splits on \n and filters — the truncated data parses cleanly.

Mode 2 gives you a confident, wrong answer with no error anywhere. Not a crash: a number that is simply too low. That is the one that ships.

What it cut

Eight distinct steps across five published plays, recorded on one laptop over two days of ordinary use. A ninth cut belongs to the synthetic three-line repro I wrote to confirm the cap, and is excluded here:

Play Step Produced Carried Never read
repo-leak-doctor history_check 1,346,888 65,536 95.1%
repo-leak-doctor tracked_files 587,958 65,536 88.9%
pr-review-verdict fetch_inline 448,889 65,536 85.4%
tech-debt scan_suppressions 155,641 65,536 57.9%
repo-onboarding-brief read_runner_scripts 146,366 65,536 55.2%
pr-review-verdict fetch_reviews 133,553 65,536 50.9%
repo-onboarding-brief read_secondary_ci 101,672 65,536 35.5%
review-gate merge_checks 82,822 65,536 20.9%

The last one is mine. review-gate answers one question before you merge: are there unresolved review threads or failing checks? Pointed at DataDog/dd-trace-js#10068 — a pull request carrying 917 checks — it returned a confident verdict computed from four fifths of the evidence.

The second outcome

A user running the audit reported that it only ever saw steps that were cut — his step had stopped being cut and started timing out instead, bypassing every degrade path he had written. He was right, and my first explanation of my own bug was wrong. I said a killed step was invisible by construction: the run fails, never reaches the presentation plane, no record is written.

That is not what happens. Forty-eight failed runs on this machine wrote presentation records, and the timed-out one was among them:

run_20260905_084310.629_1   status: failed
step evaluate -> status: "failed"
  output.diagnostic.response_id: 2
  output.diagnostic.exit: { kind: "timed_out", timeout_ms: 30000 }

The rule, measured across all 896 step outcomes on disk:

Outcome body diagnostic Steps
completed yes — 787
failed — yes 55
blocked — — 50
failed (no observation) — — 4

A completed step files its observation under outcome.output.body; a failed step files it under outcome.output.diagnostic. A killed step is a failed step. My audit read bodies, so it reported zero timeouts on a machine that had one — and the record had been sitting there the whole time, step name attached.

It was never structural. It was a place I had not looked, and I had called it a hole in the runtime.

Truncation and timeout are two outcomes of one condition: the payload exceeded what the step could carry. Only the first sets a flag anyone can read.

How much of the registry

302 packages pulled and statically analysed — 298 published, plus local working copies of my own four plays, excluded from every published figure below so my own work is not double-counted.

Class n %
consumes stdout, never reads the flag 221 73.2
consumes stdout, reads the flag 32 10.6
no stdout read 22 7.3
no process steps at all 16 5.3
legacy no-steps body 11 3.6

Of the 249 published packages that consume process stdout, 221 — 89% — never read the truncation flag. The 28 that do belong to seven owners.

A methodology correction worth recording

My first denominator was wrong, and the way it was wrong is instructive.

rote registry play list has no list-all — it requires an owner, org or community. So I enumerated by sweeping search queries: 30 queries surfaced 292 plays, 45 surfaced 532. It was tempting to call 532 "the registry."

It is not. Per-owner listing is exhaustive per owner, so I listed all 122 owners directly and got 697. Every play search had found was in that set; 165 were not. Search had missed 24% of the registry — and no number of additional queries would have revealed that, because a search sweep cannot report what it failed to surface.

If you want to enumerate that registry: list owners, do not sweep queries.

The fix

The instinct is to ask for a bigger buffer. That is the wrong shape of fix — it moves the cliff without removing it. The rule that generalises: don't move the payload, move the answer.

Aggregate where the rows still exist, inside the step, before stdout is ever written. gh's --jq and --template run in-process, so filtering there costs nothing. Carry out a bounded number of detail rows plus an explicit retained-vs-total count, so any list you print is a labelled lower bound rather than a confidently wrong number.

Applied to the step that started this, filtering passing checks server-side instead of client-side:

before   82,822 bytes   truncated
after       173 bytes   clean — check count still exact

A 99.8% reduction with no loss of the actual answer. The payload was never the point.

  • per_page=100 on a REST call is safe.
  • An unpaginated GraphQL body, a gh pr view --json status rollup, and raw git log -p or git ls-files piped into the body are not.
  • Detect it with body.stdout?.truncated === true — it is on the typed ProcessExecStream the presentation SDK hands you.

What I got wrong

The finding changed five times, and the tool caused every correction. Each wrong answer was confident.

  1. "Four authors handle this." Wrong. All four were matching their own unrelated clamping logic that happened to use the word truncated.
  2. "Only my play handles it." True across the 26 plays I had installed. False across 302.
  3. "45 are guarded." Wrong again. The proximity heuristic — looking for truncated near stdout — had a 29% false-positive rate, 12 of 41, matching things like findings_truncated beside stdout_bytes. A false "this is handled" is the worst possible output for a tool like this, so the detector now requires member access on the stream itself, with frontmatter and comments stripped first.
  4. "Search found 532, so listing finds 31% more." The share of the registry actually missed is 165/697 = 24%. 31% was 165/532 — a different question, silently answered.
  5. "The audit reads 14% of the available evidence." It reads about 84%. The script I wrote to verify the play walked only single-shape step outputs and skipped 3,328 fan-out items — so I trusted a throwaway checker over the audited tool without ever running one against the other. This one reached another person before I caught it.

Three of the five corrections came from other people running the play and reporting what it got wrong. Every one was a failure mode I had not personally hit — which is the argument for being wrong in public quickly rather than careful in private.

Using the tool at scale also found defects in the tool itself, including the one I like most: its own output reached 41 KB against the 64 KB cap at 302 plays. The truncation auditor was on course to truncate its own evidence. Compacting the index to positional rows brought it back to 29.5 KB.

What's in this repo

Path What it is
index.html the rendered writeup, served at the Pages URL above
audit/scan_runs.py run-history scanner: reads the response log and the presentation plane, joins them on response_id, reports cuts and timeouts apart
audit/scan_plays.py static scanner: reads installed play sources and separates plays that check the flag from those that consume stdout without it

Both scanners aggregate inside the step and emit a bounded summary, for the reason the writeup describes — a naive version of this audit would truncate on its own evidence.

Plays

All four are public, read-only, and need nothing beyond gh:

Play What it answers
truncation-audit which of your plays silently drop data, and which already have
flake-finder which CI jobs are flaky, grouped by job across runs rather than by run
reviewer-finder who should review this diff, by breadth of prior contact, flagging bus-factor-1 files
review-gate unresolved threads and failing checks, before merge

Built for the Rote Playoffs, September 2026.

About

Cut at 65,536 — a silent 64 KiB stdout cap in rote, what it cut on one machine, and how much of the public play registry stands in the same place.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages