Skip to content

feat(testing): CSD flows run on the five-platform matrix; state: now asserts something - #97

Merged
emooreatx merged 23 commits into
mainfrom
feat/csd-flows-on-the-matrix
Sep 29, 2026
Merged

emooreatx merged 23 commits into
mainfrom
feat/csd-flows-on-the-matrix

Conversation

@emooreatx

Copy link
Copy Markdown
Contributor

CSD flows now run on this repo's five-platform matrix. That's what a CSD needs to reach testable (CSD.md §1: "floor flips off unreleased; flow runs on the matrix"). Until now the runner (testing/gate/flow_spec.py) was vendored here with no flows and no step that ran it; CIRISAgent's gate ran the only four flows (CSD-001–004).

How a flow links to its CSD

  • testing/flows/*.yaml; each flow names csd: CSD-0NN, which must match exactly one FSD/CSD/CSD-0NN-*.md.
  • The runner loads that CSD's shows: (field id → tag, so relation works on field ids) and states: (state → tag).
  • The blocks are parsed with check_csd_v3's own pattern, so the checker and the runner can't disagree.
  • Load errors, raised before anything boots: a missing, ambiguous or unparseable CSD; any proposed: tag in requires, expect or do (a flow can't assert a tag that doesn't exist); a state: whose tag is missing or proposed; a relation operand that isn't a shows: field; a flow with no csd:.

A fix in the vendored runner: state: did nothing, so every state: assertion passed. It now requires that state's tag on screen and every other state's tag off screen. Recorded in VENDORED.md as a local delta, with the csd: key, both to go upstream to CIRISAgent.

Runner (run_platform --flows testing/flows, or run_flows.py standalone):

  • Flows load before bring-up, run after a clean smoke walk in the same app, and sign in once through the existing session fixture.
  • Verdicts: pass · refused (the client: floor is above the client under test; reported, leg stays green) · cannot-start (floor met but the first precondition never held; red, because a flow that never started silently never ran) · fail (red).
  • Results go to reports/<leg>.json under flows, screenshots to shots/flows-<leg>/.
  • Wired into all five legs of five-platform-live-qa.yml.

Seeded flow: people.yaml (CSD-005, client: ">=0.5.224"): landing on Contacts with the add card, a search that matches nothing reaching state: empty, clearing it bringing the add card back. Every tag is on main, and a test checks each exists in the client source.

Checks: 35 new tests (168 in testing/); every red path planted (11 defects, table in the commit) · packaging/gates.sh 0 · nothing under client/ changed.

Live proof: the five-platform run dispatched on this branch; its result is linked in a comment below.

Known risks on a live run: if the app shows Add Federation ID after sign-in, or doesn't land on Contacts within 30 s, the result is cannot-start. The session fixture has never run in CI before. Flows for other surfaces will need nav_map hops.

🤖 Generated with Claude Code

https://claude.ai/code/session_01XXpYDXz2XUwUCG1ePtmVMK

A CSD reaches `testable` when its flow runs on the matrix (CSD.md §1).
Nothing ran one here: flow_spec.py was vendored, but no leg drove a flow.

- testing/flows/: the home for flows. Each names its CSD (`csd: CSD-005`);
  the runner reads that CSD's `shows:` (field ids for `relation`) and
  `states:` (tags for `state:`). A missing or unparseable CSD, or a flow
  naming a `proposed:` tag, is a load error before anything boots.
- flow_spec.py local delta (VENDORED.md): the `csd:` key, CSD binding at
  load, and `state:` actually checked. Upstream's body was `pass`.
- run_flows.py: pass / fail / refused (by `client:` floor) / cannot-start.
  Refused stays green; cannot-start is red, since with the floor met it
  means the flow silently never ran.
- run_platform --flows: after a clean smoke walk, in the same app, one
  sign-in (session_fixture). Flows load before bring-up. All five legs in
  five-platform-live-qa.yml pass it, and PyYAML is installed per job.
- Seeded flow: people.yaml (CSD-005, client >=0.5.224). On a bare node it
  checks landing with the add card, a search with no match (state: empty),
  and a cleared search.
- testing/test_flows.py (35 tests), wired into build.yml. Every rule
  was mutation-checked red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XXpYDXz2XUwUCG1ePtmVMK
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, you can upgrade your account or add credits to your account and enable them for code reviews in your settings.

emooreatx and others added 2 commits September 25, 2026 18:48
…es what's on screen when a step won't advance

The first iOS run stopped at "wizard did not advance past 'you'": the four
fields were typed back to back, and iOS needs ~2 s per field for the value to
reach the ViewModel's StateFlow (client/CLAUDE.md). A step that still won't
advance now reports the tags on screen, so a required field the fixture
doesn't fill shows in the failure.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XXpYDXz2XUwUCG1ePtmVMK
@emooreatx

Copy link
Copy Markdown
Contributor Author

First live run (36199588428): the People flow (CSD-005) passed 3/3 on Linux, Android, macOS and Windows, with a screenshot per step. On iOS it never ran: the session fixture stopped at "wizard did not advance past 'you'". The fixture typed four fields back to back, and iOS needs about 2 s per field for the value to reach the StateFlow (client/CLAUDE.md). Fixed: it now settles between fields, and a step that won't advance reports what's on screen. main is merged in; a fresh full run is dispatched.

🤖 Generated with Claude Code

emooreatx added a commit that referenced this pull request Sep 26, 2026
…start

PR #97 (`feat/csd-flows-on-the-matrix`, open) defines the shape a flow here must
have, and it is stricter than main's vendored `flow_spec.py` in three ways. Two
were easy to meet; the third decides where these files live.

1. `csd:` is required and BINDING — it resolves to the one FSD/CSD/CSD-NNN-*.md
   and the runner reads that card's `shows:` and `states:`.
2. Naming a `proposed:` tag in `requires`, `expect` or `do` is a LOAD ERROR, as
   is `state: X` whose tag the CSD still marks proposed. Run against #97's real
   binder, 14 of the 15 flows loaded clean and one did not:
   `csd-090-duty-conferral` asserted `state: empty`, whose tag CSD-090 correctly
   marks `proposed:duty_source_absent` — the screen renders "this node knows no
   accord family yet" through `duty_source_error`, so the one state that is not a
   failure is the one that looks most like one. The step now asserts which rows
   are absent and says in a comment why it may not claim more.
3. THE RUNNER DOES NOT NAVIGATE, and `cannot-start` is RED on purpose. Sign-in
   lands on Contacts; every other surface is unreachable. Landing 13 flows that
   each stop at their first `requires` would turn all five matrix legs red for a
   reason that has nothing to do with the client.

So `csd-005-people.yaml` is DELETED — #97 already ships `testing/flows/people.yaml`
for CSD-005, and two flows for one CSD is the drift this repo exists to measure —
and the other 14 move to `testing/flows/drafts/`. `flow_spec.discover` globs
`p.glob("*.yaml")` and not `rglob`, so a subdirectory is staged and never run;
`discover(["testing/flows"])` returns [] here, confirmed against #97's own code.
Each draft is complete, loads clean, and promotes by `git mv` alone — the hop is
never written into a flow, it is derived from nav_map (CSD.md §2.0).

Six of the fourteen wait on navigation and nothing else: CSD-025, 036, 057, 066,
068 and 087. Navigation in the runner unblocks more flows at once than any tag or
route in this audit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XXpYDXz2XUwUCG1ePtmVMK
emooreatx added a commit that referenced this pull request Sep 26, 2026
…t flows

`test_every_tag_a_seeded_flow_names_exists_in_the_client` builds its literal set
with `re.findall(r'"([a-z][a-z0-9_]+)"', ...)`. There is no `$` in that character
class, so a tag the client composes by interpolation is invisible to it.

Seven such tags are named by four of the staged flows, and all seven are real at
run time: `age_band_adult` (`"age_band_$token"`, SetupScreen.kt:2661),
`trace_consent_yes`/`_no` (`"trace_consent_$token"`, :1095), `radio_cohort_self`
(`"radio_cohort_$value"`, ClaimNodeScreen.kt:383) and
`chk_duty_box_{moderate,review,consent_revocation}` (`"chk_duty_box_$verb"`,
DutyConferralScreen.kt:301). Promoting csd-082, csd-083, csd-085 or csd-090 as
written would redden that test against a client carrying every tag it names.

The fix belongs in the test: keep the part before the first `$` of any literal
that contains one, and accept a tag that starts with such a prefix.
`opt_run_with_ai` is the control — written as a whole literal
(SetupScreen.kt:2567) and seen today.

Recorded in testing/flows/drafts/README.md rather than worked around in the
flows, because a flow bent to fit a test's blind spot stops describing the app.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XXpYDXz2XUwUCG1ePtmVMK
emooreatx and others added 2 commits September 25, 2026 19:28
…g check sees interpolated tags

Sign-in lands on Contacts, so a flow for any other surface ended
cannot-start. Before step one, run_one now walks nav_map's derived hop
for the flow's first `requires: screen:` on this build (node or agent,
from /state). It waits for each tag, then clicks it.
- A missing hop tag is cannot-start and names the tag and its position.
- A screen with no hop that is not flow-only is cannot-start: "no nav
  hop for Screen.X".
- A flow-only screen (screen_atlas.flow_only) is waited for, not walked to.

test_every_tag_a_seeded_flow_names_exists_in_the_client read only whole
literals, so interpolated tags ("age_band_$token", "trace_consent_$token",
"radio_cohort_$value", "chk_duty_box_$verb") looked absent. A tag now
counts if it equals a literal or starts with the pre-`$` prefix of an
interpolated one. Only prefixes with two segments count, so "btn_$x"
cannot vouch for every button.

Every new rule was mutation-checked red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XXpYDXz2XUwUCG1ePtmVMK
…-matrix

# Conflicts:
#	.github/workflows/build.yml
emooreatx added a commit that referenced this pull request Sep 28, 2026
…reviewed, the 0.5.218 wait

Where we are: 72 CSDs on main (64 building, 7 sketched, 1 envisioned), the
new Network, 0.5.218, communities and key-verification cards, and the cards
retired by the dedup rule (001 into 055, 105/106 into 039, 107/108 into
054/020, 034 unused). Merged since #90 in three shapes: cards and fixes,
the checks that keep CSDs honest (route gate 121/26 -> 14/23), and the two
integration merges #115 and #124. Three PRs in review (#77 conflicting,
#97, #122); batch 2 of the gap-closing review is next.

The dependency graph stops pretending the circles are a chain: B5, B6 and
B7 are partly built ahead of B3 and B4, so the graph shows what still has
to come before what. CIRISAgent#1213 is closed and removed; the 0.5.218
cut takes its place as the hexagon four merged cards wait on. Persist#907
and #910 are closed; the households hexagon now names #686/#687/#916.

Upstream is regrouped by lane (people/chats/files, consent/data,
households/communities, accord/trust root, node ops, agent surfaces,
setup/identity) from the open issue lists in the four upstream repos.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
emooreatx and others added 11 commits September 28, 2026 13:47
…clickable before clicking it

On iOS the typed fields reach the ViewModel a beat late, so Next is on
screen but disabled, and a disabled control has no click handler: the
click was a bare 404 that named no step. Wait (bounded) for can_click,
and if it never enables say which step and what was on screen.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…ore Next, and retries once

JOIN_FEDERATION advances only when traceConsentAnswered; a Next clicked in
the same instant as the answer raced it on Windows and the step reported
'did not advance' with the question still on screen.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…k=false — retry, bounded, then name the step

The iOS automation tree omits canClick when false, so the clickable wait
passes; the click then 404s with 'No click handler'. Treat that answer as
'not yet' for up to 30 s and, if it stays, say which step and what was
on screen instead of a bare 404.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…rge left the second on its own line, run as a command)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
… asks for it

On iOS the YOU step keeps Next disabled until the federation-ID label is
valid (desktop mints the identity itself); the fixture never filled it, so
the walk stopped at 'you'. The bounded wait added before named it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…capability

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
… and the gate driver raises on success:false

iOS answered a failed text input with HTTP 200 {success:false}; the driver
checked only the status, so every field the session fixture typed on iOS
silently took nothing and the wizard's Next stayed disabled. Desktop
already answered 404. Android had the same 200. The driver now treats a
success:false body as a failure whatever the status (test red on the old
driver, green now).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…e typing or clicking it

On the iPhone the first-run wizard is taller than the screen: the account
fields sit below the age band and the AI choice, and /input refuses a
field the person could not see (CIRISClient#33). With #97's 404-on-failure
fix the refusal finally surfaced. Scroll down in steps, bounded, on an
off-screen refusal; everything else raises as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…reports what each scroll answered

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
emooreatx and others added 2 commits September 28, 2026 20:19
…tep, before calling a wizard field unreachable

A phone with the keyboard up has a small viewport; a fixed four steps down
turned around before reaching the device-name field.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…rd asks

The legs run against a bare node with no LLM; iOS asks the question and the
walk stopped at the AI step, whose Next waits for a usable LLM choice.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
emooreatx and others added 5 commits September 28, 2026 21:43
… of stopping on CircleTab

nav_map drops the row hop for a one-card tab because the wide layout opens
it directly; phones list it first, so iOS landed on CircleTab. Open the only
nav_epistemic_* row; with several, name them instead of guessing. Test added.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
iOS landed on a People list with the community roster and Contacts: the
compact layout was still on Neighbours after the circle hop. Pick the row
named for the target screen when the list has it; otherwise say which rows
were there and ask whether the circle hop applied. Test added.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
…ll marks proposed

CSD-005's trust chip shipped on main (people review), so the test that named
it went red for the product doing its job. Derive the tag instead (the same
fix the two-node branch made).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155SkTGdbwnSnR6tWBqJvUM
@emooreatx
emooreatx merged commit fdf0169 into main Sep 29, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant