Skip to content

Manta cache - #111

Draft
rjanalik wants to merge 16 commits into
mainfrom
manta-cache-stage-1-2
Draft

rjanalik wants to merge 16 commits into
mainfrom
manta-cache-stage-1-2

Conversation

@rjanalik

Copy link
Copy Markdown
Collaborator

No description provided.

@rjanalik rjanalik changed the title Manta cache stage 1 & 2 Manta cache Aug 10, 2026
@rjanalik
rjanalik force-pushed the manta-cache-stage-1-2 branch from 8c25a16 to 3fc1a23 Compare August 20, 2026 16:47
rjanalik and others added 15 commits August 20, 2026 20:31
Implement the ROADMAP Stage 1 core library as a standalone workspace
crate: SiteDescriptor, an Index with synchronous group->site and
xname->site lookups (plus groups/sites/group_members), and an async
refresh() that fans out two calls per site (GET /groups/available +
unfiltered GET /groups/nodes) via futures::try_join_all.

Built directly as a crate (collapsing Stage 2) since the cache talks to
manta-server purely over HTTP and has no compile-time dependency on it.
Wire structs are self-contained to avoid pulling manta-shared (and
transitively manta-backend-dispatcher) in for two fields. Errors use a
crate-local CacheError (thiserror); publish = false until a published
crate consumes it.

Includes 10 unit tests and an env-gated integration smoke test, and
updates the crate README/ROADMAP status banners to record the decisions.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- ci: include manta-cache in the rustfmt gate (check + auto-fix legs)
- refresh: add connect (10s) and request (300s) timeouts to the shared
  client so one hung site cannot stall the whole try_join_all fan-out;
  capture a bounded response-body snippet in CacheError::Status so
  refresh failures carry the server's diagnosis, not just a status code
- index: make group_members follow the owning site on cross-site
  collisions (never mix xnames from two sites) and document the actual
  precedence rules (listing outranks hsm-mention); test both
- site: redact the bearer token from SiteDescriptor's Debug output
- ROADMAP: correct the fixture's endpoint attribution (GET /groups, not
  /groups/available), record that it is not yet wired into tests, and
  add open questions from the review: single-call GET /groups refresh
  (verified access-scoped identically to /groups/available),
  partial-failure tolerance vs the Stage-4 blocking-startup lifecycle,
  and an in-proc population path (public snapshot constructor)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011jmoqPba44fnY3pTiQb46p
Promote SiteSnapshot, NodeMembership, and Index::from_snapshots
(né pub(crate) Index::build) to the public API. This gives embedding
processes an HTTP-free population path — the in-proc deployment shape
cannot refresh against its own not-yet-listening router at startup, so
an embedding manta-server needs to build the index from its own service
layer — and lets tests drive the index from captured data.

Wire the previously unused testdata/groups-prealps.json capture into a
new offline test (tests/fixture.rs) that folds the GET /groups payload
into a SiteSnapshot and asserts the ROADMAP's documented fixture
properties: umbrella group, empty groups, overlapping membership. The
refresh path itself still uses the two-call source; that open question
stands.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011jmoqPba44fnY3pTiQb46p
Replace the all-or-nothing try_join_all fan-out with join_all plus a
RefreshOutcome { index, failures } return: the index covers every site
that answered, and each failed site contributes its CacheError (also
logged at warn). Err is reserved for ClientBuild, the only failure that
precedes the fan-out.

This defuses the availability trap recorded in the ROADMAP: under the
Stage-4 "refresh before the listener starts" lifecycle, one unreachable
site would have cost the entire index and blocked manta-server startup.
The policy for an incomplete outcome (serve partial vs refuse) stays
with the caller via RefreshOutcome::is_complete, decided at Stage 4.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011jmoqPba44fnY3pTiQb46p
Add wiremock as a dev-dependency and tests/mock_server.rs, giving the
refresh() wire behaviour offline coverage it previously only had via
the env-gated live test: per-site X-Manta-Site and bearer headers
(two sites sharing one server URL, told apart by headers alone), the
two-call fan-out with call-count verification, tolerance of unknown
JSON fields, hsm splitting end-to-end, and the mapping of failures into
RefreshOutcome::failures (Status with body snippet, empty-body marker,
decode errors as Request) while healthy sites stay indexed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011jmoqPba44fnY3pTiQb46p
Deployment shape: standalone shared service (manta-cache-server) —
planned reuse from other projects (OpenCHAMI) rules out in-process
endpoints on manta-server; the sidecar remains a possible deployment
of the same binary but not the design target. Auth model:
service-account-style token per site, shared index, per-user
authorisation staying in the downstream manta-server handler.

Both open-questions rows are marked resolved, and the shape choice
surfaces two sub-decisions for the implementer: service-account token
sourcing (config/env/Vault + rotation) and securing the cache's own
endpoints (TLS + caller auth for a shared network service).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011jmoqPba44fnY3pTiQb46p
New workspace binary crate wrapping the manta-cache index in HTTP, per
the recorded deployment-shape decision (standalone shared service) and
auth model (service-account token per site):

- GET /api/v1/sites, /api/v1/lookup/group/{label}, and
  /api/v1/lookup/nodes?xnames=… (unanimous site or null + per-xname
  resolutions + unknowns, for the Stage-4 caller to police), plus an
  open GET /health for probes.
- cache-server.toml in the shared manta config dir
  (MANTA_CACHE_SERVER_CONFIG override); bootstrap mirrors manta-server:
  hand-rolled flags, fail-closed TLS unless allow_http, graceful
  SIGTERM drain, startup summary that never echoes tokens. Default
  ports 8444/8081 sit one above manta-server's so colocated pairs
  don't collide.
- Sub-decisions taken: per-site `token` or `token_file` (the latter
  re-read every refresh, so secrets rotate without restart), and an
  optional [server] api_token bearer guard on /api/v1/*.
- Refresh lifecycle: initial refresh before the listener binds
  (per-site failures tolerated and warned), optional
  refresh_interval_secs periodic loop that swaps the index wholesale.

12 unit tests cover config validation and every route incl. auth;
smoke-tested end-to-end with curl against a mock manta-server
(all endpoints, 401/404/400 paths, SIGTERM drain). manta-cache-server
joins the CI rustfmt gate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011jmoqPba44fnY3pTiQb46p
Split the refresh internals: fetch_snapshots() returns the raw
per-site SiteSnapshots + failures (SnapshotOutcome), and refresh()
becomes a thin wrapper folding them into an Index. A caller that keeps
its own snapshot store — the cache server's upcoming per-site
management refresh, which re-fetches one site and rebuilds via
Index::from_snapshots with the other sites' stored snapshots — needs
the snapshots that refresh() previously discarded.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011jmoqPba44fnY3pTiQb46p
POST /api/v1/refresh (full re-sync) and POST /api/v1/refresh/{site}
(single site). Both require a configured [server] api_token and answer
403 when none is set: a refresh triggers a cross-site HTTP fan-out and
must not be an open amplification lever, unlike the read-only lookups
which stay open-by-default inside the deployment perimeter.

The state now keeps the last-good SiteSnapshot per site, so a
single-site refresh re-fetches one site and rebuilds the index from
stored siblings (previous state, including that site's last-good
snapshot, keeps serving on failure). Full refresh and the
startup/periodic path share the wholesale-replace semantics.

ROADMAP Stage 4 rewritten to the post-Stage-3 reality: the in-proc
lifecycle text is superseded by the standalone service's own lifecycle;
the integration is decided as CLI-side pre-resolution (per-site
Keycloak tokens make server-side resolution self-defeating — the
request would carry the wrong site's credential); the security stance
is recorded (lookups = non-sensitive read-only routing metadata behind
optional token + network placement, TLS valued for answer integrity);
site CRUD is descoped to an open question (declarative/read-only config
deployments sit badly with a service rewriting its own TOML).

17 route/config tests incl. a wiremock success path; smoke-tested live
with curl (401/403 gating, full + per-site refresh, 404, lookups).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011jmoqPba44fnY3pTiQb46p
Stage-4 CLI-side pre-resolution: when a command targets an HSM group
or a plain xname list and neither --site nor cli.toml names a site,
the CLI asks the manta-cache-server configured by the new cli.toml
keys `cache_url` + `cache_api_token`, then proceeds exactly as if
--site <resolved> had been passed — same per-site token cache, same
X-Manta-Site header, no manta-server changes.

The clap tree has no uniform target-arg id, so extraction is a closed
per-command table (power group/nodes, get nodes/group-nodes/sessions,
apply boot group/nodes, apply boot-parameters, apply
kernel-parameters, console node); hostlist bracket expressions and
NIDs are skipped (the cache indexes xnames only). Failure semantics:
transport failures warn and degrade to the lazy "No site selected"
error — the cache is an accelerator, not a dependency — while
definitive answers (unknown group, split xname list, rejected
cache_api_token) abort with a specific message, since silently
guessing a site could aim a destructive command at the wrong cluster.
A successful resolution prints a one-line stderr notice.

Covered by unit tests (target table incl. bracket/NID guards, reply
interpretation) and three assert_cmd end-to-end tests driving the real
binary against a mock cache (resolve, degrade, split); smoke-tested
through the full real chain (manta CLI → manta-cache-server → mock
manta-server). cli.toml keys documented in the root README; ROADMAP
Stage 4 records the delivery and its scope. No clap surface change,
so man pages / autocompletions are untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011jmoqPba44fnY3pTiQb46p
Three-terminal walkthrough for the full site-resolution chain (manta
CLI -> manta-cache-server -> manta-server -> CSM) against one real
site: config snippets for cache-server.toml + the cli.toml additions,
startup ordering, and numbered test scenarios matching the ROADMAP's
table (resolution, unknown group, explicit-site bypass, degradation,
management refresh), plus the operational gotchas (token expiry,
cache-before-server startup). The token_file trick points the cache at
the CLI's own cached JWT so re-auth propagates automatically — noted
as test-only convenience, not the production setup.

Verified 2026-07-19 end-to-end against prealps.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011jmoqPba44fnY3pTiQb46p
One pointer sentence in the configuration section so the root docs are
not silent about the optional manta-cache-server component. The full
README/ARCHITECTURE sections are deliberately deferred until the
maintainer review settles the remaining open questions (conflict
policy, site CRUD, refresh source) — the crate-local docs under
crates/manta-cache/ are the current source of truth.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011jmoqPba44fnY3pTiQb46p
GET /api/v1/dump renders everything the cache holds as JSON, for humans
and jq. It exists to answer "what is actually cached, and is it stale" —
it is not consumed by manta-server or the manta CLI, and nothing in the
lookup path depends on it. No filters and no pagination: it is
hand-called by an operator, and an Alps-scale payload is a few MB.

groups + xnames are a bulk mirror of the lookup endpoints — same owners,
same member lists, no second resolution path. That is the design rule
the endpoint is built around: the dump must be a trustworthy oracle for
"what would the CLI have resolved", answerable without running a
command. Both read the derived Index, exactly as lookup_group and
lookup_nodes do.

sites is keyed by *configured* site rather than indexed site: a site
whose refresh failed is absent from the index, and that absence is
precisely what a reader is trying to explain. Each entry carries
manta_server_url, token_source (inline vs the file path, never the
token), in_index, refreshed_at + age_seconds, and last_error.

Gated behind the management middleware, for a different reason than the
refreshes: this is the only endpoint that serves group member lists,
which the recorded security stance excludes from the open-by-default
set. Operator-only by construction also keeps it clear of any future
per-user filtering question — no end user ever sees it.

conflicts is the one part that is deliberately not a mirror. Index
resolves collisions at build time and keeps only the winner, so a group
whose members were discarded when another site claimed the label is
indistinguishable from a genuinely small group. The server holds the
snapshots anyway, so they are re-scanned per request to report contested
labels and xnames with:

- claimed_by  — every site that listed the label or has a node in it;
- listed_by   — the subset that listed it in /groups/available. Separate
                because the tie-break runs over listings, not
                contributions: stating the rule over claimed_by would
                name the wrong site whenever a third site merely mentions
                the label through a node's hsm field;
- owner + reason — the winner and which documented rule produced it;
- discarded_members — the xnames missing from the served member list.
                Listed, not counted: the dump is otherwise index-only, so
                these xnames appear nowhere else in the payload, and a
                count would report that a node had gone missing without
                saying which one — leaving the reader to go ask the
                losing site's manta-server, the round trip this endpoint
                exists to save.

This is observability, not policy — the conflict-policy question stays
open — but nodes_free is a conventional pool name, so collisions are
realistic rather than hypothetical, and the dump is what makes them
visible.

Supporting changes:

- CacheState gains per-site status, written by both refresh paths. The
  semantics differ by path and the dump reads differently after each: a
  per-site fetch failure goes through refresh_all's wholesale replace, so
  the site loses its snapshot and its timestamp; an unreadable token_file
  returns before anything is replaced, so every site keeps its data and
  its timestamp with the error recorded alongside — as does a failed
  single-site refresh, since the previous snapshot is still what serves.
  A site carrying both a last_error and an intact refreshed_at is
  therefore expected.
- record_failure takes the refreshed_at observed before the fetch as a
  guard, so a failing refresh cannot stamp a stale error onto a site a
  concurrent refresh_all just refreshed successfully. The bias — drop a
  possibly-stale error rather than risk a false alarm — is documented,
  including its cost.
- refresh_all records every unreadable token_file against its own site
  and joins the messages in site-name order, rather than blaming an
  arbitrary one of them (config.sites is a HashMap) and a different one
  each tick.
- manta-cache gains Index::xnames(), since xname_to_site could not be
  enumerated at all, and CacheError::site(), which the type's docs
  already promised but never exposed.
- chrono with the clock feature, which this crate's dependency graph does
  not otherwise enable.

Tested: seven route tests covering the management gate, lookup parity,
site freshness and secret-absence, cross-site collisions, and the
listing-outranks-mention rule in both directions (a listing site that
sorts first and wins anyway, and one that sorts last and takes the label
off an earlier mention-owner); three refresh unit tests pinning the race
guard, including the never-refreshed site an operator retries by hand.
Smoke-tested against a mock manta-server with three sites — two sharing
a nodes_free label and an xname, one unreachable — plus a live run with
two vanishing token files to confirm each site is blamed by name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UVBjobqRG7iwrSfRf4DAwN
The cache was developed and demoed against a personal credential:
`token_file` pointed at the CLI's own token cache, which made the live
test a one-liner but is not a deployment. A shared service must not
refresh as a named human — the index silently inherits that person's
group roles, dies when their token expires, and attributes every
backend call to them.

Move the refresh onto one Keycloak service account per site. Because
the two credential kinds are byte-for-byte indistinguishable to the
code that uses them, make the difference observable rather than
enforced: pointing `token_file` at a personal token still works (which
keeps the LIVE-TEST demo usable), but the server now says so loudly.

**Credential storage.** `token_file` becomes the only per-site form;
inline `token` is removed. A real Keycloak credential does not belong
in a config file that gets templated, backed up, or committed, and a
path is the one form every deployment target can produce — a k8s
projected Secret, a systemd `LoadCredential=`, or a bind mount. Adds
`[server] api_token_file` for the cache's own inbound bearer. The
schema is `deny_unknown_fields` throughout, so a half-migrated site
that leaves `token` beside `token_file` fails to load rather than
silently parking a live credential in the file, and parse errors
naming either key carry the migration. Secret files readable beyond
their owner draw a startup warning — never a refusal, since k8s
projected secrets default to 0644.

**Credential advisory.** A new `credential` module decodes the token's
own claims locally (no signature verification — diagnostics only) for
principal, kind, expiry and `pa_admin`. Service accounts are
identified by Keycloak's own convention, `preferred_username ==
"service-account-" + azp`, so the check is an equality that cannot
misfire on a human. Surfaced in the startup summary and as a
`credential` block on `GET /dump`. Two details in there are
deliberate and non-obvious:

- `expires_in_days` floors rather than truncating. `num_days` divides
  toward zero, reporting a credential six hours past `exp` as `0` —
  indistinguishable from one lapsing six hours hence, and enough to
  make the obvious `expires_in_days < 0` alert miss the first day of
  an outage. Warnings classify on the timestamp itself, so a lapsed
  credential can never be described as "expires in 0 day(s)".
- The log dedup compares against the warnings **previously logged**,
  kept on `SiteStatus`, not against `previous.warnings(now)` —
  re-deriving both sides at one instant makes them equal by
  construction for an unchanged file, which would suppress every
  message after the first and never report the lapse. The store runs
  immediately after the logging loop so no exit can get between them,
  and `credential_state_changed` takes no clock, so neither hazard is
  expressible.

**Empty snapshots.** A site returning zero group labels is now a
per-site failure rather than a success. Such a site used to show
`in_index: true` with a fresh timestamp and no error while resolving
nothing. Partial visibility is expected (it is just the account's role
set), but empty means the credential resolved to no HSM roles —
expired, wrong account, or roles removed. `refresh_all` drops the
site; `refresh_site` keeps the previous snapshot serving, per its
existing contract.

Design, verified claim table for a real service-account token, and the
coverage trade-off between `pa_admin` and scoped roles are recorded in
ROADMAP Stage 5; production guidance is in the cache README.

BREAKING CHANGE: `[sites.<name>] token` is no longer accepted in
cache-server.toml, and an unknown key anywhere in the file is now an
error rather than being ignored. Write the credential to a file and
point `token_file` at it instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GthHz4MejULhVuzWSi7VBo
manta-server's routes moved to /v2 (the /api/v2 -> /v2 rename), and the
old prefix is a hard 404. The cache's outbound base URL still appended
/api/v1, so every refresh against a current manta-server would 404.

normalize_base now emits /v2, matching the CLI's MantaClient. Both sets
of test mocks move in lockstep, since each stands in for manta-server
and would otherwise keep agreeing with the client about the wrong path:
manta-cache/tests/mock_server.rs, and the wiremock matchers in
manta-cache-server/src/routes.rs's test module. Doc comments in lib.rs,
site.rs, tests/fixture.rs, and Cargo.toml follow.

manta-cache-server's own API surface -- /api/v1/{sites,lookup,refresh,
dump} -- is a separate namespace, unaffected by manta-server's rename,
and is deliberately left alone; the CLI resolves the cache by appending
/api/v1 to cache_url. Only the outbound /api/v1/groups/* mocks inside
that crate's tests changed. The two directions share a spelling, so
classify by call direction before touching either.

Not caught by CI as written: the mocks agree with the client either way,
and the only test that would notice, refresh_against_live_server, is
env-gated and skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01La5XHfUVQ6okz8nSj7yAi7
@rjanalik
rjanalik force-pushed the manta-cache-stage-1-2 branch from 3fc1a23 to 7f77058 Compare August 20, 2026 18:31
b8755ac removed the server-side socks5_proxy field and corrected the
"two schemas are disjoint" paragraph, but the near-identical sentence
20 lines below kept listing "per-site SOCKS proxies" among the details
that live in server.toml. The section now contradicted itself: one
paragraph says the server has no per-site SOCKS5 proxy knob, the next
says server.toml holds per-site SOCKS proxies.

Fixed here rather than on main because this branch already edits that
same section -- the manta-cache pointer paragraph sits directly above
the corrected sentence -- so the contradiction lands in the middle of
this branch's own change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01La5XHfUVQ6okz8nSj7yAi7
Masber pushed a commit that referenced this pull request Aug 26, 2026
* docs: sync CLAUDE/ARCHITECTURE/README/GUIDE/SECURITY with current code

The reference docs had drifted from the code in ways that cost time
rather than merely being untidy -- several described layers that were
deliberately removed, so they sent a reader off to rebuild something
that had been torn out on purpose.

CLAUDE.md:
- The "Error codes" section documented the MANTA_* catalog, ERRORS.md,
  and the ErrorResponse::new chokepoint as current fact. None of it is
  on main -- it lives on feature/custom-error-codes (PR #110, draft).
  Moved to a new "Work in flight" table, together with manta-cache
  (PR #111), each row carrying a one-command check that proves it is
  not merged yet.
- Sibling pins said mbd beta.13 / csm-rs beta.19 / ochami-rs beta.13;
  the manifest says beta.15 / beta.23 / beta.17. Rather than re-pin
  numbers that go stale, point at [workspace.dependencies] as the
  source of truth and say so explicitly.
- "doctests -- NOT run by plain cargo test" is false. `cargo test
  -p manta-shared` prints "Doc-tests manta_shared" and runs 8 of them.
- Added the toolchain caveat: rust-toolchain.toml's 1.92.0 pin only
  applies when cargo resolves through rustup's shim, so a Homebrew
  cargo silently yields a newer clippy than CI.
- Added that openapi.json is a build input, not documentation.

ARCHITECTURE.md -- seven factual corrections:
- manta-shared no longer holds the backend dispatcher (this
  contradicted line 29 of the same file).
- The CLI has no handlers/ directory; dispatch/process.rs routes.
- http_client/ has no query.rs, no QueryBuilder, and no per-resource
  sub-modules -- endpoints are reached via the progenitor client.
- infra_backend/ is not a wrapper layer. The per-domain wrappers were
  deliberately deleted; service code calls infra.backend.<method>(...)
  directly. The old text invited reintroducing the removed layer.
- AppContext is 13 fields, not 10 (read_only, token, session missing).
- Entry point is dispatch::process::process_cli.
- service/ and handlers/ module lists were missing 5 and 3 modules.

...and four things that were absent entirely: a "Generated code"
section (with the regeneration step added to the add-a-command
recipe), the CLI read-only gate, jwt_ops moving to manta-shared, and
the fact that JWT signatures are never verified locally -- which
imposes a real rule on future code and was only in a rustdoc comment.

README.md: applied the same SOCKS fix already made on
manta-cache-stage-1-2 (1265d19) -- one paragraph said the server has
no per-site SOCKS knob and the next said server.toml holds per-site
SOCKS proxies. Also added the cli.toml keys the example omitted
(hsm_group, read_only, poll knobs) so it matches examples/cli.toml.

GUIDE.md: the `manta config show` sample printed a "SOCKS5 proxy" row.
push_local_section emits exactly four rows and never that one --
fabricated output in a ```text block reads as literally observed.

SECURITY.md: the no-local-signature-verification note pointed at
server::common::jwt_ops "documented inline at the top of that module",
but that module is now an 8-line re-export. Repointed at the canonical
manta_shared::common::jwt_ops where the caveat actually lives.

Cargo.toml (comment only): dropped the claim that building manta
requires the sibling checkouts because image_pipeline.rs needs
unreleased SatTrait additions. Verified false on all counts -- no
.cargo/config.toml, no adjacent checkouts, image_pipeline.rs
references neither SatTrait nor manta_backend_dispatcher, and
`cargo check --workspace` is clean from crates.io.

Verified: cargo test --workspace (684 passed), fmt, rustdoc with the
CI RUSTDOCFLAGS, anyhow boundary, openapi.json sync, and man/shell
completion sync all clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ldf1GrsA11XsLDJ188arZ7

* docs(cli,shared): name the real cli.toml key in group-default docstrings

Three docstrings for the "operator default group" field named a
cli.toml key that does not exist. main.rs reads
`settings.get_string("hsm_group")`, so an operator following any of
these and writing the documented key into cli.toml gets no error and
no warning -- get_string(...).ok() swallows the miss and the default
is silently absent. A wrong key name fails more quietly than a wrong
type does.

  app_context.rs:59         `parent_group`       -> `hsm_group`
  kernel_parameters.rs:117  `parent_group_group` -> `hsm_group`
  group.rs:70               `group_name`         -> `hsm_group`

The correct value is not a guess: six sibling *Params structs
document the same field and already say `hsm_group`
(configuration.rs, boot_parameters.rs, hardware.rs, cluster.rs,
template.rs). All three of these trace back to the parent_hsm_group
removal in MIGRATING.md 5.7, where the key was deleted and these
docstrings were half-renamed to plausible-looking names instead of to
hsm_group. The group.rs one was the most misleading: it named
`group_name`, which is the sibling field on the same struct, so it
read as if the config key and the wire field were the same thing.

Two of the three are in manta-shared, which is published, so these
ship to docs.rs and to the GitHub Pages site that docs.yml builds
from target/doc.

No OpenAPI regeneration needed: both edited structs derive only Debug,
not ToSchema, and neither string appears in openapi.json. Confirmed by
regenerating the spec and diffing -- byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ldf1GrsA11XsLDJ188arZ7

* ci: correct the doctest comment on the doctest step

The comment claimed "Plain `cargo test --workspace` does NOT run
these." It does. `cargo test -p manta-shared` prints "Doc-tests
manta_shared" and runs 8 of them, and a full `cargo test --workspace`
prints Doc-tests sections for both manta_shared and manta_server.

The claim is load-bearing beyond this file: it was copied into
CLAUDE.md, where it told future readers they needed an extra command
they already had.

The step itself is unchanged and still worth keeping -- it reports a
doctest failure as a doctest failure rather than burying it in the
combined run, and it stays correct if the test step ever gains
--all-targets, which does suppress doctests. The comment now says
that instead of asserting something false.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ldf1GrsA11XsLDJ188arZ7

* docs: purge the removed POST /sat-file/validate and GET /summary endpoints

Found by diffing every `METHOD /path` mentioned in API.md against the
paths in the CI-gated openapi.json. Two endpoints were documented as
live across five files while returning 404.

POST /sat-file/validate was removed in f9160d2 ("feat(sat)!: remove
POST /sat-file/validate endpoint chain") along with the backend's
SatTrait::validate_sat_file. That commit touched zero .md files and
deferred the follow-ups, so the endpoint's ghost survived in:

  API.md         full endpoint section, apply-flow bullet, sequence
                 diagram round-trip, 2 server-config requirement rows
  CLI.md         "Pre-flight server-side validation" callout
  GUIDE.md       2 prose paragraphs + a mermaid flowchart node
  MIGRATING.md   section 5.10, which announces the feature as new
  ARCHITECTURE.md  the apply-flow description

Nothing replaced it, so the docs now say so plainly: validation is
client-side only (dangling image_ref, cycles), nothing checks the file
against live backend state, and an apply can therefore fail partway
through with earlier elements already applied. That last consequence
was not stated anywhere and is the part an operator needs.

MIGRATING.md 5.10 is rewritten rather than deleted -- section 5 is a
chronological within-v2 delta list, and operators who upgraded through
the intervening betas did see the pre-flight behaviour.

GET /summary was renamed twice (summary -> cache -> analysis;
fd5e801, a182763). BackendSummary still exists in
types/api/analysis.rs, so the endpoint lives on as GET
/analysis/images -- which API.md already documents in its own
section. The stale "## Summary" section was a duplicate of it under
the dead path.

Also fixes two pre-existing broken anchor links found while checking
that the deleted sections left no dangling references:

  GUIDE.md     -> MIGRATING.md 5.11 (slug was missing two hyphens
                  that the backticked `--dry-run` contributes)
  MIGRATING.md -> #2-server-setup, a heading that does not exist
                  (it is "2. Site operators (deploying the stack)")

Every internal anchor link across the 13 tracked docs now resolves.

Verified: cargo test --workspace (684 passed), fmt, openapi.json still
byte-identical to the emitted spec.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ldf1GrsA11XsLDJ188arZ7

* docs: state the real scope of the request-timeout layer

ARCHITECTURE.md said the TimeoutLayer is "applied to every API route".
It is not: build_router applies it to the `api` sub-router only, and
`/v2/auth/*` is a separate sub-router carrying the rate limiter and body
redaction with no timeout. The file already contradicted itself -- the
middleware diagram twelve lines above shows the split correctly.

Consequence worth naming: a hung upstream Keycloak on POST /v2/auth/token
is not bounded by [server].request_timeout_secs.

This is the documentation half of REVIEW_FINDINGS.md BUG-1. The code gap
is deliberately untouched -- that finding stays open. Making the prose
honest makes the gap more visible, not less.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ldf1GrsA11XsLDJ188arZ7

* docs: finish killing the "manta-shared holds the backend dispatcher" myth

The sync commit corrected ARCHITECTURE.md's crate tree but left the
identical wrong line in README.md and in manta-server's own Cargo.toml
comment -- the two places a reader most likely lands first. Both now
say what manta-shared actually carries; StaticBackendDispatcher and the
csm-rs/ochami-rs impls are server-only.

Two more accuracy fixes in the same vein:

ARCHITECTURE.md's "adding a command" recipe, step 5, sent the reader to
backend_dispatcher/mod.rs to add a dispatch arm. mod.rs holds only the
shared imports, the dispatch! macro and the `mod` lines; every
`impl <Trait> for StaticBackendDispatcher` block lives in that trait's
sibling file. The same document says so correctly two sections earlier.
Following the recipe meant editing the wrong file.

CLAUDE.md claimed the CLI "no longer links the backend traits at all".
No CLI code calls one, which was the point being made -- but
manta-backend-dispatcher is still compiled for `cargo build -p
manta-cli`, because manta-shared depends on it for the type re-exports
in types::dto. `cargo tree -p manta-cli` shows the edge at depth 2. The
distinction matters when a bad sibling beta breaks the CLI build and
the docs say that cannot happen. csm-rs and ochami-rs really are absent
from that tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TwtnVriwmTHHRZD8eYi2i6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant