Skip to content

feat: add Appendix F, provider conformance (TCK) - #439

Open
aepfli wants to merge 1 commit into
feat/provider-tck-appendixfrom
feat/provider-tck-appendix-doc
Open

aepfli wants to merge 1 commit into
feat/provider-tck-appendixfrom
feat/provider-tck-appendix-doc

Conversation

@aepfli

@aepfli aepfli commented Oct 1, 2026

Copy link
Copy Markdown
Member

Part of #417. Stacked on #423, which adds the assets this describes. Review that one first — it is the executable half and can land independently.

These were one pull request until now. The split lets the assets land while this prose is still being reviewed, because four language implementations are pinned to the branch and every rebase invalidates them.

What

Appendix F: Provider Conformance (TCK) — the normative description of the conformance assets in #423. What a TCK implementation must do, what each capability asserts, the control-API contract, the extension point, and the run-integrity checks.

Marked experimental and non-normative.

How the requirements are expressed

As normative blockquotes (58 of them), not numbered requirements — following Appendix A, since specification.json is built from the numbered sections and no appendix contributes to it.

Whether any of this should become numbered normative sections is a TSC decision, recorded as an open question rather than assumed. That is the main thing I would like a view on.

Design decisions worth reviewing

Each is argued in the appendix; this is the index.

  • Capabilities. Optional behaviour is tagged; a provider declares what it supports and the rest are skipped with a reason, never passed. Tags with no scenarios yet are reserved and must not be declarable.
  • @lifecycle and @events are separate, because a provider can initialise against its backend without emitting events, and vice versa.
  • @numeric-coercion is borrowed, not specified. The rule comes from a flagd ADR, not from the specification. An earlier draft claimed the opposite and called withholding the tag "an admission of a known bug" — wrong, and corrected. Gap: #430.
  • @string-typing and @fully-typed-values are two tags because one hid a bug. Three providers over one Flagsmith backend: two report TYPE_MISMATCH for a boolean and an integer asked as strings, one stringifies them, and all three stringify the float and the structure. The latter is the backend's shape, the former a provider defect; one tag reported the defect as a permitted absence. Gap: #433.
  • No container restarts. Outages are produced through the control API, because orchestrators reassign host ports on restart and every provider pointed at the old one then fails looking flaky.
  • In-process control is a narrow carve-out for providers with no backend. A provider with a backend must use the control API, or the suite passes while proving nothing.
  • An extension point is required, not optional — raised in review on the original PR. An extension may add questions but never replace one: it cannot occupy a canonical path, cannot satisfy a canonical scenario, and a run that skipped part of the canonical set must fail. That last rule was found by accident — a selector matching one scenario name produced a green suite and a well-formed report covering one scenario.
  • Run-integrity checks. Three failures are invisible in results and must fail the run; the one that matters is a scenario carrying a tag the implementation does not know, which all four reference implementations ignored. An unknown tag gates nothing, so its scenarios stay mandatory and a suite that has not learned a capability silently keeps demanding the old behaviour. Two further rules are about the checks themselves: a check that cannot be performed must fail rather than skip, and the revision check must be in force where the scenarios execute. Both came from one adoption producing entirely plausible numbers against the previous revision's scenarios.

One home for the vocabulary

#423 carries the capability table in assets/provider-tck/README.md so the tags are defined while this is in review. This PR moves it here and leaves a pointer behind — the appendix owns it, and two copies would be able to disagree.

Open questions

  • Evaluation context passthrough — the targeting key is covered, no other attribute is; needs an echo operation on the control API.
  • Accessor width — @large-integers covers up to 2^53 − 1; 32- versus 64-bit accessors are not modelled. Part of #430.
  • Finer-grained flag manipulation would need new control-API endpoints.
  • Flag metadata has one scenario and is blocked on a testbed fixture, not on the suite.
  • Hooks and caching are not covered; @caching is reserved.

Answers #417's Q1 (directory layout, forced rather than chosen) and Q4 (numeric coercion stays a capability). Q2, Q3, Q5 and Q6 stay open; Q7 is split into #424.

The normative description of the assets added alongside: what a TCK
implementation must do, what each capability asserts and when withholding one is
a sanctioned choice rather than a defect, the control-API contract, the
extension point, and the run-integrity checks.

Stated as normative blockquotes rather than numbered requirements, following
Appendix A -- specification.json is built from the numbered sections and no
appendix contributes to it. Whether any of this should become numbered is a TSC
decision, recorded as an open question.

The assets README hands the capability vocabulary back here, so it has one home
rather than two that can disagree.

Experimental and non-normative. Tracked in #417.

Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
@aepfli
aepfli requested a review from a team as a code owner October 1, 2026 07:54
@coderabbitai

coderabbitai Bot commented Oct 1, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: d3774265-e706-472a-a5cf-2e2c2231857d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant