Conversation
…providers
A conformance suite any Python provider can adopt to verify it implements the
provider contract of the specification, and the Python implementation of the
cross-language suite defined in Appendix F. It runs the same Gherkin, the same
canonical flag set and the same control API as the Go and Java implementations.
It uses pytest-bdd, the runner the flagd provider and the flagd testkit already
use, so an adopting package gains no new test framework.
Adoption is one fixture and one call. The step definitions ship as a pytest
plugin registered through a pytest11 entry point, so there is no conftest.py to
write and nothing to import for the vocabulary - pytest-bdd resolves steps
through the fixture system, and fixtures from an installed plugin are visible
everywhere. The feature files and flag set are packaged with the distribution,
so adopting needs no git submodule.
Capability gating uses pytest.skip from an autouse fixture, so a scenario whose
capability was not declared is reported as skipped with the reason attached
rather than silently passing. The gate keys off the node's markers rather than
its requested fixtures: pytest-bdd resolves a step's fixtures lazily, so
tck_config is not in request.fixturenames at setup time, and guarding on that
silently disabled the gate.
Two self-test suites, plus unit tests for what the Gherkin cannot assert about
itself: the SDK's InMemoryProvider, and the TCK's own updatable one. The second
exists because the first cannot exercise the configuration-change path at all.
Findings, both confirmed by running the suite:
* A boolean satisfies an Integer request. The client type-checks with
isinstance(value, int) and bool subclasses int in Python, so boolean-flag
requested as an Integer returns True with reason STATIC and no error code.
This is Python-specific - the identical scenario passes in every other
language - which is a fair argument for having more than one
implementation. Tracked as open-feature/python-sdk#619, and marked
xfail(strict=True) so it stays visible and un-hides itself once fixed.
* InMemoryProvider cannot update its flag set, which Appendix A requires of
an SDK in-memory provider. Only half the machinery is missing, since
AbstractProvider already supplies emit_provider_configuration_changed, so
ControllableInMemoryProvider is a small subclass rather than a
reimplementation and should port back as a method. Tracked as
open-feature/python-sdk#620.
Verified locally: 56 passed, 7 skipped, 2 xfailed; ruff and mypy --strict clean.
Part of open-feature/spec#417
Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
This was referenced Aug 24, 2026
Two things CI caught that local verification did not. `ruff format` is a separate pre-commit hook from `ruff check`, and only the latter was run locally. Nine files needed reformatting; the changes are cosmetic line-wrapping only. More importantly, the package was not being tested in CI at all. The build matrix is gated on dorny/paths-filter and its filter list had no entry for tools/openfeature-provider-tck, so no change under that path expanded the matrix and the suite never ran. The locally reported 56 passed / 7 skipped / 2 xfailed was local-only. Adding the filter block, mirroring the one for tools/openfeature-flagd-core, turns it on. Verified after formatting: 56 passed, 7 skipped, 2 xfailed; ruff check and mypy --strict still clean. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
`uv sync --frozen` in the build workflow validates the lockfile against the manifests, and the previous commit added openfeature-provider-tck to the workspace root's dependencies and [tool.uv.sources] without regenerating the lock. That breaks the build job for *every* package, not just this one. It was latent until now only because the paths-filter had no entry for this package, so no build job ran at all. Enabling the filter in the previous commit would have surfaced it as a red build. The regeneration also picks up openfeature-provider-flagd 0.5.1 -> 0.5.2, which the lock had missed when that release landed. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #409 +/- ##
==========================================
- Coverage 95.64% 89.12% -6.53%
==========================================
Files 24 44 +20
Lines 1057 2197 +1140
==========================================
+ Hits 1011 1958 +947
- Misses 46 239 +193 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
TckConfig.ready_timeout was documented but never read by anything, so a provider that hung while connecting would hang the whole pytest session with no useful message, and the documented knob did nothing. api.set_provider initialises synchronously and has no timeout of its own, so the bound comes from running it on a worker thread and giving up on the result. The worker is deliberately not cancelled -- Python cannot interrupt a thread blocked in a socket call -- and is left to finish or die with the process, which is acceptable because a timeout already means the scenario is failing. A config field that claims to do something it does not is exactly the kind of quiet untruth this suite exists to catch, so it is fixed rather than removed. Verified: 56 passed, 7 skipped, 2 xfailed; ruff and mypy --strict clean. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
…he tests
TckConfig.capabilities was annotated frozenset[Capability], but the README
tells adopters to write `capabilities={Capability.EVENTS, ...}` -- a set
literal. Anyone copying the documented example and running mypy got an
incompatible-argument error from the suite's own documentation. It is now
annotated Collection[Capability], which is what __post_init__ already accepted:
a set, a list or a generator all normalise to a frozenset on construction.
The reason this was invisible is the second half of the fix. mypy was
configured `files = "src"`, so the tests were never checked -- and the tests
are the reference adoption, the thing an adopting provider copies. They are now
in scope, which is what would have caught the annotation in the first place.
Verified: mypy clean over src and tests (17 files), ruff format and check
clean, 56 passed / 7 skipped / 2 xfailed.
Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
aepfli
force-pushed
the
feat/provider-tck
branch
from
August 24, 2026 12:13
a8e003a to
15a57bb
Compare
…dule The feature files, the canonical flag set and the control-API document are owned by open-feature/spec, not by this repository. Committing copies of them here forks the definition of conformance -- the one thing this suite exists to prevent -- and leaves no machine-checkable record of which spec revision the copies came from. Replace them with a git submodule at tools/openfeature-provider-tck/spec, pinned at dfa16586 (spec#423), plus a build-time copy. The copies are gitignored and carry a DO-NOT-EDIT marker, so the pin is now the only record of the revision and the two cannot drift apart unnoticed. An adopter installing this package still needs no submodule: the copies are force-included into the wheel and the sdist, and the sdist excludes the submodule itself so it carries the four assets rather than the whole spec repository. Only a contributor to this package needs the submodule, and `poe test` syncs it first. This mirrors what openfeature-flagd-api-testkit already does for the flagd test harness. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
lifecycle.feature was gated by @events, which was wrong in both directions. An SDK dispatches PROVIDER_READY around initialize for any provider (openfeature/provider/_registry.py), so a provider that declares @events passes the readiness scenario without demonstrating anything -- a NoOpProvider passes it identically. The gate made the scenario vacuous for exactly the providers it admitted. Conversely a stateless provider such as OFREP has a real initialisation to verify but no event stream of its own to declare @events for, so the gate shut it out of a scenario it should be held to. The spec revision pinned by the submodule retags the feature to @lifecycle and adds the capability to Appendix F. Add the matching enum member; plugin.py registers the marker by iterating the enum, so nothing else changes. Neither in-memory self-test declares it. They have no backend to reach, so their readiness scenario was passing vacuously too, and a skip with a reason is the honest outcome. 54 passed, 9 skipped, 2 xfailed. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
This was referenced Aug 30, 2026
…rcion The capability vocabulary and the spec submodule pin belong to this PR, so the rename does too. It was written on the report branch, which is a sibling of the flagd and OFREP adoptions rather than an ancestor -- so the adoptions could not see it, and renaming their references there would have broken them against this base. Moving it down is what lets every branch above share one vocabulary. `Capability.STRICT_NUMERIC_TYPING` becomes `Capability.NUMERIC_COERCION`, marker `numeric-coercion`, and the submodule moves to dc4d7ae8 so the executed feature files carry the renamed tag. That bump also brings two unrelated spec changes: the lifecycle readiness scenario is renamed, and control-api.yaml gains the requirement that POST /start not return until the seeded state is served. The framing is corrected at the same time, because it was wrong rather than merely stale. Both the README and the capability's own docstring asserted that "the specification requires TYPE_MISMATCH when the requested type cannot be satisfied" and concluded that not declaring the capability was "an admission of a known bug". OpenFeature defines one numeric type deliberately -- `number` is "of unspecified type or size", and differentiating integers from floats is an optional language idiom -- so no requirement governs this, and the second claim followed from the first. The rule tested here is borrowed from flagd's numeric coercion ADR, which is scoped to flagd's own implementations; a provider behaving differently is not violating the specification. The gap in the provider contract is open-feature/spec#430, and flagd's own instance is open-feature/flagd#1996. That also makes the capability genuinely optional rather than a concession to a defect, which is the opposite of what the old text said. The report branch's own files stay with it: test_report.py does not exist here, and neither do the `not_applicable` and `known_deviations` configuration fields the rename also touched. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
A TckConfig is two things. It is the configuration a run needs, and it is the set
of claims an adopter makes about their provider -- which is what turns a skipped
scenario from a hole in the run into a recorded answer. The second role was
incomplete, and the gaps were all the same shape: something an adopter has to be
able to say that the vocabulary gave them no way to say.
`not_applicable={Capability.X: "why"}` is a capability that *cannot* hold rather
than one the adopter chose not to declare. The suite treats the two identically,
because the scenarios are skipped either way, but collapsing them misrepresents
whole languages: @numeric-coercion is unsatisfiable in JavaScript, which has no
integer type, and recording that as a choice shows every JavaScript provider as
declining something none of them can have.
`known_deviations` acknowledges a gap against something the specification does
not treat as optional, with somewhere it is tracked. An acknowledgement and not
an excuse: the scenario still fails and the suite still fails with it. What it
adds is that the gap was known rather than a surprise.
And a capability is now either declarable or reserved. @targeting and @caching
gate no scenario, so declaring one cannot be verified, cannot produce a skip, and
says only that something was claimed and nothing examined -- so declaring one is
refused at construction, where the adopter's own code is still on the stack to
say which line to fix. The default is DECLARABLE_CAPABILITIES rather than the
whole enum, because "declare everything, then narrow it" is the advice and
therefore the one place a reserved tag gets declared by accident: that is how one
implementation's published report came to assert both of them.
`capability_for_tag` and `control_api` are here for a consumer that is not. A
reporter outside this package has to tell a capability-gating tag from a merely
organisational one, and has to be able to ask a control how it drove the backend.
Nothing in this commit calls either; that is the point. Two branches sit on this
one, and neither should be able to change what the other compiles against.
Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
A provider is rarely only a provider. flagd has `fractional` targeting, another
vendor has a proprietary rollout rule, and pinning those used to mean a second
harness beside the conformance suite: a second backend lifecycle, a second set
of fixtures, a second thing to keep working.
An adopter's scenarios now run inside the canonical suite instead -- same
session, same provider registration, same backend control. Almost nothing was
needed to make that happen, because pytest already scans: it collects
`conftest.py` on its own and pytest-bdd resolves step definitions through the
fixture system, so a step an adopter writes beside their test module is already
in scope for the scenarios generated into it. The only thing pytest cannot find
by itself is the feature files, because the canonical ones live inside the
installed distribution. `feature_paths()` returns both -- the packaged assets,
and a `tck-extensions` directory beside the calling module if there is one --
so an adoption gains one call and no configuration:
scenarios(*feature_paths())
An extension must never be able to stand in for a canonical scenario. Java's
suite found that a same-named feature file in a second classpath root replaced
the canonical one outright and the run went green having asked the adopter's
questions; Python has a narrower route to the same place, because pytest-bdd
names a feature file by its parent directory joined to its own name and
`tck-extensions/features/errors.feature` therefore arrives under the uri the
canonical `errors.feature` already occupies.
So the uri a feature file is identified by is derived from where the file is:
`features/` for the packaged assets and nothing else, `extensions/` for anything
below a `tck-extensions` directory -- the same prefix the Go and JavaScript
suites mount extensions under, so a consumer holding reports from several
languages applies one rule. The two cases the derivation cannot rule out are
reported rather than raised, because the scenarios are the adopter's to run and
it is publishing them as the specification's that has to be refused: a file of
the adopter's own that would reach the reserved `features/` prefix, and two
feature files that would share one uri, which nothing recording a run can hold
because it keeps one copy of a feature file per uri.
Whatever refuses to publish is not here. The derivation and both problems are
public, and the self-test reads a generated adoption back through pytest's own
JUnit XML rather than through a conformance report, which this package does not
write.
Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
…e steps Follow the conformance assets to open-feature/spec@15fe861, which adds metadata.feature, three shutdown scenarios to lifecycle.feature, the falsy-value and integer-precision scenarios to evaluation.feature, the lossless half of @numeric-coercion to errors.feature, and six flags to the canonical set. Steps: - "the error message should be empty" reads the last evaluation's error_message and accepts None or "". - "the provider is shut down" and "the provider is initialized again" call the registered provider's own shutdown() and initialize() directly, not through the SDK, so a scenario can shut down twice and an evaluation afterwards reaches the instance that was brought back. Each call is recorded as a LifecycleRecord with its duration and anything it raised; "no exception should have been thrown" now reads those records alongside the evaluation's, so there is one mechanism rather than two. A call that outlasts ready_timeout is given up on and recorded as a TimeoutError. - "the shutdown should have completed within {int}ms" bounds the most recent shutdown, parsed the way the event step's bound is. - "the provider metadata name should not be empty" asks the provider for get_metadata() and requires a non-blank string. Capabilities and flags: - @large-integers is a declarable capability. Python's int is unbounded, so both in-memory self-tests declare it. - The six new flags are transcribed into canonical_flag_set(), and a test checks the transcription against canonical-flags.json value for value and Python type for Python type, so 10.0 stays a float and false, 0 and "" stay values. - The in-memory self-tests stop declaring @numeric-coercion: the SDK's InMemoryProvider hands values back untouched and the client's type check is isinstance-based, so 10.0 requested as an integer is a TYPE_MISMATCH rather than 10. The lossless scenarios exist to catch exactly that, and the capability is optional, so the honest declaration is to leave it out. Recorded as finding 3 in the README. The lifecycle steps have no canonical scenario running them here, because neither in-memory suite declares @lifecycle, so test_lifecycle_steps drives them against a recording provider instead. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
…mes them Moves the spec submodule to ba002ce8, which renames the canonical set's three falsy flags -- false-flag, zero-flag and empty-string-flag become boolean-zero-flag, integer-zero-flag and string-zero-flag -- and follows the rename through the in-process control's flag set. The names the TCK invented were its own. Appendix B's SDK suite already had names for these three, flagd-testbed serves that vocabulary, and a provider suite that asks for a different one gets FLAG_NOT_FOUND four times over for no reason other than the disagreement. The three entries are now byte-identical to specification/assets/gherkin/test-flags.json on spec main, so a backend seeded for the SDK suite is already seeded for this one. The variant names move with the keys, from on/off, one/zero and greeting/empty to zero/non-zero throughout. That is not cosmetic: the falsy scenarios assert the variant as well as the value, so a fixture that kept the old variant names would fail on the assertion rather than the lookup. canonical_flag_set() is a transcription of the asset, so it has to move in the same commit: the type-for-type mirror test compares the two, and a commit that moved only the pin would leave the in-process fixture answering FLAG_NOT_FOUND to every falsy scenario -- the same failure the rename exists to remove, in the opposite direction. Nothing generated needed committing. Both the copied assets and spec_revision.json are gitignored and rebuilt by hatch_build_sync.py, which keeps the submodule pin the single record of the revision this package targets. The scenario count is unchanged at 40, as a pure rename should leave it. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
``a stable provider`` and ``an unavailable provider`` both registered through ``api.set_provider``, which initialises on a worker thread and returns immediately. Both steps needed ``api.set_provider_and_wait``, which is the variant that passes ``wait_for_init=True`` down to the registry. The stable case is the damaging one. The step's docstring already claimed that registration "initialises synchronously and dispatches PROVIDER_READY, so by the time this step returns the provider is ready" -- the claim the whole suite rests on, and it was not true of the call being made. Every scenario therefore ran its first evaluation against a provider still coming up and got ``PROVIDER_NOT_READY``, which looks precisely like a provider that cannot resolve anything. Against a real flagd backend that is 58 of 80 tests failing for a reason that has nothing to do with flagd. The unavailable case was wrong in the mirror image. Its comment reasons that "the SDK's registry already converts a raising initialize into PROVIDER_ERROR", which only happens if ``initialize`` is actually called; with the non-waiting variant registration returned before the provider had tried to reach its backend, so the ``@unavailable`` scenarios asserted an error state that had not happened yet. The ``contextlib.suppress`` around it stays: with the waiting variant a raising ``initialize`` can propagate, and that must not take down a scenario whose subject is the observable error state rather than how registration returned. Nothing in this package's own tests could catch it. Both self-hosted suites drive ``InMemoryProvider``, which initialises in microseconds, so the race was always won and the step's assumption held by accident. It took a backend that takes a moment to come up -- flagd behind a container -- to show the difference, which is also why it survived until the conformance suite could be run for real. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
Bumps the spec pin to fc99d5ac, which gates the scenario "A provider that was shut down can be initialized again" behind a new @reinitialization tag, and adds that capability to the vocabulary. Requirement 2.5.2 says a provider SHOULD revert to its uninitialized state after shutdown, and its supporting text adds that "some providers MAY allow reinitialization from this state". Reuse is permitted, not required, so asserting it unconditionally reported a permitted choice as a conformance failure -- the mirror image of a vacuous pass, and on its way to being filed as a defect against an implementation that was exercising a choice the specification offers it. The capability is declarable without further work: DECLARABLE_CAPABILITIES is the enum minus the reserved set, so it picks the new member up, and the declaration guard in test_declaration.py confirms the coupling -- running the new enum against the old pin fails it, because no scenario there carries the tag. Neither in-repo adoption declares it, and deliberately so. The scenario lives in lifecycle.feature, which carries @lifecycle at the feature level, so it inherits that tag and carries both; the gate skips a scenario when any capability gating it is undeclared. Neither in-memory adoption declares LIFECYCLE -- there is no backend to reach during initialisation -- so the scenario was skipped at the previous pin too, and declaring reuse would be claiming a property nothing has observed. The skip breakdown shows the move exactly: @lifecycle went from 6 skips to 4, with 2 now reported against @reinitialization. That the tag narrows @lifecycle rather than standing beside it is the trap worth naming, so it is called out in both the capability docstring and the README rather than left for an adopter to infer from a skip reason. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
The suite calls `api.set_provider_and_wait` so that a scenario evaluates a
flag only once the provider has initialised. That function arrived in
openfeature-sdk 0.10.0, but the package still declared `>=0.8.2`, so the
workspace lock resolved 0.8.4 and every scenario died on
AttributeError: module 'openfeature.api' has no attribute
'set_provider_and_wait'
CI runs `uv sync --frozen`, so it installed the locked 0.8.4 and saw the
same failure rather than the green suite the branch claims.
Raise the floor to the release that actually carries the function and
relock. Only openfeature-sdk moves, 0.8.4 -> 0.10.0.
Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
…guages do
An audit across the four language suites found three different canonical uri
forms for the same file:
gherkin/errors.feature Go
specification/assets/provider-tck/gherkin/errors.feature JavaScript
features/errors.feature this suite
A consumer joining two languages' results keys on the uri and the scenario
name, so the partition Appendix F's rule exists to guarantee is precisely the
thing that did not survive it. Appendix F now states the form rather than
implying it: a canonical feature is identified by its path *relative to the
asset directory* -- gherkin/errors.feature -- and an extension mounts under
extensions/.
The structural cause is a single local rename. Go consumes the assets as a Go
module whose root *is* the asset directory, so its embed keys are
gherkin/*.feature and it gets the right uri for nothing. The other three vendor
the assets into a locally-named directory, and the uri inherits that local name.
Here the name was CANONICAL_DIRECTORY = "features", which the sync copied
gherkin/ into and which uri_for() then reported. Point it at the name the assets
already have and the copy, the reserved prefix and the emitted uri agree again.
The reported uri looked right in one place and was wrong in another, which is
worth recording: uri_for() takes precedence over pytest-bdd's rel_filename at
the emitter's call site, so the pytest-bdd path is only a fallback. Reading the
fallback alone suggests this suite was already correct.
EXTENSIONS_DIRECTORY moves from "tck-extensions" to "extensions" in the same
commit, because Java's TCK is being renamed the same way in the same round and
the docstring's parity claim -- that an adopter shipping a provider in both
languages puts the same directory in both repositories -- is only true if both
move. The other half of that argument still holds and is now stated rather than
assumed: an extensions directory must not share the canonical name, because a
directory sharing it is how an extension comes to occupy a canonical file's
identity, and "gherkin" and "extensions" are distinct.
EXTENSIONS_URI_PREFIX stays a constant of its own even though it now equals
EXTENSIONS_DIRECTORY. The directory this suite scans and the prefix a report is
keyed by are two facts, and only the second is fixed by Appendix F. One name
doing both jobs is exactly what went wrong on the canonical half.
reserved_prefix_problem() and collision_problem() build their messages from the
constants, so they follow the rename rather than policing a stale string; the
same is true of is_canonical_uri() and uri_collisions(). The gitignore entry and
the packaged-wheel artifact list name the copied directory literally and move
with it.
One test is added. Every existing assertion is written against the constants, so
it holds whatever they say -- renaming one would leave the suite green while the
uris stopped joining with another language's, which is the failure that
happened. The two strings Appendix F fixes are now pinned as literals.
The spec assets themselves did not change, so the submodule pin does not move.
Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
The control API is normative -- Appendix F defines it as an HTTP surface a backend under test MUST expose -- so every adoption that drives a real backend needs a client for it. With the client in the flagd adoption the suite shipped the contract and not the thing that speaks it, and an adopter taking this package had to write their own. It went unnoticed because the only other adoption, OFREP, is stacked on flagd and inherited it. A third-party adopter is the case nobody was standing in for. Nothing about it was flagd-specific: urllib.request only, so the suite still gains no HTTP client dependency and no container dependency, and DEFAULT_CONFIGURATION is the configuration name Appendix F requires of every backend. Its documentation and its eighteen tests come with it; orchestrating a stack and discovering its mapped ports stays with the adopter, which is the part that is genuinely vendor-specific. Java already shipped its equivalent on the suite side; Go moved in the same round. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
The pull_request filter matches the BASE branch, so only the suite PR -- the one targeting main -- was ever checked. The report and adoption PRs stacked on it have never run CI, which is why their green ticks meant nothing: the checks on display belong to the base PR. One line, and temporary for the duration of review. The workflow is taken from the head branch, so it has to sit on the base and reach the children by rebase. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
…ason "Not declared" and "not applicable" are both skips. Giving them separate representations asks an adopter to learn more vocabulary without telling a reader anything the skip's reason does not already say: the scenario's tags say what was asked, the declaration says whether it was claimed, and the reason says why it was skipped. The gate never distinguished them, and neither did the results payload. The field's own docstring made the argument for removing it. It reserved itself for provider-specific impossibility, on the grounds that an impossibility which is a property of the language belongs in the capability documentation rather than in every report -- and both motivating cases are exactly that. @numeric-coercion cannot hold where the language has a single numeric type; @large-integers cannot hold on a 32-bit accessor. Neither is a fact about a provider, and both are now stated once in Appendix F. Nothing under providers/ ever populated the field. The report schema dropped declaration.notApplicable in open-feature/spec 7f03f672, and Appendix F records where language-level impossibility lives in 600ef9fd. Gone with it: the rule refusing a capability named in both capabilities and not_applicable. It was a rule about holding two claims at once, and with one claim left there is nothing for it to be a rule about -- a capability is declared or it is not, which the dataclass already enforces by having one field. The reserved-capability refusal and the unknown-capability refusal both still apply to what remains. The capabilities docstring and the README now say where a capability that cannot hold in a language at all is recorded, so an adopter who looks for the field finds the answer rather than its absence. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
A provider that talks to a backend needs that backend running, its
dynamically mapped host ports discovered, and an `HttpControl` built
against its control API. Every adoption needs the same three things and
until now every adoption wrote them: this package shipped the control
client and left orchestration to the adopter, so the flagd adoption alone
carried a 122-line `conftest.py` and a 170-line `suite.py` of container
wiring, and the next adopter would have paid for it again. It is the
single largest adoption cost in three of the four languages.
An adopter now names a Compose file, says which ports the provider
connects to, and builds a provider from an endpoint they are handed:
@pytest.fixture(scope="session")
def compose_backend() -> ComposeBackend:
return ComposeBackend(
compose_file="tests/tck/docker-compose.yaml",
backend_ports=[8013],
)
@pytest.fixture(scope="session")
def tck_config(tck_backend: RunningBackend) -> TckConfig:
return TckConfig(
name="my-provider",
control=tck_backend.control,
new_provider=lambda: MyProvider(
host=tck_backend.endpoint.host,
port=tck_backend.endpoint.port(8013),
),
)
`ComposeBackend` carries the concepts and defaults Java's
`ContainerizedProviderTckTest` fixed -- backend service `backend`, control
port 8080, config `default`, startup timeout 60s, plus `additional_ports`
for a stack with more than one service -- so a provider shipped in two
languages writes one Compose file and two declarations against it.
`tck_backend` is a session-scoped fixture of this package's plugin, and it
is lazy: an in-memory adopter never requests it, never defines
`compose_backend`, and never needs Docker.
The stack starts once per suite and is never restarted, because
orchestrators cannot reliably preserve dynamically mapped host ports
across a restart and a restart would silently invalidate every provider
already pointed at the old one. Unavailability stays simulated inside the
running stack through the control API. The Compose path is additional
rather than a replacement: a provider with no backend keeps supplying its
own `BackendControl`.
Startup is a readiness check rather than a pause. `docker compose up
--wait` brings the containers up, every declared port is then waited on
until it accepts a connection -- the guarantee Java gets from a
Testcontainers listening-port wait strategy, which `--wait` alone does not
give for a service with no healthcheck -- and `HttpControl.await_ready()`
probes `GET /healthz` until the control API answers, treating 404 as ready
because `control-api.yaml` defines it that way and the reference launchpad
serves no such path.
There is deliberately no settle *after* a control call. Java sleeps 50ms
after every one; flagd-testbed#394 makes `POST /start` block until the
flags are evaluable, so the sleep covers a window that no longer exists,
and a suite that sleeps instead of holding the control API to its promise
stops being able to detect when the promise breaks.
`testcontainers` is the optional `compose` extra rather than a dependency,
imported lazily, so an in-memory adopter does not install container
tooling to run a suite that never starts a container. It is in the dev
group so the lazy import is type-checked, with a mypy override because
testcontainers 4.14 ships no py.typed for `testcontainers.compose`. The
port resolution is checked against a `ComposeStack` protocol instead,
which is what lets it be tested without Docker at all -- the stack
lifecycle is proved by the flagd adoption.
`__init__.py` and the README also carry the export list and the
documentation for the two preceding commits, because the public surface
and the docs move together.
Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
The pin was parked at 009afe06 while the specification's changes were
prose-only. control-api.yaml has now changed, and it is one of the three
normative artifacts this package consumes, so the pin is due.
What moved, in the assets:
- every state-changing endpoint -- /start, /change, /reset -- must now not
return until the new state is actually being served. /start already said
so; the other two did not;
- for /change the promise is spelled out as being about the *backend*, with
the provider's own detection latency explicitly the business of the event
timeout instead;
- /restart is demoted to [OPTIONAL], because its own description falsely
claimed the TCK used it for the disconnect/reconnect scenarios.
The vendored copies under src/.../tck are gitignored and rebuilt by
hatch_build_sync.py, so the pin is the whole of the change here. The feature
files are untouched by this revision apart from a comment on @disabled-flags;
no scenario and no canonical flag changed, so the suite's tally is unaffected.
Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
control_api graduates from a comment block describing an optional duck-typed property to a required member of the BackendControl protocol, typed ControlApi -- Literal["http", "in-process"] -- with no default and no inference from the control's concrete type. The comment it replaces argued for omission on two grounds and both are void. "Adding one would make every existing control incomplete": nothing is published, and after the compose harness an adopter with a real backend writes no control at all -- the HTTP one comes with the harness. The only person who writes a control by hand is the one writing a custom one, which is precisely the case where the value cannot be inferred. "There is nothing useful the TCK can do with a control that has not said": correct, and it is the argument for requiring the control to say rather than for omitting the field. It is the one fact that decides what everything else in a report is worth. The same scenarios passing over the normative control API and passing through in-process manipulation of a provider that does have a backend are not the same claim, and this is the only field that separates them. Every run is one or the other, so an absent value is not "no claim made" but an unfalsifiable one. Closed rather than a bare str, so "HTTP" or "grpc" is a type error here instead of a conformance report that fails schema validation somewhere with nothing local to point at. ControlApi is exported for a custom control to annotate with. BackendControl is runtime-checkable, so a control that will not say is refused at the seam as well as by the type checker, which is what reaches an adopter whose test module is untyped. Appendix F states the same rule normatively at spec@93eb1a58: the control states the path, the harness must not guess, and an omitted value is not neutral. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
ComposeBackend.configuration becomes backend_configuration, and HttpControl's keyword argument with it. The word was already taken. provider.configuration in the conformance report schema is "which configuration of the provider was tested, when a provider has more than one materially different mode" -- flagd RPC versus in-process -- and it is the field TckConfig.name feeds. That is the provider's mode, not the backend's config file, so one word for both made a report's configuration mean opposite things depending on which language's TCK produced it. Java distinguished them correctly from the start; Go, Python and JavaScript all took the word for the backend one. Settled as backendConfiguration in all four languages, spelled backend_configuration here. configuration keeps its schema meaning everywhere. DEFAULT_CONFIGURATION keeps its name: it names the default *value*, "default", which is the only configuration name every backend under test must support, and Go's DefaultConfiguration is spelled the same. Neither adoption passed the field -- both take the default -- so nothing outside this package had to change. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
HttpControl.restart is deleted. POST /restart is [OPTIONAL] in control-api.yaml as of spec@93eb1a58, because its own description falsely claimed the TCK used it for the disconnect/reconnect scenarios. It does not: the @Stale scenario is written as an *unbounded* outage -- "the connection is lost", then "the connection is restored" -- which is /stop followed by /start, so the scenario ends the outage when it is ready rather than guessing in advance how long a provider needs to notice one. No step in any language reaches it, and Go dropped its binding deliberately. The docstring being removed also asserted the endpoint "is required of every backend", which was true of the document it was written against and is now false. A binding nothing can call is dead surface that misreports the contract. What would bring it back is recorded in its place: a @caching scenario asserting what a stale provider serves *during* an outage needs the flag-state preservation /restart has and /stop + /start does not. Also aligns HttpControl's documentation with the other two changes at that revision -- every state-changing endpoint owes the caller that the new state is being served before it returns, and for /change that promise is about the backend, with the provider's detection latency the business of the event timeout. Neither adds a wait: the class now says why a settle would be the wrong instrument, and where one genuinely belongs when a backend breaks the promise. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
Adds the policy item 8 settled: a containerised adoption suite is excluded
from the default build and a maintainer runs it by hand before merge. Written
down because an exclusion nobody wrote down is indistinguishable from an
oversight, which is exactly how this went unnoticed. Two reasons, and the
second decides it: Docker, and the fact that a conformance suite's honest
output includes failures that are not the provider's -- a canonical flag the
backend does not seed yet -- so a gate that must be green cannot hold it and an
xfail would blame the provider for the backend.
The corrections:
- control_api is documented as required, closed and stated rather than
guessed, in both the control section and the no-backend one;
- the compose table's configuration row is backend_configuration;
- there is no /restart binding, and the reason, replacing a bullet that
presented the endpoint as part of the contract every backend owes;
- the settle paragraph no longer claims flagd-testbed#394 closed the window.
It did not -- #394 is open and unmerged, the launchpad still returns from
/start as soon as /readyz answers, and the ~40 ms window is real and
measured. The reason not to sleep is that control-api.yaml now requires
every state-changing endpoint to serve before it returns, and that a fixed
delay is un-tunable and hides the defect from the one consumer positioned
to notice it. Where an adopter is stuck with such a backend the wait
belongs in that adoption, named and citing the defect -- which is what the
OFREP adoption's SettledControl is, and what Appendix F now prescribes.
The self-test tally was also stale: 190, not 163.
Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
Appendix F gains "Running the suite in CI" and the scenario-authoring constraint on the caching gap. No scenario, no canonical flag and no control API change: re-running the asset sync leaves all nine copied files byte-identical, checked by hash before and after. The vendored assets are gitignored and rebuilt by hatch_build_sync.py, so the pin bump is the whole of the commit. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
A reserved tag is a name held open for scenarios that do not exist yet, and it is only ever temporary: the specification writes them, the tag starts gating something, and the capability becomes declarable. Until this package follows, TckConfig refuses to let anyone declare it -- so the new scenarios skip, for a capability no adopter is allowed to claim, and the report shows a gap the provider may not have. Appendix F calls that the unclaimable capability, and it has no local symptom: a handful of extra skips in a run that is otherwise green. @targeting was reserved until spec revision 26362f85 gave it three scenarios, so this is not hypothetical. This package already checked the property -- test_declaration asserted that no packaged feature file carries a reserved tag -- but a self-test is read by whoever changes this package, and a reservation expires somewhere else. An adoption that re-pinned the assets and ran the suite saw nothing. So the check moves into the plugin, where it fails an adopter's run, and the self-test keeps the converse half: every declarable capability must be carried by some canonical scenario, which catches a tag added to the enum and never wired to anything. extensions.canonical_tags() reads the tags off the packaged feature files. Tag lines only, which matters more than it sounds: events.feature mentions @caching in a comment, saying where those scenarios will go once they exist, so a scan that read whole files would fail every adoption over a sentence. Recursive, because the shape of that directory is the specification's to change. capability.expired_reservations() is the pure comparison, matching JavaScript's expiredReservations, and neither name is exported from the package -- the adopter-facing answer is RESERVED_CAPABILITIES. The tags come from the packaged assets rather than from what was collected, for two reasons. Only the canonical set can expire a reservation: an adopter's own feature reaching for a reserved tag is a mistake in that file, not news about the specification. And a narrowed run cannot then select its way past the check. What is read off the collection is whether the session runs the conformance suite at all -- the plugin is installed for every pytest run in the environment, and an unrelated test suite has no business failing over the contents of these feature files. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
"Running it in CI" was four paragraphs of reasoning about why a conformance suite must not be a required gate, written here in this README's own words -- and written again, differently, in three other languages'. That is precisely how the known-deviation and control-path decisions came to have three answers, so the reasoning now lives once in Appendix F and this section points at it. What stays is the mechanism, which is Python's and belongs with Python: the two poe tasks build.yml reaches, the --ignore that excludes the suite from them, the test-tck task beside them, and the comment above them. Two notes stay too, because they are facts about this repository rather than restatements of the appendix: - Docker is not what decides the exclusion here. tests/e2e needs Docker too, has for years, and still runs in the default build on ubuntu-latest. - --ignore does not import the suite, and mypy in these packages is configured over src alone, so nothing in the default build would notice the suite failing to import against the harness. The appendix asks that an excluded suite keep compiling; pytest --collect-only is the form Python has for that, and the adoption branches now run it. The "Known gaps" entry for caching points at the appendix's list rather than naming the gap and stopping, since that list now carries the constraint a @caching scenario has to be written against. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
The check landed canonical-only, on the reasoning that only the specification can expire a reservation -- an adopter's own feature reaching for a reserved tag is a mistake in that file rather than news about the specification. That reasoning is sound and it answers the wrong question. A reserved capability cannot be declared, because TckConfig refuses it, so the capability gate skips every scenario carrying its tag whichever file the tag is in. An extension scenario tagged @caching is therefore unclaimable from the day it is written: it can never run, it can never be claimed, and the report shows a gap the provider may not have. That is the unclaimable-capability failure Appendix F describes, and it is the failure this check exists to surface. It arrives from the adopter's side instead of from upstream and the consequence is identical, so one check covers both causes as long as the message names both remedies -- drop the tag from RESERVED_CAPABILITIES because the specification has given it scenarios, or rename a tag of your own that reached for a reserved name. Go and Java check every scenario the run collected. This is Python coming into line rather than a local judgement: the divergence was put up for a ruling and the broader rule is what it settled on. Two things stay as they were and both are Python's alone. The canonical half is still read off the packaged feature files rather than off the collected items, because an adopter's selection must not be able to narrow a run past the specification's half; an extension has no equivalent source, since its directory is found from the adopter's own test module, so that half comes from the collection as it does in every other language. And the check still waits for a canonical scenario to be collected before it looks at anything: this package is a pytest11 plugin, so the hook fires for every pytest run in an environment that merely has it installed, and an unrelated test suite has no business failing over the contents of these feature files. Tags are read only from scenarios the uri derivation in extensions claims, so a pytest-bdd suite of the adopter's own sharing the session is left alone too. The tags are the markers pytest-bdd derives from the parsed Gherkin, which is what the capability gate itself reads, so the check and the gate cannot disagree about which scenarios are unclaimable -- and a tag inherited from the feature or the rule is included, where Scenario.tags would miss it. Still never the file text: gherkin/events.feature names @caching inside a Gherkin comment, and a text scan would fail every adoption over a sentence. Verified against a real collection rather than only the stubbed items. An adopter suite whose extensions/vendor.feature carries @caching aborts at collection with exit 4 and that message; renaming the tag collects the same 58 scenarios clean; moving @caching into a Gherkin comment collects clean. Narrowing the check back to the canonical set fails exactly one test, the one that covers the widening. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
Re-pin the spec submodule to c342461a, which moves every resolution-reason assertion out of evaluation.feature, errors.feature and lifecycle.feature -- all thirteen of them -- and into a new gherkin/reason.feature gated as a whole on @standard-reasons. Requirement 2.2.5 is a SHOULD that lets a provider populate `reason` with one of the listed values "or some other string indicating the semantic reason for the returned flag value". Asserting an exact reason in thirteen places narrowed that into a MUST for every adopter, and bought very little: every canonical flag resolves to a value distinct from the caller's default, so a provider that silently falls back was already caught by the value. So @standard-reasons is a claim rather than an exemption -- a provider saying "I use the standard vocabulary with the standard meanings", and reason.feature is what checks it. A provider that does not declare it loses nothing; its values, variants and error codes are asserted everywhere else, on MUSTs. Add the capability, declarable, beside the other twelve. It is the first whose tag is carried at the feature level rather than per scenario, which the gate handles because it reads pytest markers and pytest-bdd marks a scenario from scenario.tags | feature.tags | rule.tags. Verified by running: with the tag undeclared all nine of its scenarios skip in each in-memory suite. Both self-tests declare it, measured before declaring. InMemoryFlag.resolve reports Reason.STATIC for every flag in the decoded set, and a missing flag and a type mismatch both arrive with reason ERROR beside their error code, so the four rule-less rows and the two error scenarios pass in each. The remaining three compose the tag with @targeting and @disabled-flags, neither declared here, so they skip. Self-tests move from 196 passed / 35 skipped to 208 passed / 41 skipped: nine new scenarios in each of the two conformance suites, six running and three skipped. The thirteen removed assertions changed no scenario count -- they were lines inside scenarios that still exist. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
…s assets A rebase moves the gitlink and not the submodule's working tree. So a checkout can hold a pin naming one revision and assets on disk from another, and nothing in the build says so. That is not hypothetical. Re-pinning the assets last pass and running the suite copied the *previous* pin's Gherkin over the capability the suite had just been given, because the rebase had moved the gitlink and `poe sync-spec-assets` copies whatever is in the submodule working tree. The only thing that noticed was a self-test comparing the capability enum against the packaged assets -- and that guard fires for exactly one symptom, a declarable capability that no canonical scenario carries. A pin that changes nothing but the content of a scenario passes every guard in this package and still runs the wrong suite. The same root cause in another language ran an entire adoption suite against stale assets and reported byte-identical numbers to the run before it, with nothing failing and nothing warning. So `sync()` now brings the submodule to the revision the superproject's index records before it copies anything, and `poe test` already depends on the sync: the suite cannot run against assets it did not just check out. `git submodule update` after a rebase stops being something an operator has to remember, which is the part that was never going to hold. The guarantee is narrower than Go's and the README says so rather than implying otherwise. These assets reach the package by being copied, so a copy can always be made wrong; what this buys is that the suite cannot run without a fresh sync and a sync cannot succeed against any revision but the pinned one. Go consumes the assets as a nested module out of a read-only, checksum-verified cache and has no second artifact to go stale at all. The two fail differently. The index rather than HEAD, because the index is what the next commit records and so what a run is about to claim it tested against. If the update does not reach the pinned revision -- the commit is not in the local object store and could not be fetched -- nothing is copied and the build stops, naming the likely cause. Nothing is written when the working tree is already at the pin, so the ordinary path touches no git state. Two ways out, both loud. OPENFEATURE_TCK_SPEC_UNPINNED=1 copies whatever is checked out, for drafting a change to the canonical assets before there is a revision to pin; it warns on every sync and names the revision it used. And where the pin cannot be read at all the sync warns and continues: building from an unpacked sdist is the ordinary case, with no repository, no pin and the assets already in the tree, but so is a linked git worktree whose .git file names a path outside the running process's filesystem namespace -- a Windows worktree driven from WSL, where git answers inside the submodule and not in the superproject. The guarantee genuinely is not in force there and the warning is what says so. tests/test_spec_assets.py pins both halves. The decision is checked against a scripted git: a working tree behind the pin is checked out and the checkout confirmed, one already at the pin is left alone, a checkout that does not reach the pin stops the build, and neither warning path issues an update. And the invariant itself is checked against this very checkout -- the submodule HEAD equals the pin -- which is the assertion that would have failed last pass. It skips where no pin is readable, because there the guarantee is off and a pass would say otherwise. `hatch_build_sync.py` joins the mypy file list, and pytest gains `pythonpath` so the tests can import it the way the build hook does. Self-tests move from 208 passed / 41 skipped to 214 passed / 42 skipped: seven new tests, one of which skips here because this checkout's pin is unreadable from the shell the suite runs in. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
…eaving it to adopters Re-pin the spec submodule to 89b1519a, which adds a fifth rule for declaring to Appendix F -- a capability the language's SDK cannot express is refused by the implementation, not left to adopters -- and corrects the canonical flag set's comments, which still told every reader that every scenario expects reason STATIC. Neither feature file changes. Two such capabilities exist anywhere: @large-integers where the integer accessor is a 32-bit Integer, and @numeric-coercion where the language has one numeric type and "a float requested as an integer" does not name two different requests. Neither says anything about a provider. Leaving it to adopters means every adopter in the language has to know a fact about their language and remember to act on it, and in one implementation three separate suites each left the same capability undeclared with its own comment restating the same property -- three places to get right, and a single wrong one puts a claim in a report that no scenario could have verified. So INEXPRESSIBLE_CAPABILITIES maps such a capability to the property of the SDK that puts the question out of reach, TckConfig refuses one at construction, the capability gate skips its scenarios with that reason, and a knownDeviations entry may not name one -- the gap would be the language's and the entry would attribute it to this provider. A mapping rather than a set because the message has to name the property: an adopter who reaches this has done nothing wrong and "the specification says you may not" is not something they can act on. The two refusals stay distinguishable, in separate predicates with separate messages and separate skip reasons. A reservation is global and temporary -- no scenario anywhere carries the tag, and it expires the moment the specification writes one. An inexpressibility is one language's and permanent: the scenarios exist and other languages run and pass them. A reader seeing a capability absent from a report has to be able to tell "this provider declined" from "no provider in this language can be asked", because only the first says anything about the provider. Where a scenario is gated by both kinds, the language-wide reason wins, because the provider's declaration could not have made that scenario run either way. **The mapping is empty in Python, and that was measured rather than assumed.** `int` is arbitrary-precision; FlagType.INTEGER and FlagType.FLOAT are separate, reach separate provider methods and are type-checked against `int` and `float` separately. All four questions the two tags ask were put through the SDK's own client against a provider implementing the borrowed coercion rule: 2^53 - 1 resolved exactly, 0.5 as an Integer gave TYPE_MISMATCH and the caller's default, 10.0 as an Integer gave 10 and 10 as a Float gave 10.0. The adoptions agree from the other direction, and this was measured too rather than reasoned from the source: declaring @numeric-coercion in both flagd suites and running it, the in-process resolver refuses 0.5 as an integer and widens 10 to a float, while the RPC resolver widens 10 and silently narrows 0.5 to 0. Two resolvers of one provider giving different answers to the same three questions is exactly what a language that could not ask them makes impossible. Both are defects in an implementation, withholding the tag is the honest report for each, and neither is anything the language prevents. The third scenario fails on both for a third reason again -- flagd-testbed seeds no integral-float-flag -- which is also why the run leaves both suites' declarations exactly as they were. So nothing here is in force, and the machinery is added anyway. The rule belongs to Appendix F rather than to this package, a future capability may hit it, and the costs are not symmetric: an unused mechanism is a few lines nobody reads, while a missing one is discovered by an adopter publishing a claim no scenario could have examined. It is exercised rather than left dead -- the tests supply an entry and drive the refusal, the deviation refusal, both skip reasons and the precedence between them, so a mechanism with no instances is still known to work. DECLARABLE_CAPABILITIES is derived from both sets rather than listing what it excludes, which matters precisely because the new one is empty here: a derivation that quietly dropped it would look right in Python forever and be wrong in the one language where an entry gets added. Self-tests move from 214 passed / 42 skipped to 221 passed / 42 skipped. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
aepfli
force-pushed
the
feat/provider-tck
branch
from
September 13, 2026 14:11
5720a9a to
691c68b
Compare
…failure The build matrix is five Python versions times every package the change touched. The point of a matrix that wide is to say which combinations fail -- a package that breaks only on 3.10, a package that breaks everywhere, a package that is fine. `fail-fast` defaults to true, so the first cell to fail cancels all the others, and a run that could have reported thirty results reports one. The information lost is the part that costs the most to recover. A cancelled cell is not a passing cell and not a failing one; it is silence, and the only way to find out what it would have said is to push again and hope the cells finish in a different order. Worse, a failure in a package the pull request did not touch is indistinguishable from a failure in the package it did -- both present as "the build job failed", with most of the evidence cancelled. Observed rather than hypothesised, across two consecutive heads of one pull request: an unrelated package's cell failed, and every cell of two other packages was cancelled with it. At the earlier head four of five cells of one package completed; at the later head none did. That package's own results were never reported at all, for reasons that had nothing to do with it. The cost of the change is runner minutes spent on cells that were going to fail anyway. That is the trade this setting exists to make, and for a matrix whose whole purpose is telling combinations apart it is the right side of it. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
Two prose corrections to Appendix F and nothing else: 045950ca rewrites the numeric-coercion note that had been teaching withhold-plus-deviate, and 4cab0320 adds a sixth declaring rule saying the unit of a declaration decision is the scenario rather than the tag. `git diff --name-only 89b1519a 4cab0320` is one file, the appendix itself, and over `specification/assets/` it is empty -- the Gherkin, the canonical flag set and the control API document are byte-identical, and the synced copies under `src/` hash the same before and after. So no scenario count moves and none should: the self-tests stay at 221 passed, 42 skipped, 2 xfailed. The pin still has to move, because it is what records which questions a run asked, and a suite documenting rules from a revision it does not name is the drift this submodule exists to prevent. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
Three paragraphs here told an adopter that a provider which narrows 0.5 to 0 should withhold @numeric-coercion and record a deviation for it -- "withholding it may be a deliberate choice as readily as a known bug. Where it is a bug, say so" in the enum, "withholding the tag is the honest report for each" twice over about flagd's two resolvers. That is withhold-plus-deviate, which the known deviation rule two sections away tells an adopter to avoid, and it is not what this repository's own flagd adoption does: it declares the tag on both resolvers and puts the deviation on the one that narrows. The wording was inherited from Appendix F, which said the same thing until spec@045950ca corrected it. The distinction the paragraphs draw is worth keeping -- an undeclared capability can be a choice or a defect, and a report that cannot tell them apart is worth less -- so the distinction stays and the illustration is replaced by the one that matches: attempts the coercion and gets a direction wrong, declare and deviate; cannot attempt it at all, withhold. The SDK's in-memory provider is the second kind and says so. The paragraph about flagd also called both resolvers defective, which is false in the other direction: in-process refuses 0.5 correctly. One resolver of one implementation is wrong, which is the finding, and it survives only because the tag was declared. Also cites the sixth declaring rule (spec@4cab0320) where an adopter narrows the default capability set, rather than restating it -- that rule was written from the wording in this repository's flagd suites, and four independent statements of one rule is what this effort keeps having to undo. Its first consequence goes next to the field it concerns: a scenario failing for a fixture the backend does not serve is not a provider defect and does not belong in knownDeviations. And the two self-test withholdings that are withholdings for a defect now name the carve-out that licenses them (spec@045950ca) instead of arguing the case again, with the condition each meets. @disabled-flags meets it exactly: the sweep in test_in_process_control asserts that all four disabled flags resolve to their own default variant, so the behaviour is pinned by a test of its own and that test turns red the day the SDK honours DISABLED. @configuration-change meets it differently and the note says so -- the scenarios are not only skipped, test_controllable_conformance runs them against the subclass that supplies what the SDK lacks -- with the one thing that pin does not do stated plainly: nothing fails on its own when the SDK gains the method. An adoption has no such licence, and both notes say that too, because these files are the nearest worked example an adopter will copy from. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
…fect split Re-pins to spec@aa2ad24f, which narrows the rule this repository's flagd suites supplied. As first written it said "declare when at least one scenario gating it can be put to the provider" with no condition, which read literally forces a declaration wherever the scenarios are reachable -- including where the specification permits declining outright, and @numeric-coercion is exactly that case. The rule now opens with the condition it always had implicitly: it applies once a provider is attempting the capability, and whether the provider owes an answer at all is the known-deviation rule's question, asked first. Assets are byte-identical across all three revisions, so nothing moves. The README's citation carries the two questions in that order rather than the second one alone, because the second alone is the over-reach. `DISABLED_FLAGS` was a fourth site teaching the shape the appendix corrected, and the grep for the numeric-coercion wording did not find it: "withholding it still needs no KnownDeviation ... that holds whether the gap is architectural or a defect; where it is a defect, the adoption's note is where to say so". The two cases are not alike. A backend that gives the provider no signal is a withholding with nothing to record; a provider one unconditional index from passing -- which is the Python OFREP provider, named two paragraphs above -- is the declare-and- deviate case, and sending both to "the adoption's note" is how a defect ends up looking deliberate. `CONFIGURATION_CHANGE` gains the illustration the numeric-coercion note used to carry badly: **one capability withheld twice in this repository for two different reasons.** The OFREP adoption withholds it by choice -- no stream, no poll, nothing watching, and no defect in building a provider that way -- and the in-memory self-test withholds it because the SDK's provider cannot update its flag set at all, which Appendix A requires of it (python-sdk#620). Identical in the results, distinguishable only from what the adoption wrote down, and both true here rather than hypothetical. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
This file is the package page on PyPI, and it was 59,708 bytes -- four times Appendix F, which is the normative document it implements. Someone arriving to adopt the suite had to read past four findings, the capability rationales, the control-API invariants and the decision history to reach a forty-line example. A base README documents how to use this library. Anything equally true of another language's binding belongs in Appendix F with a pointer, so it goes: the per-capability reasoning, the rules for declaring, the two shapes of a known deviation, the control-API invariants, the reserved/inexpressible comparison, and the findings, which are recorded in their filed issues and in the findings table on open-feature/spec#417. What stays is what is Python's. The adoption surface leads, because it is the shortest of the four -- two fixtures, one scenarios() call, and a step vocabulary arriving through a pytest11 entry point, so there is no conftest.py at all. Then the full option surface as two reference tables, the capability vocabulary, the dedicated poe task, the extension point, and the gotchas that are genuinely this language's: bool subclassing int, the three steps that reach the provider directly, the SDK's InMemoryProvider, and the submodule sync. Two corrections on the way past. The poe snippet was stale -- it showed `test` and `test-cov` as direct pytest commands and omitted `test-tck-collect`, so an adopter copying it got the exclusion without the compile check, which is the opposite of what the prose fifteen lines below it described. It is now the real one, from the two adopting pyprojects. And finding 4 said the DISABLED defect was unfiled; it is python-sdk#627. No behaviour change: 221 passed, 42 skipped, 2 xfailed, unmoved. 59,708 bytes to 17,464. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
The extension tree was rooted at `tests/`, while the task block ten lines above it excludes `tests/tck` and `poe test-tck` runs that path. An adopter following both sections got a conformance module the exclusion does not reach, which is the documented-command-that-does-not-work failure this effort started from. The example module is renamed with it. `test_conformance.py` inside `tests/tck` says conformance twice, and the module name selects nothing -- the directory does -- so the example now shows what both adoptions here do, with a sentence saying why. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
…achable Two things the README asserted without saying why they hold. It cannot pass vacuously: pytest exits 5 when a path collects nothing and 4 when the path does not exist, so a suite that moved out from under the task fails the step instead of skipping it. That is the property the other languages are being pointed at this mechanism for, and it was worth writing down rather than leaving a reader to trust it. And it only holds if the step runs. poe aborts a sequence at its first failing subtask, so the compile check sat behind the default suite's result -- which on the flagd package is red on this branch stack, so the check had in fact never run there. The task block gains `ignore_fail = "return_non_zero"`, and both adoptions' `pyproject.toml` gain it with it. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
capability.py was 537 lines of prose to 52 of code, and most of the prose was
Appendix F restated: the @standard-reasons section almost paragraph for
paragraph, the numeric-coercion note, how @reinitialization came to be gated,
the reserved-versus-inexpressible distinction, the declare-and-fail rule. A
second copy of a rule is a second place for it to drift, and one of them had
already drifted -- the @disabled-flags docstring spent fourteen lines disputing
a claim ("a provider whose backend decides, such as one speaking OFREP, cannot")
that the appendix no longer makes.
So each docstring now says what the tag gates *here* -- which scenarios run,
what a withholding skips, what Python's SDK makes of the question -- and links
for the rest. What survives is what this repository owns: the two withholdings
of @configuration-change that mean different things, the pytest-bdd
feature-level marker mechanism that @standard-reasons depends on, the
InMemoryProvider facts behind three of the self-tests' declarations, the
measurement behind an empty INEXPRESSIBLE_CAPABILITIES, and the
@reinitialization composition trap.
Also drops one cross-language anecdote from extensions.py -- a same-named
feature file on a second classpath root replacing the canonical one in another
language. The rule it illustrates is stated right above it from this
implementation's own uri derivation, and the anecdote is recorded where it
happened.
No behaviour change: 221 passed, 42 skipped, 2 xfailed, unmoved, and the
module's code is identical line for line.
Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
…irement Re-pins the conformance assets to spec d47a66eb and registers the `@string-typing` capability the pin introduces. Four scenarios moved out of the mandatory mismatch matrix upstream: `boolean-flag`, `integer-flag` and `float-flag` asked for as strings, and `object-flag` asked for as a string. Every value has a string representation, so a backend that stores flag values as strings satisfies the string accessor for every flag and has no mismatch to report -- and nothing in the specification contradicts that, since the only normative statement about value type is Requirement 1.3.4, a SHOULD on the client. Appendix F carries the reasoning; the enum member links rather than restates it, as the other capabilities here do. Registering it is not optional on this side. An unrecognised tag gates nothing, so leaving the enum alone would have kept the four scenarios mandatory under a tag the suite could not see -- the pin would have moved and nothing would have changed. It also makes the capability declarable, which `test_a_reserved_capability_is_one_no_canonical_scenario_carries` then holds to the assets: a declarable tag no canonical scenario carries fails that test. Both self-tests declare it, measured rather than assumed: the SDK's `InMemoryProvider` hands values back untouched and the client's check is `isinstance(value, str)`, so a non-string flag asked for as a string is a `TYPE_MISMATCH` without the provider having to notice. All four scenarios passed as mandatory at the previous pin and pass as gated at this one, so `poe test` still reports 221 passed, 42 skipped, 2 xfailed. `INEXPRESSIBLE_CAPABILITIES` stays empty: `get_string_details` is its own accessor reaching its own provider method, so Python can put the question. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
One tag over all four string-representation cases hid a defect inside a permitted absence, which is measurement rather than taste: over one Flagsmith backend Go and Java report TYPE_MISMATCH for boolean-flag and integer-flag asked as strings where JavaScript returns "true" and "10", while all three stringify float-flag and object-flag because that store records no native float or structure type. Under a single tag the third provider withholds and its bug is published as an absence the specification permits. So @string-typing keeps the two rows a partially typed backend can still answer, and @fully-typed-values gates the float and object scenarios. The new capability narrows the old one rather than standing beside it -- every scenario carrying it carries both -- so declaring it alone runs nothing, which is the same shape as @reinitialization under @lifecycle. Both in-memory self-tests declare both. The property the split divides is undivided there: the backend is json.loads, and a decoded 0.5 is a float exactly as a decoded true is a bool. Declaring only the first would skip two scenarios that pass and publish a gap neither provider has. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
The reserved check's own direction reversed, and Appendix F calls this the one that is easy to leave out -- all four reference implementations ignored an unknown tag rather than failing. A reserved tag is one this package knows and holds shut, and its scenarios skip. An unknown tag gates nothing, so its scenarios do the opposite and stay mandatory for every adopter: a suite that has not learned a new capability does not report a new capability, it goes on demanding the old behaviour, and a provider that legitimately withholds it goes red with no reason recorded anywhere. This was not hypothetical. Re-pinning at bda599f1 put @fully-typed-values in the canonical assets, and without the enum member the float and object scenarios stayed mandatory for two adoptions whose backends can answer them and for any that cannot. Proven by experiment, with the member renamed so the tag no longer resolved: the run stopped with the message rather than the scenarios running. Canonical tags only, and that is the substance rather than a simplification. A tag resolving to nothing is a problem exactly where every tag is meant to be a capability, and the canonical assets are the only such source. An adopter's extension tags its own scenarios for its own purposes and those gate nothing by design; reading them here would turn every extension into a failing run. The reserved check does read them, and the asymmetry holds: "is this one of my reserved names" is answerable about any tag and "is this a capability I do not know" is not. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
…n more places The check compared the assets on disk against the submodule pin and skipped where the pin was unreadable, on the argument that the guarantee is genuinely not in force there, so a skip is the honest report and a pass would not be. Both halves are true and they do not reach the conclusion: a skip is honest to a reader who reads it, and the environment where the pin was unreadable -- a linked worktree, read from another filesystem namespace -- is exactly the one where a rebase leaves the assets stale, because it moves the gitlink and not the working tree. So the check went quiet precisely where it was load-bearing, which is how an adoption in another language ran a full suite against the previous revision's scenarios and published plausible numbers. Appendix F now states the rule: a run-integrity check that cannot be performed fails. The better half of the fix is upstream of that choice, and it is why this is no longer much of a dilemma. The pin is readable in a linked worktree now: a worktree records its git directory as an absolute path that another namespace cannot follow, while a submodule's is relative, so the superproject is named explicitly using git's own layout -- a submodule's git directory is <superproject>/modules/<path>, and the ancestor called "modules" has the answer as its parent. Only silence is retried; an entry that is present and is not a gitlink is a definite answer. One environment stays exempt and it is not the same kind of thing. An unpacked sdist has no submodule, no pin, and no working tree that could have drifted from one: the assets are distribution content produced by a sync that ran this check at build time. That is a check with nothing to check rather than one that cannot be performed, and failing it would accuse a downstream packager of a defect they cannot hold or fix. Appendix F names an unpacked distribution and a linked worktree in one breath; they are alike only in that git says nothing, and this distinguishes them the way checkout_pinned_spec already does for its warning. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
Assets are byte-identical across the move -- the appendix prose changed and the branch was rebased onto main. No scenario, flag or control-API change, so no tally moves. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
The spec PR was split: the assets are open-feature/spec#423 and the appendix is #439. This pin follows the assets, so appendix review no longer moves it. The embedded artifacts are byte-identical across the move -- only the assets README changed -- so no tally moves. Signed-off-by: Simon Schrottner <simon.schrottner@flagsmith.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #410
Part of open-feature/spec#417 — the
cross-language tracking issue. The language-agnostic artifacts are in
spec#423, their normative description in
spec#439; the report envelope in
spec#425. Java is the reference
(java-sdk-contrib#1830).
The PR chain
The same four-branch shape in every language, so understanding one is understanding all four:
reportandflagdare siblings, not a sequence: the follow-ups fill gaps without changing theadopter-facing API, so neither blocks the other and either can land first.
What this is
A conformance suite any Python provider can adopt to check that it implements the provider contract
in the specification — the Python implementation of Appendix F, running the same
Gherkin scenarios, against the same canonical flag set, driven through the same control API as every
other language's TCK. That shared basis is the point: "conformant" only means something if the
question is identical everywhere.
It uses pytest-bdd, the same runner
openfeature-provider-flagdandopenfeature-flagd-api-testkitalready depend on, so an adopting package gains no new testframework.
The adoption surface
Two fixtures and one call:
No
conftest.py, and nothing to import for the steps. The vocabulary ships as a pytest pluginregistered through a
pytest11entry point. That is the nicest adoption of the four, and it is notan accident — pytest-bdd resolves steps through the fixture system, and fixtures from an installed
plugin are visible to every test.
The suite owns the container stack. You name a Compose file and the ports the provider connects
to; starting Compose, discovering dynamically mapped host ports, building the HTTP control, waiting
until it accepts commands and tearing down are the suite's job. Conventions with working defaults:
backend_service(backend),control_port(8080),additional_ports,backend_configuration(
default),startup_timeout(60s). This is now required by Appendix F, after three of fourlanguages originally shipped only the control-API client and left orchestration to the adopter — at
which point every adoption hand-rolled the same wrapper.
A provider with no backend supplies its own
BackendControlinstead and needs no Compose fileand no container tooling:
testcontainersis an optional extra, imported lazily. The feature filesand canonical flag set are packaged with the distribution, so adopting needs no git submodule.
Declaring what a provider can and cannot do
An unrecognised canonical tag now fails collection. Appendix F requires it. An unknown tag gates
nothing, so every scenario carrying one stays mandatory for every adopter — a package that has not
learned a new capability does not report a new capability, it silently keeps demanding the old
behaviour. Only canonical tags are read; an adopter's extension tags gate nothing by design, and
reading them as a vocabulary would fail every run that has an extension.
The revision check no longer skips when it cannot read the pin. It used to, and the environment
where the pin was unreadable was exactly the one where a rebase leaves the assets stale — so the check
went quiet where it was load-bearing. The cause turned out to be resolvable rather than fundamental: a
worktree records its git directory as an absolute path, which WSL cannot follow, while a submodule's
is relative, so the pin reads fine from the submodule side. The condition is gone rather than merely
reported. An unpacked distribution is deliberately still exempt: it has no pin and no working tree
that could have drifted, which is a check with nothing to check rather than one that cannot run.
capabilitiesnames the optional parts of the contract this provider supports, from sixteenmembers —
EVENTS,LIFECYCLE,STALE,CONFIGURATION_CHANGE,OBJECT,VARIANTS,DISABLED_FLAGS,UNAVAILABLE_INIT,NUMERIC_COERCION,LARGE_INTEGERS,REINITIALIZATION,TARGETING,STANDARD_REASONS,STRING_TYPING,FULLY_TYPED_VALUES, andCACHING, which isreserved and cannot be declared. A
scenario gated on an undeclared capability is reported skipped with its reason, never passed:
known_deviationsis the separate statement that the provider fails something it is required to do,in either a
trackedor anuntrackedform. Appendix F settles which of its two shapes to prefer:declare the capability, let the scenario fail, and record the deviation beside the failure — rather
than withdrawing the capability so the scenario skips, which hides a defect behind something that
looks deliberate.
reasonis asserted in two places only, and they mean different things. Requirement 2.2.5 is aSHOULDpermitting "some other string", so the scenarios that resolve a value no longer assert areason at all — an exact-match assertion there fails providers the specification permits. What
replaced it is a capability.
STANDARD_REASONSdeclares that this provider speaks thespecification's own vocabulary, and a feature file then checks the whole of it:
STATIC,TARGETING_MATCH,DEFAULT,DISABLED. Undeclared, those scenarios skip with their reason, andnothing else in the suite cares what a provider calls a reason.
The error cases are not part of that bargain and are asserted unconditionally. Requirement 1.4.8
makes an error code mandatory in abnormal execution, so
ERRORthere is not a dialect — a reasonthat disagrees with the error code shipped beside it is a contradiction. The assertion is on the two
agreeing, not on which of them is authoritative.
Four things probed rather than assumed
Each verified against pytest-bdd 8.1 before the design depended on it:
@configuration-change?getattr(pytest.mark, tag)handles thempytest.skip()from an autouse fixture report skipped with the reason?scenarios()accept an absolute path into an installed package?conftest.pyTwo things that only showed up by running it, both fixed and commented:
PytestUnknownMarkWarning— noise at best, a hard failure under-W error. The plugin registersthem in
pytest_configure.request.fixturenames. pytest-bddresolves a step's fixtures lazily as each step runs, so
tck_configis not infixturenamesatsetup time, and
@unavailablescenarios ran against a config that never declared it. The gate nowkeys off the node's markers, which are on the item itself. That is exactly the failure mode the
suite exists to prevent — a gate that looks right and quietly passes everything — so it is pinned
by a test.
Self-tests, no Docker
test_in_memory_conformanceInMemoryProvider— reference adoption for a backend-less providertest_controllable_conformanceControllableInMemoryProvider— the configuration-change and lifecycle pathstest_in_process_controlInProcessControl— pins what the Gherkin cannot assert about itselftest_http_control/resetpreferred,/startfallback cached once per suite,/startafter a disconnectThere is no multi-provider suite because Python has no multi-provider — worth noting as its own gap.
Findings against the SDK
A boolean satisfies an Integer request, and this is Python-only.
boolean-flagthroughget_integer_detailsreturnsTrue, reasonSTATIC, no error code, where the specificationrequires the code default and
TYPE_MISMATCH.boolsubclassesint, soisinstance(True, int)is
Trueand the type check passes. The identical scenario passes in Java and Go — no suite inanother language could have caught this, which is a fair advertisement for the
multiple-implementations argument.
python-sdk#619, fixed on main and not yet
released, so the row stays
xfail(strict=True): it remains visible in the report and fails themoment it starts passing, which forces the marker's removal when a release carries the fix.
The in-memory provider cannot update its flag set. Appendix A requires it. Same
class of gap as go-sdk#530, found independently
in a second SDK. Only half the machinery is missing —
AbstractProvideralready suppliesemit_provider_configuration_changed— soControllableInMemoryProvideris a small subclassrather than a reimplementation. python-sdk#620,
also fixed on main and unreleased.
A pattern worth naming, because it decides what this suite can currently assert: three of the
findings across two SDKs are fixed on main and carried by no release. #619 and #620 were closed on
2026-08-30; the newest python-sdk release is v0.10.0 from 2026-06-01, so this package resolves a
pre-fix SDK. go-sdk#552 is in the same state.
Each will need a pin bump and a marker removed, not further investigation.
InMemoryFlag.stateis never read, so a flag constructedDISABLEDresolves as though enabled.Filed as a question rather than a bug, since carrying the field for configuration compatibility is a
defensible answer: python-sdk#627. The
in-memory suites withhold
@disabled-flagson the strength of it.Two findings in the flagd provider, filed as questions
Both surfaced in this repository's own provider rather than the SDK, and both are filed for triage
rather than asserted as defects:
shutdown()leaves the gRPC connectivity watcher running, so a clean shutdown emitsCannot invoke RPC: Channel closed!after the call has returned(#419). Worth noting the
scenario passes — shutdown completes within its bound — so the results record no failure and the
report has nowhere to put this. It was found only because the
@lifecyclecapability had neverbeen declared in this adoption, so those six scenarios had never run; declaring it was a one-line
change.
(#420): in-process returns
TYPE_MISMATCHfor0.5requested as an integer, RPC returns0with no error code. Bothresolvers declare
@numeric-coercion; the known deviation is recorded against the RPC suite only,because in-process does not have the defect. Calibration worth having: Java's and Go's flagd
providers narrow in both resolvers, so this provider's in-process path is ahead of the other
implementations rather than behind them.
Running it
The self-tests run in the default build with no Docker. The adoption suites live in
tests/tck/,a sibling of
tests/e2e/— Python was already the shape the other three moved to — and the directoryis the whole of the selection:
--ignore=tests/tckexcludes them frompoe test/poe cov,pytest tests/tckruns them. No-k, nopython_filesoverride, no name-shaped glob anywhere inthe chain, which is why the three adoption modules could drop the word
conformancefrom their nameswithout a single count moving.
The default build still typechecks the adoption, which Appendix F requires:
--ignoreleaves thefiles in the tree and in the package, and
poe testends with a second pytest invocation thatcollects
tests/tckwithout running it. That step cannot pass having checked nothing — pytest exits5 on an empty collection and 4 on a path that does not exist, measured four ways — so no count
assertion is needed on top of it. It was, however, not being reached on flagd: poe aborts a
sequence at its first failing subtask, and flagd's default suite is red on this stack for the SDK
reason below, so the check the README described had never run in the package where it matters most.
Both adoptions now set
ignore_fail = "return_non_zero", which runs every subtask and stillpropagates a non-zero exit.
Appendix F carries the reasoning: an adoption suite's honest output is red while real gaps remain,
so making it a required gate forces someone to silence it — and the cheapest way to silence a
conformance suite is to stop asking the question. Note Docker is not the reason:
tests/e2eneeds Docker too and still runs in the default build.
So read the tally below rather than the green check, on the adoption PRs. Worth stating plainly:
this suite previously ran in the default build with nothing excluding it, red, unnoticed. All four
languages had a version of that, each defeated by a different mechanism.
Verification
At
01bb657f:poe test(this package)ruff check,mypyviapoeThe flagd failures are documented backend fixture gaps plus the
isinstance(True, int)consequencein the flagd provider's own type check; the OFREP ones are
python-sdk-contrib#418 —
variantand
reasonindexed unconditionally on a response where the OFREP schema makes both optional.A red flagd job on this PR, which is not this PR's suite
Worth reading before the CI result: the flagd job is red on this branch and
tests/tckis notwhy.
openfeature-tckrequiresopenfeature-sdk>=0.10.0. Auvworkspace resolves one version permember, so
uv.lockmoves from 0.8.4 to 0.10.0 across the workspace — and 20 of the flagdprovider's existing tests fail on 0.10.0. Measured both ways in the same venv:
20 failed, 9 passedat 0.10.0,29 passedafter installing 0.8.4 back.build.yml's path filter listsuv.lockfor every package, so this PR runs the flagd job and that job goes red on a branch thattouches no flagd source.
There is a second casualty, and an earlier revision of this description wrongly said there was
not.
unleashfailspoe mypyon 0.10.0: itstrack()istrack(self, event_name, event_details)where the SDK'sAbstractProvider.trackis nowtrack(self, tracking_event_name, evaluation_context, tracking_event_details), so the override isincompatible. It surfaces only on the 3.14 job because that is
TARGET_PYTHON_VERSION, the only onethat runs
poe mypy— which is why a test-only sweep missed it. The remaining packages — aws-ssm,env-var, flipt, flagd-api, flagd-api-testkit and flagd-core — do pass.
Both casualties are the same kind of change and neither is fixed here: they are provider migrations
to a newer SDK, and they want their own PRs and their own reviewers.
Deliberately not fixed here: it is a provider migration to a newer SDK, and it wants its own PR
and its own reviewer. Flagging it because a reviewer seeing red on this branch would reasonably
assume the conformance suite caused it.
Known gaps
targeting-key-flagcatches aprovider that drops the context entirely; a provider that forwards the targeting key and silently
discards every other attribute still passes. Closing it needs an echo operation on the control
API, or a canonical flag whose rule keys on a custom attribute.
@cachingis reserved, no scenarios yet.Open questions
branch checks fields by hand and validates out of band. The schema is what makes reports
comparable, so this is the weakest link in the design and it wants one answer rather than four.
ControllableInMemoryProviderlive here at all, or should the SDK fix land first andthis package depend on it?