Conversation
…dotnet # Conflicts: # Datadog.Trace.Build.g.sln # Datadog.Trace.sln
Execution-Time Benchmarks Report ⏱️Execution-time results for samples comparing This PR (9351) and master. ✅ No regressions detected |
BenchmarksBenchmark execution time: 2026-10-02 01:55:47 Comparing candidate commit 7584f5a in PR branch Found 0 performance improvements and 10 performance regressions! Performance is the same for 62 metrics, 0 unstable metrics, 75 known flaky benchmarks, 48 flaky benchmarks without significant changes.
|
Summary of changes
flagevaluationEVP events for instrumented Datadog OpenFeature evaluations, including unsuccessful evaluations. These events show how flags are evaluated without sending raw targeting keys or evaluation context by default.masterindependently of #9235. Replaces the approach in the stale #9139 draft; it does not import or require Leo’s transport work.Reason for change
Provide evaluation visibility while protecting customer data by default and keeping telemetry failures separate from flag evaluation behavior. Follows the merged implementations in Go, Java, Python, and Ruby, with PII guidance and the JS implementation as additional references.
Implementation details
Key decisions / assumptions
observeFullEvaluationDatavalue and evaluation timestamp when evaluating the flag. Later configuration changes must not change that observation’s privacy treatment. Only booleantrueopts into full data.Measured performance overhead
These are local engineering measurements, not final-head measurements or a claim of performance equivalence. Baseline:
2514c73b48; measured implementation:4de2e6c67c(same production code in3cef2ec0d2). They predate the two review fixes and latest-master integration.Native ARM64, Linux/musl, .NET 10.0.12; real OpenFeature/native-instrumentation path and a responsive local Agent fixture. Each burst run measured 100,000 evaluations, 64 subjects and eight attributes after warm-up. Values are medians of three serial runs. CPUs mean affinity slots, not container quotas. Caller latency excludes final flush; CPU and allocation include all process threads and final flush.
AutoResetEventdesign stays. Provider hook-list reuse is a separate general optimization tracked in FFL-3356.Test coverage
7160ad3b03: 788 targeted tracer tests + 74 provider tests passed; one platform-specific skip. All four tracer and four provider framework builds, plus master’s startup-hook project, passed without warnings/errors. These checks ran on the resolved tree before its signed merge commit.02b6a03698: full tracer suite 13,390 passed / 70 skipped / zero failed; regression coverage for worker request-context retention and metadata-only error codes.3cef2ec0d2: system-tests 43 passed / zero failed / eight existing XPASS / 53 deselected, including all 12 Agent EVP cases. Identical baseline harness had 31 passes and 12 expected failures for absent emission. Direct-routing cases were outside scope, not reported as passing.3cef2ec0d2, stagingdev, Agent 7.84.0: baseline/candidate results matched for 81 live and 24 controlled requests, plus six no-Agent checks. Verified emitted observations, SDK identity, strict consent, and unchanged evaluation behavior.Reviewer QA
truemay include bounded full data.DD_FLAGGING_EVALUATION_COUNTS_ENABLED=false, or disconnect the Agent: evaluations must still work. Disabled emission must produce no newflagevaluationevents. Check that existing exposure/span-enrichment behavior remains unchanged.Other details
Blast radius
DD_FLAGGING_EVALUATION_COUNTS_ENABLED=falsedisables them. Existing feature-flag enable/source settings still apply.Remaining before ready for review / merge
7160ad3b03), then push. No PR merge will be performed by the agent.Note to selfresults, and replace historical evidence above with links to those comments. Publish/coordinate the separately signed test-harness changes needed for reproducibility.