Skip to content

ci: scope expensive tests to affected changes - #20727

Draft
mabdinur wants to merge 3 commits into
mainfrom
codex/ci-path-select-expensive-jobs
Draft

mabdinur wants to merge 3 commits into
mainfrom
codex/ci-path-select-expensive-jobs

Conversation

@mabdinur

@mabdinur mabdinur commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Scope

This PR path-selects several expensive non-suitespec CI families. System Tests are the largest part, but not the only part.

It does not change dd-trace-py's unit/integration selection in tests/suitespec.yml or scripts/needs_testrun.py. Suitespec tests keep their existing behavior.

flowchart TD
    Change["Change"] --> Suitespec["ddtrace suitespec tests<br/>unchanged"]
    Change --> Select["Non-suitespec CI selection"]
    Select --> System["System Tests"]
    Select --> Serverless["Serverless build / test"]
    Select --> Smoke["Package and multi-OS smoke"]
    Select --> Native["Native tests and lint"]
    Select --> Compat["Compatibility jobs"]
    Change -->|"Merge queue / full pipeline"| Full["All automated coverage"]
Loading

What changes

  • System Tests split into SSI smoke, Profiling, AppSec, and SSI environment groups with feature triggers.
  • Serverless wheel builds, uploads, validation, Lambda tests, and benchmarks share one trigger.
  • Package smoke and multi-OS wheel tests run for relevant artifact inputs.
  • Rust CI, profiling clang-tidy, profiling-native, and lib-injection tests retain focused triggers.
  • Debugging exploration and profiling correctness remain selective on ordinary PRs.
  • Merge queues, main, releases, nightly, and scheduled pipelines run everything. FULL_TEST_SUITE=true is the escape hatch.

Expected impact

Based on representative durations and 874 main commits over 90 days:

Measure Relative change (absolute)
Ordinary-PR runner time across affected jobs 87% lower [18.2 → 2.3 hours]
System Test runner time 90% lower [13.9 → 1.4 hours]
Serverless runner time 89% lower [3.23 → 0.35 hours]
Other compatibility/smoke jobs 73% lower [0.72 → 0.19 hours]
One PR plus one merge-queue run 39% lower [36.4 → 22.4 hours]

At equal PR and queue volume, modeled cost is 39% lower [about 7,000 runner-hours saved per 500 of each].

Validation

YAML, targeted lint/spelling, and profiling-native coverage passed. This CI-changing PR intentionally runs full coverage.

@mabdinur mabdinur added the changelog/no-changelog A changelog entry is not required for this PR. label Oct 1, 2026
@datadog-official

datadog-official Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Tests

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

🔄 Datadog retried 3 tests - 3 passed on retry View in Datadog

🚧 26 tests that failed were ignored due to quarantine View in Datadog

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 97239fb | Docs | View more details | Give us feedback!

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codeowners resolved as

Resolved from the full PR diff against main using the target branch CODEOWNERS file.
CODEOWNERS team requests not listed below are not required by the current file set.

.gitlab-ci.yml                                                          @DataDog/python-guild @DataDog/apm-core-python
.gitlab/benchmarks/serverless.yml                                       @DataDog/apm-ecosystems-performance @DataDog/apm-core-python
.gitlab/debugging-exploration.yml                                       @DataDog/python-guild @DataDog/apm-core-python
.gitlab/multi-os-tests.yml                                              @DataDog/python-guild @DataDog/apm-core-python
.gitlab/native.yml                                                      @DataDog/python-guild @DataDog/apm-core-python
.gitlab/package.yml                                                     @DataDog/python-guild @DataDog/apm-core-python

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

Circular import analysis

⚠️ Existing circular imports

There are 1 circular imports that already exist on the base branch and have not been changed by this PR.

ddtrace.errortracking._handled_exceptions.bytecode_injector -> ddtrace.errortracking._handled_exceptions.callbacks -> ddtrace.errortracking._handled_exceptions.collector -> ddtrace.errortracking._handled_exceptions.bytecode_reporting -> ddtrace.errortracking._handled_exceptions.bytecode_injector

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Dependency direction analysis

⚠️ Existing dependency direction violations

There are 201 dependency direction violations that already exist on the base branch and have not been changed by this PR.

Show existing violations (showing 5 of 201 highest severity)
ddtrace.internal.tracemethods -×-> ddtrace.trace  (internal-core -> product:tracing, score=132)
ddtrace.llmobs._utils -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)
ddtrace.appsec._listeners -×-> ddtrace.trace  (product:appsec -> product:tracing, score=130)
ddtrace.llmobs._integrations.google_adk -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)
ddtrace.appsec._contrib.django -×-> ddtrace.trace  (product:appsec -> product:tracing, score=130)

To see all violations, download the layers-base.json and layers-pr.json artifacts from this CI job and run:

uv run --script scripts/import-analysis/layers.py compare layers-base.json layers-pr.json

@mabdinur
mabdinur force-pushed the codex/ci-path-select-expensive-jobs branch from c57cff6 to 98099b4 Compare October 1, 2026 18:38
@pr-commenter

pr-commenter Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Benchmarks

Benchmark execution time: 2026-10-01 20:13:44

Comparing candidate commit 97239fb in PR branch codex/ci-path-select-expensive-jobs with baseline commit 9d50cdf in branch main.

📊 Benchmarking dashboard

Found 0 performance improvements and 7 performance regressions! Performance is the same for 598 metrics, 10 unstable metrics, 8 known flaky benchmarks, 16 flaky benchmarks without significant changes.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:httppropagationextract-empty_headers

  • 🟥 execution_time [+101.702ns; +126.026ns] or [+13.814%; +17.118%]

scenario:iastaspects-format_noaspect

  • 🟥 execution_time [+38.382µs; +43.699µs] or [+12.812%; +14.587%]

scenario:iastaspectsremodule-re_expand_aspect

  • 🟥 execution_time [+82.171µs; +89.622µs] or [+13.153%; +14.345%]

scenario:msgpackencoderscenario-simple_one_span

  • 🟥 execution_time [+512.949ns; +572.605ns] or [+12.487%; +13.939%]

scenario:openfeatureflagevaluation-hook-enqueue-typical

  • 🟥 execution_time [+1.672µs; +1.810µs] or [+7.649%; +8.279%]

scenario:otelspan-start

  • 🟥 execution_time [+1.794ms; +2.635ms] or [+7.147%; +10.500%]

scenario:recursivecomputation-shallow

  • 🟥 execution_time [+50.522µs; +54.987µs] or [+7.144%; +7.775%]

Unstable benchmarks

These benchmarks have a confidence interval too wide to call a change; treat them as noise rather than signal.

scenario:coreapiscenario-context_with_data_listeners

  • unstable execution_time [-681.189ns; +822.100ns] or [-6.521%; +7.870%]

scenario:coreapiscenario-core_dispatch_1_listener

  • unstable execution_time [-39.546ns; +38.746ns] or [-5.934%; +5.813%]

scenario:coreapiscenario-core_dispatch_50_listeners

  • unstable execution_time [-1944.682ns; +1929.761ns] or [-9.715%; +9.641%]

scenario:coreapiscenario-core_dispatch_exception_listeners

  • unstable execution_time [-1714.028ns; +1935.450ns] or [-9.103%; +10.279%]

scenario:coreapiscenario-core_dispatch_listeners

  • unstable execution_time [-385.268ns; +379.606ns] or [-8.979%; +8.847%]

scenario:coreapiscenario-core_dispatch_no_args_listeners

  • unstable execution_time [-227.931ns; +232.769ns] or [-8.415%; +8.594%]

scenario:coreapiscenario-core_dispatch_with_results_1_listener

  • unstable execution_time [-94.892ns; +94.272ns] or [-6.846%; +6.801%]

scenario:coreapiscenario-core_dispatch_with_results_50_listeners

  • unstable execution_time [-4693.324ns; +4627.567ns] or [-9.716%; +9.580%]

scenario:coreapiscenario-core_dispatch_with_results_listeners

  • unstable execution_time [-968.440ns; +930.912ns] or [-9.459%; +9.092%]

scenario:packagesupdateimporteddependencies-import_many_stdlib_cached

  • unstable execution_time [-61.297µs; +47.950µs] or [-10.729%; +8.393%]

Known flaky benchmarks

These benchmarks are marked as flaky and will not trigger a failure. Modify FLAKY_BENCHMARKS_REGEX to control which benchmarks are marked as flaky.

scenario:httppropagationinject-ids_only

  • 🟥 execution_time [+3.019µs; +3.138µs] or [+21.177%; +22.013%]

scenario:iastaspects-lower_aspect

  • 🟥 execution_time [+99.649µs; +104.479µs] or [+41.383%; +43.389%]

scenario:iastaspects-upper_noaspect

  • 🟥 execution_time [+35.144µs; +39.130µs] or [+20.672%; +23.016%]

scenario:iastaspectsospath-ospathbasename_aspect

  • 🟥 execution_time [+148.117µs; +153.713µs] or [+39.792%; +41.296%]

scenario:iastaspectssplit-rsplit_aspect

  • 🟥 execution_time [+27.353µs; +32.570µs] or [+17.762%; +21.149%]

scenario:span-start

  • 🟥 execution_time [+1.160ms; +1.585ms] or [+9.073%; +12.400%]

scenario:telemetryaddmetric-1-count-metric-1-times

  • 🟥 execution_time [+223.473ns; +266.427ns] or [+11.344%; +13.524%]

scenario:tracer-small

  • 🟥 execution_time [+41.123µs; +42.698µs] or [+16.426%; +17.055%]

Known flaky benchmarks without significant changes:

  • scenario:errortrackingflasksqli-baseline
  • scenario:flasksimple-iast-get
  • scenario:iastaspects-casefold_aspect
  • scenario:iastaspects-casefold_noaspect
  • scenario:iastaspects-index_aspect
  • scenario:iastaspects-ljust_noaspect
  • scenario:iastaspects-replace_aspect
  • scenario:iastaspects-rstrip_aspect
  • scenario:iastaspects-swapcase_aspect
  • scenario:iastaspects-title_noaspect
  • scenario:iastaspects-translate_aspect
  • scenario:iastaspects-translate_noaspect
  • scenario:packagespackageforrootmodulemapping-cache_off
  • scenario:packagespackageforrootmodulemapping-cache_on
  • scenario:sethttpmeta-all-enabled
  • scenario:telemetryaddmetric-record-100-metrics

@mabdinur
mabdinur force-pushed the codex/ci-path-select-expensive-jobs branch from 98099b4 to 0433583 Compare October 1, 2026 19:26
@mabdinur
mabdinur force-pushed the codex/ci-path-select-expensive-jobs branch from 0433583 to 97239fb Compare October 1, 2026 19:46
@mabdinur mabdinur changed the title ci: scope expensive tests to affected changes ci: scope expensive system tests to affected changes Oct 1, 2026
@mabdinur mabdinur changed the title ci: scope expensive system tests to affected changes ci: scope system tests to affected changes Oct 1, 2026
@mabdinur mabdinur changed the title ci: scope system tests to affected changes ci: scope expensive tests to affected changes Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/no-changelog A changelog entry is not required for this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant