Skip to content

fix(serverless): drop stale traces on MicroVM identity refresh - #20089

Closed
litianningdatadog wants to merge 0 commit into
tianning.li/3-5-runtime-metrics-identity-refreshfrom
tianning.li/4-buffer-reset-upon-microvm-start
Closed

litianningdatadog wants to merge 0 commit into
tianning.li/3-5-runtime-metrics-identity-refreshfrom
tianning.li/4-buffer-reset-upon-microvm-start

Conversation

@litianningdatadog

Copy link
Copy Markdown
Contributor

Description

When a Lambda MicroVM refreshes its runtime identity, discard active traces and writer-buffered entities that may still contain the previous runtime ID. The reset is enabled only by the MicroVM identity-refresh path; normal configuration and post-fork recreation paths retain their existing behavior.

Testing

  • 34 runtime-ID tests passed.
  • 8 focused tracer/reset/writer tests passed.
  • Targeted formatting, style, typing, and diff checks passed.

Risks

None beyond the intended discard of pre-refresh trace data.

Additional Notes

This is a stacked draft PR on top of #19822.

@litianningdatadog
litianningdatadog force-pushed the tianning.li/4-buffer-reset-upon-microvm-start branch from bb98270 to 6e93b2f Compare September 4, 2026 21:40
@cit-pr-commenter-54b7da

Copy link
Copy Markdown

Codeowners resolved as

Resolved from the full PR diff against tianning.li/3-5-runtime-metrics-identity-refresh using the target branch CODEOWNERS file.
CODEOWNERS team requests not listed below are not required by the current file set.

ddtrace/_trace/processor/__init__.py                                    @DataDog/apm-sdk-capabilities-python
ddtrace/_trace/tracer.py                                                @DataDog/apm-sdk-capabilities-python
ddtrace/internal/writer/writer.py                                       @DataDog/apm-core-python
releasenotes/notes/aws-lambda-microvm-identity-refresh-3a672cd6bcbad16d.yaml  @DataDog/apm-python
tests/tracer/runtime/test_runtime_id.py                                 @DataDog/apm-sdk-capabilities-python
tests/tracer/test_writer.py                                             @DataDog/apm-sdk-capabilities-python

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

Circular import analysis

⚠️ Existing circular imports

There are 3 circular imports that already exist on the base branch and have not been changed by this PR.

ddtrace.errortracking._handled_exceptions.bytecode_injector -> ddtrace.errortracking._handled_exceptions.callbacks -> ddtrace.errortracking._handled_exceptions.collector -> ddtrace.errortracking._handled_exceptions.bytecode_reporting -> ddtrace.errortracking._handled_exceptions.bytecode_injector
ddtrace.llmobs -> ddtrace.llmobs._evaluators -> ddtrace.llmobs._evaluators.format -> ddtrace.llmobs._experiment -> ddtrace.llmobs
ddtrace.appsec._asm_request_context -> ddtrace.appsec._iast._iast_request_context_base -> ddtrace.appsec._iast._iast_env -> ddtrace.appsec._iast.reporter -> ddtrace.appsec._exploit_prevention.stack_traces -> ddtrace.appsec._asm_request_context

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

Dependency direction analysis

⚠️ Existing dependency direction violations

There are 231 dependency direction violations that already exist on the base branch and have not been changed by this PR.

Show existing violations (showing 5 of 231 highest severity)
ddtrace.internal.tracemethods -×-> ddtrace.trace  (internal-core -> product:tracing, score=134)
ddtrace.internal.ci_visibility.filters -×-> ddtrace.trace  (product:ci_visibility -> product:tracing, score=132)
ddtrace.llmobs._evaluators.runner -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=132)
ddtrace.llmobs._integrations.mcp -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=132)
ddtrace.llmobs._integrations.vllm -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=132)

To see all violations, download the layers-base.json and layers-pr.json artifacts from this CI job and run:

uv run --script scripts/import-analysis/layers.py compare layers-base.json layers-pr.json

@datadog-prod-us1-6

datadog-prod-us1-6 Bot commented Sep 4, 2026 •

Copy link
Copy Markdown

Pipelines  Tests

✨ Unblock PR with BitsAI

⚠️ Warnings

❌ Your PR has failed checks. Please review the issues below and take necessary action before merging.

🚦 3 Pipeline jobs failed

DataDog/apm-reliability/dd-trace-py | prechecks — 🔧 Needs a code fix, caused by this PR

View more details · View in GitLab

Changelog | Validate changelog

View more details · View in GitHub Actions

Release note not found. Add a new note or label 'changelog/no-changelog' to skip validation.

DataDog/apm-reliability/dd-trace-py | download_dependency_wheels: [3.15.0rc1, 3.15]

View more details · View in GitLab

ℹ️ Info

No other issues found (see more)

🧪 All tests passed
❄️ No new flaky tests detected

Useful? React with 👍 / 👎

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 6e93b2f | Docs | View more details | Give us feedback!

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Three critical stale-data and synchronization issues must be resolved before approval.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

Prevents stale runtime-ID trace data from surviving AWS Lambda MicroVM identity refreshes.

Changes:

  • Resets active trace and writer buffers during identity refresh.
  • Adds writer-level buffer-discard support and focused tests.
  • Updates the release note.

Three critical issues remain: in-flight trace completion can race the reset, live pre-refresh spans can repopulate the aggregator, and native exporter shutdown can emit old-identity data.

File summaries
File Description
tests/tracer/test_writer.py Tests native writer buffer clearing.
tests/tracer/runtime/test_runtime_id.py Tests identity-refresh trace cleanup.
releasenotes/notes/aws-lambda-microvm-identity-refresh-3a672cd6bcbad16d.yaml Documents stale-trace prevention.
ddtrace/internal/writer/writer.py Adds buffer discard, but normal exporter shutdown can still emit buffered native stats with the old runtime ID.
ddtrace/_trace/tracer.py Enables reset during refresh, but live pre-refresh spans can later produce stale or orphaned trace fragments.
ddtrace/_trace/processor/__init__.py Coordinates reset behavior, but trace completion and writer replacement are not atomic.
Review details

Suppressed comments (1)

ddtrace/_trace/tracer.py:437

  • reset_buffer=True clears the aggregator map but leaves the current execution's active Span/Context intact. If the MicroVM snapshot restores an active old span, the first post-refresh request will still be parented to it; start_span() only adds runtime-id when there is no local parent, so that trace can be written by the new writer without the refreshed runtime ID. Clear the active context as part of this identity reset.
        self._recreate(reset_buffer=True, drop_buffered_traces=True)
  • Files reviewed: 6/6 changed files
  • Comments generated: 3
  • Review effort level: Balanced

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +590 to +591
elif drop_buffered_traces:
self.writer.drop_buffered_traces()
Comment thread ddtrace/_trace/tracer.py

def _refresh_runtime_identity(self, _runtime_id: str) -> None:
self._recreate(reset_buffer=False)
self._recreate(reset_buffer=True, drop_buffered_traces=True)
Comment on lines +1166 to +1168
def drop_buffered_traces(self) -> None:
for client in self._clients:
getattr(client.encoder, "flush")()
@pr-commenter

pr-commenter Bot commented Sep 4, 2026 •

Copy link
Copy Markdown

Benchmarks

Benchmark execution time: 2026-09-07 20:59:14

Comparing candidate commit 057569d in PR branch tianning.li/4-buffer-reset-upon-microvm-start with baseline commit 2b77209 in branch main.

📊 Benchmarking dashboard

Found 0 performance improvements and 7 performance regressions! Performance is the same for 577 metrics, 10 unstable metrics, 3 known flaky benchmarks, 15 flaky benchmarks without significant changes.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:httppropagationinject-ids_only

  • 🟥 execution_time [+2.722µs; +2.882µs] or [+16.515%; +17.489%]

scenario:iastaspects-join_aspect

  • 🟥 execution_time [+44.479µs; +48.868µs] or [+20.456%; +22.474%]

scenario:iastaspects-title_aspect

  • 🟥 execution_time [+62.443µs; +68.216µs] or [+22.732%; +24.833%]

scenario:iastaspectsospath-ospathbasename_aspect

  • 🟥 execution_time [+121.914µs; +128.130µs] or [+29.486%; +30.989%]

scenario:iastaspectssplit-rsplit_aspect

  • 🟥 execution_time [+17.381µs; +21.044µs] or [+12.385%; +14.995%]

scenario:telemetryaddmetric-1-count-metric-1-times

  • 🟥 execution_time [+617.346ns; +662.952ns] or [+22.562%; +24.229%]

scenario:tracer-small

  • 🟥 execution_time [+35.080µs; +37.930µs] or [+10.206%; +11.035%]

Unstable benchmarks

These benchmarks have a confidence interval too wide to call a change; treat them as noise rather than signal.

scenario:coreapiscenario-context_with_data_listeners

  • unstable execution_time [-822.433ns; +655.617ns] or [-7.355%; +5.863%]

scenario:coreapiscenario-core_dispatch_1_listener

  • unstable execution_time [-32.549ns; +34.054ns] or [-5.274%; +5.518%]

scenario:coreapiscenario-core_dispatch_50_listeners

  • unstable execution_time [-1547.500ns; +1756.142ns] or [-9.052%; +10.272%]

scenario:coreapiscenario-core_dispatch_exception_listeners

  • unstable execution_time [-1178.072ns; +1269.955ns] or [-9.040%; +9.746%]

scenario:coreapiscenario-core_dispatch_listeners

  • unstable execution_time [-331.757ns; +320.578ns] or [-8.971%; +8.668%]

scenario:coreapiscenario-core_dispatch_no_args_listeners

  • unstable execution_time [-249.758ns; +253.239ns] or [-8.475%; +8.593%]

scenario:coreapiscenario-core_dispatch_with_results_1_listener

  • unstable execution_time [-76.430ns; +72.202ns] or [-6.675%; +6.306%]

scenario:coreapiscenario-core_dispatch_with_results_50_listeners

  • unstable execution_time [-3795.126ns; +4312.635ns] or [-9.207%; +10.462%]

scenario:coreapiscenario-core_dispatch_with_results_listeners

  • unstable execution_time [-660.924ns; +893.116ns] or [-8.192%; +11.070%]

scenario:packagesupdateimporteddependencies-import_many_stdlib_cached

  • unstable execution_time [-52.574µs; +58.457µs] or [-9.084%; +10.101%]

Known flaky benchmarks

These benchmarks are marked as flaky and will not trigger a failure. Modify FLAKY_BENCHMARKS_REGEX to control which benchmarks are marked as flaky.

scenario:iastaspects-casefold_noaspect

  • 🟥 execution_time [+37.264µs; +43.418µs] or [+15.072%; +17.561%]

scenario:iastaspects-ljust_noaspect

  • 🟥 execution_time [+31.884µs; +40.686µs] or [+10.795%; +13.775%]

scenario:span-start

  • 🟥 execution_time [+1.825ms; +1.990ms] or [+12.928%; +14.103%]

Known flaky benchmarks without significant changes:

  • scenario:errortrackingflasksqli-baseline
  • scenario:flasksimple-iast-get
  • scenario:iastaspects-casefold_aspect
  • scenario:iastaspects-index_aspect
  • scenario:iastaspects-lower_aspect
  • scenario:iastaspects-replace_aspect
  • scenario:iastaspects-swapcase_aspect
  • scenario:iastaspects-title_noaspect
  • scenario:iastaspects-translate_aspect
  • scenario:iastaspects-translate_noaspect
  • scenario:iastaspects-upper_noaspect
  • scenario:packagespackageforrootmodulemapping-cache_off
  • scenario:packagespackageforrootmodulemapping-cache_on
  • scenario:sethttpmeta-all-enabled
  • scenario:telemetryaddmetric-record-100-metrics

@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-5-runtime-metrics-identity-refresh branch from 8bb966d to eb6e4ff Compare September 7, 2026 20:23
@litianningdatadog
litianningdatadog force-pushed the tianning.li/4-buffer-reset-upon-microvm-start branch from 6e93b2f to eb6e4ff Compare September 7, 2026 20:23
@litianningdatadog

Copy link
Copy Markdown
Contributor Author

Absorbed into the amended and rebased stack: the stale trace and writer-buffer reset changes now belong to PR #19820, with PR #19821 and #19822 rebased on top. The final tree is preserved in tianning.li/4-buffer-reset-upon-microvm-start.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants