Skip to content

chore(runtime): refresh runtime metric identity and collector state - #19822

Open
litianningdatadog wants to merge 1 commit into
tianning.li/3-4-telemetry-identity-refreshfrom
tianning.li/3-5-runtime-metrics-identity-refresh
Open

litianningdatadog wants to merge 1 commit into
tianning.li/3-4-telemetry-identity-refreshfrom
tianning.li/3-5-runtime-metrics-identity-refresh

Conversation

@litianningdatadog

@litianningdatadog litianningdatadog commented Aug 23, 2026 •

Copy link
Copy Markdown
Contributor

Stacked PRs:

Description

After a MicroVM identity refresh, runtime metrics must use the new runtime identity and
must not mix a pre-refresh measurement interval with post-refresh process activity. This
PR covers both halves of that:

Platform tags. When runtime-ID tagging is enabled, runtime metrics recollect platform
tags during the periodic runtime-metrics flush instead of keeping the tag list built before
/run, so a MicroVM identity refresh is reflected on the next flush. When runtime-ID
tagging is disabled, platform tags stay cached (no behavior change).

Collector state. A MicroVM /run transition rotates the runtime ID mid-process without
a fork. Left alone, the CPU time, context-switch, wall-clock, and GC pause/collections
collectors would report a delta spanning both the old and new identity on their next sample.
RuntimeWorker now resets those collectors' baselines through the existing
on_runtime_identity_refresh callback, registered only when the process is a MicroVM and
runtime-ID tagging is enabled; it's unregistered on stop(). Non-MicroVM processes and
MicroVMs without runtime-ID tagging register nothing and pay no extra cost.

Refresh synchronization. The worker shares the fork-safe MicroVM refresh lock from #19939.
It holds that lock across collector reset and the full runtime-metrics flush, including tag
collection, sampling, and sending, so identity rotation cannot overlap an in-flight export.
The lock is not used by non-MicroVM or runtime-ID-disabled paths.

This does not add a second identity-refresh callback/guard, and does not touch trace,
writer, or telemetry-worker lifecycle code.

Reference

Testing

Added focused subprocess/unit coverage for:

  • replacing the old runtime-ID tag after identity refresh
  • retaining the current runtime-ID on subsequent flushes
  • keeping platform tags cached when runtime-ID tagging is disabled
  • serializing MicroVM identity rotation with an in-flight runtime-metrics flush
  • a MicroVM identity refresh resetting collector state (RuntimeMetrics.reset())
  • non-MicroVM and MicroVM-without-runtime-ID processes never registering the reset hook
  • RuntimeWorker.stop() unregistering the hook
  • GCRuntimeMetricCollector/NativeProcessMetricCollector.reset() discarding pre-refresh
    pause/collections and CPU-time/context-switch/wall-clock state

The previous test that forced refreshes to interleave inside a flush was replaced because the
shared lock makes that interleaving impossible; the new regression test verifies the intended
serialization boundary directly.

Validation:

  • scripts/run-tests --venv 16cc321 -- -k 'test_runtime_metrics_microvm_flush_serializes_identity_rotation or test_runtime_metrics_refresh_identity_updates_runtime_id_tag or test_runtime_metrics_microvm_identity_refresh_resets_collectors or test_runtime_metrics_non_microvm_identity_refresh_does_not_reset_collectors' tests/runtime/test_runtime_metrics_api.py (4 passed)
  • scripts/lint fmt -- ddtrace/internal/runtime/runtime_metrics.py tests/runtime/test_runtime_metrics_api.py
  • scripts/lint style -- ddtrace/internal/runtime/runtime_metrics.py tests/runtime/test_runtime_metrics_api.py
  • scripts/lint typing -- ddtrace/internal/runtime/runtime_metrics.py tests/runtime/test_runtime_metrics_api.py
  • git diff --check

Risks

Limited to runtime-metrics platform-tag collection and the two metric collectors' interval
state. The shared lock adds synchronization only for MicroVMs with runtime-ID tagging; the
runtime-ID-disabled and non-MicroVM paths are unchanged. A flush now waits for an in-flight
MicroVM identity transition, and a transition waits for an in-flight metrics flush, preventing
partial identity attribution.

@litianningdatadog litianningdatadog added changelog/no-changelog A changelog entry is not required for this PR. aws-microvm Work related to AWS MicroVM onboarding labels Aug 23, 2026
@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Aug 23, 2026 •

Copy link
Copy Markdown

Circular import analysis

⚠️ Existing circular imports

There are 1 circular imports that already exist on the base branch and have not been changed by this PR.

ddtrace.errortracking._handled_exceptions.bytecode_injector -> ddtrace.errortracking._handled_exceptions.callbacks -> ddtrace.errortracking._handled_exceptions.collector -> ddtrace.errortracking._handled_exceptions.bytecode_reporting -> ddtrace.errortracking._handled_exceptions.bytecode_injector

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Aug 23, 2026 •

Copy link
Copy Markdown

Dependency direction analysis

⚠️ Existing dependency direction violations

There are 201 dependency direction violations that already exist on the base branch and have not been changed by this PR.

Show existing violations (showing 5 of 201 highest severity)
ddtrace.internal.tracemethods -×-> ddtrace.trace  (internal-core -> product:tracing, score=132)
ddtrace.llmobs._telemetry -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)
ddtrace.llmobs._integrations.llama_index -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)
ddtrace.llmobs._integrations.mcp -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)
ddtrace.internal.test_visibility.api -×-> ddtrace.trace  (product:ci_visibility -> product:tracing, score=130)

To see all violations, download the layers-base.json and layers-pr.json artifacts from this CI job and run:

uv run --script scripts/import-analysis/layers.py compare layers-base.json layers-pr.json

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Aug 23, 2026 •

Copy link
Copy Markdown

Codeowners resolved as

Resolved from the full PR diff against tianning.li/3-4-telemetry-identity-refresh using the target branch CODEOWNERS file.
CODEOWNERS team requests not listed below are not required by the current file set.

ddtrace/internal/runtime/collector.py                                   @DataDog/apm-sdk-capabilities-python
ddtrace/internal/runtime/metric_collectors.py                           @DataDog/apm-sdk-capabilities-python
ddtrace/internal/runtime/runtime_metrics.py                             @DataDog/apm-sdk-capabilities-python
releasenotes/notes/fix-runtime-metrics-microvm-identity-refresh-57f069db02623f9d.yaml  @DataDog/apm-python
tests/runtime/test_runtime_metrics_api.py                               @DataDog/apm-sdk-capabilities-python
tests/tracer/runtime/test_metric_collectors.py                          @DataDog/apm-sdk-capabilities-python

@litianningdatadog litianningdatadog changed the title fix(runtime): refresh runtime metric identity tags chore(runtime): refresh runtime metric identity tags Aug 23, 2026
@datadog-prod-us1-4

datadog-prod-us1-4 Bot commented Aug 23, 2026 •

Copy link
Copy Markdown
Contributor

Tests

✅ All CI checks and tests passed.

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 8a7385c | Docs | View more details | Give us feedback!

@pr-commenter

pr-commenter Bot commented Aug 23, 2026 •

Copy link
Copy Markdown

Benchmarks

Benchmark execution time: 2026-10-01 18:48:52

Comparing candidate commit 8a7385c in PR branch tianning.li/3-5-runtime-metrics-identity-refresh with baseline commit 8529b92 in branch tianning.li/3-4-telemetry-identity-refresh.

📊 Benchmarking dashboard

Found 0 performance improvements and 4 performance regressions! Performance is the same for 365 metrics, 9 unstable metrics, 4 known flaky benchmarks, 4 flaky benchmarks without significant changes.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:httppropagationextract-b3_headers

  • 🟥 execution_time [+924.829ns; +1003.602ns] or [+12.238%; +13.280%]

scenario:httppropagationextract-empty_headers

  • 🟥 execution_time [+74.267ns; +90.642ns] or [+9.892%; +12.073%]

scenario:msgpackencoderscenario-simple_one_span

  • 🟥 execution_time [+551.415ns; +606.959ns] or [+13.693%; +15.073%]

scenario:recursivecomputation-shallow

  • 🟥 execution_time [+51.885µs; +55.252µs] or [+7.229%; +7.698%]

Unstable benchmarks

These benchmarks have a confidence interval too wide to call a change; treat them as noise rather than signal.

scenario:coreapiscenario-context_with_data_listeners

  • unstable execution_time [-807.748ns; +652.060ns] or [-7.720%; +6.232%]

scenario:coreapiscenario-core_dispatch_1_listener

  • unstable execution_time [-35.735ns; +42.649ns] or [-5.412%; +6.459%]

scenario:coreapiscenario-core_dispatch_50_listeners

  • unstable execution_time [-1946.277ns; +1880.869ns] or [-9.803%; +9.473%]

scenario:coreapiscenario-core_dispatch_exception_listeners

  • unstable execution_time [-1629.789ns; +1992.013ns] or [-8.666%; +10.592%]

scenario:coreapiscenario-core_dispatch_listeners

  • unstable execution_time [-398.550ns; +362.735ns] or [-9.359%; +8.518%]

scenario:coreapiscenario-core_dispatch_no_args_listeners

  • unstable execution_time [-242.213ns; +217.215ns] or [-9.046%; +8.112%]

scenario:coreapiscenario-core_dispatch_with_results_1_listener

  • unstable execution_time [-110.614ns; +75.264ns] or [-8.127%; +5.530%]

scenario:coreapiscenario-core_dispatch_with_results_50_listeners

  • unstable execution_time [-4241.453ns; +5054.602ns] or [-8.863%; +10.562%]

scenario:coreapiscenario-core_dispatch_with_results_listeners

  • unstable execution_time [-947.416ns; +942.669ns] or [-9.286%; +9.239%]

Known flaky benchmarks

These benchmarks are marked as flaky and will not trigger a failure. Modify FLAKY_BENCHMARKS_REGEX to control which benchmarks are marked as flaky.

scenario:httppropagationinject-ids_only

  • 🟥 execution_time [+3.015µs; +3.112µs] or [+21.509%; +22.201%]

scenario:span-start

  • 🟥 execution_time [+1.284ms; +1.729ms] or [+9.516%; +12.813%]

scenario:telemetryaddmetric-1-count-metric-1-times

  • 🟥 execution_time [+236.132ns; +273.462ns] or [+11.829%; +13.700%]

scenario:tracer-small

  • 🟥 execution_time [+42.908µs; +44.186µs] or [+16.404%; +16.893%]

Known flaky benchmarks without significant changes:

  • scenario:errortrackingflasksqli-baseline
  • scenario:flasksimple-iast-get
  • scenario:sethttpmeta-all-enabled
  • scenario:telemetryaddmetric-record-100-metrics

@litianningdatadog
litianningdatadog force-pushed the tianning.li/2-flask-web-request-starting-event branch from d59e112 to 16a5332 Compare August 24, 2026 02:23
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-5-runtime-metrics-identity-refresh branch from d6a1546 to 335f6ac Compare August 24, 2026 02:24
@litianningdatadog
litianningdatadog force-pushed the tianning.li/2-flask-web-request-starting-event branch from 16a5332 to 8dd7e8e Compare August 24, 2026 02:31
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-5-runtime-metrics-identity-refresh branch from 335f6ac to 2e625cf Compare August 24, 2026 02:31
@litianningdatadog
litianningdatadog force-pushed the tianning.li/2-flask-web-request-starting-event branch 5 times, most recently from cde3045 to a0e3c42 Compare August 24, 2026 23:56
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-5-runtime-metrics-identity-refresh branch from 2e625cf to 7d36a71 Compare August 25, 2026 13:23
@litianningdatadog
litianningdatadog requested a lite review from Copilot August 25, 2026 13:36

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR ensures runtime-metrics platform tags that include runtime identity (notably runtime-id) are refreshed when the runtime identity changes (e.g., after a MicroVM /run-triggered identity refresh), so runtime metrics emitted afterward carry the new identity.

Changes:

  • Register a runtime-id change listener in RuntimeWorker to rebuild cached platform tags when identity changes.
  • Refactor platform-tag construction into a helper and add an identity-refresh callback method.
  • Add a subprocess test asserting runtime-id: tags update after runtime.refresh_identity().

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 3 comments.

File Description
ddtrace/internal/runtime/runtime_metrics.py Builds platform tags via a helper and hooks runtime-id changes to refresh cached _platform_tags.
tests/runtime/test_runtime_metrics_api.py Adds a subprocess test covering platform-tag refresh behavior on identity refresh.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread ddtrace/internal/runtime/runtime_metrics.py Outdated
Comment thread ddtrace/internal/runtime/runtime_metrics.py Outdated
Comment thread ddtrace/internal/runtime/runtime_metrics.py Outdated
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch 3 times, most recently from f7876cb to b84f21b Compare September 24, 2026 15:37
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-5-runtime-metrics-identity-refresh branch from b009343 to 7b3934b Compare September 24, 2026 15:45
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch 2 times, most recently from 156be45 to c8b45b6 Compare September 24, 2026 17:57
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-5-runtime-metrics-identity-refresh branch from 7b3934b to 707de15 Compare September 24, 2026 18:02
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch from c8b45b6 to c641a12 Compare September 24, 2026 18:29
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-5-runtime-metrics-identity-refresh branch from 707de15 to 75336e0 Compare September 24, 2026 18:30
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch 4 times, most recently from 5c23fca to 30e6c88 Compare September 24, 2026 20:41
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-5-runtime-metrics-identity-refresh branch from 75336e0 to 4fed6d7 Compare September 24, 2026 20:44
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch from 30e6c88 to e484bf7 Compare September 25, 2026 15:10
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-5-runtime-metrics-identity-refresh branch from 4fed6d7 to c037394 Compare September 25, 2026 15:12
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-4-telemetry-identity-refresh branch from e484bf7 to fb28096 Compare September 25, 2026 16:01
@litianningdatadog
litianningdatadog force-pushed the tianning.li/3-5-runtime-metrics-identity-refresh branch from c037394 to e7d5175 Compare September 25, 2026 16:02
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-01T18:23:11.780793Z 8a7385c New commits
🔒 Security Review ✅ Completed 2026-10-01T18:24:32.689403Z 8a7385c New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e7d51759d3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread ddtrace/internal/runtime/runtime_metrics.py Outdated
Comment thread ddtrace/internal/telemetry/writer.py
Comment thread ddtrace/internal/telemetry/writer.py
Comment thread ddtrace/internal/runtime/runtime_metrics.py Outdated
Comment thread ddtrace/_trace/tracer.py
Comment thread ddtrace/internal/writer/writer.py Outdated

@datadog-prod-us1-4 datadog-prod-us1-4 Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bits Code Review: FAIL

The two most critical issues are: (1) a retry-idempotency gap in the identity-refresh coordinator — when any subscriber fails, already-completed callbacks (tracer writer recreation and runtime metrics reset) re-execute for the same runtime ID, dropping buffered traces and discarding legitimate metric deltas; and (2) a race between reset() and flush() in the runtime metrics collector that can emit negative or cross-identity deltas when an identity refresh overlaps a periodic flush.

Open Bits AI session

🤖 Bits Code Review · Commit 3d4a374 · @DataDog review to ask questions

@emmettbutler emmettbutler left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Deferring review since the base branch is not main

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: fd2f295d2c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +185 to +188
runtime_id = get_runtime_id()
self._platform_tags = self._collect_platform_tags()
runtime_metrics = list(self._runtime_metrics)
if runtime_id == get_runtime_id():

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reset collectors when the refresh callback remains pending

When the MicroVM hook rotates the runtime ID but an identity-refresh callback raises before this worker's callback runs—or this worker's own collector reset fails partway—the coordinator leaves callbacks pending for a later request while the new ID is already visible. A periodic flush between those attempts reaches this block with the old collector baselines, recollects the new runtime-id, and sends a mixed pre/post-refresh interval under that new ID. The before/after comparison cannot detect this because both reads occur while holding the refresh lock; track the ID associated with the collector baselines and reset before collection when it differs.

Useful? React with 👍 / 👎.

Comment on lines +175 to +176
with self._identity_refresh_lock:
self._runtime_metrics.reset()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reset collectors at a deterministic refresh boundary

On every successful MicroVM refresh, the runtime ID is rotated before the callbacks are invoked, and those callbacks come from an unordered set. If the tracer or telemetry rebuild runs before this callback, its CPU, context switches, and GC activity occur under the new identity but are discarded here; if this callback happens first, the same activity is retained. Runtime metrics therefore vary with callback iteration order and can omit initialization work from the new runtime, so the collector reset needs a defined position immediately at the identity boundary rather than an unordered subscriber callback.

Useful? React with 👍 / 👎.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e16f6204f3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +110 to +111
monitor.reset()
self._reset_state()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep the GC baseline and pause reset atomic

When a collection starts after monitor.reset() but before gc.get_stats() is sampled by _reset_state()—which can itself occur while allocating the returned stats/list—the monitor records that pause while the new collections baseline already includes the collection. The next flush therefore reports a positive GC pause with zero corresponding collection delta. Reset the pause window and reseed the collection baseline under one synchronization boundary so activity during the identity transition is either retained or discarded consistently.

Useful? React with 👍 / 👎.

Comment on lines +107 to +109
self._identity_refresh_lock = get_runtime_identity_refresh_lock() if self._identity_refresh_enabled else None
if self._identity_refresh_enabled:
on_runtime_identity_refresh(self._on_identity_refresh)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Register refresh handling before seeding collectors

When the public RuntimeMetrics.enable() is called concurrently with a MicroVM /run refresh, the collectors are constructed and seed their CPU/GC baselines before the callback is registered here, without holding the refresh lock. If identity rotation lands in that window, the new worker misses the transition entirely; its first flush recollects the new runtime-id tag but reports deltas whose baselines came from the previous identity. Construct and register the worker state under the shared refresh lock, or verify the identity after registration and reseed when it changed.

Useful? React with 👍 / 👎.

When runtime-ID tagging is enabled, recollect platform tags during the
periodic runtime-metrics flush instead of keeping the pre-refresh list, so
an AWS Lambda MicroVM identity refresh is reflected on the next flush.
Platform tags stay cached when runtime-ID tagging is disabled.

A MicroVM /run transition also rotates the runtime ID mid-process without
a fork, which left the CPU time, context-switch, wall-clock, and GC pause
collectors mid-interval: their next sample would mix pre- and post-refresh
activity, or (for a fresh process) reports the whole process lifetime as
one delta. Reset those collectors' baselines on the existing runtime
identity refresh callback, registered only for MicroVM processes with
runtime-ID tagging enabled. Non-MicroVM processes and MicroVMs without
runtime-ID tagging register nothing and are unaffected.

This does not add another identity-refresh callback beyond the existing
one, and does not change trace, writer, or telemetry-worker lifecycle
behavior.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

aws-microvm Work related to AWS MicroVM onboarding changelog/no-changelog A changelog entry is not required for this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants