Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
99 commits
Select commit Hold shift + click to select a range
8c03217
fix(test-optimization): cache settings request in the filesystem cach…
juan-fernandez Aug 28, 2026
d005afa
fix(test-optimization): restore pre-defer suite reporting (#10017)
juan-fernandez Aug 28, 2026
cc21079
fix(test-optimization): use a 10-second flush interval (#10018)
juan-fernandez Aug 28, 2026
d9d117c
fix(test-optimization): retry transient HTTP responses (#10039)
juan-fernandez Aug 28, 2026
c13243e
chore(deps): bump test frameworks (#10042)
juan-fernandez Aug 28, 2026
69461e4
chore(deps): bump the cloud-and-messaging group across 1 directory wi…
dependabot[bot] Aug 28, 2026
5594d5f
chore(deps): bump hono (#10036)
dependabot[bot] Aug 28, 2026
a012ec7
chore(deps): bump the testing-and-build group across 1 directory with…
dependabot[bot] Aug 28, 2026
044163c
chore(deps): bump the test-versions group across 1 directory with 4 u…
dependabot[bot] Aug 28, 2026
5f1e8d0
fix(test-optimization): extend final flush timeout (#10040)
juan-fernandez Aug 28, 2026
77756f2
fix(test-optimization): increase payload request concurrency (#10038)
juan-fernandez Aug 28, 2026
5ff181c
feat(dogstatsd): enable client self-telemetry (#9698)
BridgeAR Aug 28, 2026
cc3a97f
fix(test-optimization): cap screenshot upload retries (#10056)
BridgeAR Aug 28, 2026
a544dcd
test(install): pin standardwebhooks for Anthropic fixtures (#10054)
bm1549 Aug 28, 2026
33cc450
docs(agents): simplify repository guidance (#9969)
pabloerhard Aug 28, 2026
cd26f09
chore(deps): bump the test-versions group across 1 directory with 5 u…
dependabot[bot] Aug 31, 2026
20a8abf
chore(deps): bump the web-frameworks group across 1 directory with 3 …
dependabot[bot] Aug 31, 2026
f8aa7d2
chore(deps): bump the cloud-and-messaging group across 1 directory wi…
dependabot[bot] Aug 31, 2026
4deae9c
chore(deps): bump the testing-and-build group across 1 directory with…
dependabot[bot] Aug 31, 2026
91167e9
chore(deps): bump the test-versions group across 1 directory with 2 u…
dependabot[bot] Aug 31, 2026
21fc888
refactor(cypress): build config wrappers directly (#10059)
BridgeAR Aug 31, 2026
600bbb3
chore(deps-dev): update eslint-plugin-mocha to 12.0.2 (#9970)
BridgeAR Aug 31, 2026
5f65c0a
feat(serverless): add AWS MicroVM support (#9709)
BridgeAR Aug 31, 2026
5982eb1
test(mariadb): limit clock stubs to synchronous calls (#10077)
BridgeAR Aug 31, 2026
8850ee8
feat(graphql): trace graphql-jit executions (#9343)
BridgeAR Aug 31, 2026
34c4cc4
fix(serverless): coalesce overlapping Vercel telemetry flushes (#9954)
BridgeAR Aug 31, 2026
f472e5f
refactor(eslint): detect direct literal joins (#9926)
BridgeAR Aug 31, 2026
d2f79ec
chore(eslint): prohibit immediate call-result invocation (#10060)
BridgeAR Aug 31, 2026
849ce92
test(dns): remove live resolver dependency (#9950)
BridgeAR Aug 31, 2026
6100227
chore(deps-dev): update eslint-plugin-unicorn to 74 (#10074)
BridgeAR Aug 31, 2026
01eac8d
ci(test-optimization): cover missing source test reports (#9931)
rochdev Aug 31, 2026
1ccd051
fix(openai): properly patch api promise prototypes for complete cover…
sabrenner Aug 31, 2026
6aac6c4
refactor(eslint): name returned functions before invocation (#10083)
BridgeAR Aug 31, 2026
548b288
fix(mongodb): extract queries from document sequences (#10046)
BridgeAR Aug 31, 2026
be664a1
perf(tracer): avoid conditional object spread allocations (#10052)
BridgeAR Aug 31, 2026
4b5a7b2
chore(deps): bump @happy-dom/jest-environment (#10093)
dependabot[bot] Sep 1, 2026
037eaee
chore(deps): bump the databases group across 1 directory with 8 updat…
dependabot[bot] Sep 1, 2026
44dbc4c
chore(deps): bump the ai-and-llm group across 1 directory with 17 upd…
dependabot[bot] Sep 1, 2026
debc2cd
chore(deps): bump the cloud-and-messaging group across 1 directory wi…
dependabot[bot] Sep 1, 2026
acc89a4
fix(test-optimization): support Mocha 12 (#10095)
juan-fernandez Sep 1, 2026
40f7458
perf(test-optimization): build formatted strings directly (#10087)
BridgeAR Sep 1, 2026
6e35162
chore(deps): normalize yarn lockfile (#10097)
juan-fernandez Sep 1, 2026
a93d009
perf(ai): build provider strings directly (#10088)
BridgeAR Sep 1, 2026
0c030af
perf(tracer): build runtime strings directly (#10089)
BridgeAR Sep 1, 2026
90bb24a
fix(appsec): normalize sampling priority in the API Security sampler …
CarlesDD Sep 1, 2026
ad4bdee
refactor(eslint): enforce direct string construction (#10090)
BridgeAR Sep 1, 2026
f3ea33f
fix(agentless): pass raw container ID to trace intake (#10105)
BridgeAR Sep 1, 2026
cdb8a38
feat(instrumentation): add async Orchestrion context callbacks (#10080)
juan-fernandez Sep 2, 2026
a66c5ed
feat(test-optimization): upload browser test failure videos (#9927)
juan-fernandez Sep 2, 2026
7f22904
chore(deps): bump protobufjs (#10107)
dependabot[bot] Sep 2, 2026
7dd0e55
chore(deps): bump bullmq (#10108)
dependabot[bot] Sep 2, 2026
c653334
chore(deps): bump @happy-dom/jest-environment (#10110)
dependabot[bot] Sep 2, 2026
8a2f164
chore(deps-dev): bump the dev-minor-and-patch-dependencies group acro…
dependabot[bot] Sep 2, 2026
038818c
chore(deps): bump the gh-actions-packages group across 2 directories …
dependabot[bot] Sep 2, 2026
923cfc7
test(cypress): support Cypress 16 (#10114)
juan-fernandez Sep 2, 2026
1ccd0cc
feat(test-optimization): upload WebdriverIO failure screenshots (#10113)
juan-fernandez Sep 2, 2026
19dc01b
fix(test-optimization): bound payload delivery lifecycle (#10044)
juan-fernandez Sep 2, 2026
7442398
feat(appsec): extract API Security schemas in Lambda (#10014)
CarlesDD Sep 2, 2026
fc99f88
ci(instrumentations): use supported Confluent Node versions (#9771)
BridgeAR Sep 2, 2026
2b603ce
fix(next): trace compiled response lifecycles (#9708)
BridgeAR Sep 2, 2026
36205fc
feat(apm): add data pipeline-backed agentless mode (#9839)
rochdev Sep 2, 2026
4b45aeb
chore(deps): bump the npm_and_yarn group across 1 directory with 2 up…
dependabot[bot] Sep 2, 2026
9ff8f3c
chore(deps): bump @humanfs/node (#10120)
dependabot[bot] Sep 2, 2026
4a60caa
chore(deps): bump the npm_and_yarn group across 1 directory with 2 up…
dependabot[bot] Sep 2, 2026
9204383
fix(esbuild): resolve imports in built-in subpath wrappers (#10102)
BridgeAR Sep 2, 2026
21d01cd
fix(graphql): record reused non-JIT resolver errors (#10082)
BridgeAR Sep 2, 2026
22d7478
test(bun): expand runtime smoke coverage (#10048)
BridgeAR Sep 2, 2026
7ab3689
fix(langgraph): do not mark graph interrupt errors as actual errors (…
sabrenner Sep 2, 2026
2f15023
perf(trace-encoder): cache stable strings across payloads (#10057)
BridgeAR Sep 2, 2026
dae4406
feat(debugger): support agentless Dynamic Instrumentation (#9881)
rochdev Sep 2, 2026
255a3b4
perf(rewriter): defer transformer load to first rewrite (#10099)
joeyzhao2018 Sep 2, 2026
8aff3cc
fix(config): prefer stable OTel deployment environment (#10050)
bm1549 Sep 2, 2026
ffadf86
test(vitest): cover webdriverio browser provider (#10116)
juan-fernandez Sep 3, 2026
d06f94f
test(debugger): replace axios with fetch (#10053)
BridgeAR Sep 3, 2026
4f875c0
fix(ai): resolve unbounded memory growth for the vercel ai sdk (#10106)
sabrenner Sep 3, 2026
5dc4f3e
chore(deps): bump claude-agent-sdk tested version (#10135)
sabrenner Sep 3, 2026
4921ceb
chore(deps): bump the test-versions group across 1 directory with 7 u…
dependabot[bot] Sep 3, 2026
b0cf209
chore(deps): bump the ai-and-llm group across 1 directory with 5 upda…
dependabot[bot] Sep 3, 2026
7fe3a6e
fix(agentless): preserve explicit telemetry configuration (#10124)
BridgeAR Sep 3, 2026
558befc
refactor(eslint): detect guarded array joins (#10140)
BridgeAR Sep 3, 2026
4caa4e9
bench(sirun): make benchmark workloads deterministic (#9715)
BridgeAR Sep 3, 2026
8a6d411
fix(llmobs): prevent duplicate tags in x-datadog-tags on injection (#…
BridgeAR Sep 3, 2026
ab018ee
fix(net): clean up synchronous connection failures (#10128)
BridgeAR Sep 3, 2026
6bb7a75
fix(tracing): return empty RUM data without an active span (#10130)
BridgeAR Sep 3, 2026
06956d7
feat(crashtracking): support agentless intake (#10121)
BridgeAR Sep 3, 2026
2b97e0e
fix(undici): prevent context retention after requests (#9406)
BridgeAR Sep 3, 2026
1ec7a6b
fix(config): preserve OTel metrics in agentless mode (#10142)
BridgeAR Sep 3, 2026
ba9949d
fix(graphql): record collapsed JIT resolver errors (#10027)
BridgeAR Sep 3, 2026
af03146
fix(tracing): preserve RegExp sampling rule matchers (#10134)
BridgeAR Sep 3, 2026
db022dd
fix(fs): preserve stream constructor errors (#10129)
BridgeAR Sep 3, 2026
57be093
chore(lint): enable eslint-plugin-regexp rules (#10141)
BridgeAR Sep 3, 2026
7444f02
ci: reduce test execution time and rework All Green processing (#9197)
rochdev Sep 3, 2026
1cd31d7
feat(knex, sequelize): add pool acquire span (#8924)
BridgeAR Sep 3, 2026
6c8a115
feat(instrumentation): await callbacks at function start (#10123)
juan-fernandez Sep 4, 2026
f6ee757
test: fix v5 release proposal failures (#10143)
BridgeAR Sep 4, 2026
e4f2839
fix(ci-visibility): support Vitest 5 (#10161)
juan-fernandez Sep 4, 2026
204f55f
feat(test-optimization): correlate WebdriverIO tests with RUM (#10049)
juan-fernandez Sep 4, 2026
24a8935
perf(graphql): reduce inactive resolver overhead (#10136)
BridgeAR Sep 4, 2026
4d60d8d
v6.14.0
BridgeAR Sep 4, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
5 changes: 3 additions & 2 deletions .agents/skills/apm-integrations/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -165,7 +165,8 @@ Follow these steps when creating or modifying an integration:
1. **Investigate** — Read the upstream library's source (see [Read Upstream Source First](#read-upstream-source-first)). Read 1-2 reference integrations of the same type (see table above). Understand the instrumentation and plugin patterns before writing code.
2. **Implement instrumentation** — Create the instrumentation in `packages/datadog-instrumentations/src/`. Use orchestrion for instrumentation.
3. **Implement plugin** — Create the plugin in `packages/datadog-plugin-<name>/src/`. Extend the correct base class.
4. **Register** — Add entries in `packages/dd-trace/src/plugins/index.js`, `index.d.ts`, `docs/test.ts`, `docs/API.md`, and `.github/workflows/apm-integrations.yml`.
4. **Register** — Add entries in `packages/dd-trace/src/plugins/index.js`, every supported public TypeScript surface,
`docs/test.ts`, `docs/API.md`, and `.github/workflows/apm-integrations.yml`.
5. **Write tests** — Add unit tests and ESM integration tests. See [Testing](references/testing.md) for templates.
6. **Run tests** — Validate with:

Expand All @@ -176,7 +177,7 @@ Follow these steps when creating or modifying an integration:
# If the plugin needs external services (databases, message brokers, etc.),
# check docker-compose.yml for available service names, then:
docker compose up -d <service>
PLUGINS="<name>" npm run test:plugins:ci
SERVICES="<service>" PLUGINS="<name>" npm run test:plugins:ci
```

7. **Verify** — Confirm all tests pass before marking work as complete.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -177,6 +177,7 @@ class MyPlugin extends DatabasePlugin {
// Orchestrion: static prefix = 'tracing:orchestrion:<npm-package>:<channelName>'
// Shimmer + tracingChannel: static prefix = 'tracing:apm:<name>:<operation>'
// Shimmer + manual channels: omit prefix — defaults to `apm:${id}:${operation}`
static prefix = '<channel-prefix>'
static peerServicePrecursors = ['db.name']

bindStart (ctx) {
Expand All @@ -196,6 +197,11 @@ class MyPlugin extends DatabasePlugin {

return ctx.currentStore
}

// Choose `end` (sync), `asyncEnd` (promise/callback), or `finish` (legacy manual channel).
asyncEnd (ctx) {
this.finish(ctx)
}
}

module.exports = MyPlugin
Expand All @@ -215,7 +221,7 @@ If multiple npm packages map to the same plugin (e.g., `redis` and `@redis/clien

## Step 4: Add TypeScript Definitions

In `index.d.ts`, add to the `plugins` namespace:
Add the plugin type to the `plugins` namespace in every supported public TypeScript surface:

```typescript
// In the Plugins interface:
Expand Down Expand Up @@ -296,7 +302,7 @@ PLUGINS="<name>" npm run test:plugins:ci
- [ ] Registered in hooks.js (required for both orchestrion and shimmer paths)
- [ ] Plugin created with correct base class
- [ ] Plugin registered in `packages/dd-trace/src/plugins/index.js`
- [ ] TypeScript definitions added to `index.d.ts`
- [ ] TypeScript definitions added to every supported public TypeScript surface
- [ ] Type check added to `docs/test.ts`
- [ ] Documentation added to `docs/API.md`
- [ ] CI job added to `.github/workflows/apm-integrations.yml`
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ The channel prefix is determined by the instrumentation type. Node.js `tracingCh

When using shimmer, prefer `tracingChannel` over manual channels — it provides `start/end/asyncStart/asyncEnd/error` events automatically, consistent with how orchestrion works internally.

This means the plugin only needs to define static properties and implement `bindStart`:
This asynchronous Orchestrion example creates the span in `bindStart` and finishes it in `asyncEnd`:

### Orchestrion Plugin (preferred)
```javascript
Expand All @@ -29,6 +29,10 @@ class MyPlugin extends TracingPlugin {
}, ctx)
return ctx.currentStore
}

asyncEnd (ctx) {
this.finish(ctx)
}
}
```

Expand Down
21 changes: 12 additions & 9 deletions .agents/skills/apm-integrations/references/testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -144,7 +144,6 @@ const { withVersions } = require('../../../dd-trace/test/setup/mocha')
describe('esm', () => {
let agent
let proc
let variants

withVersions('<name>', '<module-name>', version => {
useSandbox([`'<module-name>@${version}'`], false, [
Expand All @@ -154,16 +153,20 @@ describe('esm', () => {
agent = await new FakeAgent().start()
})

before(async function () {
variants = varySandbox('server.mjs', '<module-name>', '<namedExport>')
const variants = varySandbox('server.mjs', {
bindingName: 'myLib',
packageName: '<module-name>',
defaultExport: true,
namedExports: ['<named-export>'],
namedExportBinding: 'namespace',
})

afterEach(async () => {
proc && proc.kill()
await agent.stop()
})

for (const variant of varySandbox.VARIANTS) {
for (const variant of Object.keys(variants)) {
it(`is instrumented ${variant}`, async () => {
const res = agent.assertMessageReceived(({ headers, payload }) => {
assert.strictEqual(headers.host, `127.0.0.1:${agent.port}`)
Expand All @@ -182,9 +185,9 @@ describe('esm', () => {

### Key ESM Test Concepts

- `varySandbox(filename, bindingName, namedExport, packageName, byPassDefault)` generates three import-style variants (default, star, destructure) to verify all ESM import patterns
- `varySandbox.VARIANTS` is `['default', 'star', 'destructure']`
- Pass `byPassDefault: true` as fifth argument when the module has no default export
- `varySandbox(filename, options)` generates the import variants supported by the package's export shape.
- Set `defaultExport`, `namedExports`, and `namedExportBinding` from the installed package's real exports.
- Iterate over `Object.keys(variants)`; the returned object maps each generated variant to its filename.
- `useSandbox` installs package versions into a temp sandbox directory
- `spawnPluginIntegrationTestProcAndExpectExit` spawns `node <script>` with `DD_TRACE_AGENT_PORT` set to FakeAgent port
- Each `it` needs generous timeout (e.g., `20000`) for sandbox setup and process spawning
Expand All @@ -201,8 +204,8 @@ PLUGINS="<name>" npm run test:plugins:ci
PLUGINS="<name>" npm run test:plugins

# With external services (e.g., databases, message brokers)
SERVICES="rabbitmq" PLUGINS="amqplib" docker compose up -d $SERVICES
PLUGINS="amqplib" npm run test:plugins:ci
docker compose up -d rabbitmq
SERVICES="rabbitmq" PLUGINS="amqplib" npm run test:plugins:ci

# Filter within plugin tests
PLUGINS="<name>" SPEC="specific.spec.js" npm run test:plugins:ci
Expand Down
86 changes: 86 additions & 0 deletions .agents/skills/architecture-review/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,86 @@
---
name: architecture-review
description: |
Use when a dd-trace-js change introduces or substantially changes a class hierarchy, module boundary, shared helper
layer, public API, or duplicated behavior across multiple types. Triggers: architecture decision, design review,
refactor shared behavior, new abstraction, composition versus inheritance, expose internals, module coupling,
public surface, hot-path architecture, score the design.
---

# Architecture Review

Use this skill before implementing a non-trivial structural change. Do not use it for a local bug fix or a small
refactor whose boundaries and contracts remain unchanged.

The score is a decision aid, not a substitute for reasoning. Explain the evidence behind each score and reject an
abstraction that adds complexity without improving the baseline.

## Workflow

1. Describe the current design as the baseline, including its duplication, coupling, contracts, tests, and hot path.
2. Describe the smallest viable proposal and at most one meaningful alternative.
3. Identify affected public APIs, package boundaries, consumers, and per-call production paths.
4. Score the baseline and proposal from 1–10 on each dimension below using `baseline → proposal`.
5. Require the proposal to score at least 8/10 on five dimensions. Treat regressions in public-surface discipline or
hot-path fitness as blockers even if the aggregate score passes.
6. Ask the user before implementation when two viable designs have meaningful trade-offs.
7. Record the selected design's contracts and cover boundaries with observable tests.

## Six Dimensions

### 1. Drift prevention

Behavior shared by multiple types should live in one place. Adding a precondition or branch should touch one site,
not require synchronized edits across implementations.

### 2. Module coupling

Cross-module access must use an intentional boundary, never another class's internals. Adding methods to npm-exported
classes such as `Span`, `Tracer`, or OpenTelemetry bridge spans is a lasting compatibility commitment. Prefer a
callback, diagnostic channel, composition, or a redesigned module boundary over exposing internal state.

### 3. Explicit contracts

Express invariants through constructor signatures, specific JSDoc types, narrow interfaces, abstract methods when
appropriate, and `#private` state. Do not rely on undocumented conventions between modules.

### 4. Testability at boundaries

Test boundaries with multiple consumers or protocol/specification contracts directly. Exercise real entry points and
observable output; do not export internals or construct impossible object states solely for tests.

### 5. Extensibility

Evaluate the likely next consumer, type, or method. A third implementation should require a localized addition rather
than edits across every existing implementation. Do not add speculative generality without a credible next case.

### 6. Hot-path fitness

Measure overhead at architectural boundaries on the actual call path. Avoid extra allocations, closures, dispatch,
parsing, and listeners per call. A performance-motivated increase in complexity requires focused, reproducible
benchmark evidence.

## Decision Rules

- Prefer composition. Use inheritance only when at most two sibling types share a complete interface contract and the
hierarchy makes that contract clearer.
- Score the baseline honestly; a `7 → 7` rewrite is not architectural progress.
- Treat test-only exports and accessors as public-surface expansion when evaluating coupling.
- Prefer the simpler implementation when scores and measured performance are effectively equal.
- Avoid new public APIs unless the use case requires a durable compatibility contract.
- If upstream owns the broken abstraction, prefer an upstream fix over a permanent local workaround.

## Review Output

Summarize the review in a compact table:

| Dimension | Baseline | Proposal | Evidence |
| --- | ---: | ---: | --- |
| Drift prevention | | | |
| Module coupling | | | |
| Explicit contracts | | | |
| Testability at boundaries | | | |
| Extensibility | | | |
| Hot-path fitness | | | |

Then state the decision, rejected alternatives, remaining risks, and the validation needed before merging.
108 changes: 81 additions & 27 deletions .agents/skills/flaky-test-fixer/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,46 +1,100 @@
---
name: flaky-test-fixer
description: >-
Use when triaging, investigating, or fixing a suspected flaky test, intermittent
failure, nondeterministic CI failure, timing race, or test-order dependency in
dd-trace-js. Classifies infrastructure and deterministic failures before the
root-cause workflow.
Use when classifying, investigating, or fixing a suspected flaky test, intermittent test result, nondeterministic
CI test failure, timing race, hang, or test-order dependency in dd-trace-js. Classifies infrastructure and
deterministic failures before reproduction or code search.
---

# Flaky test fixer

Establish whether a failure is caused by the current change, deterministic, infrastructure, or genuinely flaky.
“Flaky”, “pre-existing”, and “unrelated” are conclusions that require evidence.
Treat a failure on the current change as caused by that change until evidence identifies another mechanism;
otherwise call it unknown.

Never make CI green by weakening or deleting assertions, filtering unexpected inputs, increasing timeouts, or adding
unexplained retries.

## Classify once

Spend one evidence pass on the failing step, the first actionable error, and whether the test process started. Do
not reproduce or search test code before this gate.
Spend one evidence pass on the exact failing command, test name, assertion or error, environment, first actionable
error, last meaningful log line, whether the test process started, and the current diff. Do not reproduce or search
test code before this gate.

- **Infrastructure:** the test process never ran because checkout, runner, registry, network, or credentials failed,
or independent evidence proves an externally owned network or service outage regardless of test-entry timing. State
the evidence and stop; ignore it for flaky-test work. Treat recurrence as a separate CI task only when asked.
- **Deterministic:** the same revision and inputs consistently fail because of a version, fixture, configuration, or
assertion mismatch. It is not a flake; handle it in the owning change.
assertion mismatch. It is not a flake; handle it in the owning change when related.
- **Genuine flake:** the same test can pass and fail at the same revision with matching relevant inputs and execution
configuration, or evidence proves nondeterministic ordering, timing, or shared state. A green rerun is evidence only
when the test ran under those matching conditions in both attempts.
- **Unknown:** evidence proves none of the above. Run one targeted reproduction or history comparison; do not promote
uncertainty to “flaky,” “infrastructure,” or “unrelated.”

## Fix the cause

1. Reproduce through the smallest real test entry point. Stress the suspected boundary and expose ordering or state;
repeated reruns without a sharper hypothesis are not diagnosis. For a hang, inspect the last error before the
leaked handle kept the process alive.
2. Write one mechanistic sentence naming the producer, consumer, state or event, and invalid ordering or lifetime.
It must explain both the pass and failure. Do not design a fix before this sentence holds.
3. Before editing, search the repository for every test with the same violated invariant and lifecycle owner.
Inventory, count, and list each member and exclusion; state the shared failure mechanism and lifecycle or
completion owner. A range is not an enumeration. Callback versus promise does not split a cohort; similar syntax
under a different contract does not join it.
- **Unknown:** evidence proves none of the above. Run one targeted reproduction or history comparison after this gate;
do not promote uncertainty to “flaky”, “infrastructure”, or “unrelated”.

Evidence for an unrelated flake can include a passing rerun plus a credible race mechanism, the same failure on the
unchanged target branch, a tracked known-flake entry, or a reproducible ordering or resource-contention dependency. A
passing rerun by itself is evidence of nondeterminism, not proof that the current change is unrelated.

## Investigate and fix the cause

1. Re-read the current diff and state a one-line candidate mechanism: “X fails because Y causes Z.” Reproduce
through the smallest real test entry point, preserving relevant environment variables and services. Stress a
suspected boundary or use focused repeated runs only when testing a concrete nondeterminism hypothesis. When safe
and practical, compare against the unchanged target branch in an isolated worktree or equivalent clean
environment. Use the actual target branch for backports, not automatically `master`.
2. Write one mechanistic sentence naming the producer, consumer, state or event, and invalid ordering or lifetime. It
must explain both the pass and failure for a flake. Do not design a fix before this sentence holds.
3. Before editing, search for every test with the same violated invariant and lifecycle owner. Inventory, count, and
list each member and exclusion; state the shared failure mechanism and lifecycle or completion owner. A range is
not an enumeration. Callback versus promise does not split a cohort; similar syntax under a different contract
does not join it.
4. Design the proof, then trace success, error, retry, cleanup, and concurrent paths. Fix the narrowest canonical
owner that restores the invariant for the whole cohort without a new public/test-only surface or production work
added solely for tests.
5. Apply that fix to the complete cohort. Do not substitute retries, skips, sleeps, timeout/tolerance increases,
filtered assertions, or broader mocks for a cause. Keep unrelated mechanisms in separate changes.
owner that restores the invariant for the whole cohort without a new public or test-only surface or production
work added solely for tests.
5. Apply the fix to the complete cohort. Stop stray requests, close leaked resources, restore hooks, and remove shared
mutable state. Use fake timers instead of real-time waits, proper resource allocation instead of arbitrary delays,
and lifecycle ownership instead of forced ordering. Do not substitute retries, skips, sleeps, timeout or tolerance
increases, filtered assertions, or broader mocks for a cause. Keep unrelated mechanisms in separate changes.
6. Verify the original failure without the fix or with a deterministic regression when practical. List every changed
sibling individually in the verification plan, then run them, a targeted repeat or stress run, the complete specs,
and the required coverage/lint from `AGENTS.md`. Report commands, iteration counts, and unproved claims.
sibling individually in the verification plan, then run them, a targeted repeat or stress run, the complete specs,
and the required coverage and lint from `AGENTS.md`. Report commands, iteration counts, and unproved claims.

## Hung jobs

Treat a hang as a potentially masked failure. Inspect the last meaningful error and leaked handles such as tracer or
remote-configuration timers, sockets, child processes, servers, and unfinished hooks before considering a timeout
increase.

## Parallel investigation

After the classification gate, when a suspected unrelated flake appears during another task, delegate its
investigation immediately if sub-agents are available and the investigation has a disjoint write scope. Continue the
main task while it runs.

Give the sub-agent:

- The exact command and failing output
- The current change summary and why the failure may be unrelated
- Relevant paths, services, runtime version, and environment variables
- A classification-first, then read-and-reproduce mandate
- A request for reproduction rate, mechanism, evidence, and the smallest proposed fix

Do not duplicate the sub-agent's investigation in the main thread or let it modify files also being changed there.

A separate branch, commit, or draft PR requires explicit user authorization. When authorized, isolate a genuine
unrelated flake fix from the feature change and base it on the appropriate clean target branch. If a confirmed
unrelated flake cannot be fixed immediately, a temporary skip is not a fix and requires a tracked reason in its own
authorized change; never silently weaken an assertion.

## Report

Report:

1. Classification: caused by the current change, deterministic pre-existing defect, genuine unrelated flake,
infrastructure, or unknown
2. Reproduction command and observed frequency
3. Evidence and one-line mechanism
4. Root-cause fix or next investigation step
5. Whether a separate tracked change is required
1 change: 1 addition & 0 deletions .claude/skills/architecture-review
1 change: 1 addition & 0 deletions .claude/skills/flaky-test-fixer
1 change: 1 addition & 0 deletions .claude/skills/serverless-integrations
9 changes: 9 additions & 0 deletions .codecov.yml
Original file line number Diff line number Diff line change
@@ -1,3 +1,12 @@
codecov:
notify:
# All Green uploads coverage per sibling workflow as each one finishes, so without this,
# Codecov would compute and post its status/patch check after the first upload lands, well
# before the rest have arrived. `scripts/all-green.mjs` calls `codecovcli send-notifications`
# once every workflow's uploads are done, which is the only thing that triggers notification
# while this is set.
manual_trigger: true

coverage:
range: 90..100
round: down
Expand Down
1 change: 1 addition & 0 deletions .cursor/skills/architecture-review
Loading
Loading