Skip to content

[APPSEC-69389] Add AI Guard sensetive data redaction - #6394

Open
Strech wants to merge 41 commits into
masterfrom
appsec-69389-add-ai-guard-sensetive-data-redaction
Open

Strech wants to merge 41 commits into
masterfrom
appsec-69389-add-ai-guard-sensetive-data-redaction

Conversation

@Strech

@Strech Strech commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

What does this PR do?

This PR introduces Sensitive Data Redaction implementation with AI Guard upgrade.

Motivation:

During chat with LLM some sensitive data might escape, so we catch it before it's too late 🛡️

Change log entry

Yes.

Additional Notes:

This PR contains many changes, during discussion we come to agreement to keep it in one PR to avoid design drift. I will highlight few adjustments worth mention.

Design changes:

  1. RubyLLM contrib now requires 2.0+ now – allowed as we are not GA and 2.0 did a massive overhaul
  2. AIGuard::APIClient renamed into AIGuard::HTTPClient
  3. AIGuard::Evaluation::Client takes a role of a glue between SDK and Backend
  4. AIGuard::ToolCall, AIGuard::Message, AIGuard::ContentPart take responsibility from AIGuard::Request of their own representation as a Hash
  5. AIGuard::ToolCall, AIGuard::Message, AIGuard::ContentPart got method with_* to generate new instance with replaced dependencies
  6. AIGuard::Request and AIGuard::Response now have clear separation of concerns

And additional to the RFC I took a step forward and introduced redaction back to the RubyLLM (not only to the model). Because of that a few missing abstractions pops up and a few limitations that goes beyond this PR.

Features outside RFC:

  1. To compute a diff between input and redaction, I used AIGuard::Contrib::RubyLLM::RedactionChangeSet
  2. A new instrumentation points for 2.0 were needed inside AIGuard::Contrib::RubyLLM::ChatInstrumentation

Bug fixes:

  1. Tool call arguments was not properly serialized as JSON, now they are
  2. Tool calls were always wrapped into separate message, now they are not

How to test the change?

CI + ST

@dd-octo-sts dd-octo-sts Bot added core Involves Datadog core libraries ai-guard labels Sep 28, 2026
@dd-octo-sts

dd-octo-sts Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Typing analysis

Note: Ignored files are excluded from the next sections.

steep:ignore comments

This PR introduces 6 steep:ignore comments, and clears 2 steep:ignore comments.

steep:ignore comments (+6-2) ❌ Introduced:
lib/datadog/ai_guard/contrib/ruby_llm/message_adapter.rb:40
lib/datadog/ai_guard/contrib/ruby_llm/redaction_change_set.rb:47
lib/datadog/ai_guard/contrib/ruby_llm/redaction_change_set.rb:58
lib/datadog/ai_guard/contrib/ruby_llm/redaction_change_set.rb:60
lib/datadog/ai_guard/contrib/ruby_llm/redaction_change_set.rb:107
lib/datadog/ai_guard/contrib/ruby_llm/redaction_change_set.rb:119
✅ Cleared:
lib/datadog/ai_guard/evaluation.rb:35
lib/datadog/ai_guard/evaluation.rb:85

Untyped methods

This PR introduces 1 partially typed method, and clears 4 partially typed methods. It increases the percentage of typed methods from 71.82% to 72.28% (+0.46%).

Partially typed methods (+1-4) ❌ Introduced:
sig/datadog/ai_guard/contrib/ruby_llm/chat_instrumentation.rbs:10
└── def provider_completion: (
            usage_recorder: untyped,
            ?stream_tracker: untyped
          ) ?{ (untyped chunk) -> untyped } -> ::RubyLLM::Message
✅ Cleared:
sig/datadog/ai_guard/api_client.rbs:13
└── def post: (::String path, body: ::Hash[::String | ::Symbol, untyped]) -> ::Hash[::String, ::String]
sig/datadog/ai_guard/api_client.rbs:19
└── def parse_response_body: (::String) -> ::Hash[::String, untyped]
sig/datadog/ai_guard/evaluation/request.rbs:32
└── def build_request_body: () -> ::Hash[::Symbol, untyped]
sig/datadog/ai_guard/evaluation/result.rbs:29
└── def initialize: (::Hash[::String, untyped] raw_response_body) -> void

If you believe a method or an attribute is rightfully untyped or partially typed, you can add # untyped:accept on the line before the definition to remove it from the stats.

@datadog-datadog-prod-us1

datadog-datadog-prod-us1 Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Tests

✅ All CI checks and tests passed.

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

🎯 Code Coverage (details)
• Patch Coverage: 99.34%
• Overall Coverage: 90.63% (+0.15%)

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 03f8e22 | Docs | View more details | Give us feedback!

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The new code breaks a marked public API and has a reproducible missing json dependency.

Review effort: Balanced
Findings: 1 High severity · 1 Medium severity · 4 Low severity

Open (6)
What changed in this PR

Adds default-on sensitive-data redaction to AI Guard evaluations and RubyLLM 2.x integration flows.

Changes:

  • Adds typed redaction parsing, application, outcomes, telemetry, and configuration.
  • Updates RubyLLM instrumentation to redact provider prompts and tool-call arguments.
  • Refactors evaluation transport and message serialization APIs.
File Description
vendor/​rbs/​ruby_llm/​0/​ruby_llm.rbs Updates RubyLLM 2.x types.
unreleased/​20260928082914.json Notes RubyLLM 2.x requirement.
unreleased/​20260928082908.json Announces redaction.
supported-configurations.json Registers the redaction setting.
spec/​datadog/​ai_guard/​redaction/​result_spec.rb Tests redaction state.
spec/​datadog/​ai_guard/​redaction/​replacements_spec.rb Tests replacement normalization.
spec/​datadog/​ai_guard/​metrics/​telemetry_spec.rb Tests telemetry metrics.
spec/​datadog/​ai_guard/​http_client_spec.rb Updates client tests.
spec/​datadog/​ai_guard/​evaluation/​result_spec.rb Tests refactored results.
spec/​datadog/​ai_guard/​evaluation/​response_spec.rb Tests response parsing.
spec/​datadog/​ai_guard/​evaluation/​request_spec.rb Tests request serialization.
spec/​datadog/​ai_guard/​evaluation/​outcome_spec.rb Tests blocking decisions.
spec/​datadog/​ai_guard/​evaluation/​message_spec.rb Tests message serialization.
spec/​datadog/​ai_guard/​evaluation/​client_spec.rb Tests evaluation orchestration.
spec/​datadog/​ai_guard/​evaluation_spec.rb Tests end-to-end redaction.
spec/​datadog/​ai_guard/​contrib/​ruby_llm/​redaction_change_set_spec.rb Tests safe RubyLLM changes.
spec/​datadog/​ai_guard/​contrib/​ruby_llm/​message_adapter_spec.rb Tests message conversion.
spec/​datadog/​ai_guard/​contrib/​ruby_llm/​chat_instrumentation_spec.rb Tests RubyLLM instrumentation.
spec/​datadog/​ai_guard/​configuration/​settings_spec.rb Tests redaction configuration.
spec/​datadog/​ai_guard/​component_spec.rb Tests HTTP client wiring.
spec/​datadog/​ai_guard_spec.rb Tests public helpers.
sig/​datadog/​core/​configuration/​settings.rbs Types the new setting.
sig/​datadog/​ai_guard/​redaction/​result.rbs Types redaction results.
sig/​datadog/​ai_guard/​redaction/​replacements.rbs Types replacement parsing.
sig/​datadog/​ai_guard/​redaction.rbs Types redaction operations.
sig/​datadog/​ai_guard/​metrics/​telemetry.rbs Types telemetry reporting.
sig/​datadog/​ai_guard/​http_client.rbs Types the renamed client.
sig/​datadog/​ai_guard/​ext.rbs Types the redaction tag.
sig/​datadog/​ai_guard/​evaluation/​tool_call.rbs Updates tool-call types.
sig/​datadog/​ai_guard/​evaluation/​result.rbs Updates result types.
sig/​datadog/​ai_guard/​evaluation/​response.rbs Adds response types.
sig/​datadog/​ai_guard/​evaluation/​request.rbs Updates request types.
sig/​datadog/​ai_guard/​evaluation/​outcome.rbs Types evaluation outcomes.
sig/​datadog/​ai_guard/​evaluation/​no_op_result.rbs Updates no-op result types.
sig/​datadog/​ai_guard/​evaluation/​message.rbs Updates message types.
sig/​datadog/​ai_guard/​evaluation/​content_part.rbs Types content serialization.
sig/​datadog/​ai_guard/​evaluation/​client.rbs Types evaluation client.
sig/​datadog/​ai_guard/​evaluation.rbs Updates evaluation signatures.
sig/​datadog/​ai_guard/​contrib/​ruby_llm/​redaction_change_set.rbs Types RubyLLM changes.
sig/​datadog/​ai_guard/​contrib/​ruby_llm/​message_adapter.rbs Types the adapter.
sig/​datadog/​ai_guard/​contrib/​ruby_llm/​chat_instrumentation.rbs Updates instrumentation types.
sig/​datadog/​ai_guard/​configuration/​ext.rbs Types the environment constant.
sig/​datadog/​ai_guard/​component.rbs Updates component types.
sig/​datadog/​ai_guard/​api_client.rbs Removes obsolete client types.
sig/​datadog/​ai_guard.rbs Updates public API signatures.
Matrixfile Updates RubyLLM test coverage.
lib/​datadog/​core/​configuration/​supported_configurations.rb Registers the environment variable.
lib/​datadog/​ai_guard/​redaction/​result.rb Implements redaction results.
lib/​datadog/​ai_guard/​redaction/​replacements.rb Normalizes backend replacements.
lib/​datadog/​ai_guard/​redaction.rb Applies replacements.
lib/​datadog/​ai_guard/​metrics/​telemetry.rb Reports evaluation metrics.
lib/​datadog/​ai_guard/​http_client.rb Renames the HTTP client.
lib/​datadog/​ai_guard/​ext.rb Adds the redaction span tag.
lib/​datadog/​ai_guard/​evaluation/​tool_call.rb Serializes tool arguments.
lib/​datadog/​ai_guard/​evaluation/​result.rb Separates result data.
lib/​datadog/​ai_guard/​evaluation/​response.rb Parses backend responses.
lib/​datadog/​ai_guard/​evaluation/​request.rb Builds request bodies.
lib/​datadog/​ai_guard/​evaluation/​outcome.rb Encapsulates blocking state.
lib/​datadog/​ai_guard/​evaluation/​no_op_result.rb Preserves unevaluated messages.
lib/​datadog/​ai_guard/​evaluation/​message.rb Supports multiple tool calls.
lib/​datadog/​ai_guard/​evaluation/​content_part.rb Serializes content parts.
lib/​datadog/​ai_guard/​evaluation/​client.rb Orchestrates evaluation and redaction.
lib/​datadog/​ai_guard/​evaluation.rb Reports redacted evaluation data.
lib/​datadog/​ai_guard/​contrib/​ruby_llm/​redaction_change_set.rb Rebuilds redacted RubyLLM messages.
lib/​datadog/​ai_guard/​contrib/​ruby_llm/​patcher.rb Loads the adapter.
lib/​datadog/​ai_guard/​contrib/​ruby_llm/​message_adapter.rb Converts RubyLLM messages.
lib/​datadog/​ai_guard/​contrib/​ruby_llm/​integration.rb Requires RubyLLM 2.0+.
lib/​datadog/​ai_guard/​contrib/​ruby_llm/​chat_instrumentation.rb Redacts provider and tool inputs.
lib/​datadog/​ai_guard/​configuration/​ext.rb Adds the environment variable.
lib/​datadog/​ai_guard/​configuration.rb Adds the redaction option.
lib/​datadog/​ai_guard/​component.rb Wires new AI Guard components.
lib/​datadog/​ai_guard.rb Updates public construction helpers.
gemfiles/​ruby_4.0_ruby_llm_min.gemfile.lock Locks RubyLLM 2.0 dependencies.
gemfiles/​ruby_4.0_ruby_llm_min.gemfile Sets RubyLLM minimum version.
gemfiles/​ruby_3.4_ruby_llm_min.gemfile.lock Locks RubyLLM 2.0 dependencies.
gemfiles/​ruby_3.4_ruby_llm_min.gemfile Sets RubyLLM minimum version.
gemfiles/​ruby_3.3_ruby_llm_min.gemfile Sets RubyLLM minimum version.
gemfiles/​ruby_3.2_ruby_llm_min.gemfile Sets RubyLLM minimum version.
gemfiles/​ruby_3.1_ruby_llm_min.gemfile Sets RubyLLM minimum version.
docs/​GettingStarted.md Documents redaction configuration.
appraisal/​ruby-4.0.rb Updates RubyLLM appraisal.
appraisal/​ruby-3.4.rb Updates RubyLLM appraisal.
appraisal/​ruby-3.3.rb Updates RubyLLM appraisal.
appraisal/​ruby-3.2.rb Updates RubyLLM appraisal.
appraisal/​ruby-3.1.rb Updates RubyLLM appraisal.
.gitlab-ci.yml Updates the system-tests pin.
.github/​workflows/​system-tests.yml Updates system-tests workflow pins.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread lib/datadog/ai_guard.rb Outdated
Comment thread lib/datadog/ai_guard/evaluation/tool_call.rb
Comment thread Matrixfile Outdated
Comment thread sig/datadog/ai_guard/evaluation/content_part.rbs
Comment thread sig/datadog/ai_guard/evaluation/message.rbs
Comment thread sig/datadog/ai_guard/evaluation/tool_call.rbs
Comment thread docs/GettingStarted.md Outdated
Comment thread lib/datadog/ai_guard/contrib/ruby_llm/chat_instrumentation.rb
Comment thread lib/datadog/ai_guard/evaluation/client.rb
Comment thread lib/datadog/ai_guard/evaluation/content_part.rb Outdated
Comment thread lib/datadog/ai_guard/evaluation/message.rb Outdated
Comment thread lib/datadog/ai_guard/evaluation/message.rb Outdated
Comment thread lib/datadog/ai_guard/evaluation/response.rb
Comment thread lib/datadog/ai_guard/evaluation/tool_call.rb
Comment thread lib/datadog/ai_guard/redaction/replacements.rb
Comment thread lib/datadog/ai_guard/redaction/replacements.rb Outdated
Comment thread lib/datadog/ai_guard/redaction/result.rb Outdated
Comment thread lib/datadog/ai_guard/redaction.rb Outdated
Comment thread lib/datadog/ai_guard/redaction.rb
Comment thread lib/datadog/ai_guard/redaction.rb Outdated
Comment thread lib/datadog/ai_guard.rb Outdated
Comment thread lib/datadog/ai_guard.rb Outdated
@Strech
Strech requested a review from y9v September 29, 2026 09:15
@pr-commenter

pr-commenter Bot commented Sep 29, 2026

Copy link
Copy Markdown

Benchmarks

Benchmark execution time: 2026-09-29 10:11:13

Comparing candidate commit b918575 in PR branch appsec-69389-add-ai-guard-sensetive-data-redaction with baseline commit ac5d9ac in branch master.

📊 Benchmarking dashboard

Found 0 performance improvements and 0 performance regressions! Performance is the same for 52 metrics, 0 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

@y9v y9v left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nice! The code looks way cleaner now

@Strech
Strech marked this pull request as ready for review October 1, 2026 07:57
@Strech
Strech requested review from a team as code owners October 1, 2026 07:57
@Strech
Strech requested review from vpellan and removed request for a team October 1, 2026 07:57
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-01T08:02:45.637584Z 2dda3a9 Draft marked ready
🔒 Security Review ✅ Completed 2026-10-01T08:09:45.318013Z 2dda3a9 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 2dda3a9cbd

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread lib/datadog/ai_guard.rb
Comment thread lib/datadog/ai_guard/configuration.rb
desired_execution_time: 300 # 5 minutes
scenarios_groups: tracer_release
ref: 70f36a2a612d0f9e98af5f4882142e941ae09717 # Automated: Updated by .github/workflows/update-system-tests.yml.
ref: 7624ea2919a72738a2e923adbdf9c2d538bc4ae2 # Automated: Updated by .github/workflows/update-system-tests.yml.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This ref is not merged yet on system-tests. We had a case yesterday of merging such a thing by accident, causing the CI to fail on every system-tests PR. Please revert this or merge it on system-tests before merging this here. There will be a job to prevent this in the future but it's not merged yet:
DataDog/system-tests#7874

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is a dead-lock, this PR heavily changing AI Guard interface and gem requirements due to non GA status, ST will be never green on prod environment because I have to change weblogs, hence I can't merge it

@Strech

Strech commented Oct 1, 2026

Copy link
Copy Markdown
Member Author

@vpellan Could you please take a minute and read through PR description please, we heavily changed the AI Guard interface as it's not GA so far and we don't want to fall into the trap. System tests are unable to pass on prod due to that refactoring, but that's OK in this case. Thanks ❤️

P.S I've resolved Codex comments as they refer to my note above

@Strech
Strech requested a review from vpellan October 1, 2026 12:20
@Strech
Strech enabled auto-merge October 1, 2026 12:42
@Strech
Strech disabled auto-merge October 1, 2026 14:27

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ai-guard core Involves Datadog core libraries

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants