Skip to content

adds missing tool level transalation for anthropic -> openai caching - #7885

Merged
akshaydeo merged 1 commit into
devfrom
claude-code-with-openai-caching-issue
Oct 3, 2026
Merged

akshaydeo merged 1 commit into
devfrom
claude-code-with-openai-caching-issue

Conversation

@akshaydeo

@akshaydeo akshaydeo commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Anthropic-style cache_control markers placed on tool_result blocks (the pattern Claude Code uses after every tool call) were being stripped without replacement when targeting gpt-5.6 and later models. Those models require an explicit prompt_cache_breakpoint field on a content part instead of cache_control. The result was that caching stopped advancing past the first tool turn: the marker was silently dropped, no breakpoint was written, and the model fell back to implicit caching anchored at the system prefix.

This PR fixes the translation on both the Responses API path and the Chat Completions path, and extends the Responses path to also handle input_image and input_file block types that OpenAI documents as valid breakpoint locations.

Changes

  • Chat Completions (chat.go): Added applyChatCacheBreakpoints and chatUsesPromptCacheBreakpoints to translate cache_control: ephemeral on text parts (including tool message content) into prompt_cache_breakpoint for gpt-5.6+ on OpenAI, Azure, and Bedrock targets. OpenRouter is deliberately excluded since it accepts cache_control verbatim. When at least one breakpoint is translated and the caller has not already set prompt_cache_options, the request is switched to explicit mode automatically. All mutations are copy-on-write so the caller's input is not modified.

  • Responses API (responses.go): Extended applyResponsesCacheBreakpoints to handle function_call_output items, including both block-form and bare-string outputs. A string output carrying a message-level cache_control marker is promoted to a single input_text block so the breakpoint has a wire-level home; isFunctionCallOutputBlocksFlattenable is updated to decline collapsing a block that already carries a prompt_cache_breakpoint. input_image and input_file blocks are now marked when targeting the OpenAI family (which documents those fields), while OpenRouter continues to receive only input_text breakpoints. responsesHasPromptCacheBreakpoint is extended to scan function_call_output blocks in addition to message content blocks.

  • E2E harness (provider-harness.json): Three new test cases cover the native Responses path (explicit-mode echo), the drop-in Anthropic messages route (cold write token count must exceed 7000 to confirm the marker reached past the tool result), and the Chat Completions path (cache write tokens must exceed 7000 under explicit mode).

Type of change

  • Bug fix

Affected areas

  • Core (Go)
  • Providers/Integrations

How to test

go test ./core/providers/openai/...

Key test functions added:

  • TestToOpenAIChatRequest_GPT56CacheBreakpoint — covers system text, tool message text, image parts (no breakpoint), unmarked text (no breakpoint), four-marker ceiling, copy-on-write safety, OpenRouter passthrough, and pre-5.6 strip-only behaviour.
  • TestToOpenAIResponsesRequest_FunctionCallOutputCacheBreakpoint — covers string output with message-level marker, block output with per-block marker, unmarked block output still flattening, and pre-breakpoint model stripping.
  • TestToOpenAIResponsesRequest_ImageAndFileCacheBreakpoints — covers input_image and input_file in user messages, OpenRouter marking only input_text, file parts inside function_call_output, and message-level marker landing on a trailing image part.

For live validation, run the Postman collection item 142 against a Bifrost instance configured with an OpenAI key and a gpt-5.6-sol deployment.

Breaking changes

  • No

Security considerations

None. No auth, secrets, or PII are involved. The change only affects how content-part metadata is translated before being forwarded to upstream providers.

Checklist

  • I read docs/contributing/README.md and followed the guidelines
  • I added/updated tests where appropriate
  • I updated documentation where needed
  • I verified builds succeed (Go and UI)
  • I verified the CI pipeline passes locally if applicable

@coderabbitai

coderabbitai Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: maximhq/bifrost/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Team
  • Run ID: 81fd11e6-ac00-4f23-a9d5-8c42d7c10ce7
📥 Commits

Reviewing files that changed from the base of the PR and between f200355 and 461e551.

📒 Files selected for processing (5)
  • core/providers/openai/responsesmarshal_test.go
  • core/providers/utils/promptcache.go
  • core/providers/utils/promptcache_test.go
  • core/schemas/provider.go
  • docs/features/prompt-caching.mdx

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 4 remain after this review.


📝 Summary

Summary by CodeRabbit

  • New Features
    • Ephemeral cache markers are translated into prompt-cache breakpoints for supported providers and models across Chat Completions and Responses requests.
    • On supported OpenAI-family models, markers can apply to eligible text, image, file, and tool-result content. Existing cache settings are preserved, and caller-supplied cache modes take precedence.
    • Breakpoints respect the shared limit, retaining the latest eligible markers when necessary.
    • Cache-marker injection can target tool results in Chat and Responses requests.

Walkthrough

Chat Completions and Responses conversion now translates eligible ephemeral cache markers into prompt-cache breakpoints. Injection-point matching can target Chat tool messages and Responses function-call outputs. Tests cover provider-specific conversion, tool outputs, and cache-write reporting.

Changes

Prompt-cache marker handling

Layer / File(s) Summary
Tool-output injection targets
core/providers/utils/promptcache.go, core/providers/utils/promptcache_test.go, core/schemas/provider.go, docs/features/prompt-caching.mdx
Responses function-call outputs and Chat tool messages can match user-role injection points. Tool-output markers are placed on the message. Tests check role and index matching, marker preservation, and input immutability. The schema comment and documentation describe the matching behavior.
Chat Completions marker conversion
core/providers/openai/chat.go, core/providers/openai/chat_test.go
Eligible text markers become breakpoints for supported providers and models. Conversion sets explicit cache mode when needed, preserves caller-supplied options, and retains the latest markers within the limit. Tests cover provider and model differences and unchanged input.
Responses message and tool-output conversion
core/providers/openai/responses.go, core/providers/openai/types.go, core/providers/openai/responsesmarshal_test.go
Responses conversion handles eligible message and function-call-output markers, including supported image and file blocks. Breakpoint-bearing output blocks are not flattened. Tests cover tool output, model support, and OpenRouter behavior.
Provider and cache-write validation
tests/e2e/api/collections/provider-harness.json
Harness tests check explicit cache mode for a Responses tool result and cache-write token counts for Anthropic drop-in and Chat Completions requests.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~30 minutes

Suggested reviewers: sammaji

Merge Risk: ⚪ Minimal · up to 461e5

No established merge-blocking issue remains. The change is ready for normal merge checks.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Issue #123 is closed and provides historical context only. No active directly linked issue supplies coding requirements for this pull request.
Out of Scope Changes check ✅ Passed The changes support the current cache-breakpoint intent. The Chat and Responses conversions, tool-result cache injection, documentation, and regression and harness tests all relate to marking tool res…
Docstring Coverage ✅ Passed Docstring coverage is 86.21% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 29 functions across 8 files. (1 skipped: 1 …
Title check ✅ Passed The title describes the main change: translating Anthropic tool-level cache markers for OpenAI caching. It is concise, though it contains spelling and capitalization errors.
Description check ✅ Passed The description is mostly complete. It explains the problem, implementation, affected areas, tests, breaking-change status, and security considerations. Screenshots are not applicable, but it does not…
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

Copy link
Copy Markdown
Contributor Author

This stack of pull requests is managed by Graphite. Learn more about stacking.

@akshaydeo
akshaydeo marked this pull request as ready for review October 3, 2026 07:12
@coderabbitai
coderabbitai Bot requested a review from TejasGhatte October 3, 2026 07:17
coderabbitai[bot]
coderabbitai Bot previously approved these changes Oct 3, 2026
@akshaydeo
akshaydeo force-pushed the claude-code-with-openai-caching-issue branch from f200355 to 461e551 Compare October 3, 2026 07:39
@coderabbitai
coderabbitai Bot requested a review from sammaji October 3, 2026 07:41

akshaydeo commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor Author

Merge activity

  • Oct 3, 7:56 AM UTC: A user started a stack merge that includes this pull request via Graphite.
  • Oct 3, 7:57 AM UTC: @akshaydeo merged this pull request with Graphite.

@akshaydeo
akshaydeo merged commit 05d3b4d into dev Oct 3, 2026
15 checks passed
@akshaydeo
akshaydeo deleted the claude-code-with-openai-caching-issue branch October 3, 2026 07:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant