adds missing tool level transalation for anthropic -> openai caching - #7885
Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (5)
Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 4 remain after this review. 📝 SummarySummary by CodeRabbit
WalkthroughChat Completions and Responses conversion now translates eligible ephemeral cache markers into prompt-cache breakpoints. Injection-point matching can target Chat tool messages and Responses function-call outputs. Tests cover provider-specific conversion, tool outputs, and cache-write reporting. ChangesPrompt-cache marker handling
Priority: ➖ Normal Estimated code review effort: 3 (Moderate) | ~30 minutes Suggested reviewers: Merge Risk: ⚪ Minimal · up to No established merge-blocking issue remains. The change is ready for normal merge checks. 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
f200355 to
461e551
Compare
Merge activity
|

Summary
Anthropic-style
cache_controlmarkers placed ontool_resultblocks (the pattern Claude Code uses after every tool call) were being stripped without replacement when targeting gpt-5.6 and later models. Those models require an explicitprompt_cache_breakpointfield on a content part instead ofcache_control. The result was that caching stopped advancing past the first tool turn: the marker was silently dropped, no breakpoint was written, and the model fell back to implicit caching anchored at the system prefix.This PR fixes the translation on both the Responses API path and the Chat Completions path, and extends the Responses path to also handle
input_imageandinput_fileblock types that OpenAI documents as valid breakpoint locations.Changes
Chat Completions (
chat.go): AddedapplyChatCacheBreakpointsandchatUsesPromptCacheBreakpointsto translatecache_control: ephemeralon text parts (including tool message content) intoprompt_cache_breakpointfor gpt-5.6+ on OpenAI, Azure, and Bedrock targets. OpenRouter is deliberately excluded since it acceptscache_controlverbatim. When at least one breakpoint is translated and the caller has not already setprompt_cache_options, the request is switched to explicit mode automatically. All mutations are copy-on-write so the caller's input is not modified.Responses API (
responses.go): ExtendedapplyResponsesCacheBreakpointsto handlefunction_call_outputitems, including both block-form and bare-string outputs. A string output carrying a message-levelcache_controlmarker is promoted to a singleinput_textblock so the breakpoint has a wire-level home;isFunctionCallOutputBlocksFlattenableis updated to decline collapsing a block that already carries aprompt_cache_breakpoint.input_imageandinput_fileblocks are now marked when targeting the OpenAI family (which documents those fields), while OpenRouter continues to receive onlyinput_textbreakpoints.responsesHasPromptCacheBreakpointis extended to scanfunction_call_outputblocks in addition to message content blocks.E2E harness (
provider-harness.json): Three new test cases cover the native Responses path (explicit-mode echo), the drop-in Anthropic messages route (cold write token count must exceed 7000 to confirm the marker reached past the tool result), and the Chat Completions path (cache write tokens must exceed 7000 under explicit mode).Type of change
Affected areas
How to test
go test ./core/providers/openai/...Key test functions added:
TestToOpenAIChatRequest_GPT56CacheBreakpoint— covers system text, tool message text, image parts (no breakpoint), unmarked text (no breakpoint), four-marker ceiling, copy-on-write safety, OpenRouter passthrough, and pre-5.6 strip-only behaviour.TestToOpenAIResponsesRequest_FunctionCallOutputCacheBreakpoint— covers string output with message-level marker, block output with per-block marker, unmarked block output still flattening, and pre-breakpoint model stripping.TestToOpenAIResponsesRequest_ImageAndFileCacheBreakpoints— coversinput_imageandinput_filein user messages, OpenRouter marking onlyinput_text, file parts insidefunction_call_output, and message-level marker landing on a trailing image part.For live validation, run the Postman collection item
142against a Bifrost instance configured with an OpenAI key and a gpt-5.6-sol deployment.Breaking changes
Security considerations
None. No auth, secrets, or PII are involved. The change only affects how content-part metadata is translated before being forwarded to upstream providers.
Checklist
docs/contributing/README.mdand followed the guidelines