Skip to content

Optimize token usage across provider pipeline - #17

Merged
apriendeau merged 6 commits into
mainfrom
austin/event-lifecycle
May 19, 2026
Merged

apriendeau merged 6 commits into
mainfrom
austin/event-lifecycle

Conversation

@apriendeau

@apriendeau apriendeau commented May 19, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Reduces token consumption by ~65% on multi-turn sessions through five optimizations:

  • Tiered tool result compression — Prior-turn tool results are replaced with one-line summaries ([read /src/main.rs: 458 lines]) before sending to the provider. Current-turn results stay full. Error results and non-zero bash exits are always kept in full to preserve diagnostic context. Full results remain on disk in SessionStore.
  • Compact line numbering — Read tool output changed from padded format ( 1\t) to compact (1:), saving ~15-20% overhead per file read.
  • Compact skill catalog — Replaced verbose XML skill catalog in system prompt with plain-text one-liner format. Cross-LLM compatible and saves 500-1000 tokens/turn.
  • Cached tool schemas — Tool schemas are built once at session start instead of regenerated every loop iteration. Prevents accidental Anthropic cache invalidation.
  • Split Anthropic cache breakpoints — system_prompt changed from Option<String> to Vec<String>. Anthropic provider gives each block its own cache_control marker so the stable base prompt stays cached when the context suffix changes. Other providers join the blocks.

Changed files

File Change
session/compress.rs New module — compress_for_provider() with per-tool summarizers
tools/read.rs Compact line number format
session/skills.rs Plain-text skill catalog
providers/traits.rs SendOptions.system_prompt: Option<String> → Vec<String>
providers/anthropic.rs Multi-block system prompt with per-block cache control
providers/{openai,google,ollama}.rs Join system prompt blocks
session/{conversation,agent_loop}.rs, tui/conversation.rs Wire compression + cached schemas + Vec system prompt

Test plan

  • All 418 tests pass
  • Clippy clean, fmt clean
  • Manual test: multi-turn session, verify prior-turn results show as summaries in provider logs
  • Manual test: bash with non-zero exit in prior turn stays full
  • Manual test: session resume loads full history, compression applies on next send

🤖 Generated with Claude Code

@apriendeau apriendeau self-assigned this May 19, 2026
@apriendeau apriendeau changed the title opitimize agent hartens and tools opitimize agent harness and tools May 19, 2026
@apriendeau apriendeau changed the title opitimize agent harness and tools Optimize token usage across provider pipeline May 19, 2026
@apriendeau
apriendeau requested a review from jberns May 19, 2026 05:08
Comment thread src/session/agent_loop.rs Outdated
apriendeau and others added 3 commits May 19, 2026 12:32
Add re-exports to config/mod.rs so consumers use crate::config::Config
instead of crate::config::schema::Config. Remove unused apply_model_choice
and add_provider from commands/model.rs. Rename ConvConfig→ConvParams and
SpawnConfig→SpawnParams since they are runtime parameter bundles, not
configuration schemas.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@apriendeau
apriendeau merged commit 6b2ba79 into main May 19, 2026
3 checks passed
@apriendeau
apriendeau deleted the austin/event-lifecycle branch May 19, 2026 18:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants