Count a request whose answer is cached as fitting, whatever the calibration - #38
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #37 (and so #36): its base is
split-compose. The change itself touches none of their code; the base only avoids a conflict on the CHANGELOG'sUnreleasedsection, where #36 adds a line too. Merge #36 and #37 first and retarget this tomain, or cherry-pick its first commit alone ontomain.The problem. Two runs of the same release on the same code can give different findings.
.jevgate/token-budget.json) is replaced after each run by the bytes per token of that run's fresh requests.I hit this while comparing branches on the corpus (see #36). One pi-fabric outline was a note in one run of 0.23.1 and clear in another, so every comparison needed each clone's calibration restored from a snapshot first. In CI, or in
check --watch, the same thing makes a gate flicker without any code change.The change. Planning counts a request as fitting when the estimate says so, or when the answer cache already holds its answer: the provider took it once.
--refreshskips the cache, so it plans by the estimate alone.TokenBudgetis unchanged. Planning now receivesLimits, which pairs the budget with that cache lookup, andFileContext.budgetholds it, so no planner changed.Measured on the corpus, 181 projects, from the cache.
scripts/claims-audit.py) and five undecided units. Under the stricter one it also asked about 42 thousand new tokens of questions.Known remainder. Those three units are in b2-bend-collections. Their whole-file rechecks fit a loose estimate, so they're sent, but the provider refuses them as too long. A refused request is not cached, and a unit that has first answers keeps its planned recheck. So the kind fallback isn't asked in that run, and the file stays undecided, while a stricter calibration asks the kind from the outline and decides it. Fixing that means remembering refusals between runs (a request the provider refused fits no more) and letting a unit whose recheck was refused take the outline-only kind in the same run. I left it for a separate change.
Test: an outline whose recheck was answered at 3.0 bytes per token is rechecked from the cache at 2.0 and keeps its status, with no request sent. It fails without the change, because the second run asks the kind from the outline alone.
Tests, clippy (also 1.98),
cargo +1.90.0 check --lockedand the self-check (no review or consider) pass.