[bench] stats: one weekly grouping; tokens by week in cli stats and the dashboard - #66
Merged
Merged
Conversation
…he dashboard Release check R3 requires the dashboard to show the release week's median tokens per edit, and neither cli stats nor the viewer exposed tokens by week: the CLI printed tokens over all entries, and the dashboard grouped only verdicts and latency by week. utils.stats.entries_by_week is now the one grouping every per-week figure is built on: entries by the ISO week of their timestamp, with unparseable timestamps in the unknown bucket. stats_by_week and latency_by_week are refactored onto it with their outputs unchanged, which the existing exact-equality tests confirm, and tokens_by_week builds on it too: per week, the same median and p90 the all-entries line prints plus the cached-rate medians, over entries that recorded usage, with empty weeks omitted rather than shown as zero. cli stats prints a Tokens by week block, one line per week, and the viewer's tokens card gains a per-week table. Tests cover the ISO week boundaries (Sunday closes W01 and Monday opens W02; the last days of 2025 fall in 2026-W01; New Year's Day 2027 falls in 2026-W53), the unknown bucket, all three weekly tables sharing the grouping, missing and malformed usage, cached-rate pricing, agreement with the all-entries figures for a single week, the empty ledger, the dashboard row, and the CLI line. The README's What It Costs section is refreshed as a fixed snapshot of the chain at 2026-09-07T06:37:19Z (2,767 entries, 2,743 with usage), stated as such rather than as a live weekly value, and the reproducibility paragraph now names the two interfaces that expose the weekly rows and the shared grouping behind them. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0122bFBUvdtXT245QFoYDkMf
Contributor
Author
|
@codex review |
SonarCloud S5906 on #66: a set-subset check written as assertTrue over a comparison reads better, and fails with the two sets shown, as assertLessEqual. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0122bFBUvdtXT245QFoYDkMf
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 13414aece7
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Codex review on #66: the weekly lines dropped the p90 that tokens_by_week computes, so cli stats could not reproduce the README table's 90th-percentile column. Each line now carries median, p90, and the cached-rate median beside the entry count. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0122bFBUvdtXT245QFoYDkMf
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Gates the v2.1.0 tag. Release check R3 requires the dashboard to show the release week's median tokens per edit, and neither
cli statsnor the viewer exposed tokens by week: the CLI printed tokens over all entries, and the dashboard grouped only verdicts and latency by week. Codex caught the README claiming otherwise on #65.What
utils.stats.entries_by_weekgroups entries by the ISO week of their timestamp, with unparseable timestamps in the unknown bucket. The verdict table and the latency table are refactored onto it with their outputs unchanged, which the existing exact-equality tests confirm.tokens_by_weekbuilds on it: per week, the same median and p90 the all-entries line prints plus the cached-rate medians, over entries that recorded usage, with empty weeks omitted rather than shown as zero.cli statsprints a "Tokens by week" block, one line per week. The viewer's tokens card gains a per-week table: entries, median, p90, median at cached rates.Tests
Receipts
On the live chain at the snapshot cutoff:
The release-week median moved from under 20,000 to just over it on the day of the snapshot as heavier edits landed; the README says it is not a settled reduction.
ruff check .clean;python -m mypy(strict, CI file set) clean.After merge
Tag the merge commit
v2.1.0(R6). Release notes are drafted at the same snapshot cutoff.