diff --git a/pkg/rds/GROWTH-FINDINGS.md b/pkg/rds/GROWTH-FINDINGS.md new file mode 100644 index 0000000..3ff5132 --- /dev/null +++ b/pkg/rds/GROWTH-FINDINGS.md @@ -0,0 +1,327 @@ +# T4 — allocated-storage growth: a refusal that finally has a size + +`pkg/rds/FINDINGS.md` §7.4 deferred trend detection over the allocated-storage +ratchet and named its seam — `Target.PriorAllocatedStorageGiB` plus +`AttributeStorageGrowth`. `cmd/RDSLIVE-FINDINGS.md` §5.2 independently listed +`AttributeStorageGrowth` as unreachable, needing "cross-run checkpoint +persistence `kilter domains` has not got". Both were describing the same hole. +This unit fills it: `growth.go` + `growth_test.go`, **979 / 943 lines, 19 tests +including one fuzz target, 5 new refusal codes and 1 advisory code.** + +Green under `gofmt -l ./pkg/rds`, `go vet ./...`, `go build ./...`, +`go test -race -count=1 ./pkg/rds/...` and `go test -race -short ./...` +(36 packages, all `ok`). `go.mod` and `go.sum` are untouched; `growth.go` +imports `fmt`, `sort`, `time` and `pkg/domain`. No existing test was edited. + +**What was NOT touched, deliberately:** `actuate*.go` and `parity*.go` are +byte-identical. This unit adds no mutating identifier, no seam an actuator +could be reached through, and no argument for making one reachable. Storage +growth is an observation and its output is a refusal; if anything it *narrows* +the case for actuation, because the thing it measures is the one RDS quantity +that provably cannot be actuated in the direction anyone would want. + +--- + +## 1. Why this is not a shrink proposal, structurally rather than by intent + +U11's central finding is that `MaxAllocatedStorage` is a **ratchet, not +headroom**. Measuring how fast the ratchet turned is therefore not a step +toward turning it back, and this unit is built so that reading it as one takes +active effort: + +| Guard | Where | +|---|---| +| The finding IS a refusal. A measured growth emits `storage-growth-is-not-reclaimable`, not an advisory with a refusal attached. | `AssessStorageGrowth` | +| No type here has a field matching `saving`, `reclaim`, `shrink`, `reduc` or `recover`. Asserted by reflection over four structs. | `TestGrowthNamesNoSaving` | +| The growth's dollar is `MonthlyUSD` on `GrowthMeasurement` — a **cost already being paid**, on the same footing as `StorageVerdict.UnusedMonthlyUSD`. There is no signed delta. | `GrowthMeasurement` | +| Nothing in `growth.go` writes `Assessment.Proposal`. The report's proposal count and savings totals stay 0 over arbitrary histories. | `FuzzGrowthNeverClaimsASaving`, `TestGrowthReachesTheReportAsARefusalAndNeverAsAProposal` | +| The refusal prose ends on AWS's own sentence, and a test rejects opportunity-shaped words in it. | `ratchetClause`, `TestGrowthIsARefusalWithASize` | + +The existing refusals are unweakened and still fire beside it: +`allocated-storage-cannot-shrink` and `storage-autoscaling-ratchet` are +asserted on the same assessment that carries the growth finding. + +## 2. The two bars, with the arithmetic + +Both are refusals with their own codes, and they are **separate codes because +they have separate remedies**: one is fixed by checkpointing more often, the +other only by waiting. + +### 2.1 `DefaultGrowthMinObservations = 4` → `storage-growth-insufficient-observations` + +n observations bound **n−1 intervals**: + +| n | intervals | what it can distinguish | +|---|---|---| +| 2 | 1 | nothing — a difference, not a trend | +| 3 | 2 | a second interval exists, but one anomalous interval is **half** the evidence and a 1–1 split has no tiebreak | +| 4 | 3 | 2-of-3: repeated behaviour outvotes a single anomaly | + +4 is the smallest n at which "it did this more than once" beats "it did this +once". Cross-check, not derivation: it equals `AutoscaleMaxModificationsPer24h` +— AWS's documented ceiling of four storage modifications per 24 hours — so +below four observations a maximally active day cannot even in principle be +resolved into separate events. + +Cross-check against the substrate: at `pkg/store`'s default 1 h cadence a +14-day span retains up to 336 snapshots, against a count cap of 768. So 4 never +binds on an hourly-checkpointing deployment; it binds exactly on a sparse one, +which is when it should. **[unverified: that `cmd/` will checkpoint at ingest +cadence rather than once per manual invocation — §6 is not written yet.]** + +### 2.2 `DefaultGrowthMinSpan = 14 days` → `storage-growth-span-too-short` + +A database's write volume has a **weekly** shape, and this package already +encodes that belief: `DefaultMinWindow = 72h` is documented as deliberately +*not* enough to contain a weekly batch job. One weekly cycle is one event, so +two is the minimum at which "it happened again" is a statement rather than an +assumption: **2 × 7 = 14 days**. + +It coincides with `DefaultFullWindow`, the span at which this package already +attaches no window caveat to a metric verdict. That coincidence is wanted: a +growth finding never claims more confidence than the CPU and memory verdicts +printed on the same line. + +What the two bars mean together, by checkpoint cadence: + +| cadence | 4 observations at | 14-day span at | first verdict | which code holds last | +|---|---|---|---|---| +| hourly | +3 h | +14 d | day 14 | span | +| daily | +3 d | +14 d | day 14 | span | +| weekly | +21 d | +14 d | day 21 | count | +| monthly | +90 d | +14 d | day 90 | count | + +`TestHistorySurvivesTheCheckpointLoopAndReachesAVerdict` walks the daily row +one day at a time and asserts which code is present on each. + +### 2.3 `DefaultGrowthMinSteps = 2` → `storage-growth-rate-not-projectable` + +This one gates the **rate only**, never the size. Storage autoscaling adds +discrete steps; one step is one event, and a per-day rate over one event +asserts a recurrence the evidence does not contain — the exact failure the +brief calls "a slope that predicts an absurd future". Two is the minimum at +which recurrence is *observed*. + +The size is stated anyway, because the size is a measured fact about the past +and does not depend on this gate at all. +`TestSingleStepGrowthStatesTheSizeAndWithholdsTheRate` pins both halves: 200 +GiB reported, no `Projection`. + +### 2.4 `DefaultGrowthMaxGapFraction = 0.5` → `storage-growth-rate-not-projectable` + +If the largest interval between two retained observations exceeds half the +span, most of the span is a single blind spot and the whole increase could have +landed inside it. At `MinObservations` evenly spaced the largest gap is +span/3 ≈ 33.3 %, so an evenly-sampled history never trips this; it trips +exactly when the retained instants are clustered. + +A policy may make every bar **stricter and never looser** — +`GrowthPolicy.normalized()` clamps up, so a caller cannot configure its way to +a two-sample trend. `TestGrowthPolicyCannotBeLoosened`. + +## 3. Every code added, and the test that pins it + +| Code | Fires when | Test | +|---|---|---| +| `storage-growth-insufficient-observations` | 0 observations (`GrowthNoHistory`) or fewer than `MinObservations` (`GrowthTooFewObservations`) | `TestGrowthRefusesTooFewObservations` | +| `storage-growth-span-too-short` | enough observations, span below `MinSpan` | `TestGrowthRefusesTooShortASpan` | +| `storage-growth-history-inconsistent` | an allocation decreased, or one instant carries two allocations | `TestGrowthRefusesAnInconsistentHistory` | +| `storage-growth-is-not-reclaimable` | **the finding**: a measured, permanent increase, with its size | `TestGrowthIsARefusalWithASize` | +| `storage-growth-rate-not-projectable` | measured and sized; the rate is withheld (single step, or irregular sampling) | `TestSingleStepGrowthStatesTheSizeAndWithholdsTheRate`, `TestIrregularSamplingWithholdsTheRateAndIsReported` | +| `allocated-storage-grew` (advisory) | alongside the refusal, carrying the permanent monthly cost | `TestGrowthReachesTheReportAsARefusalAndNeverAsAProposal` | + +`TestGrowthCodesAreDistinctFromEveryExistingCode` re-runs U11's +`TestReasonCodesAreDistinct` discipline across the union of the old 31 codes +and the new 6 — two codes sharing a value silently merge two findings. + +**Two states share one code, on purpose.** `GrowthNoHistory` and +`GrowthTooFewObservations` both emit +`storage-growth-insufficient-observations`. The *code* is the refusal ("I will +not state a growth"); the *state* is the diagnosis ("your persistence is not +wired" vs. "your persistence works and is young"). Those are different +sentences to a human and the same decision to a caller filtering by code. + +## 4. "Not enough history" is not "flat", and cannot be collapsed into it + +This is the trap the unit is shaped around, and it is closed structurally +rather than by documentation: + +- `GrowthVerdict.Measurement` is a **pointer**, `nil` on every refusal state. + A short history carries no number, so there is no zero to misread as flat. +- `GrowthState.Measured()` is the only correct test, and + `Measurement != nil ⟺ State.Measured()` is a biconditional pinned over a + 9-case table (`TestGrowthMeasurementExistsExactlyWhenMeasured`) and re-checked + on every fuzz input. +- `GrewGiB()` and `Rate()` return `(value, ok)`. A caller that drops the + boolean gets 0 — and had to drop it deliberately. +- The zero value `GrowthUnevaluated` is a **third** state: "the question was + never put", which is what an excluded instance carries. Three states, no two + reachable from one another by reading a number. + +`TestNotEnoughHistoryIsNotFlat` asserts the three are pairwise distinct and +that only the measured one answers `GrewGiB`. This is the same rule the +substrate applies at 422 rather than 500: a request that cannot be answered is +not answered with a default. + +## 5. Irregular sampling: reported, not smoothed + +`pkg/store` retains at most one snapshot per cadence bucket and then prunes by +count and by bytes, so the retained instants are irregular **by construction**. +Nothing here divides by an observation count: + +- **Span** is `Last − First` from the actual instants. `SpanDays` is that in + days, and the projected `GiBPerDay` is `GrewGiB ÷ SpanDays`. + `TestGrowthRateUsesTheActualInstants` builds a front-loaded history where + growth÷count and growth÷span differ, and asserts the rate equals the second + and **not** the first. +- **`GrowthSampling`** ships `MinGap`, `MaxGap`, `MeanGap` and + `LargestGapFraction` on every verdict, refusals included. `MeanGap` is + present precisely so the distance between it and `MinGap`/`MaxGap` is + visible; no arithmetic reads it. +- **`Truncated`** says when the retention bound dropped older observations, so + a reader knows `Span` is the retained history and not the instance's + lifetime. +- The prose says it out loud: every rendered sampling line contains *"the + instants are thinned by the store's retention policy and are not evenly + spaced — gaps run X to Y against a Z mean, and the largest single gap is N% + of the span"*. + +Above `MaxGapFraction` the rate is withheld rather than caveated, because past +that point the number is arithmetic rather than evidence. + +## 6. What `cmd/` must do — exactly two calls, plus the persistence + +`cmd/RDSLIVE-FINDINGS.md` §5.2 is right that this needs cross-run persistence, +and right that `cmd/` cannot reach it alone today. What it needs turned out to +be smaller than a new store: the history is a field on `Target`, and +`Domain.Checkpoint`/`Restore` already serialize `Target` whole. **No new bucket, +no new key, no new codec.** + +The one obstacle is that `Domain.Observe` *replaces* a target wholesale (it +must — the collector re-queries a window each tick and accumulating would +double-count), so a restored history would be discarded by the next collection. +Two additive helpers close that without touching `domain.go`: + +```go +d, _ := krds.NewDomain(sc) // sc.Growth defaults; no flag needed +if blob, err := st.LoadRDSCheckpoint(scope); err == nil { + _ = d.Restore(blob) // ← the missing read side +} + +snap, err := collector.Collect(ctx) +rds.RecordStorageHistory(snap, d.StorageHistories(), sc.Growth) // ← call 1 +_ = d.Observe(snap) // unchanged + +blob, _ := d.Checkpoint() +_ = st.SaveRDSCheckpoint(scope, blob) // ← the missing write side +``` + +`RecordStorageHistory` is idempotent — re-recording an instant replaces it +rather than doubling it, the same rule `SaveSnapshotAt` applies — so a retried +run does not manufacture a step. `StorageHistories` returns copies, so a caller +cannot reach into the domain through them. Both are asserted in +`TestHistorySurvivesTheCheckpointLoopAndReachesAVerdict`, which runs the full +Restore → Record → Observe → Checkpoint loop for twenty simulated days with two +autoscaling events in it, and never constructs a store. + +Three things `cmd/` owns and this unit deliberately did not do: + +1. **`SaveRDSCheckpoint`/`LoadRDSCheckpoint` do not exist.** `pkg/store` is out + of scope here. `SaveEvidenceCheckpoint`/`LoadEvidenceCheckpoint` is the + shape to copy — a scope-keyed, gzip-framed, version-refusing envelope — and + `WIRING-FINDINGS.md` §6.2/§6.3 want the same write side, so the three jobs + still share most of their work exactly as §5.2 predicted. +2. **The instant is the snapshot's**, falling back to `Window.End`. `cmd/` + must set `Snapshot.Timestamp`; the collector already does. A snapshot with + neither is skipped rather than stamped, because this package reads no clock + and a history keyed by when the process happened to run is not replayable. +3. **No CLI surface.** There is no `--rds-growth` flag and no way to loosen the + policy from a file, which is the `--rds-parity-rates` precedent inverted: + provenance is the gate between a magnitude and a claim, and a two-sample + trend is exactly the claim §7.4 deferred. A caller may pass a stricter + `Config.Growth` in code. + +Until that persistence lands, every modelled instance carries +`Assessment.Growth.State == GrowthNoHistory` with its reason filled, and **no +suppression** — see §7. + +## 7. Statefulness: what was traded, and what it costs + +§7.4's argument for a stateless package was "a stateless read-only report has +nothing to decay and nothing to forecast". That is now false by one field, and +the honest accounting: + +**What stayed pure.** `AssessStorageGrowth` is a pure function: no clock, no +I/O, no package state. `TestNoClockReads`, `TestNoForeignImports` and +`TestNoUnexpectedPackageState` all still pass unmodified. The *sizer* is still +pure — the state is on the input, not in the package. + +**What became stateful.** `Target.StorageHistory` is a growing slice that must +survive between runs for the finding to exist. Concretely: + +- **Report size.** Each measured assessment gains a `GrowthVerdict`, including + an echoed `GrowthPolicy` (~120 bytes/instance of duplication against + `Report.Config`). Kept anyway, for the reason `Report.Config` is echoed at + all: a verdict handed around alone must be able to state the bar it failed, + and a UI needs "3 of 4" numerically, not in prose. +- **Checkpoint size.** Bounded at `DefaultGrowthMaxObservations = 768` + observations per instance — `pkg/store`'s `MaxPerCluster`, stated + independently because `pkg/rds` imports nothing. At ~40 bytes per observation + that is ~30 KiB per instance worst case, against `pkg/store`'s 32 MiB + per-cluster budget. A 1,000-instance account at the bound is ~30 MiB of + checkpoint. **[unverified: no such account has been measured.]** A tighter + `Config.Growth.MaxObservations` is the knob. +- **A new failure mode.** The finding now depends on the caller's persistence + being correct. A checkpoint restored from another scope, or a history + attached to a reused identifier, produces a decrease — which is why + `storage-growth-history-inconsistent` refuses the whole history rather than + repairing it. Data faults surface as refusals, not as confident numbers. +- **What did not decay.** Nothing here forecasts a workload or ages a + confidence. The only forward-looking value is `GiBPerDay`, it carries + `[unverified]` in its own `String()`, and it carries no money — see §8. + +**What was reused rather than reinvented.** `AttributeStorageGrowth` is the +engine: it is called pairwise across the series, and its three outcomes drive +`Steps`, `LargestStepGiB` and the inconsistency refusal. §7.4's ledger rule +still has one implementation; this unit gave it a second arity, not a second +copy. + +## 8. The rate is never money + +The fourth trap, closed three ways: + +1. `GrowthProjection` has **no** field matching `usd`, `cost`, `price`, + `saving`, `dollar` or `monthly`. Reflection test. +2. `Claimable()` is a **method** returning `false` unconditionally — the + `Advisory.Actuatable()` precedent — so no serialized form and no future + struct literal can claim otherwise. +3. `GrowthProjection.String()` carries the literal `[unverified]` marker, so a + caller formatting the number gets the marker whether it wanted it or not. + +The only dollar the finding produces is `GrowthMeasurement.MonthlyUSD`: the +monthly cost of storage **already allocated**, priced by the same +`RateCard.StorageMonthlyUSD` that prices the floor, carrying the same +`RateProvenance`. It flows into `Assessment.WorstRateProvenance` and so into +`unverified-rate` exactly like every other RDS dollar, and into +`Report.StorageGrowth().UnrecoverableGrowthMonthlyUSD` — a field named after a +cost — and never into a savings column. +`Report.Validate` needed no change: it gates proposals, and this unit produces +none. + +## 9. Things that would falsify parts of this unit + +1. **If a real fleet's growth is dominated by single large steps**, `MinSteps` + makes `storage-growth-rate-not-projectable` the common case and the rate is + near-dead weight. That is the *intended* failure mode — the size still ships + — but if it holds universally, the projection should be deleted rather than + defended. +2. **If `cmd/` ends up checkpointing only on manual invocation**, a weekly + operator waits 21 days for a first verdict (§2.2 table) and `MinObservations` + binds where `MinSpan` was meant to. The fix is the cadence, not the bar. +3. **If AWS ever ships an API that reduces allocated storage**, §1's entire + argument inverts, `storage-growth-is-not-reclaimable` becomes a false + statement, and this file is the first thing to delete. Nothing about that is + subtle, which is why the refusal is named after the claim it makes. +4. **If `MaxGapFraction = 0.5` proves too permissive** — a history with two + clusters 45 %/55 % apart passes and shouldn't — the constant moves and + nothing else does. It is read in exactly one place. diff --git a/pkg/rds/growth.go b/pkg/rds/growth.go new file mode 100644 index 0000000..8b530c4 --- /dev/null +++ b/pkg/rds/growth.go @@ -0,0 +1,979 @@ +package rds + +import ( + "fmt" + "sort" + "time" + + "github.com/agenticode/kilter/pkg/domain" +) + +// Allocated-storage growth — FINDINGS.md §7.4's "the one genuinely +// time-series-shaped RDS finding", built on the seam that section names: +// [AttributeStorageGrowth], applied across a series of observations instead of +// across a pair. +// +// # What this produces, and what it does not +// +// The product is A REFUSAL WITH A SIZE. "This instance's allocated storage +// grew 240 GiB over 21 days, here is the evidence, and no reduction is +// available." [ReasonGrowthIsNotReclaimable] is that sentence. Trap 8 is +// unchanged and unweakened by any of it: MaxAllocatedStorage is a RATCHET, not +// headroom, so measuring how fast the ratchet turned is not a step toward +// turning it back. There is no such step. [GrowthMeasurement] therefore has no +// field named after a saving, a reduction or a reclaim, and +// TestGrowthNamesNoSaving asserts that by reflection rather than by comment. +// +// # Two samples are not a trend, and this file is mostly about that +// +// Storage autoscaling adds "the greater of 10 GiB, 10 % of the current +// allocation, or the growth predicted for the next 7 hours" [verified, see +// [AutoscalingTrigger]] in DISCRETE STEPS, up to four times in 24 hours. A +// single scale-up inside the observation window is one event; a slope fitted +// through it predicts that the event recurs, which is exactly the claim one +// event cannot support. So the gates below come first and the arithmetic +// second: +// +// - fewer than [GrowthPolicy.MinObservations] instants ⇒ +// [ReasonGrowthInsufficientObservations], and NOTHING is measured; +// - a span below [GrowthPolicy.MinSpan] ⇒ [ReasonGrowthSpanTooShort], +// and nothing is measured; +// - an allocation that DECREASED, or two contradicting observations at one +// instant ⇒ [ReasonGrowthHistoryInconsistent] — no RDS API reduces +// allocated storage, so such a history is two instances under one +// identifier and is not averaged into a trend; +// - measured, but the growth arrived in a single step or behind one +// unobserved gap covering most of the span ⇒ the SIZE is stated and the +// RATE is withheld with [ReasonGrowthRateNotProjectable]. +// +// "Not enough history" and "measured, and flat" are different facts and this +// file makes them structurally different values, not two readings of one zero. +// See [GrowthVerdict.Measurement]. +// +// # The retained history is thinned, so instants are not evenly spaced +// +// pkg/store retains at most one snapshot per cadence bucket and prunes by +// count and by bytes, so a history is a sequence of IRREGULAR instants and a +// slope computed as (last−first)/count would be a fiction. Everything here +// works from the actual instants: the span is Last−First and the gaps between +// consecutive observations are reported as [GrowthSampling] rather than +// smoothed into a mean. [GrowthSampling.LargestGapFraction] is the number that +// says how much of the "observed" span nobody actually observed. +// +// # A rate is a claim about the future +// +// [GrowthProjection] carries no dollar field of any kind, its Claimable method +// returns false unconditionally, and nothing in this file writes to +// [Assessment.Proposal]. A growth projection cannot become a cost claim +// through a side door because there is no door: [Report.Validate] gates +// proposals, and this file produces none. + +// Growth policy defaults. Each is a POLICY choice, not an AWS fact, and the +// arithmetic behind each is in GROWTH-FINDINGS.md §2. +const ( + // DefaultGrowthMinObservations is the fewest distinct instants this + // package will state anything from. n observations bound n−1 intervals: + // 2 gives one interval, which is a difference rather than a trend; 3 + // gives two, where a single anomalous interval is half the evidence and + // ties cannot be broken; 4 gives three, the smallest n at which repeated + // behaviour outvotes one anomaly. It is also + // [AutoscaleMaxModificationsPer24h], the documented ceiling on storage + // modifications in a day — below four observations a maximally active day + // cannot even in principle be resolved into separate events. + DefaultGrowthMinObservations = 4 + // DefaultGrowthMinSpan is the shortest history this package will state + // anything over. A database's write volume has a WEEKLY shape — this + // package already encodes that belief in [DefaultMinWindow], documented as + // deliberately not long enough to contain a weekly batch job — and one + // weekly cycle is one event, so two are the minimum at which "it happened + // again" is a statement rather than an assumption. 2 × 7 days. It + // coincides with [DefaultFullWindow], the span at which this package + // already attaches no window caveat, so a growth finding never claims more + // confidence than the metric verdicts printed beside it. + DefaultGrowthMinSpan = 14 * 24 * time.Hour + // DefaultGrowthMinSteps is the fewest strictly-increasing transitions + // before a RATE is projected from them. One step is one event; a rate + // derived from one event asserts that the event recurs, which is the one + // thing a single event contains no evidence for. Two is the minimum at + // which recurrence is observed. The size of the growth is reported either + // way — it is a measured fact and does not depend on this gate. + DefaultGrowthMinSteps = 2 + // DefaultGrowthMaxGapFraction bounds how much of the span may sit inside a + // single unobserved interval before the rate is withheld. At + // DefaultGrowthMinObservations evenly spaced, the largest gap is span/3 ≈ + // 0.333, so an evenly-sampled history never trips this; it trips exactly + // when the retained instants are clustered and the "span" is mostly blind. + // Above half the span the whole growth could have arrived in one interval + // nobody looked at, and a per-day rate over it is arithmetic rather than + // evidence. + DefaultGrowthMaxGapFraction = 0.5 + // DefaultGrowthMaxObservations bounds how much history is read. It is + // pkg/store's DefaultSnapshotRetention().MaxPerCluster, so this package + // refuses to hold more than the substrate beneath it retains — the two + // bounds are stated independently because pkg/rds imports nothing and must + // not learn pkg/store's number by importing it. + DefaultGrowthMaxObservations = 768 +) + +// SourceCheckpointHistory names where a growth observation came from: not one +// CloudWatch call and not one DescribeDBInstances call, but a sequence of them +// persisted across runs. Evidence carrying this source is evidence that +// depends on the caller's persistence being wired — see GROWTH-FINDINGS.md §6. +const SourceCheckpointHistory = "checkpointed-observation-history" + +// Growth reason codes. All five are refusals: this domain is modelling the +// instance and declining to state something about it. None is an exclusion. +const ( + // ReasonGrowthInsufficientObservations: the history holds fewer distinct + // instants than [GrowthPolicy.MinObservations]. Two samples are not a + // trend and neither are three. NOTHING is measured — the verdict carries + // no [GrowthMeasurement] at all, so "we have not looked long enough" + // cannot be misread as "we looked and it is flat". + ReasonGrowthInsufficientObservations = "storage-growth-insufficient-observations" + + // ReasonGrowthSpanTooShort: enough observations arrived, but they cover + // less than [GrowthPolicy.MinSpan]. Four snapshots an hour apart describe + // an hour, and a database's growth shape is weekly. Distinct from + // [ReasonGrowthInsufficientObservations] because the remedies differ: one + // needs a denser checkpoint cadence, the other needs time to pass. + ReasonGrowthSpanTooShort = "storage-growth-span-too-short" + + // ReasonGrowthHistoryInconsistent: the history is not one instance's. + // Either an allocation DECREASED — which no RDS API can do, so the two + // observations are a replaced instance reusing an identifier — or one + // instant carries two different allocations. Averaging such a history into + // a trend would turn a data fault into a confident number, so the whole + // history is refused rather than repaired. See [StorageShrankImpossible]. + ReasonGrowthHistoryInconsistent = "storage-growth-history-inconsistent" + + // ReasonGrowthIsNotReclaimable: the finding itself, stated as a refusal + // because that is what it is. The allocation grew by a measured amount + // over a measured span, that increase is permanent, and no RDS API and no + // proposal in this package can return it. "You can't reduce the amount of + // storage for a DB instance after storage has been allocated" [verified]. + // A reader who takes this line as an opportunity has read it backwards. + ReasonGrowthIsNotReclaimable = "storage-growth-is-not-reclaimable" + + // ReasonGrowthRateNotProjectable: the growth was measured and sized, and + // the per-day RATE is withheld. Either every GiB arrived in a single step + // (fewer than [GrowthPolicy.MinSteps] increases) or more than + // [GrowthPolicy.MaxGapFraction] of the span sits inside one unobserved + // interval. A rate is a claim about the future; these histories do not + // contain one, and the size is reported without it. + ReasonGrowthRateNotProjectable = "storage-growth-rate-not-projectable" +) + +// AdvisoryStorageGrew reports a measured, permanent allocated-storage +// increase and its monthly cost. Like every advisory in this package the +// MonthlyUSD is a MAGNITUDE — here, the recurring cost the fleet acquired +// while nobody was asked — and never a saving. +const AdvisoryStorageGrew = "allocated-storage-grew" + +// GrowthState is the typed answer to "what does the history support?". +// +// The refusal states and the measured states are deliberately not +// distinguishable by a zero: a caller cannot fall into reading +// [GrowthTooFewObservations] as [GrowthFlat] by testing a number against zero, +// because a non-measured verdict carries no number to test. Use +// [GrowthState.Measured]. +type GrowthState string + +const ( + // GrowthUnevaluated is the zero value: the growth question was never put. + // A target that carried no history at all is [GrowthNoHistory], not this. + GrowthUnevaluated GrowthState = "" + // GrowthNoHistory: not one observation was supplied. Operationally this + // means the caller is not persisting checkpoints between runs, which is a + // property of the WIRING and not of the instance — so it is recorded on + // the assessment and, alone among these states, adds no suppression. See + // [Sizer.growthFindings]. + GrowthNoHistory GrowthState = "no-history" + // GrowthTooFewObservations: some history, below MinObservations. + GrowthTooFewObservations GrowthState = "too-few-observations" + // GrowthSpanTooShort: enough observations, too little elapsed time. + GrowthSpanTooShort GrowthState = "span-too-short" + // GrowthInconsistent: the observations are not one instance's history. + GrowthInconsistent GrowthState = "inconsistent-history" + // GrowthFlat: MEASURED, and the allocation never moved. This is a finding, + // not an absence of one. + GrowthFlat GrowthState = "flat" + // GrowthGrew: MEASURED, and the allocation ratcheted up. + GrowthGrew GrowthState = "grew" +) + +// Measured reports whether a number was produced. It is the only correct way +// to tell "flat" from "we could not say", and [GrowthVerdict.Measurement] is +// non-nil for exactly the states it answers true for. +func (s GrowthState) Measured() bool { return s == GrowthFlat || s == GrowthGrew } + +// GrowthPolicy is the evidence bar. Every field is a refusal threshold; see +// the Default* constants for the arithmetic behind each. +type GrowthPolicy struct { + MinObservations int `json:"minObservations"` + MinSpan time.Duration `json:"minSpan"` + MinSteps int `json:"minSteps"` + MaxGapFraction float64 `json:"maxGapFraction"` + MaxObservations int `json:"maxObservations"` +} + +// DefaultGrowthPolicy returns the shipped bar. +func DefaultGrowthPolicy() GrowthPolicy { + return GrowthPolicy{ + MinObservations: DefaultGrowthMinObservations, + MinSpan: DefaultGrowthMinSpan, + MinSteps: DefaultGrowthMinSteps, + MaxGapFraction: DefaultGrowthMaxGapFraction, + MaxObservations: DefaultGrowthMaxObservations, + } +} + +// normalized clamps a policy to the shipped floor. A caller may make the bar +// STRICTER and may not make it looser: a MinObservations of 2 would ship the +// exact claim §7.4 deferred this finding to avoid, so it is clamped up rather +// than honoured. +func (p GrowthPolicy) normalized() GrowthPolicy { + if p.MinObservations < DefaultGrowthMinObservations { + p.MinObservations = DefaultGrowthMinObservations + } + if p.MinSpan < DefaultGrowthMinSpan { + p.MinSpan = DefaultGrowthMinSpan + } + if p.MinSteps < DefaultGrowthMinSteps { + p.MinSteps = DefaultGrowthMinSteps + } + if p.MaxGapFraction <= 0 || p.MaxGapFraction > DefaultGrowthMaxGapFraction { + p.MaxGapFraction = DefaultGrowthMaxGapFraction + } + if p.MaxObservations <= 0 || p.MaxObservations > DefaultGrowthMaxObservations { + p.MaxObservations = DefaultGrowthMaxObservations + } + return p +} + +// StorageObservation is one instant's view of an instance's allocated storage: +// what [DescribeDBInstances] reported, and when it reported it. +type StorageObservation struct { + At time.Time `json:"at"` + AllocatedGiB int64 `json:"allocatedGiB"` + // MaxAllocatedGiB is the autoscaling ceiling at that instant, carried so a + // reader can see the ceiling move too. It takes no part in the arithmetic. + MaxAllocatedGiB int64 `json:"maxAllocatedGiB,omitempty"` +} + +// StorageHistory is one instance's observations, oldest first. It is the +// STATE this finding introduces — see GROWTH-FINDINGS.md §5 for what that +// costs — and it is carried on [Target] so it round-trips through the domain +// checkpoint with everything else. +type StorageHistory []StorageObservation + +// Append folds one observation into the history and returns the result, +// oldest first, bounded to max observations by dropping the oldest. +// +// Re-appending an instant already present REPLACES it, which is the same +// idempotence pkg/store's SaveSnapshotAt gives: re-ingesting a history a +// second time is a no-op rather than a doubling. An observation with a zero +// instant or a non-positive allocation is DROPPED — "we did not look" is not +// an observation of zero storage, and admitting it would put a fake step into +// every history whose first record predates the field. +// +// max ≤ 0 means [DefaultGrowthMaxObservations]. The receiver is not mutated. +func (h StorageHistory) Append(obs StorageObservation, max int) StorageHistory { + if max <= 0 || max > DefaultGrowthMaxObservations { + max = DefaultGrowthMaxObservations + } + out := make(StorageHistory, 0, len(h)+1) + replaced := false + for _, o := range h { + if o.At.IsZero() || o.AllocatedGiB <= 0 { + continue + } + if o.At.Equal(obs.At) { + if obs.At.IsZero() || obs.AllocatedGiB <= 0 { + continue + } + out = append(out, normalizeObs(obs)) + replaced = true + continue + } + out = append(out, normalizeObs(o)) + } + if !replaced && !obs.At.IsZero() && obs.AllocatedGiB > 0 { + out = append(out, normalizeObs(obs)) + } + sortHistory(out) + if len(out) > max { + out = out[len(out)-max:] + } + return out +} + +// normalizeObs puts an observation's instant in UTC so two histories that +// travelled through different time zones compare and sort identically. +func normalizeObs(o StorageObservation) StorageObservation { + o.At = o.At.UTC() + if o.MaxAllocatedGiB < 0 { + o.MaxAllocatedGiB = 0 + } + return o +} + +// sortHistory orders by instant, then by allocation so a same-instant +// contradiction has one canonical order and is DETECTED rather than resolved +// by whichever record happened to arrive first. +func sortHistory(h StorageHistory) { + sort.SliceStable(h, func(i, j int) bool { + if !h[i].At.Equal(h[j].At) { + return h[i].At.Before(h[j].At) + } + return h[i].AllocatedGiB < h[j].AllocatedGiB + }) +} + +// clean returns the history in canonical form and reports whether two +// observations at one instant disagreed. +// +// Identical (instant, allocation) pairs collapse silently — that is a replayed +// checkpoint, not a fault. Two different allocations at one instant do NOT +// collapse: that is a contradiction, and a contradiction is the caller's to +// explain rather than this package's to average away. +func (h StorageHistory) clean(max int) (StorageHistory, bool) { + if max <= 0 || max > DefaultGrowthMaxObservations { + max = DefaultGrowthMaxObservations + } + out := make(StorageHistory, 0, len(h)) + for _, o := range h { + if o.At.IsZero() || o.AllocatedGiB <= 0 { + continue + } + out = append(out, normalizeObs(o)) + } + sortHistory(out) + conflict := false + dedup := out[:0] + for i, o := range out { + if i > 0 && o.At.Equal(dedup[len(dedup)-1].At) { + if o.AllocatedGiB != dedup[len(dedup)-1].AllocatedGiB { + conflict = true + } + continue + } + dedup = append(dedup, o) + } + if len(dedup) > max { + dedup = dedup[len(dedup)-max:] + } + return dedup, conflict +} + +// GrowthSampling describes the SHAPE of the retained history, and is filled +// for every state including the refusals: a refusal that does not say how +// short the history was is not actionable. +// +// The gaps are reported rather than averaged. pkg/store thins its history to +// at most one snapshot per cadence bucket and prunes by count and by bytes, so +// the retained instants are irregular by construction and a mean gap would +// describe a history that was never recorded. +type GrowthSampling struct { + Observations int `json:"observations"` + First time.Time `json:"first,omitzero"` + Last time.Time `json:"last,omitzero"` + // Span is Last−First, from the ACTUAL instants. It is not + // Observations × any assumed cadence. + Span time.Duration `json:"span,omitempty"` + // MinGap and MaxGap bound the intervals between consecutive observations. + MinGap time.Duration `json:"minGap,omitempty"` + MaxGap time.Duration `json:"maxGap,omitempty"` + // MeanGap is Span ÷ (Observations−1): the interval an evenly-spaced + // history WOULD have had. It is reported beside MinGap/MaxGap precisely so + // the distance between them is visible, and no arithmetic here uses it. + MeanGap time.Duration `json:"meanGap,omitempty"` + // LargestGapFraction is MaxGap ÷ Span — the share of the span that sits + // inside a single interval nobody observed. The headline irregularity + // number, and the one [ReasonGrowthRateNotProjectable] tests. + LargestGapFraction float64 `json:"largestGapFraction,omitempty"` + // Regular is LargestGapFraction ≤ the policy's MaxGapFraction. + Regular bool `json:"regular,omitempty"` + // Truncated reports that the supplied history was longer than the policy + // reads, so Span is the retained span and not the instance's lifetime. + Truncated bool `json:"truncated,omitempty"` +} + +// Sampling measures a history's shape under a policy, without judging it. +func (h StorageHistory) Sampling(p GrowthPolicy) GrowthSampling { + p = p.normalized() + clean, _ := h.clean(p.MaxObservations) + s := GrowthSampling{Observations: len(clean), Truncated: len(clean) < countUsable(h)} + if len(clean) == 0 { + return s + } + s.First, s.Last = clean[0].At, clean[len(clean)-1].At + s.Span = s.Last.Sub(s.First) + if len(clean) < 2 { + return s + } + s.MinGap, s.MaxGap = s.Span, 0 + for i := 1; i < len(clean); i++ { + gap := clean[i].At.Sub(clean[i-1].At) + if gap < s.MinGap { + s.MinGap = gap + } + if gap > s.MaxGap { + s.MaxGap = gap + } + } + s.MeanGap = s.Span / time.Duration(len(clean)-1) + if s.Span > 0 { + s.LargestGapFraction = float64(s.MaxGap) / float64(s.Span) + } + s.Regular = s.LargestGapFraction <= p.MaxGapFraction + return s +} + +// countUsable is how many observations survive the "we did not look is not a +// zero" filter, before the retention bound is applied. +func countUsable(h StorageHistory) int { + n := 0 + for _, o := range h { + if !o.At.IsZero() && o.AllocatedGiB > 0 { + n++ + } + } + return n +} + +// GrowthProjection is the per-day rate, and it is a CLAIM ABOUT THE FUTURE. +// +// It has no dollar field, no saving, and no reclaim, and Claimable is a method +// returning false unconditionally so no serialized form and no future struct +// literal can say otherwise. TestGrowthProjectionCarriesNoMoney asserts the +// absence by reflection. The growth's monthly COST lives on +// [GrowthMeasurement], where it describes money already being spent rather +// than money predicted. +type GrowthProjection struct { + // GiBPerDay is the measured growth divided by the ACTUAL span in days — + // Last−First, never Observations × an assumed cadence. [unverified]: it + // extrapolates a step function whose steps are driven by write volume this + // package does not observe. + GiBPerDay float64 `json:"giBPerDay"` + // Steps and SpanDays are the two numbers the rate was divided out of, so a + // reader can reconstruct it rather than trust it. + Steps int `json:"steps"` + SpanDays float64 `json:"spanDays"` +} + +// Claimable is false for every projection, always. A growth rate may motivate +// a question and may never become a number in a savings column. +func (GrowthProjection) Claimable() bool { return false } + +// String renders the projection with its unverified marker attached, so the +// marker cannot be dropped by a caller formatting the number itself. +func (g GrowthProjection) String() string { + return fmt.Sprintf("[unverified] ~%.2f GiB/day (%d increases over %.1f days)", + g.GiBPerDay, g.Steps, g.SpanDays) +} + +// GrowthMeasurement is what the history actually showed. It exists only when +// [GrowthState.Measured] is true, which is what keeps "not enough history" +// from being readable as "flat". +// +// Read the ABSENT fields as the specification, the same way [StorageVerdict] +// does: there is a MonthlyUSD and there is no SavingMonthlyUSD, no +// ReclaimableGiB and no RecommendedGiB, because allocated storage is a monotone +// ratchet and a field named after a reduction would be a promise no RDS API +// can keep. +type GrowthMeasurement struct { + FirstGiB int64 `json:"firstGiB"` + LastGiB int64 `json:"lastGiB"` + // GrewGiB is LastGiB−FirstGiB and is never negative: a decrease is + // [GrowthInconsistent] and never reaches here. + GrewGiB int64 `json:"grewGiB"` + // Steps is how many transitions strictly increased, classified by + // [AttributeStorageGrowth] — the same ledger rule §7.4 names, applied + // pairwise across the series. Steps == 1 means the whole growth is one + // event. + Steps int `json:"steps"` + // LargestStepGiB is the biggest single increase. A growth dominated by one + // step is a scale-up, not a trend, whatever the span says. + LargestStepGiB int64 `json:"largestStepGiB,omitempty"` + // MonthlyUSD is what the GROWTH now costs every month at the card's + // storage rate. It is a magnitude and a permanent one: this is money the + // account started spending without being asked and cannot stop spending. + MonthlyUSD float64 `json:"monthlyUSD,omitempty"` + RateProvenance RateProvenance `json:"rateProvenance,omitempty"` + // Projection is non-nil only when the rate survived the step and gap + // gates. Nil with a positive GrewGiB is the normal, correct outcome for a + // single autoscaling event. + Projection *GrowthProjection `json:"projection,omitempty"` +} + +// GrowthVerdict is this package's complete answer about one instance's +// allocated-storage history. +type GrowthVerdict struct { + State GrowthState `json:"state"` + Policy GrowthPolicy `json:"policy,omitzero"` + Sampling GrowthSampling `json:"sampling,omitzero"` + // Measurement is non-nil if and only if State.Measured(). That biconditional + // is the whole defence against a caller collapsing "we have not looked long + // enough" into "we looked and it is flat", and + // TestGrowthMeasurementExistsExactlyWhenMeasured pins it. + Measurement *GrowthMeasurement `json:"measurement,omitempty"` + // Code and Reason are the refusal this verdict carries. Every state except + // [GrowthFlat] has one — a flat, well-observed history is a measurement + // with nothing to refuse. + Code string `json:"code,omitempty"` + Reason string `json:"reason,omitempty"` + // RateCode and RateReason are the SECOND refusal a measured growth can + // carry: the size was stated and the rate was withheld. They are separate + // fields because a caller must be able to see that a growth was measured + // AND that its rate was refused, which one code slot cannot say. + RateCode string `json:"rateCode,omitempty"` + RateReason string `json:"rateReason,omitempty"` +} + +// Measured is true when a number was produced. Prefer it to any test against a +// field: on a refusal there is no field to test. +func (v GrowthVerdict) Measured() bool { return v.State.Measured() } + +// GrewGiB returns the measured increase and whether one was measured at all. +// A caller that ignores the boolean gets 0, which is indistinguishable from a +// flat instance — the two-return form exists so that mistake has to be made on +// purpose. +func (v GrowthVerdict) GrewGiB() (int64, bool) { + if v.Measurement == nil { + return 0, false + } + return v.Measurement.GrewGiB, true +} + +// Rate returns the projected growth rate and whether one was projectable. +// The projection is [unverified] and carries no money; see [GrowthProjection]. +func (v GrowthVerdict) Rate() (GrowthProjection, bool) { + if v.Measurement == nil || v.Measurement.Projection == nil { + return GrowthProjection{}, false + } + return *v.Measurement.Projection, true +} + +// AssessStorageGrowth measures one instance's allocated-storage history. +// +// It is pure: no clock, no I/O, no package state. The instants come from the +// history the caller persisted, and every threshold comes from p. card may be +// a zero RateCard, in which case the growth is reported in GiB with no dollar +// beside it — a weaker finding, and still the right one. +// +// It never proposes anything. There is no return path to [Proposal] from here. +func AssessStorageGrowth(d DBInstance, h StorageHistory, p GrowthPolicy, card RateCard) GrowthVerdict { + p = p.normalized() + v := GrowthVerdict{Policy: p, Sampling: h.Sampling(p)} + clean, conflict := h.clean(p.MaxObservations) + name := d.DisplayName() + + if conflict { + v.State = GrowthInconsistent + v.Code = ReasonGrowthHistoryInconsistent + v.Reason = fmt.Sprintf( + "the allocated-storage history of %s records two different allocations at one instant, so it "+ + "is not one instance's history — a replaced instance reusing an identifier is the common "+ + "cause. No growth is stated from it: averaging a contradiction into a trend turns a data "+ + "fault into a confident number. Nothing here implies a reduction is available; %s", + name, ratchetClause) + return v + } + if len(clean) == 0 { + v.State = GrowthNoHistory + v.Code = ReasonGrowthInsufficientObservations + v.Reason = fmt.Sprintf( + "no allocated-storage history was supplied for %s, so no growth is stated. Storage autoscaling "+ + "moves the floor on its own and \"autoscaling operations aren't logged by AWS CloudTrail\", "+ + "so the only way to see the floor move is to compare persisted observations across runs — "+ + "which requires the caller to checkpoint this domain between invocations. This is a "+ + "property of the wiring, not of the instance", name) + return v + } + if len(clean) < p.MinObservations { + v.State = GrowthTooFewObservations + v.Code = ReasonGrowthInsufficientObservations + v.Reason = fmt.Sprintf( + "%s has %d allocated-storage observation(s), below the %d this package will state a growth "+ + "from. %d observations bound %d interval(s); two samples are a difference rather than a "+ + "trend, and storage autoscaling adds lumpy discrete steps, so a slope drawn through too "+ + "few of them predicts a future the evidence does not contain. This is NOT a report that "+ + "the allocation is flat — nothing was measured. Wait for more observations", + name, len(clean), p.MinObservations, len(clean), max(len(clean)-1, 0)) + return v + } + if v.Sampling.Span < p.MinSpan { + v.State = GrowthSpanTooShort + v.Code = ReasonGrowthSpanTooShort + v.Reason = fmt.Sprintf( + "%s has %d allocated-storage observations but they span only %s, below the %s minimum. A "+ + "database's write volume has a weekly shape, and one weekly cycle is one event; %s is the "+ + "shortest history in which \"it happened again\" is a statement rather than an assumption. "+ + "This is NOT a report that the allocation is flat — nothing was measured. Denser sampling "+ + "does not fix it; only elapsed time does", + name, len(clean), fmtSpan(v.Sampling.Span), fmtSpan(p.MinSpan), fmtSpan(p.MinSpan)) + return v + } + + // The ledger rule from trap 8, applied pairwise across the series rather + // than across a single pair. One implementation of the rule, one test. + m := &GrowthMeasurement{FirstGiB: clean[0].AllocatedGiB, LastGiB: clean[len(clean)-1].AllocatedGiB} + for i := 1; i < len(clean); i++ { + attr, why := AttributeStorageGrowth(clean[i-1].AllocatedGiB, clean[i].AllocatedGiB) + switch attr { + case StorageShrankImpossible: + v.State = GrowthInconsistent + v.Code = ReasonGrowthHistoryInconsistent + v.Reason = fmt.Sprintf( + "the allocated-storage history of %s is not monotone and therefore is not one instance's: "+ + "%s. No growth rate is drawn through it, and no reduction is inferred from the "+ + "decrease — %s", name, why, ratchetClause) + return v + case StorageGrewUnattributed: + step := clean[i].AllocatedGiB - clean[i-1].AllocatedGiB + m.Steps++ + if step > m.LargestStepGiB { + m.LargestStepGiB = step + } + } + } + m.GrewGiB = m.LastGiB - m.FirstGiB + if m.GrewGiB <= 0 { + v.State = GrowthFlat + v.Measurement = m + return v + } + + v.State = GrowthGrew + if usd, prov, ok := card.StorageMonthlyUSD(d.StorageType, m.GrewGiB); ok { + m.MonthlyUSD, m.RateProvenance = zeroIfNotFinite(usd), prov + } + + // The rate, and the two ways it is withheld. The SIZE above does not + // depend on either gate: it is a measured fact about the past. + days := v.Sampling.Span.Hours() / 24 + switch { + case m.Steps < p.MinSteps: + v.RateCode = ReasonGrowthRateNotProjectable + v.RateReason = fmt.Sprintf( + "%s grew %d GiB across %s, and no growth RATE is stated from it: the whole increase arrived in "+ + "%d step(s), below the %d this package projects from. Storage autoscaling adds \"the "+ + "greater of %d GiB, %s of the current allocation, or the growth predicted for the next "+ + "%s\" in discrete steps, so one step is one event and a per-day rate over it asserts a "+ + "recurrence the evidence does not contain. The size above is measured; the future is not", + name, m.GrewGiB, fmtSpan(v.Sampling.Span), m.Steps, p.MinSteps, + AutoscaleMinIncrementGiB, fmtPct(AutoscaleIncrementFraction), AutoscalePredictionHorizon) + case !v.Sampling.Regular: + v.RateCode = ReasonGrowthRateNotProjectable + v.RateReason = fmt.Sprintf( + "%s grew %d GiB across %s, and no growth RATE is stated from it: the largest gap between two "+ + "retained observations is %s, %s of the span, above the %s bound. The retained history is "+ + "thinned rather than evenly sampled, and when most of a span sits inside one interval "+ + "nobody observed, the whole increase could have arrived inside it. The size above is "+ + "measured; the rate would be arithmetic rather than evidence", + name, m.GrewGiB, fmtSpan(v.Sampling.Span), fmtSpan(v.Sampling.MaxGap), + fmtPct(v.Sampling.LargestGapFraction), fmtPct(p.MaxGapFraction)) + case days > 0: + m.Projection = &GrowthProjection{ + GiBPerDay: zeroIfNotFinite(float64(m.GrewGiB) / days), + Steps: m.Steps, + SpanDays: days, + } + } + + v.Code = ReasonGrowthIsNotReclaimable + v.Reason = fmt.Sprintf( + "the allocated storage of %s grew %d GiB — from %d GiB to %d GiB — across %s of retained history, "+ + "in %d observation(s) and %d increase(s), the largest of them %d GiB%s. This domain has no "+ + "actuator, so it was not Kilter; storage autoscaling is the usual cause and its operations are "+ + "not logged by CloudTrail, so there is no event to correlate against, and the increase is "+ + "recorded as unattributed. None of it is recoverable: %s. This is the measured size of a "+ + "permanent cost, not a reduction that is available", + name, m.GrewGiB, m.FirstGiB, m.LastGiB, fmtSpan(v.Sampling.Span), v.Sampling.Observations, + m.Steps, m.LargestStepGiB, growthCostClause(m), ratchetClause) + v.Measurement = m + return v +} + +// ratchetClause is the one sentence every growth refusal ends on, quoted so +// the reader gets AWS's words rather than this package's paraphrase +// [verified: USER_PIOPS.Autoscaling.html]. +const ratchetClause = "\"You can't reduce the amount of storage for a DB instance after storage has been " + + "allocated\", and the documented alternatives are a blue/green deployment or a migration to a new " + + "instance, neither of which Kilter performs" + +func growthCostClause(m *GrowthMeasurement) string { + if m.MonthlyUSD <= 0 { + return "" + } + return fmt.Sprintf(", which now costs %s/mo at a %s storage rate", fmtUSD(m.MonthlyUSD), m.RateProvenance) +} + +// fmtSpan renders a duration in days-and-hours, because "504h0m0s" is not a +// span a reader can weigh against a weekly cycle. +func fmtSpan(d time.Duration) string { + if d <= 0 { + return "0h" + } + days := int64(d / (24 * time.Hour)) + rem := d - time.Duration(days)*24*time.Hour + hours := int64(rem / time.Hour) + switch { + case days == 0 && hours == 0: + return fmt.Sprintf("%dm", int64(d/time.Minute)) + case days == 0: + return fmt.Sprintf("%dh", hours) + case hours == 0: + return fmt.Sprintf("%dd", days) + } + return fmt.Sprintf("%dd%dh", days, hours) +} + +// growthFindings turns the growth verdict into report output: evidence for the +// sampling shape, a refusal carrying the size, and an advisory carrying the +// permanent cost. +// +// The one state that produces no suppression is [GrowthNoHistory]. An empty +// history means the CALLER is not persisting checkpoints, which is one fact +// about the wiring and not N facts about N instances; stamping the same +// sentence on every row of a fleet report would bury the findings that are +// about instances. The verdict is still recorded on the assessment, typed and +// with its reason filled, so a caller can see it and GROWTH-FINDINGS.md §6 +// says what to do about it. +func (s *Sizer) growthFindings(a *Assessment, t Target, snap *Snapshot) { + v := AssessStorageGrowth(a.Instance, t.StorageHistory, s.cfg.Growth, s.cfg.Rates) + a.Growth = v + if v.State == GrowthNoHistory { + return + } + + a.Evidence = append(a.Evidence, domain.Evidence{ + Metric: "allocated-storage-history", Value: describeSampling(v.Sampling), + Window: fmtSpan(v.Sampling.Span), Samples: v.Sampling.Observations, + Source: SourceCheckpointHistory, At: v.Sampling.Last, + }) + + if v.Code != "" { + a.suppress(v.Code, v.Reason) + } + if v.RateCode != "" { + a.suppress(v.RateCode, v.RateReason) + } + + m := v.Measurement + if m == nil { + return + } + if v.State == GrowthFlat { + a.Evidence = append(a.Evidence, domain.Evidence{ + Metric: "allocated-storage-growth", + Value: fmt.Sprintf("flat: %d GiB unchanged across %s and %d observations", + m.LastGiB, fmtSpan(v.Sampling.Span), v.Sampling.Observations), + Window: fmtSpan(v.Sampling.Span), Samples: v.Sampling.Observations, + Source: SourceCheckpointHistory, At: v.Sampling.Last, + }) + return + } + + a.Evidence = append(a.Evidence, domain.Evidence{ + Metric: "allocated-storage-growth", + Value: fmt.Sprintf("grew %d GiB (%d → %d) in %d step(s) across %s", + m.GrewGiB, m.FirstGiB, m.LastGiB, m.Steps, fmtSpan(v.Sampling.Span)), + Window: fmtSpan(v.Sampling.Span), Samples: v.Sampling.Observations, + Source: SourceCheckpointHistory, At: v.Sampling.Last, + }) + if p, ok := v.Rate(); ok { + a.Evidence = append(a.Evidence, domain.Evidence{ + Metric: "allocated-storage-growth-rate", Value: p.String(), + Window: fmtSpan(v.Sampling.Span), Samples: v.Sampling.Observations, + Source: SourceCheckpointHistory, At: v.Sampling.Last, + }) + } + + a.advise(Advisory{ + Code: AdvisoryStorageGrew, + Message: fmt.Sprintf("%s acquired %d GiB of allocated storage (%d → %d GiB) across %s, in %d "+ + "increase(s) nothing in this account requested and no CloudTrail event records", + a.Instance.DisplayName(), m.GrewGiB, m.FirstGiB, m.LastGiB, + fmtSpan(v.Sampling.Span), m.Steps), + Caveat: "this is a permanent cost that has already been incurred, not a saving and not an " + + "opportunity: no RDS API reduces allocated storage, so the increase cannot be undone at any " + + "price. " + describeSampling(v.Sampling) + growthRateCaveat(v), + MonthlyUSD: m.MonthlyUSD, + RateProvenance: m.RateProvenance, + }) +} + +// growthRateCaveat states, inside the advisory's own caveat, whether a rate +// was projected — so the advisory cannot be read as carrying a forecast when +// it does not, or as carrying a verified one when it does. +func growthRateCaveat(v GrowthVerdict) string { + if p, ok := v.Rate(); ok { + return " Growth rate " + p.String() + ": a claim about the future from a step function whose steps " + + "this package does not observe the cause of, and never a dollar figure." + } + return " No growth rate is projected from this history; the size above is a measured fact about the past." +} + +// describeSampling renders the irregularity rather than smoothing it, because +// a reader weighing a growth number needs to know how much of the span was +// actually looked at. +func describeSampling(s GrowthSampling) string { + if s.Observations == 0 { + return "no retained observations" + } + if s.Observations == 1 { + return fmt.Sprintf("1 retained observation at %s", s.Last.UTC().Format(time.RFC3339)) + } + out := fmt.Sprintf("%d retained observations spanning %s (%s → %s); the instants are thinned by the "+ + "store's retention policy and are not evenly spaced — gaps run %s to %s against a %s mean, and the "+ + "largest single gap is %s of the span", + s.Observations, fmtSpan(s.Span), s.First.UTC().Format(time.RFC3339), + s.Last.UTC().Format(time.RFC3339), fmtSpan(s.MinGap), fmtSpan(s.MaxGap), fmtSpan(s.MeanGap), + fmtPct(s.LargestGapFraction)) + if s.Truncated { + out += ". Older observations were dropped by the retention bound, so the span is the retained " + + "history and not the instance's lifetime" + } + return out +} + +// GrowthTotals is the fleet-level roll-up of this finding, for a caller that +// wants the headline without walking every assessment. +// +// Note what is counted and what is not: there is a MonthlyUSD naming a cost +// and there is no saving, because none of this is recoverable. Refused counts +// the instances this package DECLINED to measure — the number that tells an +// operator their history is too young, which no total of measured growth +// could. +type GrowthTotals struct { + // Measured is instances with enough history to state something. + Measured int `json:"measured"` + // Grew and Flat partition Measured. Both are findings. + Grew int `json:"grew"` + Flat int `json:"flat"` + // Refused is instances with a history that did not clear the bar. It + // excludes instances with no history at all, which are NoHistory. + Refused int `json:"refused"` + // NoHistory is instances for which nothing was persisted — a count of the + // wiring gap, not of the fleet. + NoHistory int `json:"noHistory"` + // RateRefused is instances measured as having grown whose per-day rate was + // withheld. A subset of Grew. + RateRefused int `json:"rateRefused"` + + GrewGiB int64 `json:"grewGiB"` + // UnrecoverableGrowthMonthlyUSD is what the fleet's measured growth now + // costs every month, permanently. It is a cost and never a saving. + UnrecoverableGrowthMonthlyUSD float64 `json:"unrecoverableGrowthMonthlyUSD"` +} + +// StorageGrowth rolls the growth verdicts up across a report. It walks +// Assessments, which [Sizer.Assess] has already sorted, so the result does not +// depend on map order or on collection order. +func (r *Report) StorageGrowth() GrowthTotals { + var t GrowthTotals + if r == nil { + return t + } + for _, a := range r.Assessments { + v := a.Growth + switch { + case v.State == GrowthUnevaluated: + continue + case v.State == GrowthNoHistory: + t.NoHistory++ + continue + case !v.Measured(): + t.Refused++ + continue + } + t.Measured++ + if v.State == GrowthFlat { + t.Flat++ + continue + } + t.Grew++ + if v.RateCode != "" { + t.RateRefused++ + } + if m := v.Measurement; m != nil { + t.GrewGiB += m.GrewGiB + t.UnrecoverableGrowthMonthlyUSD += m.MonthlyUSD + } + } + return t +} + +// --- The persistence seam, and the only two calls cmd/ has to make ---------- + +// StorageHistories returns every instance's allocated-storage history as this +// domain currently holds it, keyed by target ID. +// +// It exists because [Domain.Observe] REPLACES a target wholesale — the +// collector re-queries a window each tick and accumulating would double-count +// the overlap — so a history restored from a checkpoint would be discarded by +// the next collection unless the caller carries it across. That is deliberate +// on Observe's part and it is why the history is the CALLER's to fold in; this +// method plus [RecordStorageHistory] are the two halves of doing so. +// +// The returned map and its slices are copies: a caller cannot reach into the +// domain's state through them. Ranging over the map is not deterministic, so +// use it as a lookup table and never as an output order. +func (d *Domain) StorageHistories() map[string]StorageHistory { + d.mu.RLock() + defer d.mu.RUnlock() + out := make(map[string]StorageHistory, len(d.targets)) + for id, t := range d.targets { + if len(t.StorageHistory) == 0 { + continue + } + out[id] = append(StorageHistory(nil), t.StorageHistory...) + } + return out +} + +// RecordStorageHistory folds each target's currently-observed allocated +// storage into the history that target carried before, and writes the result +// back onto the snapshot. +// +// Call it on a freshly collected snapshot BEFORE [Domain.Observe] or +// [Domain.Learn], with prior taken from [Domain.StorageHistories] after +// [Domain.Restore]. The history then rides the existing checkpoint — it is a +// field on [Target], which [Domain.Checkpoint] already serializes whole — so +// persisting this finding needs no new store, no new key and no new codec. +// +// The observation instant is the SNAPSHOT's, falling back to the end of its +// window: this package reads no clock, and a history keyed by when the process +// happened to run is not replayable. A snapshot with neither is skipped rather +// than stamped with a zero instant, because "we did not look" is not an +// observation. +// +// It is idempotent. Re-recording a snapshot already folded in replaces that +// instant rather than doubling it, the same rule pkg/store's SaveSnapshotAt +// applies, so a retried run does not manufacture a step. +func RecordStorageHistory(snap *Snapshot, prior map[string]StorageHistory, p GrowthPolicy) { + if snap == nil { + return + } + at := snap.Timestamp + if at.IsZero() { + at = snap.Window.End + } + if at.IsZero() { + return + } + max := p.normalized().MaxObservations + for i := range snap.Targets { + t := &snap.Targets[i] + h := t.StorageHistory + if len(h) == 0 { + h = prior[t.Ref.ID] + } + t.StorageHistory = h.Append(StorageObservation{ + At: at, + AllocatedGiB: t.Instance.AllocatedStorageGiB, + MaxAllocatedGiB: t.Instance.MaxAllocatedStorageGiB, + }, max) + } +} diff --git a/pkg/rds/growth_test.go b/pkg/rds/growth_test.go new file mode 100644 index 0000000..ed4b052 --- /dev/null +++ b/pkg/rds/growth_test.go @@ -0,0 +1,943 @@ +package rds + +import ( + "encoding/json" + "math" + "strings" + "testing" + "time" + + "github.com/agenticode/kilter/pkg/domain" +) + +// The growth finding's tests are mostly about the REFUSALS, because the +// refusals are the product. Every gate in growth.go has a test that pins its +// code, and the two collapses the design exists to prevent — "not enough +// history" read as "flat", and a growth read as a shrink opportunity — each +// have a test of their own. + +// obs builds one observation d before testEnd. +func obs(before time.Duration, gib int64) StorageObservation { + return StorageObservation{At: testEnd.Add(-before), AllocatedGiB: gib} +} + +// evenHistory lays n observations evenly across span, ending at testEnd, with +// allocation taken from alloc(i). Evenly spaced so the irregularity gate is +// not the thing under test unless a test makes it so. +func evenHistory(n int, span time.Duration, alloc func(i int) int64) StorageHistory { + h := make(StorageHistory, 0, n) + if n < 2 { + for i := 0; i < n; i++ { + h = append(h, obs(span, alloc(i))) + } + return h + } + // Instants are computed as span·i/(n−1) rather than i·(span/(n−1)) so the + // endpoints are EXACT: a test of the minimum-span boundary must not be + // decided by a truncated division. + for i := 0; i < n; i++ { + off := time.Duration(int64(span) * int64(i) / int64(n-1)) + h = append(h, obs(span-off, alloc(i))) + } + return h +} + +func flatHistory(n int, span time.Duration, gib int64) StorageHistory { + return evenHistory(n, span, func(int) int64 { return gib }) +} + +// growInstance is the instance every growth test measures: gp2, so the shipped +// storage rate can put a dollar on the finding. +func growInstance(gib int64) DBInstance { + return DBInstance{ + ARN: "arn:aws:rds:us-east-1:1234:db:grow", Identifier: "grow", + Class: "db.m6i.large", Engine: "postgres", LicenseModel: LicenseGPL, + StorageType: StorageGP2, AllocatedStorageGiB: gib, + } +} + +func growAssess(h StorageHistory, alloc int64) GrowthVerdict { + return AssessStorageGrowth(growInstance(alloc), h, DefaultGrowthPolicy(), DefaultRates()) +} + +// --- The two gates, and the arithmetic behind them -------------------------- + +// Fewer than MinObservations instants ⇒ nothing is measured, whatever the +// span. Three snapshots a month apart are still three snapshots. +func TestGrowthRefusesTooFewObservations(t *testing.T) { + for n := 0; n < DefaultGrowthMinObservations; n++ { + h := evenHistory(n, 60*24*time.Hour, func(i int) int64 { return 100 + int64(i)*50 }) + v := growAssess(h, 250) + if v.Measured() { + t.Fatalf("%d observations produced a measurement; two samples are not a trend", n) + } + if v.Measurement != nil { + t.Fatalf("%d observations produced a non-nil Measurement", n) + } + wantState := GrowthTooFewObservations + if n == 0 { + wantState = GrowthNoHistory + } + if v.State != wantState { + t.Fatalf("%d observations gave state %q, want %q", n, v.State, wantState) + } + if v.Code != ReasonGrowthInsufficientObservations { + t.Fatalf("%d observations gave code %q, want %q", n, v.Code, ReasonGrowthInsufficientObservations) + } + if _, ok := v.GrewGiB(); ok { + t.Fatalf("%d observations reported a growth", n) + } + } + // One more clears it, over the same span. + h := evenHistory(DefaultGrowthMinObservations, 60*24*time.Hour, func(i int) int64 { return 100 + int64(i)*50 }) + if v := growAssess(h, 250); !v.Measured() { + t.Fatalf("%d observations still refused: %s", DefaultGrowthMinObservations, v.Code) + } +} + +// Enough observations, too little elapsed time. A dense hour of snapshots +// describes an hour; the remedy is time, not sampling, and the code says so +// separately from the count gate. +func TestGrowthRefusesTooShortASpan(t *testing.T) { + h := evenHistory(48, DefaultGrowthMinSpan-time.Hour, func(i int) int64 { return 100 + int64(i) }) + v := growAssess(h, 147) + if v.Measured() || v.Measurement != nil { + t.Fatal("a span below the minimum produced a measurement") + } + if v.State != GrowthSpanTooShort || v.Code != ReasonGrowthSpanTooShort { + t.Fatalf("state %q code %q, want %q / %q", v.State, v.Code, GrowthSpanTooShort, ReasonGrowthSpanTooShort) + } + if v.Code == ReasonGrowthInsufficientObservations { + t.Fatal("a short span was reported as too few observations; the two have different remedies") + } + // Exactly the minimum span clears it. + h = evenHistory(48, DefaultGrowthMinSpan, func(i int) int64 { return 100 + int64(i) }) + if v := growAssess(h, 147); !v.Measured() { + t.Fatalf("a span of exactly %s was refused: %s", DefaultGrowthMinSpan, v.Code) + } +} + +// THE collapse this design exists to prevent. "We have not looked long enough" +// and "we looked and it never moved" must not be reachable from one another by +// reading a zero. +func TestNotEnoughHistoryIsNotFlat(t *testing.T) { + short := growAssess(flatHistory(2, 60*24*time.Hour, 200), 200) + brief := growAssess(flatHistory(48, 24*time.Hour, 200), 200) + measured := growAssess(flatHistory(48, 30*24*time.Hour, 200), 200) + + if measured.State != GrowthFlat || !measured.Measured() { + t.Fatalf("a well-observed unchanging allocation gave state %q, want %q", measured.State, GrowthFlat) + } + for _, v := range []GrowthVerdict{short, brief} { + if v.State == GrowthFlat { + t.Fatalf("state %q reported an unmeasured history as flat", v.State) + } + if v.Measured() { + t.Fatalf("state %q claims to be measured", v.State) + } + if v.Measurement != nil { + t.Fatalf("state %q carries a Measurement a caller could read a zero out of", v.State) + } + // The two-return accessor is the structural half: a caller that + // ignores ok gets 0, and has to have ignored ok on purpose. + if _, ok := v.GrewGiB(); ok { + t.Fatalf("state %q answered GrewGiB", v.State) + } + } + // And flat DOES answer, with a real zero that means something. + got, ok := measured.GrewGiB() + if !ok || got != 0 { + t.Fatalf("a measured flat history gave (%d, %v), want (0, true)", got, ok) + } + // The three states are pairwise distinct, so no caller can merge them. + if short.State == brief.State || short.State == measured.State || brief.State == measured.State { + t.Fatalf("states collide: %q / %q / %q", short.State, brief.State, measured.State) + } +} + +// Measurement is non-nil if and only if the state is a measured one. This +// biconditional is what the rest of the package is allowed to rely on. +func TestGrowthMeasurementExistsExactlyWhenMeasured(t *testing.T) { + cases := []StorageHistory{ + nil, + {}, + flatHistory(1, 0, 200), + flatHistory(3, 60*24*time.Hour, 200), + flatHistory(48, time.Hour, 200), + flatHistory(48, 30*24*time.Hour, 200), + evenHistory(48, 30*24*time.Hour, func(i int) int64 { return 200 + int64(i)*10 }), + // Non-monotone: refused, so no measurement. + {obs(40*24*time.Hour, 400), obs(30*24*time.Hour, 300), obs(20*24*time.Hour, 300), obs(0, 300)}, + // Same instant, two allocations. + {obs(40*24*time.Hour, 100), obs(40*24*time.Hour, 200), obs(20*24*time.Hour, 200), obs(0, 200)}, + } + for i, h := range cases { + v := growAssess(h, 200) + if v.State.Measured() != (v.Measurement != nil) { + t.Fatalf("case %d: state %q Measured()=%v but Measurement!=nil is %v", + i, v.State, v.State.Measured(), v.Measurement != nil) + } + if v.State == GrowthUnevaluated { + t.Fatalf("case %d: AssessStorageGrowth returned the zero state; it always answers", i) + } + if !v.State.Measured() && v.Code == "" { + t.Fatalf("case %d: state %q carries no refusal code", i, v.State) + } + if !v.State.Measured() && v.Reason == "" { + t.Fatalf("case %d: state %q carries no reason", i, v.State) + } + } +} + +// A decreasing allocation, or one instant carrying two allocations, is not one +// instance's history. It is refused whole rather than repaired, and no +// reduction is inferred from the decrease. +func TestGrowthRefusesAnInconsistentHistory(t *testing.T) { + shrank := evenHistory(10, 30*24*time.Hour, func(i int) int64 { + if i == 7 { + return 100 // an allocation no RDS API can produce + } + return 200 + int64(i)*10 + }) + conflict := flatHistory(10, 30*24*time.Hour, 200) + conflict = append(conflict, StorageObservation{At: conflict[4].At, AllocatedGiB: 999}) + + for name, h := range map[string]StorageHistory{"shrank": shrank, "same-instant": conflict} { + v := growAssess(h, 300) + if v.State != GrowthInconsistent { + t.Fatalf("%s: state %q, want %q", name, v.State, GrowthInconsistent) + } + if v.Code != ReasonGrowthHistoryInconsistent { + t.Fatalf("%s: code %q, want %q", name, v.Code, ReasonGrowthHistoryInconsistent) + } + if v.Measurement != nil { + t.Fatalf("%s: a contradictory history was averaged into a measurement", name) + } + if !strings.Contains(v.Reason, "You can't reduce the amount of storage") { + t.Fatalf("%s: the refusal does not quote the ratchet: %s", name, v.Reason) + } + } +} + +// --- The finding itself: a refusal with a size ------------------------------ + +// A measured growth produces a REFUSAL carrying the size, an advisory carrying +// the permanent cost, and nothing that could be read as an available +// reduction. +func TestGrowthIsARefusalWithASize(t *testing.T) { + h := evenHistory(30, 30*24*time.Hour, func(i int) int64 { return 200 + int64(i/6)*40 }) + v := growAssess(h, 360) + if v.State != GrowthGrew { + t.Fatalf("state %q, want %q", v.State, GrowthGrew) + } + if v.Code != ReasonGrowthIsNotReclaimable { + t.Fatalf("code %q, want %q", v.Code, ReasonGrowthIsNotReclaimable) + } + grew, ok := v.GrewGiB() + if !ok || grew != 160 { + t.Fatalf("grew %d (ok=%v), want 160", grew, ok) + } + if v.Measurement.Steps != 4 { + t.Fatalf("steps = %d, want 4", v.Measurement.Steps) + } + if v.Measurement.LargestStepGiB != 40 { + t.Fatalf("largest step = %d, want 40", v.Measurement.LargestStepGiB) + } + if v.Measurement.MonthlyUSD <= 0 { + t.Fatal("the refusal carries no size in dollars; the magnitude IS the finding") + } + // The prose must read as a permanent cost, never as an opportunity. + if !strings.Contains(v.Reason, "not a reduction that is available") { + t.Fatalf("the refusal does not say the reduction is unavailable: %s", v.Reason) + } + for _, banned := range []string{"could be reduced", "recommend shrink", "reclaim", "downsize the storage"} { + if strings.Contains(strings.ToLower(v.Reason), banned) { + t.Fatalf("the growth refusal contains %q; it must not read as a shrink proposal: %s", + banned, v.Reason) + } + } +} + +// --- The rate, and the two ways it is withheld ------------------------------ + +// One autoscaling step inside a long, dense window is one event. The size is +// still stated; the rate is not. +func TestSingleStepGrowthStatesTheSizeAndWithholdsTheRate(t *testing.T) { + h := evenHistory(60, 30*24*time.Hour, func(i int) int64 { + if i < 30 { + return 200 + } + return 400 + }) + v := growAssess(h, 400) + if v.State != GrowthGrew { + t.Fatalf("state %q, want %q", v.State, GrowthGrew) + } + if v.Measurement.Steps != 1 { + t.Fatalf("steps = %d, want 1", v.Measurement.Steps) + } + if grew, _ := v.GrewGiB(); grew != 200 { + t.Fatalf("grew %d, want 200 — the SIZE does not depend on the rate gate", grew) + } + if v.RateCode != ReasonGrowthRateNotProjectable { + t.Fatalf("rate code %q, want %q", v.RateCode, ReasonGrowthRateNotProjectable) + } + if _, ok := v.Rate(); ok { + t.Fatal("a rate was projected from a single step; one event is not a recurrence") + } + if v.Measurement.Projection != nil { + t.Fatal("Projection is non-nil with the rate refused") + } + // Two steps over the same span clears the gate. + h2 := evenHistory(60, 30*24*time.Hour, func(i int) int64 { return 200 + int64(i/20)*100 }) + v2 := growAssess(h2, 400) + if v2.RateCode != "" { + t.Fatalf("two steps were still refused a rate: %s", v2.RateReason) + } + if _, ok := v2.Rate(); !ok { + t.Fatal("two steps over a full span produced no rate") + } +} + +// The retained history is THINNED, so the instants cluster. When most of the +// span sits inside one gap nobody observed, the growth could have arrived +// entirely inside it and no rate is drawn through it — but the irregularity is +// reported rather than smoothed. +func TestIrregularSamplingWithholdsTheRateAndIsReported(t *testing.T) { + // Two dense clusters 24 days apart: plenty of observations, plenty of + // span, three increases, and 80 % of the span unobserved. + var h StorageHistory + for i := 0; i < 6; i++ { + h = append(h, obs(30*24*time.Hour-time.Duration(i)*time.Hour, 200+int64(i)*10)) + } + for i := 0; i < 6; i++ { + h = append(h, obs(5*24*time.Hour-time.Duration(i)*time.Hour, 400+int64(i)*10)) + } + v := growAssess(h, 450) + if v.State != GrowthGrew { + t.Fatalf("state %q, want %q", v.State, GrowthGrew) + } + if v.Sampling.Regular { + t.Fatalf("a history with an %s gap across a %s span was called regular", + v.Sampling.MaxGap, v.Sampling.Span) + } + if v.Sampling.LargestGapFraction <= DefaultGrowthMaxGapFraction { + t.Fatalf("largest gap fraction = %v, want > %v", v.Sampling.LargestGapFraction, + DefaultGrowthMaxGapFraction) + } + if v.RateCode != ReasonGrowthRateNotProjectable { + t.Fatalf("rate code %q, want %q", v.RateCode, ReasonGrowthRateNotProjectable) + } + if !strings.Contains(v.RateReason, "thinned") { + t.Fatalf("the rate refusal does not name the thinning: %s", v.RateReason) + } + // The size survives the rate refusal, and the irregularity is stated. + if grew, _ := v.GrewGiB(); grew != 250 { + t.Fatalf("grew %d, want 250", grew) + } + if !strings.Contains(describeSampling(v.Sampling), "not evenly spaced") { + t.Fatal("the sampling description smooths the irregularity away") + } +} + +// The rate divides by the ACTUAL span between the first and last instants, not +// by an assumed cadence times the observation count. The history below is +// deliberately front-loaded so the two answers differ. +func TestGrowthRateUsesTheActualInstants(t *testing.T) { + // 20 observations across 30 days: 14 crammed into the first 4½ days at 8 h + // apart, then 6 stragglers. The count and the span-in-days differ, so + // growth÷count and growth÷span are distinguishable answers. + day := 24 * time.Hour + var h StorageHistory + for i := 0; i < 14; i++ { + gib := int64(200) + if i >= 7 { + gib = 230 + } + h = append(h, obs(30*day-time.Duration(i)*8*time.Hour, gib)) + } + for i, before := range []time.Duration{24 * day, 20 * day, 16 * day, 12 * day, 6 * day, 0} { + gib := int64(260) + if i >= 3 { + gib = 290 + } + h = append(h, obs(before, gib)) + } + v := growAssess(h, 290) + if v.State != GrowthGrew { + t.Fatalf("state %q (%s), want %q", v.State, v.Code, GrowthGrew) + } + p, ok := v.Rate() + if !ok { + t.Fatalf("no rate projected: %s", v.RateReason) + } + grew, _ := v.GrewGiB() + wantDays := v.Sampling.Span.Hours() / 24 + if math.Abs(p.SpanDays-wantDays) > 1e-9 { + t.Fatalf("SpanDays = %v, want the actual span %v", p.SpanDays, wantDays) + } + if want := float64(grew) / wantDays; math.Abs(p.GiBPerDay-want) > 1e-9 { + t.Fatalf("GiBPerDay = %v, want %v", p.GiBPerDay, want) + } + // And it is NOT the naive per-sample figure. + if naive := float64(grew) / float64(v.Sampling.Observations); math.Abs(p.GiBPerDay-naive) < 1e-9 { + t.Fatal("the rate equals growth ÷ observation count; the instants were not used") + } + // MinGap and MaxGap bracket the mean rather than equalling it, which is + // the fact a reader needs in order to weigh the rate at all. + if !(v.Sampling.MinGap < v.Sampling.MeanGap && v.Sampling.MeanGap < v.Sampling.MaxGap) { + t.Fatalf("gaps min=%s mean=%s max=%s do not describe an irregular history", + v.Sampling.MinGap, v.Sampling.MeanGap, v.Sampling.MaxGap) + } +} + +// --- Structural guarantees -------------------------------------------------- + +// A rate is a claim about the future and must never carry money. Asserted by +// reflection so the next person to add a field argues with a test. +func TestGrowthProjectionCarriesNoMoney(t *testing.T) { + for _, name := range structFieldNames(t, GrowthProjection{}) { + lower := strings.ToLower(name) + for _, banned := range []string{"usd", "cost", "price", "saving", "dollar", "monthly"} { + if strings.Contains(lower, banned) { + t.Errorf("rds.GrowthProjection has a field %q. A growth rate is a forecast over a step "+ + "function whose steps this package does not observe the cause of; attaching money to "+ + "it would let an unverified projection become a claim", name) + } + } + } + if (GrowthProjection{GiBPerDay: 1000}).Claimable() { + t.Error("a growth projection claims to be claimable") + } + if !strings.Contains((GrowthProjection{GiBPerDay: 1.5, Steps: 3, SpanDays: 30}).String(), "[unverified]") { + t.Error("a rendered projection does not carry its [unverified] marker") + } +} + +// The measurement names a COST and never a reduction — the same discipline +// TestStorageVerdictNamesNoSaving applies to the floor, applied to the ratchet +// over time. +func TestGrowthNamesNoSaving(t *testing.T) { + for _, v := range []any{GrowthMeasurement{}, GrowthVerdict{}, GrowthTotals{}, GrowthSampling{}} { + for _, name := range structFieldNames(t, v) { + lower := strings.ToLower(name) + for _, banned := range []string{"saving", "reclaim", "shrink", "reduc", "recover"} { + if strings.Contains(lower, banned) && !strings.Contains(lower, "unrecoverable") { + t.Errorf("%T has a field %q; allocated storage is a one-way ratchet and a field named "+ + "after a reduction would be a promise no RDS API can keep", v, name) + } + } + } + } +} + +// A caller may demand MORE evidence than this package requires and may never +// demand less: a two-sample policy is exactly the claim §7.4 deferred. +func TestGrowthPolicyCannotBeLoosened(t *testing.T) { + loose := GrowthPolicy{MinObservations: 2, MinSpan: time.Minute, MinSteps: 1, + MaxGapFraction: 0.99, MaxObservations: 1 << 20} + got := loose.normalized() + if got != DefaultGrowthPolicy() { + t.Fatalf("a loosened policy survived normalization: %+v", got) + } + if (GrowthPolicy{}).normalized() != DefaultGrowthPolicy() { + t.Fatal("the zero policy does not normalize to the shipped bar") + } + // Stricter survives. + strict := GrowthPolicy{MinObservations: 30, MinSpan: 90 * 24 * time.Hour, MinSteps: 5, + MaxGapFraction: 0.1, MaxObservations: 100} + if strict.normalized() != strict { + t.Fatalf("a stricter policy was relaxed: %+v", strict.normalized()) + } + // And a stricter policy actually refuses what the default accepts. + h := evenHistory(10, 20*24*time.Hour, func(i int) int64 { return 200 + int64(i)*10 }) + if v := growAssess(h, 290); !v.Measured() { + t.Fatalf("the default policy refused a clean history: %s", v.Code) + } + v := AssessStorageGrowth(growInstance(290), h, strict, DefaultRates()) + if v.Measured() { + t.Fatal("a stricter policy measured a history it should have refused") + } +} + +// --- The history container -------------------------------------------------- + +func TestStorageHistoryAppendIsIdempotentBoundedAndOrdered(t *testing.T) { + var h StorageHistory + for i := 0; i < 10; i++ { + h = h.Append(obs(time.Duration(10-i)*24*time.Hour, 100+int64(i)), 0) + } + if len(h) != 10 { + t.Fatalf("len = %d, want 10", len(h)) + } + // Re-appending the same instants replaces rather than doubles: re-ingesting + // a history is a no-op, the same rule pkg/store's SaveSnapshotAt applies. + again := h + for _, o := range h { + again = again.Append(o, 0) + } + if len(again) != 10 { + t.Fatalf("re-appending doubled the history to %d", len(again)) + } + if b1, _ := json.Marshal(h); string(b1) != mustJSON(t, again) { + t.Fatal("re-appending an identical history changed it") + } + // Out-of-order arrival sorts. + shuf := StorageHistory{obs(0, 300), obs(20*24*time.Hour, 100), obs(10*24*time.Hour, 200)} + var built StorageHistory + for _, o := range shuf { + built = built.Append(o, 0) + } + for i := 1; i < len(built); i++ { + if !built[i-1].At.Before(built[i].At) { + t.Fatal("Append did not order the history by instant") + } + } + // Zero instants and non-positive allocations are dropped: "we did not + // look" is not an observation of zero storage. + junk := built.Append(StorageObservation{AllocatedGiB: 500}, 0). + Append(StorageObservation{At: testEnd.Add(-time.Hour)}, 0) + if len(junk) != len(built) { + t.Fatalf("Append accepted a junk observation: %d → %d", len(built), len(junk)) + } + // The bound drops the OLDEST. + bounded := h + for i := 0; i < 5; i++ { + bounded = bounded.Append(obs(time.Duration(i)*time.Hour, 500), 4) + } + if len(bounded) != 4 { + t.Fatalf("bounded len = %d, want 4", len(bounded)) + } + if !bounded[len(bounded)-1].At.Equal(h[len(h)-1].At.UTC()) && bounded[0].At.Before(h[0].At) { + t.Fatal("the bound dropped the newest rather than the oldest") + } + // A truncated history says so. + long := flatHistory(20, 30*24*time.Hour, 200) + s := long.Sampling(GrowthPolicy{MaxObservations: 5}) + if !s.Truncated { + t.Fatal("a history longer than the read bound did not report itself truncated") + } +} + +func mustJSON(t *testing.T, v any) string { + t.Helper() + b, err := json.Marshal(v) + if err != nil { + t.Fatal(err) + } + return string(b) +} + +// Input ORDER must not reach the verdict: a history delivered backwards, or +// interleaved, measures the same. +func TestGrowthIsShuffleInvariant(t *testing.T) { + base := evenHistory(24, 30*24*time.Hour, func(i int) int64 { return 200 + int64(i/6)*25 }) + want := mustJSON(t, growAssess(base, 275)) + for pi, order := range permutations(len(base)) { + got := mustJSON(t, growAssess(permute(base, order), 275)) + if got != want { + t.Fatalf("permutation %d produced a different verdict; input order reached the output", pi) + } + } +} + +// --- Report integration ----------------------------------------------------- + +// withHistory attaches a history to every target in a snapshot. +func withHistory(snap *Snapshot, h func(id string) StorageHistory) *Snapshot { + for i := range snap.Targets { + snap.Targets[i].StorageHistory = h(snap.Targets[i].Instance.Identifier) + } + return snap +} + +// A fleet with no persisted history produces NO growth suppression: an unwired +// caller is one fact about the wiring, not N facts about N instances. The +// verdict is still recorded, typed, on every assessment. +func TestNoHistoryAddsNoSuppression(t *testing.T) { + rep := assess(t, collect(t, mixedFixture()), nil) + for _, code := range []string{ + ReasonGrowthInsufficientObservations, ReasonGrowthSpanTooShort, + ReasonGrowthHistoryInconsistent, ReasonGrowthIsNotReclaimable, + ReasonGrowthRateNotProjectable, + } { + if n := rep.Totals.RefusedByCode[code]; n != 0 { + t.Fatalf("an unwired fleet emitted %q ×%d; the wiring gap is not a per-instance finding", + code, n) + } + } + modelled := 0 + for _, a := range rep.Assessments { + if a.Excluded() { + // Exclusions fire alone and never reach the growth gate. + if a.Growth.State != GrowthUnevaluated { + t.Fatalf("%s is excluded but carries a growth verdict %q", a.Target.ID, a.Growth.State) + } + continue + } + modelled++ + if a.Growth.State != GrowthNoHistory { + t.Fatalf("%s: growth state %q, want %q", a.Target.ID, a.Growth.State, GrowthNoHistory) + } + if a.Growth.Reason == "" { + t.Fatalf("%s: the no-history verdict states no reason", a.Target.ID) + } + } + if got := rep.StorageGrowth(); got.NoHistory != modelled || got.Measured != 0 || got.Refused != 0 { + t.Fatalf("totals %+v, want NoHistory=%d and nothing measured or refused", got, modelled) + } +} + +// The end-to-end shape: a growing instance produces a refusal, an advisory +// carrying the permanent cost, a valid report, and NOT ONE PROPOSAL or dollar +// of savings anywhere. +func TestGrowthReachesTheReportAsARefusalAndNeverAsAProposal(t *testing.T) { + f := &Fixture{ + Instances: []DBInstanceRecord{ + rec("grew", "db.m6i.large", "postgres", withStorage(600, 2000, StorageGP2)), + rec("still", "db.m6i.large", "postgres", withStorage(200, 0, StorageGP2)), + rec("young", "db.m6i.large", "postgres", withStorage(300, 0, StorageGP2)), + }, + Metrics: mergeMetrics( + metricsFor("grew", 30, 10, 8<<30, 100*GiB), + metricsFor("still", 30, 10, 8<<30, 50*GiB), + metricsFor("young", 30, 10, 8<<30, 50*GiB), + ), + } + snap := withHistory(collect(t, f), func(id string) StorageHistory { + switch id { + case "grew": + return evenHistory(30, 30*24*time.Hour, func(i int) int64 { return 200 + int64(i/6)*100 }) + case "still": + return flatHistory(30, 30*24*time.Hour, 200) + default: + return flatHistory(2, 30*24*time.Hour, 300) + } + }) + rep := assess(t, snap, nil) + + grew := must(t, rep, "grew") + wantCode(t, grew, ReasonGrowthIsNotReclaimable) + ad := wantAdvisory(t, grew, AdvisoryStorageGrew) + if ad.MonthlyUSD <= 0 { + t.Fatal("the growth advisory carries no magnitude") + } + if ad.Actuatable() || ad.Action() != "advisory" { + t.Fatal("the growth advisory claims to be actuatable") + } + if !strings.Contains(ad.Caveat, "permanent cost") || !strings.Contains(ad.Caveat, "not a saving") { + t.Fatalf("the growth advisory's caveat does not say it is unrecoverable: %s", ad.Caveat) + } + // The ratchet refusal that was already there is still there, unweakened. + wantCode(t, grew, ReasonStorageCannotShrink) + wantCode(t, grew, ReasonStorageAutoscalingRatchet) + + still := must(t, rep, "still") + if still.Growth.State != GrowthFlat { + t.Fatalf("still: growth state %q, want %q", still.Growth.State, GrowthFlat) + } + for _, code := range []string{ReasonGrowthIsNotReclaimable, ReasonGrowthInsufficientObservations, + ReasonGrowthSpanTooShort, ReasonGrowthRateNotProjectable} { + wantNoCode(t, still, code) + } + + young := must(t, rep, "young") + wantCode(t, young, ReasonGrowthInsufficientObservations) + if young.Growth.Measured() { + t.Fatal("young: two observations were measured") + } + + // Nothing proposed, nothing claimed, anywhere. + if rep.Totals.Proposals != 0 { + t.Fatalf("%d proposals; a growth measurement must never become one", rep.Totals.Proposals) + } + if rep.Totals.GrossSavingsMonthlyUSD != 0 || rep.Totals.NetSavingsMonthlyUSD != 0 { + t.Fatalf("the report claims savings: gross %v net %v", + rep.Totals.GrossSavingsMonthlyUSD, rep.Totals.NetSavingsMonthlyUSD) + } + for _, a := range rep.Assessments { + if a.Proposal != nil { + t.Fatalf("%s carries a proposal", a.Target.ID) + } + } + // The roll-up counts refusals separately from measurements, so an operator + // cannot read "1 measured" as "3 instances checked". + tot := rep.StorageGrowth() + if tot.Measured != 2 || tot.Grew != 1 || tot.Flat != 1 || tot.Refused != 1 { + t.Fatalf("totals %+v, want measured 2 (1 grew, 1 flat) and 1 refused", tot) + } + if tot.GrewGiB != 400 || tot.UnrecoverableGrowthMonthlyUSD <= 0 { + t.Fatalf("totals %+v, want 400 GiB with a cost", tot) + } + // The text report renders without depending on anything unordered. + var b strings.Builder + if err := rep.WriteText(&b); err != nil { + t.Fatal(err) + } + if !strings.Contains(b.String(), ReasonGrowthIsNotReclaimable) { + t.Fatal("the growth refusal does not reach the text report") + } +} + +// Every growth code is a stable, distinct, storable string, and none collides +// with a code this package already ships. +func TestGrowthCodesAreDistinctFromEveryExistingCode(t *testing.T) { + existing := []string{ + ReasonAuroraNotSupported, ReasonClusterMemberNotSupported, ReasonModeOff, ReasonUnknownEngine, + ReasonUnknownInstanceClass, ReasonEngineNotPriced, ReasonUnknownDeployment, ReasonUnverifiedRate, + ReasonInstanceClassIsAFailover, ReasonFreeableMemoryIsPageCache, ReasonBufferPoolScalesWithClass, + ReasonMemorySemanticsUnencoded, ReasonStorageCannotShrink, ReasonStorageAutoscalingRatchet, + ReasonReplicaIsFailoverCapacity, ReasonMultiAZIsAvailabilityPosture, ReasonInsufficientWindow, + ReasonNoMetricEvidence, ReasonTruncatedMetrics, ReasonSizeFlexibilityExcluded, + ReasonInstanceStateUnstable, ReasonNoStoragePerformanceModel, ReasonCommitmentNegative, + ReasonCommitmentNeutral, AdvisoryIdleInstance, AdvisoryIdleReadReplica, AdvisoryStorageFloor, + AdvisoryStorageAutoscaling, AdvisoryMultiAZMultiplier, AdvisoryUnverifiedRate, + AdvisoryReservationStranding, + } + seen := map[string]bool{} + for _, c := range existing { + seen[c] = true + } + added := map[string]string{ + "ReasonGrowthInsufficientObservations": ReasonGrowthInsufficientObservations, + "ReasonGrowthSpanTooShort": ReasonGrowthSpanTooShort, + "ReasonGrowthHistoryInconsistent": ReasonGrowthHistoryInconsistent, + "ReasonGrowthIsNotReclaimable": ReasonGrowthIsNotReclaimable, + "ReasonGrowthRateNotProjectable": ReasonGrowthRateNotProjectable, + "AdvisoryStorageGrew": AdvisoryStorageGrew, + } + for name, code := range added { + if code == "" { + t.Errorf("%s is empty", name) + continue + } + if code != strings.ToLower(code) || strings.ContainsAny(code, " _") { + t.Errorf("%s = %q: codes are lower-case and hyphenated so they are safe to store and group on", + name, code) + } + if seen[code] { + t.Errorf("%s = %q collides with a code this package already ships; two codes with one value "+ + "silently merge two findings", name, code) + } + seen[code] = true + } + // The states are distinct too, and none is the zero value by accident. + states := map[GrowthState]bool{} + for _, s := range []GrowthState{GrowthNoHistory, GrowthTooFewObservations, GrowthSpanTooShort, + GrowthInconsistent, GrowthFlat, GrowthGrew} { + if s == GrowthUnevaluated { + t.Errorf("state %q equals the unevaluated zero value", s) + } + if states[s] { + t.Errorf("state %q is declared twice", s) + } + states[s] = true + } +} + +// FuzzGrowthNeverClaimsASaving drives arbitrary histories — including +// nonsensical, decreasing, zero-span and enormous ones — through the whole +// path and asserts the invariants that must hold over every one of them. +func FuzzGrowthNeverClaimsASaving(f *testing.F) { + f.Add(4, int64(100), int64(50), int64(3600), int64(0)) + f.Add(0, int64(0), int64(0), int64(0), int64(0)) + f.Add(50, int64(1), int64(-1), int64(1), int64(1)) + f.Add(9, int64(1<<40), int64(1<<40), int64(1), int64(1<<40)) + + f.Fuzz(func(t *testing.T, n int, start, step, gapSec, jitter int64) { + if n < 0 { + n = -n + } + n %= 200 + if gapSec < 0 { + gapSec = -gapSec + } + gapSec %= 90 * 24 * 3600 + h := make(StorageHistory, 0, n) + at := testEnd.Add(-200 * 24 * time.Hour) + alloc := start % (1 << 20) + for i := 0; i < n; i++ { + h = append(h, StorageObservation{At: at, AllocatedGiB: alloc}) + at = at.Add(time.Duration(gapSec+int64(i)*(jitter%3600)) * time.Second) + alloc += step % (1 << 16) + } + inst := growInstance(alloc) + v := AssessStorageGrowth(inst, h, DefaultGrowthPolicy(), DefaultRates()) + + if v.State == GrowthUnevaluated { + t.Fatal("AssessStorageGrowth returned the unevaluated zero state; it always answers") + } + if v.State.Measured() != (v.Measurement != nil) { + t.Fatalf("state %q: Measured()=%v but Measurement!=nil is %v", + v.State, v.State.Measured(), v.Measurement != nil) + } + if !v.State.Measured() && v.Code == "" { + t.Fatalf("state %q carries no refusal code", v.State) + } + if m := v.Measurement; m != nil { + if m.GrewGiB < 0 { + t.Fatalf("a negative growth %d escaped; a decrease is inconsistent, never a reduction", + m.GrewGiB) + } + if m.MonthlyUSD < 0 || !finite(m.MonthlyUSD) { + t.Fatalf("growth cost %v is not a finite non-negative magnitude", m.MonthlyUSD) + } + if (m.GrewGiB == 0) != (v.State == GrowthFlat) { + t.Fatalf("state %q with GrewGiB %d: flat and grew must partition the measured states", + v.State, m.GrewGiB) + } + if p := m.Projection; p != nil { + if !finite(p.GiBPerDay) || p.GiBPerDay < 0 { + t.Fatalf("projected rate %v is not a finite non-negative number", p.GiBPerDay) + } + if p.Claimable() { + t.Fatal("a projection claims to be claimable") + } + if p.Steps < DefaultGrowthMinSteps { + t.Fatalf("a rate was projected from %d step(s)", p.Steps) + } + if v.Sampling.Span < DefaultGrowthMinSpan { + t.Fatalf("a rate was projected across %s", v.Sampling.Span) + } + } + } + if v.Sampling.Observations < DefaultGrowthMinObservations && v.State.Measured() { + t.Fatalf("%d observations were measured", v.Sampling.Observations) + } + + // And through the sizer: still no proposal, still a valid report. + snap := &Snapshot{Domain: Kind, Window: testWindow(), Timestamp: testEnd, Targets: []Target{{ + Ref: domain.TargetRef{Domain: Kind, Scope: "1234/us-east-1", ID: inst.ARN, Name: inst.Identifier}, + Instance: inst, + StorageHistory: h, + }}} + s, err := NewSizer(DefaultConfig()) + if err != nil { + t.Fatal(err) + } + rep := s.Assess(testNow, snap, nil) + if err := rep.Validate(); err != nil { + t.Fatalf("a growth history produced an invalid report: %v", err) + } + if rep.Totals.Proposals != 0 || rep.Totals.GrossSavingsMonthlyUSD != 0 { + t.Fatalf("a growth history produced %d proposals and %v of savings", + rep.Totals.Proposals, rep.Totals.GrossSavingsMonthlyUSD) + } + }) +} + +// The persistence loop, end to end and without a store: Restore → fold the +// prior histories in → Observe → Checkpoint, repeated until the bar is +// cleared. This is exactly the sequence GROWTH-FINDINGS.md §6 asks cmd/ for, +// and it is the reason `AttributeStorageGrowth` was unreachable until now. +func TestHistorySurvivesTheCheckpointLoopAndReachesAVerdict(t *testing.T) { + f := &Fixture{ + Instances: []DBInstanceRecord{ + rec("ratchet", "db.m6i.large", "postgres", withStorage(200, 2000, StorageGP2)), + }, + Metrics: metricsFor("ratchet", 30, 10, 8<<30, 20*GiB), + } + cfg := DefaultConfig() + var blob []byte + // Twenty daily runs. Storage autoscales twice along the way — nothing + // Kilter did, nothing CloudTrail recorded. + for day := 0; day < 20; day++ { + d, err := NewDomain(cfg) + if err != nil { + t.Fatal(err) + } + if err := d.Restore(blob); err != nil { + t.Fatalf("day %d: Restore: %v", day, err) + } + switch day { + case 7: + f.Instances[0].AllocatedStorage = 260 + case 14: + f.Instances[0].AllocatedStorage = 330 + } + snap := collect(t, f) + snap.Timestamp = testEnd.Add(-time.Duration(19-day) * 24 * time.Hour) + + // The two calls cmd/ owes, and the only two. + RecordStorageHistory(snap, d.StorageHistories(), cfg.Growth) + if err := d.Observe(snap); err != nil { + t.Fatalf("day %d: Observe: %v", day, err) + } + if blob, err = d.Checkpoint(); err != nil { + t.Fatalf("day %d: Checkpoint: %v", day, err) + } + + rep := d.Report(testNow, nil) + if err := rep.Validate(); err != nil { + t.Fatalf("day %d: %v", day, err) + } + a := must(t, rep, "ratchet") + switch { + case day < DefaultGrowthMinObservations-1: + wantCode(t, a, ReasonGrowthInsufficientObservations) + if a.Growth.Measured() { + t.Fatalf("day %d: %d observations were measured", day, day+1) + } + case day < 14: + // Enough observations from day 3 on; the SPAN gate holds alone + // until day 14, which is the point of having two codes. + if a.Growth.Measured() { + t.Fatalf("day %d: measured across only %s", day, a.Growth.Sampling.Span) + } + wantCode(t, a, ReasonGrowthSpanTooShort) + default: + wantNoCode(t, a, ReasonGrowthSpanTooShort) + wantNoCode(t, a, ReasonGrowthInsufficientObservations) + if a.Growth.State != GrowthGrew { + t.Fatalf("day %d: state %q, want %q", day, a.Growth.State, GrowthGrew) + } + } + } + + // The final verdict: 130 GiB across 19 days in two steps, refused with its + // size, and a rate that survives both gates. + d, err := NewDomain(cfg) + if err != nil { + t.Fatal(err) + } + if err := d.Restore(blob); err != nil { + t.Fatal(err) + } + a := must(t, d.Report(testNow, nil), "ratchet") + wantCode(t, a, ReasonGrowthIsNotReclaimable) + grew, ok := a.Growth.GrewGiB() + if !ok || grew != 130 { + t.Fatalf("grew %d (ok=%v), want 130", grew, ok) + } + if a.Growth.Measurement.Steps != 2 { + t.Fatalf("steps = %d, want 2", a.Growth.Measurement.Steps) + } + if p, ok := a.Growth.Rate(); !ok { + t.Fatalf("no rate from two steps across 19 days: %s", a.Growth.RateReason) + } else if p.SpanDays < 18 || p.SpanDays > 20 { + t.Fatalf("rate span %v days, want ~19", p.SpanDays) + } + // A rerun of the same day is idempotent: it must not manufacture a step. + snap := collect(t, f) + snap.Timestamp = testEnd + RecordStorageHistory(snap, d.StorageHistories(), cfg.Growth) + if err := d.Observe(snap); err != nil { + t.Fatal(err) + } + again := must(t, d.Report(testNow, nil), "ratchet") + if again.Growth.Sampling.Observations != a.Growth.Sampling.Observations { + t.Fatalf("re-recording the same instant changed the observation count: %d → %d", + a.Growth.Sampling.Observations, again.Growth.Sampling.Observations) + } + if g, _ := again.Growth.GrewGiB(); g != grew { + t.Fatalf("re-recording changed the growth: %d → %d", grew, g) + } + // StorageHistories hands out copies; mutating one cannot reach the domain. + hs := d.StorageHistories() + for id := range hs { + hs[id][0].AllocatedGiB = 999999 + } + if g, _ := must(t, d.Report(testNow, nil), "ratchet").Growth.GrewGiB(); g != grew { + t.Fatal("StorageHistories leaked the domain's own slices") + } +} diff --git a/pkg/rds/rds.go b/pkg/rds/rds.go index ee5fb7a..e3c7e95 100644 --- a/pkg/rds/rds.go +++ b/pkg/rds/rds.go @@ -709,6 +709,19 @@ type Target struct { // never by the collector, because only the domain has a previous // observation to remember. PriorAllocatedStorageGiB int64 `json:"priorAllocatedStorageGiB,omitempty"` + // StorageHistory is every allocated-storage observation of this instance + // that survived the caller's retention, oldest first. It is the seam + // FINDINGS.md §7.4 named: PriorAllocatedStorageGiB answers "did it move + // since last time", and this answers "how far has it moved, over how long, + // and in how many steps" — which is the only version of the question a + // lumpy, thinned, one-way ratchet can be asked honestly. + // + // It is filled by the CALLER, from persisted checkpoints, and never by the + // collector: a collector sees one instant and a growth is a claim about + // several. Empty is the normal state of an unwired deployment and produces + // [GrowthNoHistory] rather than a fleet of identical refusals. See + // [StorageHistory.Append] and GROWTH-FINDINGS.md §6. + StorageHistory StorageHistory `json:"storageHistory,omitempty"` } // SeriesFor returns the named series, and false when it was not delivered. diff --git a/pkg/rds/sizer.go b/pkg/rds/sizer.go index 0f6729b..f9e4085 100644 --- a/pkg/rds/sizer.go +++ b/pkg/rds/sizer.go @@ -97,6 +97,14 @@ type Config struct { // Parity is the U13 seam. Nil in U11 — see [StorageParity]. Parity StorageParity `json:"-"` + + // Growth is the evidence bar the allocated-storage growth finding must + // clear before it states anything. The zero value means + // [DefaultGrowthPolicy], and normalized() clamps a supplied policy to that + // floor rather than honouring one that is looser: a caller must be able to + // demand MORE evidence than this package requires and never less. See + // [AssessStorageGrowth]. + Growth GrowthPolicy `json:"growth"` } // DefaultConfig returns the shipped policy. @@ -107,6 +115,7 @@ func DefaultConfig() Config { FullWindow: DefaultFullWindow, StorageOverprovisionThreshold: DefaultStorageOverprovisionThreshold, IdleCPUPercent: DefaultIdleCPUPercent, + Growth: DefaultGrowthPolicy(), } } @@ -129,6 +138,7 @@ func (c Config) normalized() Config { if c.IdleCPUPercent <= 0 { c.IdleCPUPercent = DefaultIdleCPUPercent } + c.Growth = c.Growth.normalized() return c } @@ -253,8 +263,13 @@ type Assessment struct { CurrentMonthlyUSD float64 `json:"currentMonthlyUSD,omitempty"` RateProvenance RateProvenance `json:"rateProvenance,omitempty"` - Memory MemoryVerdict `json:"memory,omitzero"` - Storage StorageVerdict `json:"storage,omitzero"` + Memory MemoryVerdict `json:"memory,omitzero"` + Storage StorageVerdict `json:"storage,omitzero"` + // Growth is the allocated-storage RATCHET over time, from the history the + // caller persisted. Its zero value is [GrowthUnevaluated] — the question + // was never put — which is a third state distinct from both "not enough + // history" and "measured, and flat". See [GrowthVerdict]. + Growth GrowthVerdict `json:"growth,omitzero"` Idle IdleVerdict `json:"idle,omitzero"` Commitment CommitmentVerdict `json:"commitment,omitzero"` @@ -488,6 +503,16 @@ func (s *Sizer) assessTarget(now time.Time, snap *Snapshot, t Target, ledger dom a.Storage = AssessStorage(inst, free, s.cfg.Rates) s.storageFindings(&a, snap) + + // Trap 8 over TIME. The floor moved on its own; this says how far, over + // how long, and refuses to say more than the history supports. Ordering: + // after storageFindings so the report reads floor-then-growth (the growth + // elaborates the ratchet the line above already refused), and before + // rateProvenanceFindings so the unverified-rate caveat lands after every + // dollar it is a caveat on. Its dollar comes from the same storage rate + // [Assessment.WorstRateProvenance] already weighs. See [AssessStorageGrowth]. + s.growthFindings(&a, t, snap) + s.rateProvenanceFindings(&a) // --- Idle, and the replica distinction (§2.5) -------------------------