Multimode batched evaluation of factorized CC (cost model + placement + dry-run) - #583
Open
evaleev wants to merge 259 commits into
Open
Multimode batched evaluation of factorized CC (cost model + placement + dry-run)#583evaleev wants to merge 259 commits into
evaleev wants to merge 259 commits into
Conversation
The static per-node walk in cost_profile() prices each node once, so its
flops/exec are order- and batching-blind and never reflect the per-occ-block
REPLAY recompute the batched evaluator does at runtime -- the reason a dry-run
could not predict occ-batching being slower than aux-only.
Split the reported cost:
- Rename CostProfile::{flops,exec_cost,n_ops} -> model_{flops,exec,n_ops} (the
static DP-model quantities, unchanged).
- Add dryrun_{flops,exec,n_ops}, tallied from the existing Trace::On replay: an
optional CostSink is attached to the shared dry-run CostModel, and every
actual product-op execution (DryRunOps::prod) folds its own SLICED-extent
flops/exec (the same numbers already computed for the per-op OpCost log) into
it. Because a sliced occ-dependent op run N times does ~1/N work per pass, its
sliced-cost sum is work-neutral; only occ-INDEPENDENT work re-executed at full
size once per block inflates -- so dryrun_* isolates exactly the recompute.
The sink is opt-in (nullptr default) and lives on the dry-run CostModel, which
is constructed only inside cost_profile(); eval.hpp / make_evaluator are
untouched, so mpqc's production evaluator path is byte-identical.
water-20 [.][dryrun-water20-overcompute], heavy occ batching: model_flops
ratio 1.0 (flat), dryrun_flops/exec ratio ~1.98, dryrun_n_ops ~55x -- the
recompute is now visible where the model walk saw nothing.
Extend [.][dryrun-water20-overcompute] to three configs -- aux-only, occ+aux
order_aware=false (MPQC production / root-level forest seed), occ+aux
order_aware=true (node-level placement) -- and report the recompute-aware
dryrun_{exec,n_ops} per config plus the OA-true-vs-false ratio.
Diagnoses the water-20 order_aware=true slowdown: under occ batching,
order_aware=true does ~2x the dryrun_exec and ~12x the op-executions of
order_aware=false, and (via its resident-scan peak model reporting higher
peaks) also tips more terms over peak_threshold so occ batching engages where
order_aware=false leaves them un-batched -- a double hit that matches the
observed runtime regression.
Flipping the default to true (previous commit on this branch) turned on the order-aware cost model AND, together with batch_spectator_indices, node-level external placement. On water-20 that path is a runtime regression -- the water-20 dryrun diagnostic measures ~2x the traffic and ~12x the op-executions of the root-level forest seed for the same term -- and, worse, was reported on Owl (job 649250) to produce a WRONG schedule: incorrect PNO-CCSD iteration energies and malformed eval ops. Restore the known-good default (false = legacy set-keyed DP + root-level forest seed) until node-level placement is root-caused and fixed. The two tests that pin order_aware off explicitly keep passing (the pin now merely matches the default).
order_aware_recompute conflated two orthogonal concerns: the order-aware recompute COST MODEL (which factorization the DP selects) and the node-level external-mode EMISSION placement (per-node External stamps vs the root-level forest seed). node_level_placement was defined as order_aware_recompute && batch_spectator_indices, so enabling the more-realistic cost model forced the node-level emission along with it. A water-8 A/B (holding the cost model fixed, toggling only emission) shows the regression is the EMISSION, not the cost model: node-level placement runs ~6x slower (~124 vs ~20 s/iter) and emits ~8x more batch scopes than the root-level seed, because it nests a batch scope at every carrying node and the batched evaluator replays each. The order-aware cost model with root-seed emission is correct and cheap. Node-level placement also produces a wrong residual on water-20 (size-dependent; not reproduced at water-8/he10). Split node_level_placement into its own BatchPolicy/CostParams/CostModel flag (default false), threaded alongside order_aware_recompute. order_aware_recompute now drives only selection; node_level_placement drives only emission. Both default false, so this is behavior-neutral. A gated SEQUANT_NODE_LEVEL_PLACEMENT env knob forces the placement for A/B diagnostics without recompiling a flag through the caller. Tests that engaged node-level placement via order_aware_recompute=true now set node_level_placement=true explicitly; the node-level correctness sweep now drives the emission via its sweep variable.
Now that node-level emission is separately gated by node_level_placement (default off), the order-aware recompute cost model is selection-only and safe to default on: it charges recompute realistically and picks better-batching factorizations, while emission stays the correct, cheap root-level forest seed. Verified no churn across the [optimize] suite (628 assertions) and the batched emission tests. The internal CostModel/oracle-helper member defaults stay false (documented seeded-probe reason); the public BatchPolicy/CostParams path threads the true default through.
…le recompute
Add a batched-evaluation schedule-visualizer pipeline and make avoidable
recompute a first-class per-node output of cost_profile().
- schedule_dump.hpp (new): per-term IR schedule-record emitter
(schedule_ir_json) and a shared cost_op_signature() join key (result index
labels + the sorted operand pair). One definition, used by three producers
so a DAG node, its runtime Build event, and cost_profile's per-node number
all carry the identical key -- no signature is reconstructed downstream.
- eval.hpp: runtime schedule-dump hooks (SCHEDULE_RUN_EVENT / RUN_GROUP),
each internal Build event stamped with the same signature and per-loop
dependent-mode flags, all gated by SEQUANT_SCHED_DUMP (production path
byte-identical when unset).
- CostProfile gains per-node avoidable_nodes {label,count,exec,flops} plus
avoidable_exec / avoidable_ops and avoidable_time(). DryRunOps::prod tallies
each build's necessary = product of block counts of its touched (dependent)
modes, so builds - necessary is the avoidable recompute (a node rebuilt once
per block of a mode it does not touch). The cost model is dense, so necessary
is exact and no empirical correction is needed. avoidable_nodes_from_sink()
is the shared rollup, reused by the schedule-dump test to emit these numbers.
- Consolidate the two dryrun avoidable witnesses (occ-veto, extmode) onto
cost_profile's structural per-node rollup, replacing the BatchGroup/BatchIter
trace-parse string-match reconstruction with a read of cp.avoidable_* (the
structural signature is relabeling-proof; the trace is still parsed only for
the scatter/group Begin markers cost_profile does not expose).
…uckets The per-node avoidable-recompute rollup keyed only by the label signature (result+operand indices), so a node built at many slice sizes landed in one bucket. Because the exec model is roofline-like (nonlinear in slice size), those builds span orders of magnitude (a 257x spread on the C60 external-occ arm); pricing avoidable as count * a single per-build exec then exceeded the whole replay's exec -- the impossible avoidable_time > 100%. Fix: key the sink by the EXTENDED signature (label + the touched modes' realized extents), so every build in a bucket ran at the same slice-context and hence the same roofline cost. Within a cost-homogeneous bucket avoidable_exec = total_exec * (builds - necessary)/builds is exact; the buckets of one DAG node aggregate back to a per-label AvoidableNode (what the visualizer joins on by hash->sig). Verified: buckets are homogeneous (count*last == total*frac per bucket), all arms bounded <= 100% (C60 external-occ 390% -> 0.005%), and the giant term reads 77%, matching an independent dependent-mode analysis (78%). NodeCost now accumulates total_exec/total_flops; min/max/last are kept only for the SEQUANT_AVOIDABLE_DEBUG homogeneity dump.
Redefine the per-value avoidable-recompute metric: it is now measured in FLOPs
against the batching-free (unlimited-memory) ideal -- the arithmetic the batched
replay repeats beyond building each value once at full extent -- rather than in
roofline exec against a within-scheme "necessary" reference.
Two problems with the exec-weighted metric drove this:
- roofline exec is nonlinear in slice size, so one value's differently-sized
builds spanned ~257x; pricing avoidable as count x a single per-build exec
exceeded the whole replay's exec (avoidable_time > 100%). The prior
cost-homogeneous slice-context bucketing removed the >100% but stayed
exec-weighted;
- referencing "necessary = distinct slices the scheme produces" is circular --
it scores an un-hoisted value's per-block rebuilds as necessary, so it cannot
see the recompute hoisting exists to avoid.
FLOPs is linear in extents, hence additive across slices: disjoint per-block
slices that tile a value sum to exactly full_flops (0 avoidable), while a value
rebuilt full per block sums to N*full ((N-1)*full avoidable). So avoidable =
max(0, total_flops - full_flops) per value, bounded in [0, dryrun_flops] by
construction, needs no slice-context bucketing, and answers "what does batching
cost vs. infinite memory". NodeCost drops to {builds, total_flops, full_flops};
CostProfile.avoidable_exec -> avoidable_flops, avoidable_time() = avoidable_flops
/ dryrun_flops.
Witnesses re-baselined (nterms=55, FLOPs): occ-veto 1.8/6.5/15.4%; the extmode
witness shows external-occ (~1.95%) and contracted-occ (~1.97%) essentially
equal -- external-mode batching is NOT a recompute fix on the C60 forest (its
original conclusion, now with honest magnitudes). The [cost_profile] giant reads
77.8%, matching an independent dependent-mode analysis (78%).
The batch-variant caching veto in cache_manager had two disjuncts: (a) a node whose own batched_here() carries a Contracted, batchable mode FREE in its own result, and (b) a non-empty cross-occurrence lifetime mask. Disjunct (a) is structurally dead: a Contracted mode is summed AT the node, so it can never be free in that node's result, and post-role-split a free index is stamped External, never Contracted -- so the condition never holds (the occ-veto test's [veto-reach] probe read 0, structurally, not by forest accident). It only ever guarded a malformed emission. Remove disjunct (a) and the `is_batchable_contracted_index` parameter from the cache_manager factory (the only functional caller, build_dryrun_cache, drops the arg; no production caller passed it), the `CacheConfig::is_batchable_index` field and the cost_profile() overwrite that fed it, and the now-obsolete [veto-reach]/[veto-hazard] probes plus the disjunct-(a) sub-test. Disjunct (b) (the cross-occurrence lifetime-mask veto -- the load-bearing F1 correctness guard) is unchanged. Behavior-preserving: [cache_manager] (200), [lifetime_mask] (76), [dryrun], and [eval] (TA production path, 457) all green. Does NOT touch BatchPolicy::is_batchable_contracted_index (the batching decision) or BatchPolicy::is_batchable_index() (the eval accept union) -- both stay.
Design note reframing batched-eval cache placement as register allocation. Three identities -- value (hash), instance (use-site), cell (a materialized copy serving a subset of instances). A value's instances partition into cells; perfect CSE = one cell/value, no CSE = one cell/instance, and the peak budget chooses the granularity in between (a partial un-CSE / materialization DAG). Cell identity = (value, home-scope, split-index): home-scope is the loop level (the axis batching adds), split-index names a same-scope peak split (the RA live-range-split rename). Placement is register allocation + loop-invariant code motion + rematerialization: hoisting a shared value lengthens its live range (peak) to save recompute; slicing adds partial-hoist granularity. Objective: minimize recompute (the rational, batching-aware reuse count W-1 times build cost) subject to the whole-forest peak profile <= peak_threshold. Peak is a placement (post-CSE, whole-forest) constraint, not a factorizer one; cost_profile()'s replay peak is the detection safety net, and a peak that survives full splitting is factorization-inherent. Includes a prior-art section (rematerialization/checkpointing -- Checkmate; electronic-structure space-time tradeoff -- Cociorva/Sadayappan PLDI 2002; register allocation; pebble games), four worked cases, and open items (group-scoped cache keying, the greedy split move, per-placement footprint, W's fixed point).
O1 (cell keying) resolved as a router + dumb stores, not a wider cache key. Add
§7a "Runtime realization": one value-keyed store per (home-scope, split-index)
-- the cache stays TreeNode-keyed unchanged -- plus an explicit router
{value, use-site} -> (home-scope, split-index) that is the placement pass's
output and replaces the implicit parent_ fall-through search. Reads route via
the map then reuse the EXISTING Enter-stage slicer, (use-scope - home-scope)
INTERSECT carried(N), fed the home scope directly instead of via hops; default
{value} -> (home, 0) is byte-identical. Standardize terminology on "home scope"
(= the code's "lifetime scope" = store scope; consumer's is "use scope"). Update
§4, §9, and O1 accordingly; residual O1 sub-items are the use-site/occurrence id,
the parent_/hops audit, and the naming standardization.
Add §7b: the placement pass as a register-allocation spill loop. Seed = perfect CSE (recompute-minimal, peak-maximal); walk up the recompute axis to walk down peak until peak <= threshold. Objective and constraint are exactly cost_profile()'s avoidable_flops and peak_bytes -- no new measurement. Moves: SHRINK (slice a carried mode a cell holds full -- the existing external-slice / node_level_placement, now driven off the true whole-forest peak) and EVICT (delay/un-hoist an invariant cell held idle, or split a long-lived cell's instances into short-lived groups -- the new CSE-aware move the per-term DP cannot see). Greedy: candidates = cells alive at the binding peak point, prefer free shrinks then max ΔPeak/ΔRecompute (the spill metric), apply, incrementally re-cost, repeat; terminate on fit or on a factorization-inherent peak. Residual sub-items O2a (incremental profile update), O2b (per-move estimator/lookahead), O2c (subsume vs run-after the DP external-slice pass).
Clarify §7b: O2 runs after the per-term min-time factorizer and takes the factorization AND batch-loop assignments (batched_here) as FIXED, deciding only the whole-forest eval/placement strategy (home-scope + router); it never adds, removes, or re-assigns a batch loop. Reframe "shrink" from "slice a carried mode" to "re-home a cell into an EXISTING carried loop" -- a placement choice on the fixed nest, not a batching change; deciding to batch an un-batched mode (adding a loop) is the factorizer's lever. Split the termination boundary into two non-O2 failure modes: factorization-inherent (a single intermediate > budget) vs. re-batch-needed (fixed batching left placement too little room, e.g. a shared cell needing slicing on a mode no single term batched) -- both detected via peak_bytes and fed back, giving the structure factorize+batch -> O2 place -> if infeasible re-batch.
Add §7c. Cell footprint is home-relative: a carried mode is sliced (block extent) iff its fixed batch loop encloses the cell's home, else held full -- the existing moment-aware memsize with home-relative extent overrides, so O2's shrink ΔPeak is just the footprint delta. The peak profile is max weighted-interval overlap: each cell is a [first-use, last-use] interval (from the router's use-sites + the static schedule order) weighted by footprint; peak = max over static points of the sum of live cells' footprints (a sweep line), and the argmax is O2's binding peak point. Because it SUMS co-resident live cells it corrects today's peak_bytes = max(scratch, cache) under-count (a lower bound per §1); the replay stays the oracle (must sum, not max, co-residency). The weighted-interval form updates incrementally under an O2 move (feeds O2a). Residual O3a-c: the sweep structure, the summed- co-residency replay oracle, composite/proto sizing.
Add §7d. Define home_scope(value) = deepest scope enclosing the loops of (sliced_modes ∪ demoted_external_modes). sliced_modes is the cross-occurrence meet (max-reuse upper bound); the demotion fold adds the External batched_here stamps the meet demoted (has_demoted_external) -- occurrences bind them to incompatible blocks, so the value can't be a single full value above those loops and its home must be inside them. The fold is exactly what unifies the current cache-veto-vs-has_demoted_external disagreement into one authority both the cache and the runtime read. Per-block is temporal (one external-loop-homed cell re-instantiated per iteration), so no split-index -- that stays reserved for O2's peak-driven same-scope splits. Structural and computed from the meet before O2, which only lowers homes further for peak; consistent with W (the demoted mode is free tiling). Residual O5a-b: confirm the exact signal / edge cases and the seed router construction. Also tie O6 to §7b's two failure modes.
O4 (W's computation order) is not a fixed point: W is a function of the current placement, well-defined at the home_scope seed and re-costed incrementally per O2 move -- seed-then-refine, subsumed by §7b/§7d. O6 (feedback) scoped to a minimal detect-and-report step (surface the binding cell + failure mode so a schedule fails loudly, not silent OOM), with the re-batch/re-factorize hint as a follow-on that the detect step precedes. All major open items (O1-O6) now designed or resolved; the spec is design-complete.
Phased plan for the placement-as-register-allocation design. Phase 1 (detailed, bite-sized TDD) corrects cost_profile()'s peak_bytes from max(scratch, cache) hwmarks to the instant-resolved co-resident SUM across the scope chain (spec 7c/O3b) -- adds CacheManager current_residency()/chain_residency(), threads the chain sum into note_working_set, simplifies the fold, and re-baselines the documented-RED peak figures from measurement. Phases 2-5 (router+home_scope seed, static peak sweep, the O2 greedy, feedback) are a roadmap, each a future plan. Global constraints: no en-dashes, clang-format, byte-identical perfect-CSE default, replay stays the peak oracle.
- cost_profile.hpp: peak_bytes doc no longer claims an EXACT co-resident peak. Document the two known deviations, both conservative (never-under): the ancestor-resident-unsliced-read over-count (alive() is local-only), and the scatter dest-local under-count. The over-count fix (chain-aware alive()) is tracked as a Phase-3 O3b prerequisite so the static peak sweep and this replay agree. - test_cache_manager.cpp: (void)-cast the residency test's [[nodiscard]] store() calls, matching the sibling tests.
…-count) The per-op peak-trace hwmark added an operand's bytes whenever the operand was not a LOCAL cache hit. But an operand read full from an ANCESTOR cache aliases the parent buffer, which chain_residency() already counts -- so it was double-counted (an over-estimate) whenever a hoisted invariant was read full inside a loop it does not carry. Replace the local-only (alive && canon_phase==1 [&& layout_is_default]) alias proxies at all five operand-guard sites with CacheManager::chain_holds(): a pointer-identity test against every alive entry on the whole scope chain. This counts each live buffer exactly once -- a value read full from a cache at any scope is counted via residency, while a sliced/permuted/phase-shifted read is a distinct buffer and is added. Pointer identity subsumes all three old proxies (a fresh apply_phase/permute buffer has a distinct pointer), so the local-case numbers are unchanged; only the ancestor double-count is removed. New CacheManager::chain_holds() + entry::holds(); unit-tested by pointer identity in test_cache_manager. Effect on the occ-veto witness aux+occ leg: 886.1 -> 525.9 GB (the ancestor double-count removed); unbatched/aux-only and both extmode legs unchanged (no ancestor-resident full read); acceptance gates unmoved (Contracted-occ=0, External-occ=244).
Occurrence-key = canonicalize_slots(node sub-expression, named_indices = batched indices) -> SlotCanonicalizationMetadata (the build_subnet_metadata pattern); router consulted by both the Enter read and the place_at_this_level store; empty router => byte-identical. Home stays derived per-node; 7d reconciliation, rational W, the O2 pass, and split-index are deferred.
…at_hops (Phase 2 T2)
… seeded empty (Phase 2 T3)
Reverts the uncommitted member-anchored mode_to_level WIP (dag_scope.hpp, eval.hpp, member_axis.hpp, ordered_executor.hpp, ordered_schedule.hpp, test_ordered_schedule.cpp) back to ddc360e content -- that approach failed on symmetric CSE intermediates (I(i1,i2) == I(i2,i1)), which a per-cell (cell, loop) -> mode map cannot represent when different use-sites bind the loop to different physical slots of the same shared value. The env-gated [PROD] diagnostic in backends/tiledarray/result.hpp is kept (unstaged) as a validation tool for the follow-on tasks. Adds tests/unit/test_sliced_canonical_layout.cpp: a minimal probe using a Symmetry::Symm 2-occ-index intermediate shared by two use-site contractions that bind the batched loop to different physical slots (slot 1 in one, slot 0 in the other). It pins today's baseline: occurrence_key correctly folds the two occurrences to one router key, yet index_position shows the loop sits at a different physical slot per occurrence for the same value -- the exact ambiguity the loop-coloring design (doc/dev/specs/2026-08-23-sliced-value-canonical-layout-loop-coloring-design.md) resolves by coloring sliced slots with their DAG-scope loop identity.
…urrence_canonical_layout
…e canonical layout
The Task-7-P2 consumer-aware seam (LoopColoredSliceSeam::by_hash_consumer + SlicedModeAssignment::occ_facts) replaced the old per-cell node->ModeToLevel map on the ordered path, and the transitional equivalence assert was already retired. Delete the whole dead chain: ValueCell::mode_to_level, compute_cell_mode_to_level / populate_cell_mode_to_level, the cache-side plumbing (cell_mode_to_level_/set_cell_mode_to_level/cell_mode_to_level/ mode_to_level_of), the executor's self-wiring block and its CellModeToLevelGuard, the eval.hpp m2l equivalence-gate diagnostic and its two inherit sites, and the now-obsolete populate_cell_mode_to_level unit test. The ModeToLevel TYPE and mode_to_level_from_signature (used by compute_sliced_mode_assignment and slicing_signature.hpp) are unrelated and stay. Also convert the manual current_consumer save/restore around run_ordered_contracted_block's two evaluate_impl call sites into an RAII guard (CurrentConsumerGuard, mirroring LoopColoredSliceSeamGuard): the manual restore leaked on a throwing evaluate_impl. Behavior on the non-throwing path is unchanged (byte-identical test results).
… deadlock) Batched slice-on-use resolved a canonical Index label against the fetched node, but a divergent (relabeled) CSE occurrence is a different index-frame that only shares SLOTS with the value's canonical frame, so the label match failed (unsliced) or mis-resolved -> a shared occ index served full (32) in one operand and sliced (16) in the other -> the ToT einsum's tiled-range TA_ASSERT is elided in Release so the DistEval deadlocked. Fix, in-frame end to end: - compute_sliced_mode_assignment stamps each occurrence in its OWN frame (iterate occ.carried, not the value's first-occurrence w.carried). - LoopColoredSliceSeam by_hash/by_hash_consumer and occ_facts carry a physical POSITION computed in each occurrence's own frame at schedule time; mode_of returns a position; slice_to_use uses it directly -- no index_position, no cross-frame label match, no first-match guess. - Remove space_mapped_slicing (it guessed a mode by space); the resident-reads guard, the [homed] escape-output trace line, and the truthful sliced=/scope= and tiled-range traces stay. Fixes the water-8 occ+aux 5-member (L*R) deadlock. A sibling-occ-loop case (two forced-split occ loops sharing one canonical axis) still mis-binds a mode's loop by SPACE in the assignment -- documented in doc/dev/specs/2026-08-24-frame-correct-slot-slicing-design.md; the fix is the next step. Env-gated investigation diagnostics (SEQUANT_UT_SMA_DIAG/HOME_DIAG/ PROD_TR, trange_annot) are retained for that pass and reverted at finalize.
Two TEST_CASEs assert the pre-2026-08-23 label seam (mode_of -> optional<Index>, by_hash keyed by Index) and no longer compile against the position-based seam, which blocks linking unit_tests-sequant. Guard them with #if 0 and a marker; the loop-open-vs-sliced-mask plan Task 4 rewrites the seam and these cases against the final per-occurrence positional semantics.
Add EvalExpr::batch_loops_opened_here() -- the subset of batched_here() for which a node is the loop-OPEN site (outermost node introducing the physical batch loop), distinct from the per-node sliced mask the DP stamps on every carrying node. Carried on NodeBatchAnnotation::opened_here and applied by binarize alongside set_batched_here. Lets a consumer reconstructing the enclosing-loop nest (peak_profile's ectx) count each physical loop once instead of once-per-carrying-node. Default empty; OFF path byte-identical. Task 1 of the loop-open-vs-sliced-mask plan.
reconstruct_batched_modes now fills NodeBatchAnnotation::opened_here -- the subset of a node's axes for which THIS node introduces the physical batch loop: - external, root-seed path: opened at the term root only (an external mode is on the final result, so the root is its outermost carrier); - external, node-level path: the injection site (new opened_at_node, NOT the D1-propagated carrying descendants that placed_at_node also records); - contracted: at the unique contraction node (aprime), mirroring the emitted Contracted axes. Lets peak_profile build its enclosing-loop nest (ectx) from opens so one physical loop counts once instead of once-per-carrying-node. axes emit unchanged; OFF path byte-identical. Observable witness is the w8 ectx (Task 3); the optimize-suite external-emission unit tests are pre-existing-red on this branch (orthogonal calibration drift, surfaced now that the suite links again). Task 2 of the loop-open-vs-sliced-mask plan.
peak_profile's visit accumulated n->batched_here() into the enclosing-loop
context, but the DP stamps an external mode's sliced mask on EVERY carrying
node, so one physical batch loop piled up once-per-carrying-node -- ectx became
the same occ index repeated (i i i ...) and no longer matched the DAG scope's
one-loop-one-level de-duplication. Read n->batch_loops_opened_here() instead,
which names each physical loop once. own_modes likewise moves to opens, matching
its documented intent ('its own node, not an ancestor's').
Verified on w8 occ+aux (SEQUANT_UT_SMA_DIAG): |ectx| drops from 6/7/3/4 to 2/1,
no duplicates; single-external occurrences now align |ectx|==|scope|==1 and
two-external ones show the two distinct loops cleanly, so dedup(ectx) & carried
yields clean per-occurrence slice positions for Task 4. OFF path byte-identical.
Task 3 of the loop-open-vs-sliced-mask plan.
…adlock) Replace compute_sliced_mode_assignment's EXACT + REGIME-2 base_key passes with a single per-occurrence positional rule. For each occurrence occ of value W consumed by C, the loops the runtime crosses fetching W are C's enclosing DAG blocks (build_scope), outermost first; dedup(occ.ectx) -- now one entry per physical loop (Task 3) -- names those loops in W's OWN frame, outermost first. Pair by nest position: block scope[k] slices occ mode nest[k], at physical position index-in(occ.carried). All in occ's own frame -- no base_key, no cross-frame label match, no first-match guess. Divergent (relabeled) and symmetric occurrences (one value sliced on different positions by different consumers) are handled uniformly by the consumer-keyed occ_facts; by_value is left empty (the ordered executor always fetches under a tracked consumer). The base_key block-find collapsed sibling occ loops (all occ share base_key 'i'), mis-slicing divergent occurrences -> two operands of one contraction disagreed on a shared occ mode's extent -> ToTxToT DistEval deadlock (TA_ASSERT elided in Release). w8 occ+aux ordered now completes: CSV-CCk Energy -1.602851115435274 vs reference -1.6028511154353591 (8.5e-14, FP summation order; lossless). SMA diagnostics retained (inert); removed in Task 5. Task 4 of the loop-open-vs-sliced-mask plan.
…-leaf assignment case peak_profile now builds OccurrenceRec::ectx from batch_loops_opened_here (Task 3), so fixtures that manually stamp batched_here must also stamp the loop-open: - test_eval_ta forest-descent equivalence cases (one Contracted block, one External block, reset-in-Contracted) and the orderedsched_stamp helper: aux Contracted opens at its contraction node, occ External opens at the root only; - test_eval_dryrun whole-scope outer-homed-aux case: both loops open at the root. Guard 'compute_sliced_mode_assignment: a DF-leaf's aux+occ modes map ...': it asserts the value-keyed loop_of/by_value API, which the per-occurrence occ_facts (consumer-keyed) seam supersedes -- by_value is now legitimately empty and loop_of returns nullopt. To be rewritten against occ_facts / by_hash_consumer. Restores the eval unit suite to its pre-existing baseline (0 new regressions); w8 occ+aux ordered stays lossless (-1.6028511154354086).
…ts gap) Because by_value is now empty (per-occurrence occ_facts is the slice source and the value-keyed by_hash fallback is unused), a fetch the schedule failed to attribute a sliced-mode fact to would silently serve the operand UNSLICED and mismatch its contraction partner -- and TA's tiled-range assert is elided in Release, so the mismatched ToTxToT DistEval deadlocks instead of erroring. Add LoopColoredSliceSeam::participates(hash, loop) -- whether the value has ANY sliced-mode fact under the loop (any consumer, either map) -- and a guard in slice_to_use: when the value PARTICIPATES in the crossed loop but this fetch got no fact, throw sequant::Exception instead of proceeding unsliced. A value with NO fact under the loop is genuinely invariant and correctly left unsliced, so the guard does not fire there. w8 occ+aux ordered stays lossless (-1.6028511154355747) -- occ_facts is complete for it, the guard never fires.
Remove the SEQUANT_UT_SMA_DIAG investigation scaffolding from compute_sliced_mode_assignment: the [SMA-ALIGN]/[SMA-PAIR] traces and the hardcoded w8 target-hash predicate (_sma_tgt). These were env-gated (inert without the var), so removal is byte-identical. The general env-gated eval traces ([HOME-SLICE], [PROD-PRE], trange_annot) are left in place -- they are not hardcoded to w8 and remain useful for the ongoing multimode-batched-eval branch. Task 5 of the loop-open-vs-sliced-mask plan.
'batched_here' read as 'batching happens at this node', but it is the per-node sliced-mode MASK -- the modes this node slices when it evaluates, which the DP stamps on EVERY carrying node. That mismatch is what made the propagated mask look like a stack of loops in peak_profile's ectx. Now that the discrete loop-open lives in batch_loops_opened_here(), the contrast made the old name actively misleading. Mechanical, behavior-neutral rename: EvalExpr::batched_here() -> node_slice_mask(), set_batched_here() -> set_node_slice_mask(), field batch_axes_ -> node_slice_mask_ (the field is private to eval_expr.hpp; the look-alikes term_batch_axes_mutex and batch_axes_indices are untouched). sliced_modes() (the cross-occurrence residency meet) keeps its name. mpqc builds, eval unit suite 0 regressions, w8 occ+aux ordered lossless (-1.6028511154358982).
Three tests validate the loop-COLORED (3-arg NamedIndexColorMap) occurrence_key as Pillar-1 value-id: depth coloring DISTINGUISHES which slot a loop slices on a non-symmetric tensor; a symmetric tensor FOLDS the two depth assignments; an empty color map is byte-identical to the 2-arg space-only key (the #1 non-regression anchor). The primitive is occurrence_key as-is; the per-scope ValueIdColoring + value_id_of lookup land in Task 2. Task 1 of the Pillar-1 slice-colored-value-identity plan.
build_ordered_schedule now records OrderedSchedule::home_mode_depth -- per value homed below the top scope, its sliced result mode -> DAG-scope depth, in the value's OWN carried (canonical) frame. Recorded ONCE at home placement by walking the realized block tree and pairing the cell's carried with the enclosing levels by space + nest position (the frame-safe occ_facts routing). NOT re-derived from ectx; keys are value-frame, never base_key. home_modes was the wrong source (home minus own-loop modes -> empty). value_id.hpp adds ValueIdColoring, value_id_coloring() (from a home_mode_depth entry), and value_id_hash() (colored occurrence_key hash when the node carries an in-scope sliced mode, else plain hash::value -- byte-identical to TreeNodeHasher). occ_facts untouched. Task 2 of the Pillar-1 slice-colored-value-identity plan.
…entity TreeNodeHasher gains an optional id_override; TreeNodeEqualityComparator an optional colored_eq_override returning optional<bool> (nullopt => fall through to the byte-identical structural compare). Injected from cache_manager (kept out of this low-level header, so no tensor-network dependency). When installed on a sub-top cache they key by the home-slice-colored value-id -- value_id_hash / occurrence_key -- distinguishing I(i,_) from I(_,i) by which slot carries the loop var (ONE per-scope coloring; the node's own slots do the distinguishing, so no per-node map that would collide on the shared node-hash). Null on the main cache => plain hash::value / structural, byte-identical. Test pins: value_id_hash distinguishes the two slots, null path == hash::value byte-for-byte, and the hasher/comparator honour the overrides. 0 new regressions (the 9 eval-tag fails are the pre-existing baseline). Task 3 of the Pillar-1 slice-colored-value-identity plan.
Add the value-keyed runtime-cache/remat-router key CachedValue{node, coloring}
and its functors:
- CachedValueHasher = value_id_hash(node, coloring) (colored occurrence_key
hash when sliced; hash::value byte-identical when not)
- CachedValueEqual = colored occurrence_key graph compare when both sliced,
else structural TreeNodeEqualityComparator
This is what lets the cache distinguish post-remat split values that share ONE
forest node (build_value_node_map is hash-keyed): the discriminating home-slice
coloring is recorded beside the node, not re-derived from the shared node. Factor
the shared slice-relevance fork into value_is_sliced(node, coloring), used by both
the hasher and the comparator.
Test (test_occurrence_key.cpp): in an unordered_map keyed by CachedValue, a
non-symmetric B sliced on slot 0 vs slot 1 -> two entries; empty coloring folds
them to one (node-id, as today) and hashes byte-identically to hash::value; a
slot-symmetric B folds the two slicings. Pre-existing [dryrun][cache] 'giant'
failure confirmed unrelated (fails identically with these changes stashed).
…pty=identity) Promote the runtime cache from node-keyed to value-keyed. CacheManager gains a key/node split: key_type stays TreeNode (the NODE type -- drivers, custom evaluator, persistence predicate, for_each_key, and every caller are unchanged), and a new cache_key_type = CachedValue<TreeNode> keys cache_map_ + the map-keying method params. CachedValue is IMPLICITLY constructible from a node, so every existing cache call site (~37 across eval/ordered/scope executors) keeps compiling and, with an empty coloring, keys the map byte-identically to the old node keying. Only the ordered executor (Task 5) will pass a genuinely colored key; until then every key is empty-colored, so value_id_hash == hash::value and CachedValueEqual == the structural TreeNodeEqualityComparator -- provably identical. The recompute tally stays node-keyed (a per-node diagnostic the eval tests consume node-by-node) via key.node; for_each_key hands callers key.node. CachedValueHasher carries the force_hash_collisions param so the collision-safety tests still force collisions. Byte-identity verified: [cache]/[value-id]/[occurrence_key] shows only the pre-existing [dryrun][cache] 'giant' failure; [placement_remat]/[placement_router] failing-assertion set is IDENTICAL (deterministic diff) between the node-keyed baseline and this value-keyed build.
…ange)
Factor evaluate_impl's innermost single-op compute into a free value-in/value-out
kernel apply_one_op(node, left, right): dispatch on node->op_type() and call the
operand Result method with the annotation computed from node -- Adjoint ->
left->adjoint({node.left()->annot(), node->annot()}); Sum -> left->sum(*right,
ann); Product -> left->prod(*right, ann, de_nest?...). No cache, no value identity.
Everything else stays exactly where it was, so the seam is a byte-identical move:
the shaped-product hook consult (caller calls apply_one_op only when the hook
declines), apply_phase + store in finish_phase_b (apply_phase is CONDITIONAL --
would not be byte-identical if moved), the in-place-Sum fast path (inline), and the
recompute tally / EvalTrace / timing / last_op_flops sentinel / force_sync fence
(wrap the call). Leaf is not an apply_one_op case (a leaf_evaluator fetch).
This is B-full stage A ('extract first'): the value/occurrence-driven ordered
executor (Task 7) will call this same kernel with home-colored operand fetches and
a colored store. evaluate_impl is otherwise untouched (whole-scope/direct paths
keep it).
Byte-identity verified: clean build; [eval] shows only the pre-existing TA
batched-ToT External-occ failure (test_eval_ta.cpp:2803, identical on HEAD); the
[dryrun] cost_profile test fails 5/43 identically in isolation on both HEAD and
this change (the full-[dryrun] delta was Catch2 test-order randomization over the
dryrun tests' shared static state, not this refactor).
…redSchedule Add OrderedSchedule::operand_vids -- per value_id, the value_ids of its DIRECT operands -- populated in build_ordered_schedule from the dep graph it already computes for the topo-sort (ordered_schedule_dep_graph(rich).depends_on, whose edges come from every OccurrenceRec's consumer_point, so a split operand resolves to the specific consumed value rather than an ambiguous node hash). This is the value/occurrence prerequisite the value-driven ordered executor (Task 7) consumes to fetch each operand by its OWN home-colored key. Pillar 2's occ_facts already own the per-occurrence use-context (slice-on-use), so only the operand VALUE_IDS are needed here. Purely additive: the field is populated from an already-computed graph (no new work) and read only by Task 7, so existing behavior is unchanged. Test (test_ordered_schedule.cpp): on the sp2 occ-outer/aux-inner fixture, operand_vids is non-empty, equals the dep graph's depends_on, and every entry keys a non-leaf value with in-range, non-self operands.
Key the ordered runtime cache by VALUE, not node, so a sliced value homed below
the top scope is found by its distinct home identity -- fixing the I(i,_) vs
I(_,i) collision. The coloring reaches evaluate_impl's bare-node store/access
through a per-SCOPE coloring context on each batched scratch:
- CacheManager gains value_coloring_ctx_ (node hash -> home coloring) + recolor:
a map-keying method colors an EMPTY-coloring key from it (an explicitly-colored
value_of(vid) key is honored as-is), so store/access are value-keyed.
recolor_registered_entries() re-keys a freshly built scratch's member entries
the same way -- the crux: make_batched_scratch pre-registers members by bare
node, so without this the colored store/read would miss them (a homed shared
sliced value 'vanished'). The scratch now genuinely holds VALUES: two values of
one node coexist as distinct colored entries.
- make_batched_scratch takes an optional coloring context, sets it on the new
scratch and recolors its registration. Null (non-ordered / whole-scope paths,
and the top-level cache where nothing is sliced) => byte-identical node keying.
- run_ordered_contracted_block builds the per-scope context (its BuildSteps +
escape outputs + their operands, from operand_vids + home_mode_depth) and hands
it to make_batched_scratch; the executor's own home stores/reads pass
value_of(vid) directly. The root (top-level) cache stays uncolored.
Validated: forest-descent ordered-vs-forest equivalence unit tests pass (lossless
batched == forest; the value-keyed-scratch fix is what makes the sliced shared
intermediate S=g*h resolvable); no new unit regressions (remaining failures are
pre-existing on HEAD); w8 occ+aux CCSD is wet-lossless (energy -1.60285111543...,
no completeness-guard fire, no vanished-value, EXIT 0).
…achedValue) Remove the id_override on TreeNodeHasher and colored_eq_override on TreeNodeEqualityComparator (added while pursuing the abandoned node-keyed K1 route). Under CachedValue keying, the concrete CachedValueHasher/CachedValueEqual own the coloring logic, so these functors return to their pre-Pillar-1 node-id form for compute_dag_boulevard/CSE and the top-level cache. Drop the now-unused <functional>/<optional> includes and the test lines that exercised the overrides (the colored keying is covered by the CachedValue map test). Byte-identical: [value-id]/[occurrence_key]/[cache] shows only the pre-existing [dryrun][cache] giant failure.
Remove the self-contained, verified-dead test-only cluster subsumed by the value-keyed cache + occ_facts: loop_colored_id, populate_occurrence_canonical_layout, and populate_canonical_layouts (ordered_schedule.hpp; no runtime caller), and the fields they alone wrote -- ValueCell::canonical_layout and OccurrenceRec::perm_to_canonical (peak_profile.hpp; no live reader). Delete test_sliced_canonical_layout.cpp (its whole subject) and its CMakeLists entry, and the two #if 0 'rewrite pending' tests in test_ordered_schedule.cpp that referenced the removed symbols. Kept (still live): SlicedModeAssignment (by_value/loop_of) and compute_sliced_mode_assignment -- the LoopColoredSliceSeam builds from them and consults loop_of_level at runtime; retiring by_value there is a separate follow-up. Byte-identical (test-only removal): builds clean; the remaining [ordered-schedule] failures (build_ordered_schedule water-20 / forced-split / executor-shape) are pre-existing known-open issues, unchanged by this removal.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Multimode batched evaluation of factorized coupled-cluster equations
Adds cost-model-driven multimode batching to the SeQuant evaluator so
large-system CSV/PNO-CC residuals can be evaluated without forming their
largest transients whole. Six squashed commits (dry-run backend, optimizer,
evaluator, supporting core, tests, docs).
What it does
-> scattered into disjoint slices) and contracted (DF aux, summed ->
accumulated). Loops nest external-outside-contracted.
DenseTimeSpace): minimizes flops withpeak_thresholdas a ceiling; role-split (contracted/external) batchability;order-aware placement over the combined nest.
(a cached intermediate fetched from an outer scope is sliced to the current
block) decouples correctness from placement; per-level placement driven by
a per-canonical lifetime mask (cross-occurrence proto-aware meet) unioned
with contracted residency; iterative (stack-safe) tree traversal.
C60 PNO-CCSD dry run (55-term residual, aux K@256, occ@8, 100 GB budget)
The DP selects the same factorization regardless of what is batchable
(flops are unchanged); batching only slices modes to lower the peak. The
roofline-time column moves because a giant intermediate executed whole is
memory-bound (
machine_balance x traffic) but compute-bound when sliced -- thecache-blocking win of the same schedule, not a cheaper one. Recompute overhead
(
avoidable_time) is 1.8% -> 6.5% -> 39.8% as slicing gets more aggressive.Validation
[eval]449,[lifetime_mask]76,[optimize]628 assertions green;OFF (order-blind) path byte-identical.
events) matches unbatched to < 1e-9, within the 1e-7 precision, no aborts.
Follow-ups (non-blocking, from the final review): dedup the proto-expansion
helper; add a real-forest hidden-tag hash-regression test; revisit the
stamp_lifetime_masksconst_cast.