Skip to content

Repository files navigation

Dynamic Execution Architecture

Dynamic Execution Architecture (DEA) is an executable research scaffold for a learned microarchitectural control plane. A global controller observes machine telemetry every eight modeled cycles and selects a bounded, cross-subsystem policy bundle while hard simulator logic preserves dependencies, resource limits, retirement correctness, and thermal enforcement.

The controller coordinates:

  • instruction issue priority;
  • DRAM request ordering, cache admission, and sequential prefetch depth;
  • eco, balanced, or turbo power/performance mode;
  • CPU/FPU versus vector-accelerator routing;
  • off, conservative, or aggressive branch speculation;
  • shortest, congestion-adaptive, or thermal-aware NoC routing.

This is the project thesis: plan how valid computation should flow through a heterogeneous machine, without rewriting program semantics or granting a learned policy control over correctness.

Architecture

workload DAG (adaptive or family generator, seeded)
          |
          v
hard execution fabric <-----------------------------------+
  DAG + resources + cache/DRAM + predictor + NoC + thermal|
          |                                                |
          v                                                |
global machine state (21-dim observation incl. active action)|
          |                                                |
          v                                                |
expert / distilled actor / PPO actor                       |
          |                                                |
          v                                                |
bounded ControlAction -------------------------------------+
  issue + memory + cache + prefetch + power + accelerator
  + speculation + NoC routing
          |
          v
safety governor may override unsafe power actions

The global controller decides every eight modeled cycles; local dispatch continues each cycle. Both checked-in learned controllers require one matrix-vector product per decision, and their inference energy is charged to the simulation.

Audit: defects found and fixed

A critical audit of the original scaffold found that several subsystems were decorative rather than causal, and that the learning pipeline had failure modes it could not observe. Each item below is pinned by a regression test (tests/test_dea_audit.py, tests/test_dea_rl.py):

  1. The observation could not see the currently applied action, so no policy — expert or learned — could express hysteresis. Fixed: six one-hot "active bundle" features plus recent-L1-hit-rate and memory-in-flight features (21-dim state).
  2. The thermal model was numerically inert: no workload ever got hot, so turbo was free and the governor never fired. Fixed: calibrated per-zone heating (0.11/cycle) so turbo genuinely trades cycles for thermal exposure and draws governor overrides under sustained use.
  3. The objective used an all-time peak temperature squared, which is path-dependent noise. Replaced with integrated squared exposure Σ_cycles max(0, T_hot − target)².
  4. The objective double-counted controller inference energy (subtracted once, penalized again). Now charged exactly once.
  5. The RL reward double-counted stalls and DRAM misses already inside cycle cost. Rewards are now strictly the incremental objective of the interval; a test asserts the sum of interval costs equals the final objective to 6 decimal places.
  6. Checkpoints accepted stale feature schemas. Both learned controllers stamp and check format_version; artifacts record a global dea_version (currently 0.3.0), echoed in benchmark protocols.
  7. Cache sampling drew from an unseeded global RNG, making whole runs irreproducible and tests flaky. The legacy chip prototype's cache is now seeded per run.
  8. Behavior cloning was trained only on the phase-changing family, so the distilled student collapsed off-distribution (it lost catastrophically on branch-heavy graphs). Distillation is now multi-family with adaptive majority.
  9. PPO did not move: a sharp argmax-equal clone has zero exploration entropy, so every gradient update re-enacted the teacher. The warm start now shrinks cloned weights by a temperature factor (×0.25); argmax is invariant to positive scaling, so the greedy initial policy is exactly the teacher while sampling becomes exploratory.
  10. Late-training collapse appeared in every long PPO run. Updates now stop early when the approximate KL exceeds a trust-region threshold (0.03), and checkpoint selection uses a 12-seed held-out gate with a recorded learning curve.
  11. Non-adaptive workload families were single fixed graphs, inviting memorization and meaningless multi-seed evaluation. All four family generators are now seeded and parameterized (length, chaos cadence, walk instance, vector width, reuse mix).
  12. The cross-family benchmark shared one consumed graph across controllers, silently scoring zero work for everyone after the first. Evaluation generates a fresh graph per controller; a regression test documents the consume-on-run contract.

Objective

cycles
+ 0.08  * (energy_units - controller_energy_units)
+ 0.6   * sum_over_cycles( max(0, hottest_zone_temperature - target)^2 )
+ 0.02  * controller_energy_units

The scalar is an experimental preference, not a physical constant. The artifact separately preserves latency, energy, thermal, cache, prefetch, speculation, NoC, accelerator, safety, and controller-overhead metrics.

Held-out result (100 disjoint adaptive workloads, seeds 10000–10099)

Training used seeds 0–319 for PPO (with per-update selection gates on later seeds) and seeds 0–79 for distillation rollouts; evaluation seeds are disjoint and the benchmark refuses overlapping ranges by default.

Controller Mean objective Cycles Energy Thermal exposure Paired Δ vs balanced
Fixed balanced 1377.6 1294.1 1042.9 0.0 baseline
Fixed streaming 1359.2 1252.1 1136.5 26.9 −18.4 ± 10.9
Independent heuristics 1368.4 1281.0 1079.9 2.0 −9.2 ± 2.7
Best fixed mode per workload 1343.1 1243.8 1104.4 18.2 −34.5 ± 6.0
Hierarchical teacher 1315.0 1237.0 979.5 0.0 −62.5 ± 4.8
Distilled actor (BC) 1315.6 1237.0 987.8 0.0 −61.9 ± 4.8
Validation-gated PPO 1287.5 1211.4 956.6 0.0 −90.0 ± 3.5
  • PPO improves the objective by 6.5% versus balanced (−90.0 ± 3.5), beats the retrospectively selected best fixed mode by 55.6 ± 5.7, and beats its own behavior-cloned initialization by 28.1 objective units (2.1%) while retiring identical work.
  • The PPO advantage is real but modest in relative terms; it comes from cycle reduction (1211.4 vs 1237.0) at lower energy. Unlike earlier rounds, all three training seeds improved over their warm start on held-out selection seeds — the previous result ("PPO is operational but does not improve on cloning") no longer holds.
  • Learned policies hold thermal exposure at 0.0: they simply do not enter turbo when headroom is missing, so the governor never overrides them.

Cross-family generalization (12 seeded instances per family)

Family Balanced Expert Distilled PPO
pointer_chase 3466 3470 (+3) 3467 (+1) 3467 (+1)
branch_heavy 616 621 (+5) 602 (−14, 11W/1L) 608 (−8, 9W/3L)
matrix_heavy 4979 4279 (−700) 4278 (−701) 4072 (−907)
mixed 2649 2559 (−90) 2557 (−92) 2560 (−90)

Honest reading:

  • pointer_chase is ILP-bound. A serialized load→address→load chain offers nothing to coordinate; no controller improves it, and we report that as a negative result rather than tuning around it.
  • branch_heavy transfer is fixed, not assumed. Before multi-family distillation the student lost catastrophically here; the distilled actor now beats balanced on 11 of 12 held-out instances and edges the expert teacher.
  • matrix_heavy shows real positive transfer, and PPO transfers best (−18% vs balanced), suggesting RL found accelerator/streaming coordination that cloning under-used.
  • mixed-family phase transitions are handled by all learned controllers roughly equally.

Training methodology

Distillation (train_dea.py): the hierarchical teacher controls 80 seeded rollouts drawn adapt-majority from five families (adaptive 4 : branch_heavy 1 : matrix_heavy 1 : mixed 1 : pointer_chase 1). Every per-interval decision becomes a labeled example (19,028 examples); a class-balanced softmax regression fits the linear student to 97.3% top-1 accuracy. Class balancing matters: without non-adaptive families the latency bundle was never observed at all.

PPO (train_ppo_dea.py): clipped surrogate, GAE(λ=0.95), γ=0.99, linear entropy anneal 0.02→0.002, Adam (policy 2e-3, value 1e-2), batches of 8 episodes, ≤8 epochs per update with approximate-KL early stopping at 0.03, rewards scaled by 0.02 for value-function conditioning. Exploration comes from the temperature-scaled warm start described above. Checkpoint selection evaluates every candidate on 12 held-out adaptive seeds and keeps the gate argmin; the full selection learning curve is stored in the checkpoint.

Three independent training seeds (80, 500, 1000; 320 episodes each) all improved over the warm start (selection-objective deltas −2.0%, −0.06%, −1.1%), and every run showed late-curve degradation that the gate rejected. The shipped checkpoint (seed 1000) won a common held-out gate (seeds 9000–9011: 1313.3 vs 1320.9 and 1331.8).

Evidence and reproducibility artifacts:

What is causally modeled

Memory hierarchy

  • 16-line L1 and 64-line L2 with exact LRU state and a rolling 32-access hit-rate window;
  • L1, L2, row-hit DRAM, and row-miss DRAM latency;
  • low-reuse cache bypass;
  • prefetch bandwidth, energy, pollution, and usefulness;
  • rolling miss rate, sequentiality, and memory-level-parallelism observations.

Compute, branch, and interconnect

  • bounded ALU, FPU, load, store, branch, and accelerator resources;
  • dependency critical paths and power-mode issue widths;
  • vector eligibility and accelerator wake/setup cost;
  • signature-indexed two-bit branch counters updated only after resolution;
  • explicit correct-prediction benefit and misprediction flush cost;
  • finite-capacity direct and alternate NoC routes with causal reservations;
  • rolling branch-miss and NoC-congestion observations.

Power, thermal, and safety

  • operating modes alter issue width, compute latency, energy, and leakage;
  • four thermal zones model core, memory, accelerator, and controller activity with calibrated heating;
  • controller inference has explicit energy cost, charged exactly once in the objective;
  • a hard governor forces recovery or disables turbo near the limit; the expert controller recovers hysteretically.

Energy, power, and temperature are normalized experimental proxies, not calibrated joules, watts, or silicon measurements.

Workloads

workloads/adaptive.py generates three serialized randomized macro-phases — streaming vectors, irregular pointer chains, control-heavy lanes — and the policy never sees phase labels. The four stress families are seeded generators whose difficulty knobs vary per seed:

  • branch_heavy: chaotic branch storms with predictor aliasing;
  • matrix_heavy: wide FMA chains over streaming rows;
  • pointer_chase: serialized unpredictable address walks;
  • mixed: gemm → barrier → irregular transitions.

Learning systems

  • learned: class-balanced behavior cloning from complete multi-family expert rollouts.
  • ppo: dense interval reward aligned exactly with the objective, GAE, clipped PPO, entropy annealing, KL-bounded updates, temperature-scaled cloning initialization, validation-gated selection.
  • expert: transparent phase-reactive teacher with hysteretic thermal recovery.
  • independent: local heuristics without a coordinated global mode.
  • balanced, streaming, irregular, latency, control: fixed bundles included in best_static.

One RL transition advances at most eight hardware cycles and is rewarded with the exact negative incremental objective of that interval. The sampled action is the action injected into the simulator; the environment never substitutes an expert label.

Reproduce

Python 3.10+ is required.

python -m pip install -r requirements.txt
python -m unittest discover -s tests -v
python train_dea.py --episodes 80 --seed 0 --epochs 600
python train_ppo_dea.py --seed 80 --output artifacts/dea_ppo_s80.json
python train_ppo_dea.py --seed 500 --output artifacts/dea_ppo_s500.json
python train_ppo_dea.py --seed 1000 --output artifacts/dea_ppo_s1000.json
python benchmark_dea.py --seeds 100 --seed-start 10000
python run_dea.py --compare --seed 10000

train_ppo_dea.py --warm-start artifacts/dea_ppo.json after copying your chosen seed checkpoint to artifacts/dea_ppo.json. Training/evaluation overlap is rejected by default; --family-seeds controls cross-family sample size. The benchmark records checkpoint hashes, DEA/python/platform versions, every raw run, paired 95% confidence intervals, and completion status.

Serve the live dashboard:

python -m uvicorn dashboard.app:app --host 127.0.0.1 --port 8000

Open http://127.0.0.1:8000. It streams active policies, cache and branch misses, NoC congestion, accelerator routing, thermal zones, safety actions, switches, and objective.

Repository map

dea/contracts.py          state, action, objective, safety contracts, dea_version
dea/memory.py             causal L1/L2/DRAM and prefetch model
dea/branch.py             two-bit predictor and speculation costs
dea/noc.py                finite-capacity adaptive interconnect
dea/controller.py         fixed, independent, expert, and distilled policies
dea/environment.py        dense control-interval RL environment (graph factory)
dea/ppo.py                clipped PPO actor-critic, KL early stopping, checkpoints
dea/simulator.py          multi-domain closed-loop machine
workloads/                seeded adaptive + four family generators
train_dea.py              multi-family expert-to-fast-actor distillation
train_ppo_dea.py          family-mixed, warm-started, gated PPO training
benchmark_dea.py          held-out paired evaluation + cross-family section
dashboard/                live control-plane telemetry
tests/test_dea_audit.py   one regression test per audited defect
tests/test_dea_rl.py      environment, PPO mechanics, family generators
chip/, policy/            retained narrow issue-scheduling prototype

Research boundary and open limitations

This is a causal multi-domain laboratory, not a validated processor model. Real architecture research still needs detailed fetch/decode/rename/commit semantics, established traces, calibrated controller latency/storage/energy, and Pareto reporting across objectives rather than a scalarization. gem5 O3CPU, gem5 Garnet 2.0, and ChampSim are the natural fidelity bridges. The PPO implementation follows Schulman et al., 2017.

Known open limitations, stated plainly:

  • the policy class is a single linear map over 21 features; it cannot represent arbitrary decision boundaries;
  • gains beyond behavior cloning are real but small (≈2% objective on the training family); most of the advantage over baselines still originates in the expert design;
  • pointer-chasing workloads are bounded by dependency structure, and no controller — learned or hand-written — can fix that inside this action space;
  • thermal parameters are calibrated for trade-off visibility, not against silicon measurements.

The repository currently supports this claim:

A hierarchical controller can coordinate causal execution, memory, speculation, interconnect, power, and accelerator knobs in an abstract phase-changing machine. A small distilled policy preserves that behavior off its training distribution, and PPO — initialized from the softened clone, gated on held-out seeds, and KL-bounded — adds a statistically significant further improvement while correctness and safety remain outside learned control.

About

A research prototype for a learned microarchitectural control plane in CPU branch prediction

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages