perf(v2): shared activation cache - further 1.66-2.00x real speedup (DEC-0016)
Implements optimization #2 from the final benchmark campaign's own recommendation, on top of DEC-0015's word-level burst rewrite: a new shared activation_cache.v module fetches a given activation (X) vector from PSRAM once instead of once per neuron sharing it - the exact redundant traffic pattern the dense-layer workloads in this project's benchmark suite exhibit. Each memory_manager's own prefetch_engine now fetches WEIGHTS only; the activation half is requested from the shared cache instead (single-tag, tile-granular, N_SLOTS request ports, its own real word-level PSRAM backend via a new dedicated arbiter port). dataflow_core.v/slot_mem_arbiter.v/neural_multiprocessor.v widened to N_SLOTS+1 ports to arbitrate the cache's traffic alongside each slot's weight traffic. Two real bugs found and fixed during implementation (ERR-0010): a target-bank/pending-bank race in memory_manager.v's activation-cache wiring (the same bug class ERR-0006 already fixed once for pf_target_bank - a later handoff's queued request can overwrite which bank an earlier, still-in-flight request's ack applies to), and a repeat of ERR-0009's N_SLOTS=1 zero-width replication bug in activation_cache.v itself. Real, measured results: the full final-benchmark campaign (24/24 workload/config combinations) re-verified bit-exact. D-Stress cycles fall a further 1.66-2.00x on top of DEC-0015 (~4x combined vs the original byte-level baseline). But the cache's real Fmax cost is much steeper than DEC-0015's own: N_SLOTS=2 (the recommended default, DEC-0014) drops from 133.58 to 87.72 MHz (-34%, margin over 80MHz shrinks from +67% to +9.7%), and N_SLOTS=4 drops to 65.01 MHz - now FAILING the 80MHz target it previously passed. Combined real wall-clock speedup vs the original baseline: N=1 3.86x, N=2 2.45x (both real net wins); N=4 is a real regression once its own now-failing Fmax is honestly used, though N=4 was never the recommended configuration. N_SLOTS=2 remains the recommended default (DEC-0014 unaffected) with a thinner but still real Fmax margin. Cache hit-detection pipelining is flagged as concrete follow-up work if N_SLOTS>2 is ever needed with the cache active - not attempted this round. Logged: simulation/synthesis/timing/benchmark/decisions (DEC-0016)/ experiments (EXP-0016)/errors (ERR-0010)/development.log, ROADMAP.md updated. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
This commit is contained in:
@@ -152,3 +152,31 @@ PASS/FAIL: M4 3/3 PASS (166/446/728 -> 84/204/322 cycles, bit-exact).
|
||||
(see benchmark.log EXP-0015 for the full table).
|
||||
errors: none found during this implementation (clean first-pass
|
||||
correctness at every regression point).
|
||||
|
||||
[2026-09-05] EXP-0016 -- regression + real improvement measurement
|
||||
after adding shared activation_cache.v (DEC-0016)
|
||||
test: hardware/v2/sim/tb_memory_manager.v (M4, updated: memory_manager
|
||||
now wired to a real activation_cache instance + a real 2-port
|
||||
arbiter, N_SLOTS=1 scope), hardware/v2/sim/tb_dataflow_core.v (M7,
|
||||
updated: sim_word_mem array widened to N_SLOTS+1, X data poked once
|
||||
into the shared cache's own backing memory instead of duplicated
|
||||
per-slot), hardware/v2/sim/tb_neural_multiprocessor.v (M8,
|
||||
UNCHANGED -- black-box), hardware/v2/sim/tb_benchmark_suite.v (final
|
||||
campaign, UNCHANGED, re-run at N_SLOTS=1/2/4/8)
|
||||
simulator: Verilator 5.050 (--binary --timing)
|
||||
PASS/FAIL: M4 3/3 PASS (bit-exact, cycle counts higher than DEC-0015
|
||||
alone for this SPECIFIC single-instance test -- expected, no sharing
|
||||
benefit possible with only one memory_manager, only the cache's real
|
||||
arbitration overhead shows up here). M7 4/4 PASS. M8 4/4 PASS
|
||||
(unchanged testbench). Final campaign 24/24 PASS bit-exact,
|
||||
D-Stress cycles reduced a further 1.66-2.00x on top of DEC-0015's
|
||||
own reduction (see benchmark.log EXP-0016 for the full table) --
|
||||
the real sharing benefit this cache targets only manifests with
|
||||
multiple neurons genuinely sharing one x_base, which only the full
|
||||
campaign's dense-layer workloads (not M4/M7/M8's own small tests)
|
||||
exercise.
|
||||
errors: 2 real bugs found and fixed during implementation (errors.log
|
||||
ERR-0010): a target-bank/pending-bank race (same class as ERR-0006,
|
||||
a new instance in the activation-cache side of memory_manager.v),
|
||||
and a repeat of ERR-0009's N_SLOTS=1 zero-width replication bug
|
||||
(this time in activation_cache.v itself).
|
||||
|
||||
Reference in New Issue
Block a user