perf(v2): shared activation cache - further 1.66-2.00x real speedup (DEC-0016)
Implements optimization #2 from the final benchmark campaign's own recommendation, on top of DEC-0015's word-level burst rewrite: a new shared activation_cache.v module fetches a given activation (X) vector from PSRAM once instead of once per neuron sharing it - the exact redundant traffic pattern the dense-layer workloads in this project's benchmark suite exhibit. Each memory_manager's own prefetch_engine now fetches WEIGHTS only; the activation half is requested from the shared cache instead (single-tag, tile-granular, N_SLOTS request ports, its own real word-level PSRAM backend via a new dedicated arbiter port). dataflow_core.v/slot_mem_arbiter.v/neural_multiprocessor.v widened to N_SLOTS+1 ports to arbitrate the cache's traffic alongside each slot's weight traffic. Two real bugs found and fixed during implementation (ERR-0010): a target-bank/pending-bank race in memory_manager.v's activation-cache wiring (the same bug class ERR-0006 already fixed once for pf_target_bank - a later handoff's queued request can overwrite which bank an earlier, still-in-flight request's ack applies to), and a repeat of ERR-0009's N_SLOTS=1 zero-width replication bug in activation_cache.v itself. Real, measured results: the full final-benchmark campaign (24/24 workload/config combinations) re-verified bit-exact. D-Stress cycles fall a further 1.66-2.00x on top of DEC-0015 (~4x combined vs the original byte-level baseline). But the cache's real Fmax cost is much steeper than DEC-0015's own: N_SLOTS=2 (the recommended default, DEC-0014) drops from 133.58 to 87.72 MHz (-34%, margin over 80MHz shrinks from +67% to +9.7%), and N_SLOTS=4 drops to 65.01 MHz - now FAILING the 80MHz target it previously passed. Combined real wall-clock speedup vs the original baseline: N=1 3.86x, N=2 2.45x (both real net wins); N=4 is a real regression once its own now-failing Fmax is honestly used, though N=4 was never the recommended configuration. N_SLOTS=2 remains the recommended default (DEC-0014 unaffected) with a thinner but still real Fmax margin. Cache hit-detection pipelining is flagged as concrete follow-up work if N_SLOTS>2 is ever needed with the cache active - not attempted this round. Logged: simulation/synthesis/timing/benchmark/decisions (DEC-0016)/ experiments (EXP-0016)/errors (ERR-0010)/development.log, ROADMAP.md updated. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
This commit is contained in:
@@ -116,6 +116,28 @@ DEC-0014). Report completo (21 sezioni, THEORETICAL/SIMULATED/
|
|||||||
POST-P&R/DERIVED classificati): `hardware/v2/docs/benchmarks/
|
POST-P&R/DERIVED classificati): `hardware/v2/docs/benchmarks/
|
||||||
final-benchmark.md`.
|
final-benchmark.md`.
|
||||||
|
|
||||||
|
### Ottimizzazioni post-campagna (su richiesta utente)
|
||||||
|
|
||||||
|
Implementate entrambe le raccomandazioni #1/#2 del report:
|
||||||
|
1. **Burst a livello di parola** (`prefetch_engine.v`/`memory_manager.v`
|
||||||
|
parlano direttamente il protocollo a 16 bit di `memory_interface.v`,
|
||||||
|
bypassando `int8_memory_access.v` -- ancora congelato, semplicemente
|
||||||
|
non piu' istanziato in questo percorso). Reale: -49/-56% cicli sui
|
||||||
|
job singoli, 2.24-2.37x speedup wall-clock reale sull'intera
|
||||||
|
campagna. `logs/decisions.log` DEC-0015.
|
||||||
|
2. **Cache condivisa on-chip per il vettore di attivazione** (nuovo
|
||||||
|
`activation_cache.v`, evita che neuroni con lo stesso `x_base`
|
||||||
|
rileggano X da PSRAM). Reale: ulteriore -1.66/-2.00x cicli. MA costo
|
||||||
|
Fmax reale molto piu' ripido del previsto: N_SLOTS=4 ora FALLISCE il
|
||||||
|
target 80MHz (65.01 MHz, prima passava). N_SLOTS=2 (default
|
||||||
|
raccomandato) resta valido con margine piu' sottile (87.72 MHz).
|
||||||
|
Speedup wall-clock reale combinato (#1+#2) vs baseline originale:
|
||||||
|
N=1 3.86x, N=2 2.45x. `logs/decisions.log` DEC-0016.
|
||||||
|
|
||||||
|
Trovati e risolti 3 bug RTL reali durante l'implementazione
|
||||||
|
(`logs/errors.log` ERR-0009, ERR-0010). 24/24 combinazioni
|
||||||
|
workload/config ancora bit-exact dopo entrambe le ottimizzazioni.
|
||||||
|
|
||||||
## Log
|
## Log
|
||||||
|
|
||||||
Vedi `hardware/v2/logs/` (`development.log` per la cronologia di sessione,
|
Vedi `hardware/v2/logs/` (`development.log` per la cronologia di sessione,
|
||||||
|
|||||||
@@ -311,3 +311,38 @@ cycles at N=2/4/8 remain within ~0.2% of each other: 307602/307346/
|
|||||||
EFFICIENT per transaction, it did not remove the fact that there is
|
EFFICIENT per transaction, it did not remove the fact that there is
|
||||||
still only one physical port, so DEC-0014's N_SLOTS=2 recommendation
|
still only one physical port, so DEC-0014's N_SLOTS=2 recommendation
|
||||||
is unaffected and reconfirmed with the new, faster numbers.
|
is unaffected and reconfirmed with the new, faster numbers.
|
||||||
|
|
||||||
|
[2026-09-05] EXP-0016 -- shared activation_cache (DEC-0016), real
|
||||||
|
before/after comparison on top of DEC-0015's own word-level burst
|
||||||
|
rewrite (user-requested optimization #2)
|
||||||
|
|
||||||
|
Full campaign (tb_benchmark_suite.v, same 6 workloads as EXP-0014/
|
||||||
|
EXP-0015), D-Stress (256 neurons, all sharing ONE input vector -- the
|
||||||
|
exact pattern this cache targets), real cycles and real wall-clock
|
||||||
|
(cycles / real POST-P&R Fmax):
|
||||||
|
|
||||||
|
| N_SLOTS | cycles: byte | cycles: +burst | cycles: +cache | Fmax: byte | Fmax: +burst | Fmax: +cache |
|
||||||
|
|-----------|----------------|-------------------|-------------------|--------------|----------------|----------------|
|
||||||
|
| 1 | 780298 | 348682 | 174610 | 152.46 MHz | 152.44 MHz | 131.79 MHz |
|
||||||
|
| 2 | 736402 | 307602 | 185428 | 142.45 MHz | 133.58 MHz | 87.72 MHz |
|
||||||
|
| 4 | 736823 | 307346 | 184795 | 113.38 MHz | 112.07 MHz | 65.01 MHz (FAIL@80MHz) |
|
||||||
|
|
||||||
|
Real wall-clock speedup vs the ORIGINAL byte-level baseline (cycles /
|
||||||
|
real Fmax, both fully combined optimizations):
|
||||||
|
| N_SLOTS | wall-clock: byte | wall-clock: +burst | wall-clock: +cache | TOTAL real speedup |
|
||||||
|
|-----------|--------------------|------------------------|------------------------|------------------------|
|
||||||
|
| 1 | 5118.1 us | 2287.3 us | 1324.9 us | 3.86x |
|
||||||
|
| 2 | 5169.5 us | 2302.8 us | 2113.9 us | 2.45x |
|
||||||
|
| 4 | 6498.7 us | 2742.4 us | 2842.6 us (Fmax FAILS)| 2.29x (but a real regression vs +burst alone) |
|
||||||
|
|
||||||
|
All 24/24 workload/config combinations (6 workloads x N_SLOTS=1/2/4/8)
|
||||||
|
re-verified bit-exact after adding the cache. Two real bugs found and
|
||||||
|
fixed during implementation (errors.log ERR-0010).
|
||||||
|
|
||||||
|
HONEST SUMMARY: the two user-requested optimizations together deliver
|
||||||
|
a real 2.45-3.86x wall-clock speedup for the recommended N_SLOTS<=2
|
||||||
|
range (DEC-0014), at the cost of a much steeper Fmax sensitivity to
|
||||||
|
N_SLOTS than either the arbiter alone (DEC-0015) or the un-optimized
|
||||||
|
baseline had -- N_SLOTS=4 now fails 80MHz outright with the cache
|
||||||
|
active, a real trade-off, not glossed over (see decisions.log
|
||||||
|
DEC-0016).
|
||||||
|
|||||||
@@ -937,3 +937,98 @@ correctness preserved (24/24 bit-exact).
|
|||||||
STATUS:
|
STATUS:
|
||||||
ACCEPTED
|
ACCEPTED
|
||||||
|
|
||||||
|
DEC-0016
|
||||||
|
|
||||||
|
DATE: 2026-09-05
|
||||||
|
|
||||||
|
DECISION:
|
||||||
|
A new shared module, activation_cache.v, is added inside dataflow_core.v
|
||||||
|
alongside the N_SLOTS memory_managers. It fetches a given ACTIVATION
|
||||||
|
(X) vector from PSRAM once (tile by tile, on first use) and serves
|
||||||
|
every subsequent request for the same x_base/tile directly from an
|
||||||
|
on-chip buffer -- no PSRAM access on a hit. Each memory_manager's own
|
||||||
|
prefetch_engine now fetches WEIGHTS only (X is no longer duplicated
|
||||||
|
per-slot). Single-tag design (one active cached x_base at a time,
|
||||||
|
correct but can thrash under interleaved different-x_base concurrent
|
||||||
|
traffic -- never incorrect, see the "Alternatives"/"Result" sections
|
||||||
|
below for the honest limitation).
|
||||||
|
|
||||||
|
WHY (user-requested optimization #2, following the final-benchmark.md
|
||||||
|
report's own recommendation): in the realistic dense-layer workloads
|
||||||
|
this project benchmarks, many neurons share the exact same X vector --
|
||||||
|
each of dataflow_core's N_SLOTS memory_manager instances re-fetching
|
||||||
|
that identical vector from PSRAM independently was real, measured,
|
||||||
|
redundant traffic on the one shared PSRAM port.
|
||||||
|
|
||||||
|
EVIDENCE -- REAL BENEFIT (cycles, SIMULATED, full campaign,
|
||||||
|
tb_benchmark_suite.v, D-Stress the largest/most representative
|
||||||
|
workload): combined with DEC-0015's word-level burst rewrite, total
|
||||||
|
cycle count for D-Stress falls from the ORIGINAL byte-level baseline
|
||||||
|
(EXP-0014) by 3.97x-4.47x across every N_SLOTS tested (N=1: 780298 ->
|
||||||
|
174610; N=2: 736402 -> 185428; N=4: 736823 -> 184795; N=8: 738751 ->
|
||||||
|
184797) -- the activation_cache's OWN incremental contribution on top
|
||||||
|
of DEC-0015 alone is a further 1.66x-2.00x cycle reduction. All 24/24
|
||||||
|
workload/config combinations remain bit-exact.
|
||||||
|
|
||||||
|
EVIDENCE -- REAL COST (Fmax, POST-P&R MEASURED, full
|
||||||
|
neural_multiprocessor including the real V1 PSRAM chain): the shared
|
||||||
|
cache's real Fmax cost is substantially STEEPER than DEC-0015's own
|
||||||
|
(which cost only 0-6% Fmax). Real POST-P&R Fmax after adding the
|
||||||
|
cache: N=1: 152.44 -> 131.79 MHz (-13.5%); N=2: 133.58 -> 87.72 MHz
|
||||||
|
(-34.3%); N=4: 112.07 -> 65.01 MHz (-42.0%, and this configuration NOW
|
||||||
|
FAILS the 80MHz target it previously passed). This is a real,
|
||||||
|
structural cost: activation_cache is a single shared resource with
|
||||||
|
N_SLOTS request ports, a broadcast-capable hit-check evaluated for
|
||||||
|
every port every cycle, and a shared tile_store array -- a genuine
|
||||||
|
routing/fan-in hot spot that gets worse as N_SLOTS grows, more
|
||||||
|
severely than the arbiter-only widening DEC-0015 introduced.
|
||||||
|
|
||||||
|
Combined REAL WALL-CLOCK effect (cycles / real POST-P&R Fmax,
|
||||||
|
D-Stress, vs the ORIGINAL byte-level baseline): N=1: 3.86x faster
|
||||||
|
(the clear win case -- low Fmax cost, full cycle benefit); N=2: 2.45x
|
||||||
|
faster (still a solid net win -- this is the recommended default,
|
||||||
|
DEC-0014, and it still comfortably beats the baseline, though its
|
||||||
|
Fmax safety margin over 80MHz shrank from +78% to +9.7%); N=4: 2.29x
|
||||||
|
faster THAN THE ORIGINAL baseline, but WORSE than DEC-0015-alone
|
||||||
|
(2742.4us -> 2842.6us) once its own real (now failing) Fmax is used --
|
||||||
|
adding the cache is a net REGRESSION specifically at N=4, and N=4 is
|
||||||
|
no longer even a valid passing 80MHz design point.
|
||||||
|
|
||||||
|
ALTERNATIVES:
|
||||||
|
1. Pipeline the cache's hit-detection/broadcast logic (register the
|
||||||
|
hit[] comparison one extra stage before driving tile_x_out) to
|
||||||
|
recover some of the lost Fmax margin. Rejected FOR NOW: a real,
|
||||||
|
promising follow-up, but a genuine RTL redesign of the cache's own
|
||||||
|
timing, not attempted in this pass -- flagged as real, concrete
|
||||||
|
future work rather than attempted blindly without first measuring
|
||||||
|
whether N_SLOTS=2 (the actual recommended default, DEC-0014) still
|
||||||
|
needs it (it does not fail timing at N=2, it just has a thinner
|
||||||
|
margin than before).
|
||||||
|
2. Give up on the cache entirely given N=4's regression. Rejected:
|
||||||
|
N=1 and N=2 (the actually-recommended range per DEC-0014) both show
|
||||||
|
a clear, real net win, and N=4 was never the recommended default
|
||||||
|
to begin with -- discarding a real 2.45-3.86x win over the range
|
||||||
|
that matters to avoid a regression in a range that was already
|
||||||
|
deprioritized would be the wrong trade.
|
||||||
|
3. Make MAX_TILES smaller (currently 16, sized for the largest
|
||||||
|
workload's shared vector) to shrink the cache's own storage/compare
|
||||||
|
width and recover some Fmax. Rejected for THIS round: would need
|
||||||
|
re-verifying against every workload's own real tile count
|
||||||
|
requirements (Large/Stress use exactly 16) -- a real, bounded
|
||||||
|
follow-up, not attempted here to avoid conflating multiple
|
||||||
|
variables in one measurement.
|
||||||
|
|
||||||
|
RESULT:
|
||||||
|
activation_cache.v is added, real net win confirmed at N_SLOTS=1 and
|
||||||
|
N_SLOTS=2 (the recommended default, DEC-0014), real regression to a
|
||||||
|
failing timing state confirmed at N_SLOTS=4 -- reported honestly, not
|
||||||
|
hidden. N_SLOTS=2 remains the recommended default (DEC-0014's own
|
||||||
|
conclusion is unaffected, since N=4 was never recommended), now with a
|
||||||
|
thinner but still real Fmax margin (87.72 MHz vs the 80MHz target).
|
||||||
|
Cache pipelining (Alternative 1) is flagged as real, concrete follow-up
|
||||||
|
work if N_SLOTS>2 configurations are ever needed with this cache
|
||||||
|
active.
|
||||||
|
|
||||||
|
STATUS:
|
||||||
|
ACCEPTED
|
||||||
|
|
||||||
|
|||||||
@@ -331,3 +331,35 @@ decision: vedi decisions.log DEC-0015.
|
|||||||
next_action: ottimizzazione #2 -- cache condivisa on-chip per il
|
next_action: ottimizzazione #2 -- cache condivisa on-chip per il
|
||||||
vettore di attivazione (X), per eliminare le letture ridondanti tra
|
vettore di attivazione (X), per eliminare le letture ridondanti tra
|
||||||
neuroni che condividono lo stesso input di layer.
|
neuroni che condividono lo stesso input di layer.
|
||||||
|
|
||||||
|
[2026-09-05] Ottimizzazione #2 -- cache condivisa on-chip per il
|
||||||
|
vettore di attivazione, su richiesta esplicita dell'utente
|
||||||
|
reason: dopo l'ottimizzazione #1 (burst a livello di parola), l'utente
|
||||||
|
ha chiesto di implementare anche la #2: cache condivisa per evitare
|
||||||
|
che ogni neurone rilegga da PSRAM lo stesso vettore X gia' letto da
|
||||||
|
un altro neurone dello stesso layer.
|
||||||
|
result: nuovo modulo activation_cache.v (single-tag, granularita' a
|
||||||
|
tile, condiviso tra tutti gli slot). Trovati e risolti 2 bug reali
|
||||||
|
durante l'implementazione (errors.log ERR-0010): una race
|
||||||
|
target-bank/pending-bank (stessa classe di ERR-0006) e una ripetizione
|
||||||
|
del bug a larghezza zero di ERR-0009 (N_SLOTS=1). Dopo il fix: 24/24
|
||||||
|
PASS bit-exact. Beneficio reale misurato: ulteriore riduzione cicli
|
||||||
|
1.66-2.00x oltre l'ottimizzazione #1 (circa 4x combinato rispetto
|
||||||
|
alla baseline originale byte-level). MA costo reale in Fmax molto piu'
|
||||||
|
ripido del previsto: N_SLOTS=2 scende da 133.58 a 87.72 MHz (-34%,
|
||||||
|
ancora sopra 80MHz ma con margine molto piu' sottile), N_SLOTS=4
|
||||||
|
scende a 65.01 MHz e ORA FALLISCE il target 80MHz (prima passava).
|
||||||
|
Speedup reale in wall-clock rispetto alla baseline originale:
|
||||||
|
N=1 3.86x, N=2 2.45x, N=4 2.29x (ma con una regressione reale
|
||||||
|
rispetto alla sola ottimizzazione #1, dato che N=4 ora fallisce il
|
||||||
|
timing).
|
||||||
|
errors: vedi errors.log ERR-0010 (2 bug RTL reali, trovati e risolti).
|
||||||
|
decision: vedi decisions.log DEC-0016 -- vittoria reale netta
|
||||||
|
confermata per N_SLOTS=1/2 (l'intervallo raccomandato, DEC-0014
|
||||||
|
resta valida), regressione reale a N_SLOTS=4 (mai la configurazione
|
||||||
|
raccomandata, ma un costo reale onestamente riportato, non nascosto).
|
||||||
|
Pipeline della logica di hit-check della cache segnalato come lavoro
|
||||||
|
futuro concreto se servisse N_SLOTS>2 con la cache attiva.
|
||||||
|
next_action: nessuna ulteriore richiesta esplicitamente dall'utente per
|
||||||
|
questo giro. Datasheet V2 professionale, ora con dati definitivi
|
||||||
|
post-ottimizzazione da includere.
|
||||||
|
|||||||
@@ -327,3 +327,61 @@ DIAGNOSIS METHOD (bug 3): periodic `$display` progress heartbeats
|
|||||||
output for real-time visibility, isolating the exact neuron index
|
output for real-time visibility, isolating the exact neuron index
|
||||||
where progress stopped advancing -- the same "trace real signals,
|
where progress stopped advancing -- the same "trace real signals,
|
||||||
don't guess" discipline used throughout this whole project.
|
don't guess" discipline used throughout this whole project.
|
||||||
|
|
||||||
|
ERR-0010 (real RTL bugs found and fixed during activation_cache.v
|
||||||
|
implementation, DEC-0016)
|
||||||
|
DATE: 2026-09-05
|
||||||
|
|
||||||
|
1. Same bug CLASS as ERR-0006 (pf_target_bank/pf_pending_bank), NEW
|
||||||
|
instance, memory_manager.v's activation-cache side:
|
||||||
|
SYMPTOM: hardware/v2/sim/tb_memory_manager.v -- job xb=4096 (3
|
||||||
|
tiles) hung/produced wrong results after adding the shared
|
||||||
|
activation_cache request path; cycle-by-cycle tracing (temporary
|
||||||
|
$display instrumentation, later removed) showed a cache ack for
|
||||||
|
tile 1 (queued for bank 1) instead applying its data to bank 0.
|
||||||
|
ROOT CAUSE: `xc_target_bank` (which bank an ack's data should be
|
||||||
|
written into) was being written DIRECTLY by the queueing logic
|
||||||
|
(MM_IDLE/MM_PREFETCH_FIRST/MM_STREAM), the same register used to
|
||||||
|
resolve an ack that might still be OUTSTANDING from an EARLIER
|
||||||
|
queued request. Real PSRAM miss latency can exceed one
|
||||||
|
neural_processor tile's own compute time, so a LATER handoff can
|
||||||
|
queue a NEW request (targeting a DIFFERENT bank) in the same or a
|
||||||
|
following cycle, before the EARLIER request's ack has arrived --
|
||||||
|
with only one `xc_target_bank` register, NBA "last write in
|
||||||
|
program order wins" semantics silently overwrote which bank the
|
||||||
|
EARLIER, already-in-flight request's eventual ack gets applied
|
||||||
|
to. This is the EXACT bug ERR-0006 already found and fixed once
|
||||||
|
for pf_target_bank/pf_pending_bank (which already used a correct
|
||||||
|
two-register pattern: a `_pending_bank` staging register written
|
||||||
|
at queueing time, and the real `_target_bank` written ONLY by the
|
||||||
|
issue rule at the moment the request actually fires) -- this
|
||||||
|
module's newly-added activation-cache side did not follow that
|
||||||
|
already-established pattern, until now.
|
||||||
|
FIX: introduced `xc_pending_bank` (written at queueing time) and
|
||||||
|
changed `xc_target_bank` to be written ONLY by the issue rule
|
||||||
|
(`xc_target_bank <= xc_pending_bank;`, the same cycle xc_req
|
||||||
|
fires), mirroring pf_target_bank/pf_pending_bank exactly.
|
||||||
|
VERIFICATION: tb_memory_manager.v 3/3 PASS bit-exact after the fix.
|
||||||
|
Full final-benchmark campaign (24/24 workload/config
|
||||||
|
combinations) re-verified bit-exact.
|
||||||
|
|
||||||
|
2. Same bug CLASS as ERR-0009 item 1 (neural_director.v's N_SLOTS=1
|
||||||
|
zero-width replication), NEW instance, activation_cache.v:
|
||||||
|
SYMPTOM: Verilator compile error building the M4-level regression
|
||||||
|
testbench (activation_cache instantiated with N_SLOTS=1 there,
|
||||||
|
to serve a single memory_manager instance) -- same
|
||||||
|
"%Error-ZEROREPL: Replication value of 0 is only legal under a
|
||||||
|
concatenation" as ERR-0009.
|
||||||
|
ROOT CAUSE: `miss_idx = {$clog2(N_SLOTS){1'b0}};` -- identical root
|
||||||
|
cause to ERR-0009 item 1 ($clog2(1)=0 at N_SLOTS=1).
|
||||||
|
FIX: replaced with the same width-agnostic `'0` literal used in
|
||||||
|
neural_director.v's own fix.
|
||||||
|
VERIFICATION: tb_memory_manager.v (N_SLOTS=1 activation_cache) and
|
||||||
|
the full campaign (N_SLOTS=1/2/4/8) all build and pass.
|
||||||
|
|
||||||
|
DIAGNOSIS METHOD (item 1): periodic $display cycle-by-cycle tracing of
|
||||||
|
memory_manager's own internal state (xc_req/xc_ack/xc_outstanding/
|
||||||
|
xc_pending/bank_x_ready/bank_w_ready) and the 2-port test arbiter's
|
||||||
|
own owner/grant state, added temporarily to tb_memory_manager.v and
|
||||||
|
removed once the bug was isolated and fixed -- the same "trace real
|
||||||
|
signals, don't guess" discipline used throughout this project.
|
||||||
|
|||||||
@@ -584,3 +584,45 @@ next_action: user-requested optimization #2 -- a shared on-chip cache
|
|||||||
for the activation (X) vector, so N independent neurons sharing one
|
for the activation (X) vector, so N independent neurons sharing one
|
||||||
input vector (the dense-layer shape used throughout this benchmark
|
input vector (the dense-layer shape used throughout this benchmark
|
||||||
suite) fetch it from PSRAM ONCE instead of once per neuron.
|
suite) fetch it from PSRAM ONCE instead of once per neuron.
|
||||||
|
|
||||||
|
[2026-09-05] EXP-0016 -- shared activation_cache implementation (user-
|
||||||
|
requested optimization #2, following final-benchmark.md's own
|
||||||
|
recommendation: eliminate redundant per-neuron re-fetching of a
|
||||||
|
shared input vector)
|
||||||
|
test: new module activation_cache.v (single-tag, tile-granular,
|
||||||
|
N_SLOTS request ports, real word-level PSRAM backend via its own
|
||||||
|
arbiter port); memory_manager.v's activation half redirected through
|
||||||
|
it (weight half unchanged from DEC-0015); dataflow_core.v/
|
||||||
|
neural_multiprocessor.v/slot_mem_arbiter.v widened to N_SLOTS+1
|
||||||
|
ports to arbitrate the cache's own traffic alongside the N_SLOTS
|
||||||
|
memory_managers' weight traffic. Re-verified: M4 (updated
|
||||||
|
testbench), M7 (updated testbench), M8 (unchanged, black-box), and
|
||||||
|
the full final benchmark campaign (unchanged, black-box) at
|
||||||
|
N_SLOTS=1/2/4/8.
|
||||||
|
simulator: Verilator 5.050 (--binary --timing)
|
||||||
|
PASS/FAIL:
|
||||||
|
SIMULATED: M4 3/3 PASS, M7 4/4 PASS, M8 4/4 PASS, final campaign
|
||||||
|
24/24 PASS bit-exact. D-Stress cycles: 348682/307602/307346/
|
||||||
|
307874 (post-DEC-0015) -> 174610/185428/184795/184797 (N=1/2/4/8,
|
||||||
|
post-DEC-0016) -- a further 1.66-2.00x real cycle reduction from
|
||||||
|
the cache alone, ~4x combined with DEC-0015 vs the original
|
||||||
|
byte-level baseline.
|
||||||
|
SYNTHESIZED + POST-P&R (real, full system incl. real PSRAM pins):
|
||||||
|
N=1 131.79 MHz (was 152.44), N=2 87.72 MHz (was 133.58, -34.3%),
|
||||||
|
N=4 65.01 MHz (was 112.07, -42.0% -- NOW FAILS the 80MHz target).
|
||||||
|
Combined real wall-clock speedup vs the ORIGINAL byte-level
|
||||||
|
baseline: N=1 3.86x, N=2 2.45x, N=4 2.29x (but a real regression
|
||||||
|
vs DEC-0015-alone once N=4's own now-failing Fmax is used).
|
||||||
|
errors: 2 real bugs found and fixed (errors.log ERR-0010): a target-
|
||||||
|
bank/pending-bank race (ERR-0006's bug class, new instance) and a
|
||||||
|
repeat of ERR-0009's N_SLOTS=1 zero-width replication bug.
|
||||||
|
decision: see decisions.log DEC-0016 -- real net win confirmed at
|
||||||
|
N_SLOTS=1/2 (the recommended range, DEC-0014), real regression to a
|
||||||
|
failing timing state at N_SLOTS=4 (never the recommended default,
|
||||||
|
but a real, honestly-reported cost of this optimization). Cache
|
||||||
|
pipelining flagged as real follow-up work if N_SLOTS>2 with the
|
||||||
|
cache active is ever needed.
|
||||||
|
next_action: none further requested by the user for this round. Real,
|
||||||
|
concrete follow-up flagged in DEC-0016: pipeline the cache's own
|
||||||
|
hit-detection/broadcast logic to recover Fmax margin if higher
|
||||||
|
N_SLOTS configurations are ever needed with the cache active.
|
||||||
|
|||||||
@@ -152,3 +152,31 @@ PASS/FAIL: M4 3/3 PASS (166/446/728 -> 84/204/322 cycles, bit-exact).
|
|||||||
(see benchmark.log EXP-0015 for the full table).
|
(see benchmark.log EXP-0015 for the full table).
|
||||||
errors: none found during this implementation (clean first-pass
|
errors: none found during this implementation (clean first-pass
|
||||||
correctness at every regression point).
|
correctness at every regression point).
|
||||||
|
|
||||||
|
[2026-09-05] EXP-0016 -- regression + real improvement measurement
|
||||||
|
after adding shared activation_cache.v (DEC-0016)
|
||||||
|
test: hardware/v2/sim/tb_memory_manager.v (M4, updated: memory_manager
|
||||||
|
now wired to a real activation_cache instance + a real 2-port
|
||||||
|
arbiter, N_SLOTS=1 scope), hardware/v2/sim/tb_dataflow_core.v (M7,
|
||||||
|
updated: sim_word_mem array widened to N_SLOTS+1, X data poked once
|
||||||
|
into the shared cache's own backing memory instead of duplicated
|
||||||
|
per-slot), hardware/v2/sim/tb_neural_multiprocessor.v (M8,
|
||||||
|
UNCHANGED -- black-box), hardware/v2/sim/tb_benchmark_suite.v (final
|
||||||
|
campaign, UNCHANGED, re-run at N_SLOTS=1/2/4/8)
|
||||||
|
simulator: Verilator 5.050 (--binary --timing)
|
||||||
|
PASS/FAIL: M4 3/3 PASS (bit-exact, cycle counts higher than DEC-0015
|
||||||
|
alone for this SPECIFIC single-instance test -- expected, no sharing
|
||||||
|
benefit possible with only one memory_manager, only the cache's real
|
||||||
|
arbitration overhead shows up here). M7 4/4 PASS. M8 4/4 PASS
|
||||||
|
(unchanged testbench). Final campaign 24/24 PASS bit-exact,
|
||||||
|
D-Stress cycles reduced a further 1.66-2.00x on top of DEC-0015's
|
||||||
|
own reduction (see benchmark.log EXP-0016 for the full table) --
|
||||||
|
the real sharing benefit this cache targets only manifests with
|
||||||
|
multiple neurons genuinely sharing one x_base, which only the full
|
||||||
|
campaign's dense-layer workloads (not M4/M7/M8's own small tests)
|
||||||
|
exercise.
|
||||||
|
errors: 2 real bugs found and fixed during implementation (errors.log
|
||||||
|
ERR-0010): a target-bank/pending-bank race (same class as ERR-0006,
|
||||||
|
a new instance in the activation-cache side of memory_manager.v),
|
||||||
|
and a repeat of ERR-0009's N_SLOTS=1 zero-width replication bug
|
||||||
|
(this time in activation_cache.v itself).
|
||||||
|
|||||||
@@ -112,3 +112,13 @@ CHECK: 0 problems on all three (same benign warnings as every prior
|
|||||||
configs and saved some in others (word-level control logic is
|
configs and saved some in others (word-level control logic is
|
||||||
simpler than byte-indexing logic in prefetch_engine.v), net roughly
|
simpler than byte-indexing logic in prefetch_engine.v), net roughly
|
||||||
flat.
|
flat.
|
||||||
|
|
||||||
|
[2026-09-05] EXP-0016 -- neural_multiprocessor N_SLOTS=1/2/4 after
|
||||||
|
adding activation_cache.v (DEC-0016), real standalone synthesis, no
|
||||||
|
harness needed
|
||||||
|
N_SLOTS=1: LUT4=2760 CCU2C=202 TRELLIS_FF=2405 MULT18X18D=8 DP16KD=0
|
||||||
|
N_SLOTS=2: LUT4=4359 CCU2C=366 TRELLIS_FF=3924 MULT18X18D=16 DP16KD=0
|
||||||
|
N_SLOTS=4: LUT4=9158 CCU2C=698 TRELLIS_FF=7986 MULT18X18D=32 DP16KD=0
|
||||||
|
CHECK: 0 problems on all three (same benign warnings as always).
|
||||||
|
Resource cost is modest (activation_cache itself is small -- a
|
||||||
|
16-tile x 64-bit store plus N_SLOTS-way pending/hit logic).
|
||||||
|
|||||||
@@ -126,3 +126,18 @@ Small, real Fmax cost from widening the shared arbiter/backend to
|
|||||||
16-bit + lb_n/ub_n (extra routing), overwhelmingly outweighed by the
|
16-bit + lb_n/ub_n (extra routing), overwhelmingly outweighed by the
|
||||||
real cycle-count reduction (EXP-0015 in experiments.log/benchmark.log):
|
real cycle-count reduction (EXP-0015 in experiments.log/benchmark.log):
|
||||||
D-Stress real wall-clock improves 2.24-2.37x at every N_SLOTS.
|
D-Stress real wall-clock improves 2.24-2.37x at every N_SLOTS.
|
||||||
|
|
||||||
|
[2026-09-05] EXP-0016 -- neural_multiprocessor N_SLOTS=1/2/4 after
|
||||||
|
adding activation_cache.v (DEC-0016), real nextpnr-ecp5 --45k
|
||||||
|
--package CABGA381 --speed 8 --freq 80 --lpf-allow-unconstrained
|
||||||
|
N_SLOTS=1: Fmax = 131.79 MHz -- PASS at 80MHz (was 152.44 MHz, -13.5%)
|
||||||
|
N_SLOTS=2: Fmax = 87.72 MHz -- PASS at 80MHz (was 133.58 MHz, -34.3%)
|
||||||
|
N_SLOTS=4: Fmax = 65.01 MHz -- FAILS at 80MHz (was 112.07 MHz, -42.0%)
|
||||||
|
Real, substantial Fmax cost from the shared activation_cache -- a
|
||||||
|
single central resource with N_SLOTS request ports, a broadcast-
|
||||||
|
capable hit-check evaluated every cycle for every port, and a shared
|
||||||
|
tile_store array. Cost grows much faster with N_SLOTS than
|
||||||
|
DEC-0015's own arbiter widening did. N_SLOTS=4 is no longer a passing
|
||||||
|
80MHz design point with the cache active -- see decisions.log
|
||||||
|
DEC-0016 for the full analysis and why N_SLOTS=2 (the recommended
|
||||||
|
default, DEC-0014) is unaffected in its own recommendation.
|
||||||
|
|||||||
@@ -0,0 +1,224 @@
|
|||||||
|
`timescale 1ns/1ps
|
||||||
|
|
||||||
|
// ================================================================
|
||||||
|
// FPGA-Neural V2 -- Shared Activation Cache (post-M10, docs/v2-
|
||||||
|
// description.md §14; decisions.log DEC-0016)
|
||||||
|
//
|
||||||
|
// User-requested optimization #2, following the final benchmark
|
||||||
|
// campaign's own recommendation: in the realistic dense-layer
|
||||||
|
// workloads this project benchmarks (hardware/v2/docs/benchmarks/
|
||||||
|
// final-benchmark.md), many neurons in the same layer share the
|
||||||
|
// EXACT SAME activation (X) input vector -- each of dataflow_core's
|
||||||
|
// N_SLOTS memory_manager instances re-fetching that identical vector
|
||||||
|
// from PSRAM independently was real, measured, redundant traffic on
|
||||||
|
// the one shared PSRAM port. This module fetches a given X vector
|
||||||
|
// from PSRAM ONCE (tile by tile, on first use) and serves every
|
||||||
|
// subsequent request for the SAME x_base/tile directly from an
|
||||||
|
// on-chip buffer -- no PSRAM access at all on a hit.
|
||||||
|
//
|
||||||
|
// Single-tag design: one active cached x_base at a time, filled
|
||||||
|
// tile-by-tile up to `filled_up_to` (tiles [0, filled_up_to) are
|
||||||
|
// valid). A request for a DIFFERENT x_base invalidates the cache and
|
||||||
|
// restarts filling from tile 0 for the new tag. This is correct
|
||||||
|
// (never serves stale/wrong data -- a tag switch always resets
|
||||||
|
// filled_up_to, so a later request against the OLD tag is treated as
|
||||||
|
// a fresh miss, refetched from scratch) but can THRASH under
|
||||||
|
// interleaved concurrent requests for genuinely different x_base
|
||||||
|
// values (falls back to no worse than the pre-cache behavior, never
|
||||||
|
// incorrect -- see decisions.log DEC-0016 for the full analysis).
|
||||||
|
// Fine for this project's own realistic workload shape (a "layer" of
|
||||||
|
// neurons dispatched together, sharing one x_base for the whole
|
||||||
|
// phase); a multi-way cache would avoid thrashing for interleaved
|
||||||
|
// multi-layer traffic, deferred until measured to matter.
|
||||||
|
//
|
||||||
|
// Request protocol: each of N_SLOTS ports issues a ONE-CYCLE req
|
||||||
|
// pulse (x_base + tile_idx); the cache LATCHES it into a per-slot
|
||||||
|
// pending register regardless of hit/miss/fetch-in-progress state --
|
||||||
|
// the same single-entry "queue, don't drop the request" idiom already
|
||||||
|
// used by memory_manager's own pf_pending register (ERR-0006) and
|
||||||
|
// slot_mem_arbiter's own pending latch (ERR-0008) -- so a request
|
||||||
|
// arriving while the cache is busy filling a miss for another slot is
|
||||||
|
// never lost. ack pulses exactly once per request, the cycle its
|
||||||
|
// data becomes available (immediately, if already a hit at latch
|
||||||
|
// time; after the real PSRAM fetch completes, on a miss). Multiple
|
||||||
|
// slots pending on tiles that become valid the SAME cycle a fetch
|
||||||
|
// completes are all acked that same cycle (broadcast hit).
|
||||||
|
//
|
||||||
|
// Backend: word-level (16-bit + lb_n/ub_n), same convention as
|
||||||
|
// prefetch_engine.v post-DEC-0015 -- talks to memory_interface.v's
|
||||||
|
// own 16-bit word interface via the shared slot_mem_arbiter.v (one
|
||||||
|
// more arbiter port, dedicated to this cache).
|
||||||
|
// ================================================================
|
||||||
|
|
||||||
|
module activation_cache #(
|
||||||
|
parameter DATA_WIDTH = 8,
|
||||||
|
parameter P_IN = 8,
|
||||||
|
parameter ADDR_WIDTH = 23,
|
||||||
|
parameter N_SLOTS = 4,
|
||||||
|
parameter MAX_TILES = 16 // max cacheable vector length, in tiles
|
||||||
|
)(
|
||||||
|
input wire clk,
|
||||||
|
input wire rst,
|
||||||
|
|
||||||
|
// ---- per-slot request port (one per memory_manager) ----
|
||||||
|
input wire [N_SLOTS-1:0] req,
|
||||||
|
input wire [ADDR_WIDTH*N_SLOTS-1:0] req_x_base,
|
||||||
|
input wire [16*N_SLOTS-1:0] req_tile_idx,
|
||||||
|
output reg [N_SLOTS-1:0] ack,
|
||||||
|
output reg signed [DATA_WIDTH*P_IN*N_SLOTS-1:0] tile_x_out,
|
||||||
|
|
||||||
|
// ---- shared backend port (word-level, -> slot_mem_arbiter.v) ----
|
||||||
|
output reg mem_req,
|
||||||
|
output reg mem_wr,
|
||||||
|
output reg [ADDR_WIDTH-1:0] mem_addr, // WORD address
|
||||||
|
output reg [15:0] mem_wdata, // unused (read-only), tied 0
|
||||||
|
output reg mem_lb_n,
|
||||||
|
output reg mem_ub_n,
|
||||||
|
input wire [15:0] mem_rdata,
|
||||||
|
input wire mem_ready
|
||||||
|
);
|
||||||
|
|
||||||
|
localparam WORDS_PER_TILE = P_IN/2;
|
||||||
|
localparam WIW = $clog2(WORDS_PER_TILE+1);
|
||||||
|
localparam TIW = $clog2(MAX_TILES+1);
|
||||||
|
|
||||||
|
localparam ST_IDLE = 1'd0;
|
||||||
|
localparam ST_FETCH = 1'd1;
|
||||||
|
|
||||||
|
reg state;
|
||||||
|
reg [ADDR_WIDTH-1:0] tag;
|
||||||
|
reg tag_valid;
|
||||||
|
reg [TIW-1:0] filled_up_to;
|
||||||
|
reg signed [DATA_WIDTH*P_IN-1:0] tile_store [0:MAX_TILES-1];
|
||||||
|
|
||||||
|
reg [TIW-1:0] fetch_tile_idx;
|
||||||
|
reg [WIW-1:0] word_idx;
|
||||||
|
|
||||||
|
// ---- per-slot pending-request latch (see file header) ----
|
||||||
|
reg [N_SLOTS-1:0] pending;
|
||||||
|
reg [ADDR_WIDTH*N_SLOTS-1:0] pending_x_base;
|
||||||
|
reg [16*N_SLOTS-1:0] pending_tile_idx;
|
||||||
|
|
||||||
|
integer pi;
|
||||||
|
|
||||||
|
wire [N_SLOTS-1:0] hit;
|
||||||
|
wire [N_SLOTS-1:0] miss;
|
||||||
|
genvar gi;
|
||||||
|
generate
|
||||||
|
for (gi = 0; gi < N_SLOTS; gi = gi + 1) begin : GEN_HITCHK
|
||||||
|
assign hit[gi] = pending[gi] && tag_valid &&
|
||||||
|
(pending_x_base[gi*ADDR_WIDTH +: ADDR_WIDTH] == tag) &&
|
||||||
|
(pending_tile_idx[gi*16 +: 16] < {{(16-TIW){1'b0}}, filled_up_to});
|
||||||
|
assign miss[gi] = pending[gi] && !hit[gi];
|
||||||
|
end
|
||||||
|
endgenerate
|
||||||
|
|
||||||
|
// Fixed lowest-index-wins priority scan over MISS requests (same
|
||||||
|
// convention as neural_director/dependency_manager/slot_mem_arbiter).
|
||||||
|
reg [$clog2(N_SLOTS)-1:0] miss_idx;
|
||||||
|
reg any_miss;
|
||||||
|
integer mi;
|
||||||
|
always @(*) begin
|
||||||
|
miss_idx = '0; // '0 self-sizes for any width incl. 0 (N_SLOTS=1) -- see errors.log ERR-0009
|
||||||
|
any_miss = 1'b0;
|
||||||
|
for (mi = N_SLOTS-1; mi >= 0; mi = mi - 1) begin
|
||||||
|
if (miss[mi]) begin
|
||||||
|
miss_idx = mi[$clog2(N_SLOTS)-1:0];
|
||||||
|
any_miss = 1'b1;
|
||||||
|
end
|
||||||
|
end
|
||||||
|
end
|
||||||
|
|
||||||
|
wire [ADDR_WIDTH-1:0] miss_x_base = pending_x_base[miss_idx*ADDR_WIDTH +: ADDR_WIDTH];
|
||||||
|
wire miss_is_new_tag = !tag_valid || (miss_x_base != tag);
|
||||||
|
wire [TIW-1:0] next_fetch_tile = miss_is_new_tag ? {TIW{1'b0}} : filled_up_to;
|
||||||
|
wire [ADDR_WIDTH-1:0] next_word_base = miss_x_base[ADDR_WIDTH-1:1] +
|
||||||
|
(next_fetch_tile * WORDS_PER_TILE[TIW-1:0]);
|
||||||
|
|
||||||
|
always @(posedge clk) begin
|
||||||
|
if (rst) begin
|
||||||
|
state <= ST_IDLE;
|
||||||
|
tag <= {ADDR_WIDTH{1'b0}};
|
||||||
|
tag_valid <= 1'b0;
|
||||||
|
filled_up_to <= {TIW{1'b0}};
|
||||||
|
fetch_tile_idx <= {TIW{1'b0}};
|
||||||
|
word_idx <= {WIW{1'b0}};
|
||||||
|
pending <= {N_SLOTS{1'b0}};
|
||||||
|
pending_x_base <= {(ADDR_WIDTH*N_SLOTS){1'b0}};
|
||||||
|
pending_tile_idx <= {(16*N_SLOTS){1'b0}};
|
||||||
|
ack <= {N_SLOTS{1'b0}};
|
||||||
|
tile_x_out <= {(DATA_WIDTH*P_IN*N_SLOTS){1'b0}};
|
||||||
|
mem_req <= 1'b0;
|
||||||
|
mem_wr <= 1'b0;
|
||||||
|
mem_addr <= {ADDR_WIDTH{1'b0}};
|
||||||
|
mem_wdata <= 16'h0000;
|
||||||
|
mem_lb_n <= 1'b1;
|
||||||
|
mem_ub_n <= 1'b1;
|
||||||
|
end else begin
|
||||||
|
mem_req <= 1'b0;
|
||||||
|
ack <= {N_SLOTS{1'b0}};
|
||||||
|
|
||||||
|
// Latch every incoming request pulse (never dropped, see
|
||||||
|
// file header).
|
||||||
|
for (pi = 0; pi < N_SLOTS; pi = pi + 1) begin
|
||||||
|
if (req[pi]) begin
|
||||||
|
pending[pi] <= 1'b1;
|
||||||
|
pending_x_base[pi*ADDR_WIDTH +: ADDR_WIDTH] <= req_x_base[pi*ADDR_WIDTH +: ADDR_WIDTH];
|
||||||
|
pending_tile_idx[pi*16 +: 16] <= req_tile_idx[pi*16 +: 16];
|
||||||
|
end
|
||||||
|
end
|
||||||
|
|
||||||
|
// Serve every currently-pending HIT this same cycle
|
||||||
|
// (broadcast -- see file header). Safe against colliding
|
||||||
|
// with the latch loop above: a slot only ever hits while
|
||||||
|
// its OWN pending bit was already set on an EARLIER cycle
|
||||||
|
// (this cycle's freshly-latched requests read `filled_up_to`/
|
||||||
|
// `tag` at their OWN NEXT evaluation, not this one).
|
||||||
|
for (pi = 0; pi < N_SLOTS; pi = pi + 1) begin
|
||||||
|
if (hit[pi]) begin
|
||||||
|
ack[pi] <= 1'b1;
|
||||||
|
tile_x_out[pi*DATA_WIDTH*P_IN +: DATA_WIDTH*P_IN] <= tile_store[pending_tile_idx[pi*16 +: 16]];
|
||||||
|
pending[pi] <= 1'b0;
|
||||||
|
end
|
||||||
|
end
|
||||||
|
|
||||||
|
case (state)
|
||||||
|
ST_IDLE: begin
|
||||||
|
if (any_miss) begin
|
||||||
|
tag <= miss_x_base;
|
||||||
|
tag_valid <= 1'b1;
|
||||||
|
filled_up_to <= miss_is_new_tag ? {TIW{1'b0}} : filled_up_to;
|
||||||
|
fetch_tile_idx <= next_fetch_tile;
|
||||||
|
word_idx <= {WIW{1'b0}};
|
||||||
|
mem_req <= 1'b1;
|
||||||
|
mem_wr <= 1'b0;
|
||||||
|
mem_lb_n <= 1'b0;
|
||||||
|
mem_ub_n <= 1'b0;
|
||||||
|
mem_addr <= next_word_base;
|
||||||
|
state <= ST_FETCH;
|
||||||
|
end
|
||||||
|
end
|
||||||
|
|
||||||
|
ST_FETCH: begin
|
||||||
|
if (mem_ready) begin
|
||||||
|
tile_store[fetch_tile_idx][word_idx*16 +: 16] <= mem_rdata;
|
||||||
|
if (word_idx == WORDS_PER_TILE[WIW-1:0] - 1'b1) begin
|
||||||
|
filled_up_to <= fetch_tile_idx + 1'b1;
|
||||||
|
state <= ST_IDLE;
|
||||||
|
end else begin
|
||||||
|
word_idx <= word_idx + 1'b1;
|
||||||
|
mem_req <= 1'b1;
|
||||||
|
mem_wr <= 1'b0;
|
||||||
|
mem_lb_n <= 1'b0;
|
||||||
|
mem_ub_n <= 1'b0;
|
||||||
|
mem_addr <= tag[ADDR_WIDTH-1:1] + fetch_tile_idx*WORDS_PER_TILE[TIW-1:0] + word_idx + 1'b1;
|
||||||
|
end
|
||||||
|
end
|
||||||
|
end
|
||||||
|
|
||||||
|
default: state <= ST_IDLE;
|
||||||
|
endcase
|
||||||
|
end
|
||||||
|
end
|
||||||
|
|
||||||
|
endmodule
|
||||||
@@ -34,19 +34,32 @@
|
|||||||
// component gluing the two together.
|
// component gluing the two together.
|
||||||
//
|
//
|
||||||
// Scope (see hardware/v2/logs/decisions.log DEC-0009):
|
// Scope (see hardware/v2/logs/decisions.log DEC-0009):
|
||||||
// - activation_buffer.v/weight_buffer.v/result_buffer.v (M3) are NOT
|
// - activation_buffer.v/weight_buffer.v/result_buffer.v (M3, BRAM-
|
||||||
// instantiated inside dataflow_core yet -- they belong on the OTHER
|
// backed FIFOs) are NOT instantiated here -- superseded by a
|
||||||
// side of the Memory Backend Interface (§15's own diagram: Memory
|
// different, measurement-driven shared cache (activation_cache.v,
|
||||||
// Manager -> Memory Backend Interface -> PSRAM Controller), and
|
// post-M10 DEC-0016, see below), not the original M3 modules
|
||||||
// each memory_manager instance already owns its own prefetch double
|
// themselves.
|
||||||
// buffer (M4) for the fast path. Wiring the M3 buffers in as a
|
// - each slot's Memory Backend Interface port is exposed SEPARATELY
|
||||||
// shared on-chip cache in front of PSRAM is real future work, not
|
// (N_SLOTS independent ports) rather than arbitrated down to one
|
||||||
// done here (no measured need for it yet, §22/§30).
|
// shared PSRAM master -- real PSRAM integration (including whatever
|
||||||
// - each slot's byte-level Memory Backend Interface port is exposed
|
// arbitration N_SLOTS>1 requires) is done one level up, in
|
||||||
// SEPARATELY (N_SLOTS independent ports) rather than arbitrated
|
// neural_multiprocessor.v (M8).
|
||||||
// down to one shared PSRAM master -- real PSRAM integration
|
//
|
||||||
// (including whatever arbitration N_SLOTS>1 requires) is explicitly
|
// Post-M10 (decisions.log DEC-0016): a single shared activation_cache
|
||||||
// M8's job, not this one's.
|
// instance sits alongside the N_SLOTS memory_managers, serving the
|
||||||
|
// ACTIVATION (X) half of each tile fetch -- in the realistic dense-
|
||||||
|
// layer workloads this project benchmarks, many neurons share the
|
||||||
|
// exact same X vector, and fetching it from PSRAM once instead of
|
||||||
|
// once per memory_manager instance is real, measured, redundant-
|
||||||
|
// traffic elimination (see hardware/v2/docs/benchmarks/
|
||||||
|
// final-benchmark.md's own recommendation #2). Each memory_manager's
|
||||||
|
// own prefetch_engine now fetches WEIGHTS only. The exposed
|
||||||
|
// slot_mem_* arrays are sized N_SLOTS+1: indices [0, N_SLOTS) are the
|
||||||
|
// per-slot memory_managers' own weight+write-back backend ports
|
||||||
|
// (unchanged in spirit from before), index [N_SLOTS] is the shared
|
||||||
|
// activation_cache's own backend port -- all N_SLOTS+1 arbitrated
|
||||||
|
// together by neural_multiprocessor.v's slot_mem_arbiter.v (N_PORTS
|
||||||
|
// widened to N_SLOTS+1 there to match).
|
||||||
// ================================================================
|
// ================================================================
|
||||||
|
|
||||||
module dataflow_core #(
|
module dataflow_core #(
|
||||||
@@ -73,19 +86,20 @@ module dataflow_core #(
|
|||||||
input wire [15:0] reg_n_tiles,
|
input wire [15:0] reg_n_tiles,
|
||||||
input wire [ADDR_WIDTH-1:0] reg_result_addr,
|
input wire [ADDR_WIDTH-1:0] reg_result_addr,
|
||||||
|
|
||||||
// ---- per-slot Memory Backend Interface (arrayed, one per slot --
|
// ---- Memory Backend Interface, arrayed N_SLOTS+1 wide (see file
|
||||||
// see file header on why arbitration to one shared PSRAM port is
|
// header: indices [0,N_SLOTS) are the per-slot memory_managers'
|
||||||
// NOT done here). WORD-level (16-bit) post-M10 (decisions.log
|
// own weight+write-back ports, index [N_SLOTS] is the shared
|
||||||
// DEC-0015) -- see memory_manager.v/prefetch_engine.v's own
|
// activation_cache's own port). WORD-level (16-bit) post-M10
|
||||||
// headers for why. ----
|
// (decisions.log DEC-0015) -- see memory_manager.v/
|
||||||
output wire [N_SLOTS-1:0] slot_mem_req,
|
// prefetch_engine.v's own headers for why. ----
|
||||||
output wire [N_SLOTS-1:0] slot_mem_wr,
|
output wire [N_SLOTS:0] slot_mem_req,
|
||||||
output wire [ADDR_WIDTH*N_SLOTS-1:0] slot_mem_addr, // WORD address
|
output wire [N_SLOTS:0] slot_mem_wr,
|
||||||
output wire [16*N_SLOTS-1:0] slot_mem_wdata,
|
output wire [ADDR_WIDTH*(N_SLOTS+1)-1:0] slot_mem_addr, // WORD address
|
||||||
output wire [N_SLOTS-1:0] slot_mem_lb_n,
|
output wire [16*(N_SLOTS+1)-1:0] slot_mem_wdata,
|
||||||
output wire [N_SLOTS-1:0] slot_mem_ub_n,
|
output wire [N_SLOTS:0] slot_mem_lb_n,
|
||||||
input wire [16*N_SLOTS-1:0] slot_mem_rdata,
|
output wire [N_SLOTS:0] slot_mem_ub_n,
|
||||||
input wire [N_SLOTS-1:0] slot_mem_ready
|
input wire [16*(N_SLOTS+1)-1:0] slot_mem_rdata,
|
||||||
|
input wire [N_SLOTS:0] slot_mem_ready
|
||||||
);
|
);
|
||||||
|
|
||||||
localparam NODE_IDW = $clog2(N_NODES);
|
localparam NODE_IDW = $clog2(N_NODES);
|
||||||
@@ -153,6 +167,14 @@ module dataflow_core #(
|
|||||||
assign dm_producer_done_valid = dir_job_out_done;
|
assign dm_producer_done_valid = dir_job_out_done;
|
||||||
assign dm_producer_done_node_id = completed_node_id_16[NODE_IDW-1:0];
|
assign dm_producer_done_node_id = completed_node_id_16[NODE_IDW-1:0];
|
||||||
|
|
||||||
|
// ---- shared activation_cache request bus (one port per slot,
|
||||||
|
// collected here for the cache instance below) ----
|
||||||
|
wire [N_SLOTS-1:0] xc_req;
|
||||||
|
wire [ADDR_WIDTH*N_SLOTS-1:0] xc_x_base;
|
||||||
|
wire [16*N_SLOTS-1:0] xc_tile_idx;
|
||||||
|
wire [N_SLOTS-1:0] xc_ack;
|
||||||
|
wire signed [DATA_WIDTH*P_IN*N_SLOTS-1:0] xc_tile_x;
|
||||||
|
|
||||||
// ---- N_SLOTS x (Memory Manager (M4) + Neural Processor (M1)) ----
|
// ---- N_SLOTS x (Memory Manager (M4) + Neural Processor (M1)) ----
|
||||||
genvar g;
|
genvar g;
|
||||||
generate
|
generate
|
||||||
@@ -177,6 +199,11 @@ module dataflow_core #(
|
|||||||
.operand_valid(mm_operand_valid), .operand_ready(mm_operand_ready),
|
.operand_valid(mm_operand_valid), .operand_ready(mm_operand_ready),
|
||||||
.input_data(mm_input_data), .weight_data(mm_weight_data), .tile_last(mm_tile_last),
|
.input_data(mm_input_data), .weight_data(mm_weight_data), .tile_last(mm_tile_last),
|
||||||
.result_valid(mm_result_valid), .result_ready(mm_result_ready), .result_data(mm_result_data),
|
.result_valid(mm_result_valid), .result_ready(mm_result_ready), .result_data(mm_result_data),
|
||||||
|
.xc_req(xc_req[g]),
|
||||||
|
.xc_x_base(xc_x_base[g*ADDR_WIDTH +: ADDR_WIDTH]),
|
||||||
|
.xc_tile_idx(xc_tile_idx[g*16 +: 16]),
|
||||||
|
.xc_ack(xc_ack[g]),
|
||||||
|
.xc_tile_x(xc_tile_x[g*DATA_WIDTH*P_IN +: DATA_WIDTH*P_IN]),
|
||||||
.mem_req(slot_mem_req[g]), .mem_wr(slot_mem_wr[g]),
|
.mem_req(slot_mem_req[g]), .mem_wr(slot_mem_wr[g]),
|
||||||
.mem_addr(slot_mem_addr[g*ADDR_WIDTH +: ADDR_WIDTH]),
|
.mem_addr(slot_mem_addr[g*ADDR_WIDTH +: ADDR_WIDTH]),
|
||||||
.mem_wdata(slot_mem_wdata[g*16 +: 16]),
|
.mem_wdata(slot_mem_wdata[g*16 +: 16]),
|
||||||
@@ -215,4 +242,21 @@ module dataflow_core #(
|
|||||||
end
|
end
|
||||||
endgenerate
|
endgenerate
|
||||||
|
|
||||||
|
// ---- shared activation_cache (M10+, DEC-0016) -- serves the
|
||||||
|
// ACTIVATION half of every slot's tile fetch, using arbiter port
|
||||||
|
// index N_SLOTS (the last one) for its own PSRAM traffic on a
|
||||||
|
// cache miss. ----
|
||||||
|
activation_cache #(
|
||||||
|
.DATA_WIDTH(DATA_WIDTH), .P_IN(P_IN), .ADDR_WIDTH(ADDR_WIDTH), .N_SLOTS(N_SLOTS)
|
||||||
|
) u_activation_cache (
|
||||||
|
.clk(clk), .rst(rst),
|
||||||
|
.req(xc_req), .req_x_base(xc_x_base), .req_tile_idx(xc_tile_idx),
|
||||||
|
.ack(xc_ack), .tile_x_out(xc_tile_x),
|
||||||
|
.mem_req(slot_mem_req[N_SLOTS]), .mem_wr(slot_mem_wr[N_SLOTS]),
|
||||||
|
.mem_addr(slot_mem_addr[N_SLOTS*ADDR_WIDTH +: ADDR_WIDTH]),
|
||||||
|
.mem_wdata(slot_mem_wdata[N_SLOTS*16 +: 16]),
|
||||||
|
.mem_lb_n(slot_mem_lb_n[N_SLOTS]), .mem_ub_n(slot_mem_ub_n[N_SLOTS]),
|
||||||
|
.mem_rdata(slot_mem_rdata[N_SLOTS*16 +: 16]), .mem_ready(slot_mem_ready[N_SLOTS])
|
||||||
|
);
|
||||||
|
|
||||||
endmodule
|
endmodule
|
||||||
|
|||||||
@@ -10,42 +10,41 @@
|
|||||||
// The processor sees only "data available" (operand_valid/ready,
|
// The processor sees only "data available" (operand_valid/ready,
|
||||||
// tile_last) -- never PSRAM request/wait cycles directly (§12).
|
// tile_last) -- never PSRAM request/wait cycles directly (§12).
|
||||||
//
|
//
|
||||||
// Post-M10 (decisions.log DEC-0015): this port talks directly to
|
// Post-M10 (decisions.log DEC-0015): the WEIGHT backend port talks
|
||||||
// memory_interface.v's own 16-bit word interface instead of routing
|
// directly to memory_interface.v's own 16-bit word interface instead
|
||||||
// through int8_memory_access.v's byte-splitting layer -- every real
|
// of routing through int8_memory_access.v's byte-splitting layer --
|
||||||
// transaction now moves a full PSRAM word (2 bytes) instead of
|
// every real transaction now moves a full PSRAM word (2 bytes)
|
||||||
// discarding half of one, halving the real transaction count for
|
// instead of discarding half of one. int8_memory_access.v itself is
|
||||||
// prefetch_engine's own reads. int8_memory_access.v itself is
|
|
||||||
// untouched (still frozen V1); V2 simply no longer instantiates it in
|
// untouched (still frozen V1); V2 simply no longer instantiates it in
|
||||||
// this datapath, reusing the lower (word-level) layer directly
|
// this datapath, reusing the lower (word-level) layer directly
|
||||||
// instead, the same "reuse what fits" precedent slot_mem_arbiter.v
|
// instead, the same "reuse what fits" precedent slot_mem_arbiter.v
|
||||||
// already set for hardware/v1/rtl/mem_arbiter.v.
|
// already set for hardware/v1/rtl/mem_arbiter.v.
|
||||||
//
|
//
|
||||||
|
// Post-M10 (decisions.log DEC-0016): the ACTIVATION (X) side is no
|
||||||
|
// longer fetched from PSRAM by this module's own prefetch_engine at
|
||||||
|
// all -- it is requested from a shared activation_cache.v instance
|
||||||
|
// (one per dataflow_core, not one per slot), which fetches a given
|
||||||
|
// X vector from PSRAM once and serves every memory_manager sharing
|
||||||
|
// that same x_base directly on-chip. Each bank therefore becomes
|
||||||
|
// ready only once BOTH its activation half (cache ack) AND its
|
||||||
|
// weight half (prefetch_engine's own pf_done, now W-only) have
|
||||||
|
// arrived -- bank_ready[b] = bank_x_ready[b] && bank_w_ready[b].
|
||||||
|
//
|
||||||
// Double-buffered prefetch (§13): while the processor consumes tile
|
// Double-buffered prefetch (§13): while the processor consumes tile
|
||||||
// N from bank "current", this module retargets the single
|
// N from bank "current", this module retargets its single
|
||||||
// prefetch_engine instance (M4) at bank "next" to fetch tile N+1
|
// prefetch_engine instance (M4, W-only) and issues a fresh
|
||||||
|
// activation_cache request at bank "next" to fetch tile N+1
|
||||||
// concurrently. On tile handoff, banks swap; if a bank isn't ready in
|
// concurrently. On tile handoff, banks swap; if a bank isn't ready in
|
||||||
// time (prefetch slower than compute for this run), operand_valid
|
// time, operand_valid simply stays low until it is -- a real stall,
|
||||||
// simply stays low until it is -- a real stall, not hidden, so its
|
// not hidden (§22). NOTE (measured characteristic, not yet optimized
|
||||||
// frequency is genuinely measurable (§22, deferred to M9). NOTE
|
// -- see decisions.log DEC-0006): the bank-swap-and-check control path
|
||||||
// (measured characteristic, not yet optimized -- see
|
// itself costs a minimum 1 idle cycle per tile handoff even when the
|
||||||
// hardware/v2/logs/decisions.log DEC-0006): the bank-swap-and-check
|
// next bank was already prefetched in time, unlike neural_processor.v's
|
||||||
// control path itself costs a minimum 1 idle cycle per tile handoff
|
// own zero-gap tile acceptance.
|
||||||
// even when the next bank was already prefetched in time, unlike
|
|
||||||
// neural_processor.v's own zero-gap tile acceptance -- a real,
|
|
||||||
// deliberately-not-hidden overhead of this first Memory Manager
|
|
||||||
// implementation, left for M10 (Optimization) to revisit with real
|
|
||||||
// stall-percentage data (§22) rather than optimized blindly now.
|
|
||||||
//
|
//
|
||||||
// One job = one neuron's worth of tiles (n_tiles), read from x_base/
|
// One job = one neuron's worth of tiles (n_tiles), read from x_base/
|
||||||
// w_base (PSRAM byte addresses), followed by writing the single
|
// w_base (PSRAM byte addresses), followed by writing the single INT8
|
||||||
// INT8 result back to result_addr. The result write only happens
|
// result back to result_addr.
|
||||||
// after the last tile has been handed off and prefetch_engine is
|
|
||||||
// idle (temporally disjoint from prefetching by construction), so no
|
|
||||||
// separate backend arbiter is needed at this milestone -- see
|
|
||||||
// decisions.log DEC-0006 for why, and what changes once multiple
|
|
||||||
// concurrent jobs/processors need to share one backend port
|
|
||||||
// (deferred, not yet needed).
|
|
||||||
// ================================================================
|
// ================================================================
|
||||||
|
|
||||||
module memory_manager #(
|
module memory_manager #(
|
||||||
@@ -56,8 +55,7 @@ module memory_manager #(
|
|||||||
input wire clk,
|
input wire clk,
|
||||||
input wire rst,
|
input wire rst,
|
||||||
|
|
||||||
// ---- job control (from a future Neural Director, M5; driven
|
// ---- job control (from Neural Director, M5) ----
|
||||||
// directly by a testbench at M4) ----
|
|
||||||
input wire job_start,
|
input wire job_start,
|
||||||
input wire [ADDR_WIDTH-1:0] x_base,
|
input wire [ADDR_WIDTH-1:0] x_base,
|
||||||
input wire [ADDR_WIDTH-1:0] w_base,
|
input wire [ADDR_WIDTH-1:0] w_base,
|
||||||
@@ -78,12 +76,19 @@ module memory_manager #(
|
|||||||
output reg result_ready,
|
output reg result_ready,
|
||||||
input wire signed [DATA_WIDTH-1:0] result_data,
|
input wire signed [DATA_WIDTH-1:0] result_data,
|
||||||
|
|
||||||
// ---- Memory Backend Interface (word-level, matches
|
// ---- shared activation_cache.v request port (post-M10 DEC-0016
|
||||||
|
// -- one per memory_manager instance, cache is shared/instantiated
|
||||||
|
// once per dataflow_core) ----
|
||||||
|
output reg xc_req, // one-cycle pulse
|
||||||
|
output reg [ADDR_WIDTH-1:0] xc_x_base,
|
||||||
|
output reg [15:0] xc_tile_idx,
|
||||||
|
input wire xc_ack, // one-cycle pulse
|
||||||
|
input wire signed [DATA_WIDTH*P_IN-1:0] xc_tile_x,
|
||||||
|
|
||||||
|
// ---- WEIGHT Memory Backend Interface (word-level, matches
|
||||||
// hardware/v1/rtl/memory_interface.v's contract exactly -- see
|
// hardware/v1/rtl/memory_interface.v's contract exactly -- see
|
||||||
// prefetch_engine.v's own header and decisions.log DEC-0015 for
|
// prefetch_engine.v's own header and decisions.log DEC-0015/
|
||||||
// why this is now word- rather than byte-level: int8_memory_access.v
|
// DEC-0016 for why this is word-level and weight-only) ----
|
||||||
// is no longer in the datapath, each transaction moves a full
|
|
||||||
// 16-bit PSRAM word instead of discarding half of it) ----
|
|
||||||
output wire mem_req,
|
output wire mem_req,
|
||||||
output wire mem_wr,
|
output wire mem_wr,
|
||||||
output wire [ADDR_WIDTH-1:0] mem_addr, // WORD address
|
output wire [ADDR_WIDTH-1:0] mem_addr, // WORD address
|
||||||
@@ -108,17 +113,55 @@ module memory_manager #(
|
|||||||
reg [15:0] tile_idx; // tile currently presented (bank `current`)
|
reg [15:0] tile_idx; // tile currently presented (bank `current`)
|
||||||
reg current_bank; // 0 or 1
|
reg current_bank; // 0 or 1
|
||||||
|
|
||||||
reg [1:0] bank_ready; // bank_ready[b] = bank b holds valid, unconsumed prefetched data
|
// ---- double-buffer storage: X half filled by the shared cache,
|
||||||
|
// W half filled by this module's own prefetch_engine -- a bank is
|
||||||
|
// usable once BOTH halves have arrived. ----
|
||||||
|
reg [1:0] bank_x_ready, bank_w_ready;
|
||||||
|
wire [1:0] bank_ready = bank_x_ready & bank_w_ready;
|
||||||
|
|
||||||
// ---- double-buffer storage (owned here, filled by prefetch_engine) ----
|
|
||||||
reg signed [DATA_WIDTH*P_IN-1:0] bank_x [0:1];
|
reg signed [DATA_WIDTH*P_IN-1:0] bank_x [0:1];
|
||||||
reg signed [DATA_WIDTH*P_IN-1:0] bank_w [0:1];
|
reg signed [DATA_WIDTH*P_IN-1:0] bank_w [0:1];
|
||||||
|
|
||||||
// ---- single prefetch_engine instance, retargeted per bank ----
|
// ---- activation_cache request bookkeeping: single-entry pending
|
||||||
|
// (same idiom as pf_pending below -- only one outstanding cache
|
||||||
|
// request at a time, one instance to serve, one bank as its target). ----
|
||||||
|
reg xc_pending;
|
||||||
|
reg [ADDR_WIDTH-1:0] xc_pending_x_base;
|
||||||
|
reg [15:0] xc_pending_tile_idx;
|
||||||
|
reg xc_pending_bank;
|
||||||
|
// xc_target_bank is the bank the CURRENTLY-outstanding (already
|
||||||
|
// issued) cache request will fill -- set ONLY by the issue rule
|
||||||
|
// below, from xc_pending_bank, at the exact moment xc_req fires.
|
||||||
|
// Queueing logic (MM_IDLE/MM_PREFETCH_FIRST/MM_STREAM) writes
|
||||||
|
// xc_pending_bank, NEVER xc_target_bank directly -- writing
|
||||||
|
// xc_target_bank directly from queueing was a real bug (found via
|
||||||
|
// simulation): a later handoff can queue a NEW request (targeting
|
||||||
|
// a DIFFERENT bank) in the same cycle an EARLIER request is being
|
||||||
|
// issued, and program-order NBA "last write wins" would silently
|
||||||
|
// overwrite which bank the EARLIER (already in-flight) request's
|
||||||
|
// eventual ack gets applied to -- the exact same class of bug
|
||||||
|
// ERR-0006 already found and fixed once for pf_target_bank/
|
||||||
|
// pf_pending_bank (which already used this two-register pattern
|
||||||
|
// correctly; this module's activation-cache side did not, until
|
||||||
|
// now).
|
||||||
|
reg xc_target_bank;
|
||||||
|
// Tracks whether THIS instance's own cache request is still
|
||||||
|
// awaiting its ack (real PSRAM miss latency can easily exceed one
|
||||||
|
// neural_processor tile's own compute time, so a later handoff's
|
||||||
|
// "queue the next request" can genuinely race an earlier request
|
||||||
|
// still in flight -- the same class of race ERR-0006 already found
|
||||||
|
// and fixed once for pf_pending/pf_busy; fixed here the same way,
|
||||||
|
// with an explicit outstanding flag this module controls directly
|
||||||
|
// rather than inferring busy-ness from a signal with its own
|
||||||
|
// latency quirk).
|
||||||
|
reg xc_outstanding;
|
||||||
|
|
||||||
|
// ---- single prefetch_engine instance (WEIGHT-only post-DEC-0016),
|
||||||
|
// retargeted per bank ----
|
||||||
reg pf_start;
|
reg pf_start;
|
||||||
reg [ADDR_WIDTH-1:0] pf_x_addr, pf_w_addr;
|
reg [ADDR_WIDTH-1:0] pf_w_addr;
|
||||||
wire pf_busy, pf_done;
|
wire pf_busy, pf_done;
|
||||||
wire signed [DATA_WIDTH*P_IN-1:0] pf_tile_x, pf_tile_w;
|
wire signed [DATA_WIDTH*P_IN-1:0] pf_tile_w;
|
||||||
|
|
||||||
reg pf_target_bank; // which bank the CURRENTLY-running (or just-launched) prefetch fills
|
reg pf_target_bank; // which bank the CURRENTLY-running (or just-launched) prefetch fills
|
||||||
|
|
||||||
@@ -132,15 +175,17 @@ module memory_manager #(
|
|||||||
// descriptor instead of touching pf_start directly; a single
|
// descriptor instead of touching pf_start directly; a single
|
||||||
// always-active rule issues pf_start once the engine is free.
|
// always-active rule issues pf_start once the engine is free.
|
||||||
reg pf_pending;
|
reg pf_pending;
|
||||||
reg [ADDR_WIDTH-1:0] pf_pending_x, pf_pending_w;
|
reg [ADDR_WIDTH-1:0] pf_pending_w;
|
||||||
reg pf_pending_bank;
|
reg pf_pending_bank;
|
||||||
|
|
||||||
// prefetch_engine drives its OWN internal backend wires; the
|
// prefetch_engine drives its OWN internal weight-backend wires;
|
||||||
// result-write FSM below drives its own. A combinational mux
|
// the result-write FSM below drives its own. A combinational mux
|
||||||
// (never both at once, by construction -- see file header)
|
// (never both at once, by construction -- MM_WRITE_RESULT/MM_DONE
|
||||||
// selects which one actually reaches the real output port,
|
// only run after every tile for this job has already been fetched,
|
||||||
// avoiding a two-driver conflict on mem_req/mem_wr/mem_addr/
|
// so prefetch_engine is guaranteed idle) selects which one actually
|
||||||
// mem_wdata/mem_lb_n/mem_ub_n.
|
// reaches the real output port, same pattern as the pre-DEC-0015
|
||||||
|
// design, just weight-only now (the activation side moved to the
|
||||||
|
// shared activation_cache.v, DEC-0016).
|
||||||
wire pf_mem_req, pf_mem_wr;
|
wire pf_mem_req, pf_mem_wr;
|
||||||
wire [ADDR_WIDTH-1:0] pf_mem_addr;
|
wire [ADDR_WIDTH-1:0] pf_mem_addr;
|
||||||
wire [15:0] pf_mem_wdata;
|
wire [15:0] pf_mem_wdata;
|
||||||
@@ -150,9 +195,9 @@ module memory_manager #(
|
|||||||
.DATA_WIDTH(DATA_WIDTH), .P_IN(P_IN), .ADDR_WIDTH(ADDR_WIDTH)
|
.DATA_WIDTH(DATA_WIDTH), .P_IN(P_IN), .ADDR_WIDTH(ADDR_WIDTH)
|
||||||
) u_prefetch (
|
) u_prefetch (
|
||||||
.clk(clk), .rst(rst),
|
.clk(clk), .rst(rst),
|
||||||
.fetch_start(pf_start), .x_addr(pf_x_addr), .w_addr(pf_w_addr),
|
.fetch_start(pf_start), .w_addr(pf_w_addr),
|
||||||
.fetch_busy(pf_busy), .fetch_done(pf_done),
|
.fetch_busy(pf_busy), .fetch_done(pf_done),
|
||||||
.tile_x(pf_tile_x), .tile_w(pf_tile_w),
|
.tile_w(pf_tile_w),
|
||||||
.mem_req(pf_mem_req), .mem_wr(pf_mem_wr), .mem_addr(pf_mem_addr), .mem_wdata(pf_mem_wdata),
|
.mem_req(pf_mem_req), .mem_wr(pf_mem_wr), .mem_addr(pf_mem_addr), .mem_wdata(pf_mem_wdata),
|
||||||
.mem_lb_n(pf_mem_lb_n), .mem_ub_n(pf_mem_ub_n),
|
.mem_lb_n(pf_mem_lb_n), .mem_ub_n(pf_mem_ub_n),
|
||||||
.mem_rdata(mem_rdata), .mem_ready(mem_ready)
|
.mem_rdata(mem_rdata), .mem_ready(mem_ready)
|
||||||
@@ -163,12 +208,6 @@ module memory_manager #(
|
|||||||
reg [15:0] wr_mem_wdata;
|
reg [15:0] wr_mem_wdata;
|
||||||
reg wr_mem_lb_n, wr_mem_ub_n;
|
reg wr_mem_lb_n, wr_mem_ub_n;
|
||||||
|
|
||||||
// wr_mem_req is SET while state==MM_WRITE_RESULT but only becomes
|
|
||||||
// valid (via NBA) the FOLLOWING cycle, i.e. while state==MM_DONE --
|
|
||||||
// the mux must select the write-back source across BOTH states,
|
|
||||||
// not just the one that issues it (an off-by-one here silently
|
|
||||||
// dropped the write request entirely -- found and fixed here, see
|
|
||||||
// hardware/v2/logs/errors.log ERR-0006).
|
|
||||||
wire wr_active = (state == MM_WRITE_RESULT) || (state == MM_DONE);
|
wire wr_active = (state == MM_WRITE_RESULT) || (state == MM_DONE);
|
||||||
assign mem_req = wr_active ? wr_mem_req : pf_mem_req;
|
assign mem_req = wr_active ? wr_mem_req : pf_mem_req;
|
||||||
assign mem_wr = wr_active ? 1'b1 : pf_mem_wr;
|
assign mem_wr = wr_active ? 1'b1 : pf_mem_wr;
|
||||||
@@ -188,45 +227,63 @@ module memory_manager #(
|
|||||||
result_ready <= 1'b0;
|
result_ready <= 1'b0;
|
||||||
pf_start <= 1'b0;
|
pf_start <= 1'b0;
|
||||||
current_bank <= 1'b0;
|
current_bank <= 1'b0;
|
||||||
bank_ready <= 2'b00;
|
bank_x_ready <= 2'b00;
|
||||||
|
bank_w_ready <= 2'b00;
|
||||||
tile_idx <= 16'h0;
|
tile_idx <= 16'h0;
|
||||||
|
pf_pending <= 1'b0;
|
||||||
|
xc_req <= 1'b0;
|
||||||
|
xc_pending <= 1'b0;
|
||||||
|
xc_outstanding <= 1'b0;
|
||||||
wr_mem_req <= 1'b0;
|
wr_mem_req <= 1'b0;
|
||||||
wr_mem_addr <= {ADDR_WIDTH{1'b0}};
|
wr_mem_addr <= {ADDR_WIDTH{1'b0}};
|
||||||
wr_mem_wdata <= 16'h0000;
|
wr_mem_wdata <= 16'h0000;
|
||||||
wr_mem_lb_n <= 1'b1;
|
wr_mem_lb_n <= 1'b1;
|
||||||
wr_mem_ub_n <= 1'b1;
|
wr_mem_ub_n <= 1'b1;
|
||||||
pf_pending <= 1'b0;
|
|
||||||
end else begin
|
end else begin
|
||||||
job_done <= 1'b0;
|
job_done <= 1'b0;
|
||||||
pf_start <= 1'b0;
|
pf_start <= 1'b0;
|
||||||
|
xc_req <= 1'b0;
|
||||||
result_ready <= 1'b0;
|
result_ready <= 1'b0;
|
||||||
|
|
||||||
// Latch a completed prefetch into its target bank.
|
// Latch a completed weight prefetch into its target bank.
|
||||||
if (pf_done) begin
|
if (pf_done) begin
|
||||||
bank_x[pf_target_bank] <= pf_tile_x;
|
|
||||||
bank_w[pf_target_bank] <= pf_tile_w;
|
bank_w[pf_target_bank] <= pf_tile_w;
|
||||||
bank_ready[pf_target_bank] <= 1'b1;
|
bank_w_ready[pf_target_bank] <= 1'b1;
|
||||||
end
|
end
|
||||||
|
|
||||||
// Issue a pending fetch request as soon as the (single)
|
// Latch a completed activation-cache fetch into its target
|
||||||
|
// bank and clear the outstanding flag (see its own
|
||||||
|
// declaration comment above).
|
||||||
|
if (xc_ack) begin
|
||||||
|
bank_x[xc_target_bank] <= xc_tile_x;
|
||||||
|
bank_x_ready[xc_target_bank] <= 1'b1;
|
||||||
|
xc_outstanding <= 1'b0;
|
||||||
|
end
|
||||||
|
|
||||||
|
// Issue a pending weight fetch as soon as the (single)
|
||||||
// prefetch engine is genuinely free. The `!pf_start` guard
|
// prefetch engine is genuinely free. The `!pf_start` guard
|
||||||
// is required, not cosmetic: pf_busy does not read 1 until
|
// is required, not cosmetic -- see hardware/v2/logs/
|
||||||
// the cycle AFTER pf_start was first observed (prefetch_
|
// errors.log ERR-0006.
|
||||||
// engine's own fetch_busy<=1 is one clock behind its own
|
|
||||||
// fetch_start sampling), so checking !pf_busy alone leaves
|
|
||||||
// a genuine one-cycle window where a second pending
|
|
||||||
// request would fire on top of the one just launched,
|
|
||||||
// silently corrupting pf_target_bank for the fetch already
|
|
||||||
// in flight (found and fixed here -- see
|
|
||||||
// hardware/v2/logs/errors.log ERR-0006).
|
|
||||||
if (pf_pending && !pf_busy && !pf_start) begin
|
if (pf_pending && !pf_busy && !pf_start) begin
|
||||||
pf_start <= 1'b1;
|
pf_start <= 1'b1;
|
||||||
pf_x_addr <= pf_pending_x;
|
|
||||||
pf_w_addr <= pf_pending_w;
|
pf_w_addr <= pf_pending_w;
|
||||||
pf_target_bank <= pf_pending_bank;
|
pf_target_bank <= pf_pending_bank;
|
||||||
pf_pending <= 1'b0;
|
pf_pending <= 1'b0;
|
||||||
end
|
end
|
||||||
|
|
||||||
|
// Issue a pending activation-cache request only once this
|
||||||
|
// instance's own PREVIOUS request has been genuinely acked
|
||||||
|
// (xc_outstanding low) -- see that flag's own declaration
|
||||||
|
// comment for why checking xc_req alone is not enough.
|
||||||
|
if (xc_pending && !xc_outstanding) begin
|
||||||
|
xc_req <= 1'b1;
|
||||||
|
xc_outstanding <= 1'b1;
|
||||||
|
xc_x_base <= xc_pending_x_base;
|
||||||
|
xc_tile_idx <= xc_pending_tile_idx;
|
||||||
|
xc_target_bank <= xc_pending_bank;
|
||||||
|
xc_pending <= 1'b0;
|
||||||
|
end
|
||||||
|
|
||||||
case (state)
|
case (state)
|
||||||
|
|
||||||
MM_IDLE: begin
|
MM_IDLE: begin
|
||||||
@@ -237,30 +294,37 @@ module memory_manager #(
|
|||||||
result_addr_reg <= result_addr;
|
result_addr_reg <= result_addr;
|
||||||
tile_idx <= 16'h0;
|
tile_idx <= 16'h0;
|
||||||
current_bank <= 1'b0;
|
current_bank <= 1'b0;
|
||||||
bank_ready <= 2'b00;
|
bank_x_ready <= 2'b00;
|
||||||
|
bank_w_ready <= 2'b00;
|
||||||
operand_valid <= 1'b0;
|
operand_valid <= 1'b0;
|
||||||
// kick off the very first fetch (tile 0 into bank 0)
|
// kick off the very first fetch (tile 0 into bank 0)
|
||||||
pf_pending <= 1'b1;
|
pf_pending <= 1'b1;
|
||||||
pf_pending_x <= x_base;
|
|
||||||
pf_pending_w <= w_base;
|
pf_pending_w <= w_base;
|
||||||
pf_pending_bank <= 1'b0;
|
pf_pending_bank <= 1'b0;
|
||||||
|
xc_pending <= 1'b1;
|
||||||
|
xc_pending_x_base <= x_base;
|
||||||
|
xc_pending_tile_idx <= 16'h0;
|
||||||
|
xc_pending_bank <= 1'b0;
|
||||||
state <= MM_PREFETCH_FIRST;
|
state <= MM_PREFETCH_FIRST;
|
||||||
end
|
end
|
||||||
end
|
end
|
||||||
|
|
||||||
MM_PREFETCH_FIRST: begin
|
MM_PREFETCH_FIRST: begin
|
||||||
if (bank_ready[0] || (pf_done && pf_target_bank == 1'b0)) begin
|
if (bank_ready[0]) begin
|
||||||
// Present tile 0; concurrently start prefetching
|
// Present tile 0; concurrently start prefetching
|
||||||
// tile 1 into bank 1, if there is one.
|
// tile 1 into bank 1, if there is one.
|
||||||
operand_valid <= 1'b1;
|
operand_valid <= 1'b1;
|
||||||
input_data <= pf_done ? pf_tile_x : bank_x[0];
|
input_data <= bank_x[0];
|
||||||
weight_data <= pf_done ? pf_tile_w : bank_w[0];
|
weight_data <= bank_w[0];
|
||||||
tile_last <= (n_tiles_reg == 16'h1);
|
tile_last <= (n_tiles_reg == 16'h1);
|
||||||
if (n_tiles_reg > 16'h1) begin
|
if (n_tiles_reg > 16'h1) begin
|
||||||
pf_pending <= 1'b1;
|
pf_pending <= 1'b1;
|
||||||
pf_pending_x <= x_base_reg + P_IN[ADDR_WIDTH-1:0];
|
|
||||||
pf_pending_w <= w_base_reg + P_IN[ADDR_WIDTH-1:0];
|
pf_pending_w <= w_base_reg + P_IN[ADDR_WIDTH-1:0];
|
||||||
pf_pending_bank <= 1'b1;
|
pf_pending_bank <= 1'b1;
|
||||||
|
xc_pending <= 1'b1;
|
||||||
|
xc_pending_x_base <= x_base_reg;
|
||||||
|
xc_pending_tile_idx <= 16'h1;
|
||||||
|
xc_pending_bank <= 1'b1;
|
||||||
end
|
end
|
||||||
state <= MM_STREAM;
|
state <= MM_STREAM;
|
||||||
end
|
end
|
||||||
@@ -269,7 +333,8 @@ module memory_manager #(
|
|||||||
MM_STREAM: begin
|
MM_STREAM: begin
|
||||||
if (operand_valid && operand_ready) begin
|
if (operand_valid && operand_ready) begin
|
||||||
// This tile consumed; free its bank, swap.
|
// This tile consumed; free its bank, swap.
|
||||||
bank_ready[current_bank] <= 1'b0;
|
bank_x_ready[current_bank] <= 1'b0;
|
||||||
|
bank_w_ready[current_bank] <= 1'b0;
|
||||||
current_bank <= ~current_bank;
|
current_bank <= ~current_bank;
|
||||||
tile_idx <= tile_idx + 16'h1;
|
tile_idx <= tile_idx + 16'h1;
|
||||||
operand_valid <= 1'b0; // re-asserted below once the new bank is ready
|
operand_valid <= 1'b0; // re-asserted below once the new bank is ready
|
||||||
@@ -279,17 +344,14 @@ module memory_manager #(
|
|||||||
state <= MM_WAIT_RESULT;
|
state <= MM_WAIT_RESULT;
|
||||||
end else if (tile_idx + 16'h2 < n_tiles_reg) begin
|
end else if (tile_idx + 16'h2 < n_tiles_reg) begin
|
||||||
// Queue a prefetch for the tile AFTER next into
|
// Queue a prefetch for the tile AFTER next into
|
||||||
// the bank we just freed (current_bank, pre-swap)
|
// the bank we just freed (current_bank, pre-swap).
|
||||||
// -- it will actually launch once the (single)
|
|
||||||
// prefetch engine is free (see the pf_pending
|
|
||||||
// issue rule above); it is very likely still
|
|
||||||
// busy with the tile-N+1 fetch kicked off on the
|
|
||||||
// PREVIOUS handoff, so this almost always queues
|
|
||||||
// rather than launching immediately.
|
|
||||||
pf_pending <= 1'b1;
|
pf_pending <= 1'b1;
|
||||||
pf_pending_x <= x_base_reg + (tile_idx + 16'h2) * P_IN[ADDR_WIDTH-1:0];
|
|
||||||
pf_pending_w <= w_base_reg + (tile_idx + 16'h2) * P_IN[ADDR_WIDTH-1:0];
|
pf_pending_w <= w_base_reg + (tile_idx + 16'h2) * P_IN[ADDR_WIDTH-1:0];
|
||||||
pf_pending_bank <= current_bank; // the one just freed
|
pf_pending_bank <= current_bank; // the one just freed
|
||||||
|
xc_pending <= 1'b1;
|
||||||
|
xc_pending_x_base <= x_base_reg;
|
||||||
|
xc_pending_tile_idx <= tile_idx + 16'h2;
|
||||||
|
xc_pending_bank <= current_bank;
|
||||||
end
|
end
|
||||||
end else if (!operand_valid) begin
|
end else if (!operand_valid) begin
|
||||||
// Waiting for the new current bank to become ready
|
// Waiting for the new current bank to become ready
|
||||||
@@ -322,8 +384,9 @@ module memory_manager #(
|
|||||||
|
|
||||||
MM_WRITE_RESULT: begin
|
MM_WRITE_RESULT: begin
|
||||||
// prefetch_engine is guaranteed idle here (no more
|
// prefetch_engine is guaranteed idle here (no more
|
||||||
// tiles to fetch for this job), so driving the shared
|
// tiles to fetch for this job), so driving the
|
||||||
// backend port directly is safe -- see file header.
|
// shared weight backend port directly is safe --
|
||||||
|
// see file header/wr_active above.
|
||||||
wr_mem_req <= 1'b1;
|
wr_mem_req <= 1'b1;
|
||||||
wr_mem_addr <= result_addr_reg[ADDR_WIDTH-1:1]; // byte -> word
|
wr_mem_addr <= result_addr_reg[ADDR_WIDTH-1:1]; // byte -> word
|
||||||
state <= MM_DONE;
|
state <= MM_DONE;
|
||||||
|
|||||||
@@ -72,12 +72,14 @@ module neural_multiprocessor #(
|
|||||||
);
|
);
|
||||||
|
|
||||||
// ---- dataflow_core (M7, control logic unmodified; per-slot
|
// ---- dataflow_core (M7, control logic unmodified; per-slot
|
||||||
// backend port widened to 16-bit + lb_n/ub_n per DEC-0015) ----
|
// backend port widened to 16-bit + lb_n/ub_n per DEC-0015, and to
|
||||||
wire [N_SLOTS-1:0] slot_mem_req, slot_mem_wr;
|
// N_SLOTS+1 ports per DEC-0016 -- the extra port is the shared
|
||||||
wire [ADDR_WIDTH*N_SLOTS-1:0] slot_mem_addr;
|
// activation_cache's own backend traffic) ----
|
||||||
wire [16*N_SLOTS-1:0] slot_mem_wdata, slot_mem_rdata;
|
wire [N_SLOTS:0] slot_mem_req, slot_mem_wr;
|
||||||
wire [N_SLOTS-1:0] slot_mem_lb_n, slot_mem_ub_n;
|
wire [ADDR_WIDTH*(N_SLOTS+1)-1:0] slot_mem_addr;
|
||||||
wire [N_SLOTS-1:0] slot_mem_ready;
|
wire [16*(N_SLOTS+1)-1:0] slot_mem_wdata, slot_mem_rdata;
|
||||||
|
wire [N_SLOTS:0] slot_mem_lb_n, slot_mem_ub_n;
|
||||||
|
wire [N_SLOTS:0] slot_mem_ready;
|
||||||
|
|
||||||
dataflow_core #(
|
dataflow_core #(
|
||||||
.DATA_WIDTH(DATA_WIDTH), .P_IN(P_IN), .ACC_WIDTH(ACC_WIDTH), .ADDR_WIDTH(ADDR_WIDTH),
|
.DATA_WIDTH(DATA_WIDTH), .P_IN(P_IN), .ACC_WIDTH(ACC_WIDTH), .ADDR_WIDTH(ADDR_WIDTH),
|
||||||
@@ -93,7 +95,9 @@ module neural_multiprocessor #(
|
|||||||
.slot_mem_rdata(slot_mem_rdata), .slot_mem_ready(slot_mem_ready)
|
.slot_mem_rdata(slot_mem_rdata), .slot_mem_ready(slot_mem_ready)
|
||||||
);
|
);
|
||||||
|
|
||||||
// ---- N_SLOTS -> 1 arbiter (M8, word-level per DEC-0015) ----
|
// ---- (N_SLOTS+1) -> 1 arbiter (M8, word-level per DEC-0015;
|
||||||
|
// widened to N_SLOTS+1 ports per DEC-0016 to also arbitrate the
|
||||||
|
// shared activation_cache's own backend traffic) ----
|
||||||
wire arb_m_req, arb_m_wr;
|
wire arb_m_req, arb_m_wr;
|
||||||
wire [ADDR_WIDTH-1:0] arb_m_addr;
|
wire [ADDR_WIDTH-1:0] arb_m_addr;
|
||||||
wire [15:0] arb_m_wdata;
|
wire [15:0] arb_m_wdata;
|
||||||
@@ -102,7 +106,7 @@ module neural_multiprocessor #(
|
|||||||
wire arb_m_ready;
|
wire arb_m_ready;
|
||||||
|
|
||||||
slot_mem_arbiter #(
|
slot_mem_arbiter #(
|
||||||
.ADDR_WIDTH(ADDR_WIDTH), .N_PORTS(N_SLOTS)
|
.ADDR_WIDTH(ADDR_WIDTH), .N_PORTS(N_SLOTS+1)
|
||||||
) u_arbiter (
|
) u_arbiter (
|
||||||
.clk(clk), .rst(rst),
|
.clk(clk), .rst(rst),
|
||||||
.s_req(slot_mem_req), .s_wr(slot_mem_wr), .s_addr(slot_mem_addr),
|
.s_req(slot_mem_req), .s_wr(slot_mem_wr), .s_addr(slot_mem_addr),
|
||||||
|
|||||||
@@ -1,43 +1,44 @@
|
|||||||
`timescale 1ns/1ps
|
`timescale 1ns/1ps
|
||||||
|
|
||||||
// ================================================================
|
// ================================================================
|
||||||
// FPGA-Neural V2 -- Prefetch Engine (M4, docs/v2-description.md §13;
|
// FPGA-Neural V2 -- Weight Prefetch Engine (M4, docs/v2-description.md
|
||||||
// word-level burst rewrite post-M10 -- see hardware/v2/logs/
|
// §13; word-level burst rewrite post-M10 DEC-0015; X-fetch moved out
|
||||||
// decisions.log DEC-0015)
|
// to a shared activation_cache.v post-M10 DEC-0016)
|
||||||
//
|
//
|
||||||
// Fetches ONE tile (P_IN activation bytes + P_IN weight bytes) from
|
// Fetches ONE tile's P_IN WEIGHT bytes from the WORD-level Memory
|
||||||
// the WORD-level Memory Backend Interface, P_IN/2 sixteen-bit
|
// Backend Interface, P_IN/2 sixteen-bit transactions instead of P_IN
|
||||||
// transactions per array instead of P_IN single-byte ones.
|
// single-byte ones (DEC-0015 -- see this rationale in full below).
|
||||||
//
|
//
|
||||||
// WHY: hardware/v1/rtl/int8_memory_access.v (the byte-level backend
|
// Historical note: this module used to ALSO fetch the P_IN
|
||||||
// this engine originally sat on) converts every 8-bit logical request
|
// ACTIVATION (X) bytes for the same tile. DEC-0016 moved that
|
||||||
// into a FULL 16-bit PSRAM word access internally (mem_addr <= addr
|
// responsibility to a new shared activation_cache.v instead: in the
|
||||||
// >> 1, one byte lane selected via lb_n/ub_n) -- so a byte-at-a-time
|
// realistic dense-layer workloads this project actually benchmarks
|
||||||
// fetch was ALREADY paying for two bytes of real PSRAM bandwidth per
|
// (hardware/v2/docs/benchmarks/final-benchmark.md), many neurons
|
||||||
// transaction while only using one. This engine now talks directly to
|
// share the exact same X vector, and each of memory_manager.v's own
|
||||||
// hardware/v1/rtl/memory_interface.v's own 16-bit word interface
|
// N_SLOTS instances re-fetching that identical vector from PSRAM
|
||||||
// (skipping int8_memory_access.v entirely -- both are frozen V1 files,
|
// independently was real, measured, redundant traffic on the one
|
||||||
// unmodified either way, §1/§34; V2 is simply choosing to reuse the
|
// shared PSRAM port -- exactly the kind of real recommendation the
|
||||||
// lower layer instead of the byte-splitting one on top of it, the
|
// benchmark campaign was built to surface. Weights (W) are NOT shared
|
||||||
// same "reuse what fits" precedent already set by slot_mem_arbiter.v
|
// across neurons (each neuron has its own trained weight vector), so
|
||||||
// not reusing hardware/v1/rtl/mem_arbiter.v verbatim). psram_controller.v's
|
// there is no equivalent caching opportunity on the W side -- this
|
||||||
// own real page-mode support (already implemented, unmodified) then
|
// engine keeps fetching W directly from PSRAM, unchanged in spirit
|
||||||
// serves consecutive same-page word reads faster than a cold access --
|
// from DEC-0015, just no longer also fetching X.
|
||||||
// this engine's job is simply to stop discarding half of every word it
|
|
||||||
// already paid for, and to halve the number of real backend
|
|
||||||
// round-trips needed per tile.
|
|
||||||
//
|
//
|
||||||
// CONSTRAINT: P_IN must be even, and x_addr/w_addr must be word-
|
// WHY word-level (DEC-0015, unchanged rationale): int8_memory_access.v
|
||||||
// aligned (even BYTE addresses) -- each 16-bit transaction covers
|
// (the byte-level backend this engine originally sat on) converts
|
||||||
// BYTE addresses {addr, addr+1} as {low byte, high byte} (matches
|
// every 8-bit logical request into a FULL 16-bit PSRAM word access
|
||||||
// int8_memory_access.v's own addr[0] convention exactly, replicated
|
// internally (mem_addr <= addr >> 1, one byte lane selected via
|
||||||
// here since that module is no longer in the datapath). A host/loader
|
// lb_n/ub_n) -- so a byte-at-a-time fetch was ALREADY paying for two
|
||||||
// placing X/W tile arrays at even byte offsets (already true of every
|
// bytes of real PSRAM bandwidth per transaction while only using one.
|
||||||
// address used in this project's own testbenches) satisfies this
|
// This engine talks directly to hardware/v1/rtl/memory_interface.v's
|
||||||
// with no special handling.
|
// own 16-bit word interface (skipping int8_memory_access.v entirely --
|
||||||
|
// both are frozen V1 files, unmodified either way, §1/§34).
|
||||||
//
|
//
|
||||||
// The double-buffering strategy itself (§13) remains memory_manager.v's
|
// CONSTRAINT: P_IN must be even, and w_addr must be word-aligned (even
|
||||||
// responsibility -- unchanged by this rewrite.
|
// BYTE address) -- each 16-bit transaction covers BYTE addresses
|
||||||
|
// {addr, addr+1} as {low byte, high byte} (matches int8_memory_access.v's
|
||||||
|
// own addr[0] convention exactly, replicated here since that module is
|
||||||
|
// no longer in the datapath).
|
||||||
// ================================================================
|
// ================================================================
|
||||||
|
|
||||||
module prefetch_engine #(
|
module prefetch_engine #(
|
||||||
@@ -49,11 +50,9 @@ module prefetch_engine #(
|
|||||||
input wire rst,
|
input wire rst,
|
||||||
|
|
||||||
input wire fetch_start,
|
input wire fetch_start,
|
||||||
input wire [ADDR_WIDTH-1:0] x_addr, // BYTE address, word-aligned
|
|
||||||
input wire [ADDR_WIDTH-1:0] w_addr, // BYTE address, word-aligned
|
input wire [ADDR_WIDTH-1:0] w_addr, // BYTE address, word-aligned
|
||||||
output reg fetch_busy,
|
output reg fetch_busy,
|
||||||
output reg fetch_done, // one-cycle pulse
|
output reg fetch_done, // one-cycle pulse
|
||||||
output reg signed [DATA_WIDTH*P_IN-1:0] tile_x,
|
|
||||||
output reg signed [DATA_WIDTH*P_IN-1:0] tile_w,
|
output reg signed [DATA_WIDTH*P_IN-1:0] tile_w,
|
||||||
|
|
||||||
// ---- word-level Memory Backend Interface (matches
|
// ---- word-level Memory Backend Interface (matches
|
||||||
@@ -69,7 +68,6 @@ module prefetch_engine #(
|
|||||||
);
|
);
|
||||||
|
|
||||||
localparam ST_IDLE = 2'd0;
|
localparam ST_IDLE = 2'd0;
|
||||||
localparam ST_READ_X = 2'd1;
|
|
||||||
localparam ST_READ_W = 2'd2;
|
localparam ST_READ_W = 2'd2;
|
||||||
localparam ST_DONE = 2'd3;
|
localparam ST_DONE = 2'd3;
|
||||||
|
|
||||||
@@ -79,7 +77,6 @@ module prefetch_engine #(
|
|||||||
reg [1:0] state;
|
reg [1:0] state;
|
||||||
reg [WIW-1:0] word_idx;
|
reg [WIW-1:0] word_idx;
|
||||||
|
|
||||||
wire [ADDR_WIDTH-1:0] x_word_base = x_addr[ADDR_WIDTH-1:1];
|
|
||||||
wire [ADDR_WIDTH-1:0] w_word_base = w_addr[ADDR_WIDTH-1:1];
|
wire [ADDR_WIDTH-1:0] w_word_base = w_addr[ADDR_WIDTH-1:1];
|
||||||
|
|
||||||
always @(posedge clk) begin
|
always @(posedge clk) begin
|
||||||
@@ -106,32 +103,10 @@ module prefetch_engine #(
|
|||||||
word_idx <= 0;
|
word_idx <= 0;
|
||||||
mem_req <= 1'b1;
|
mem_req <= 1'b1;
|
||||||
mem_wr <= 1'b0;
|
mem_wr <= 1'b0;
|
||||||
mem_addr <= x_word_base;
|
mem_addr <= w_word_base;
|
||||||
mem_lb_n <= 1'b0; // both byte lanes -- fetch the whole word
|
mem_lb_n <= 1'b0; // both byte lanes -- fetch the whole word
|
||||||
mem_ub_n <= 1'b0;
|
mem_ub_n <= 1'b0;
|
||||||
state <= ST_READ_X;
|
state <= ST_READ_W;
|
||||||
end
|
|
||||||
end
|
|
||||||
|
|
||||||
ST_READ_X: begin
|
|
||||||
if (mem_ready) begin
|
|
||||||
tile_x[word_idx*16 +: 16] <= mem_rdata;
|
|
||||||
if (word_idx == WORDS_PER_TILE[WIW-1:0] - 1'b1) begin
|
|
||||||
word_idx <= 0;
|
|
||||||
mem_req <= 1'b1;
|
|
||||||
mem_wr <= 1'b0;
|
|
||||||
mem_addr <= w_word_base;
|
|
||||||
mem_lb_n <= 1'b0;
|
|
||||||
mem_ub_n <= 1'b0;
|
|
||||||
state <= ST_READ_W;
|
|
||||||
end else begin
|
|
||||||
word_idx <= word_idx + 1'b1;
|
|
||||||
mem_req <= 1'b1;
|
|
||||||
mem_wr <= 1'b0;
|
|
||||||
mem_addr <= x_word_base + word_idx + 1'b1;
|
|
||||||
mem_lb_n <= 1'b0;
|
|
||||||
mem_ub_n <= 1'b0;
|
|
||||||
end
|
|
||||||
end
|
end
|
||||||
end
|
end
|
||||||
|
|
||||||
|
|||||||
@@ -96,11 +96,21 @@ module tb;
|
|||||||
reg [ADDR_WIDTH-1:0] reg_x_base, reg_w_base, reg_result_addr;
|
reg [ADDR_WIDTH-1:0] reg_x_base, reg_w_base, reg_result_addr;
|
||||||
reg [15:0] reg_n_tiles;
|
reg [15:0] reg_n_tiles;
|
||||||
|
|
||||||
wire [N_SLOTS-1:0] slot_mem_req, slot_mem_wr;
|
// Arrays sized N_SLOTS+1 post-M10 (decisions.log DEC-0016) -- index
|
||||||
wire [ADDR_WIDTH*N_SLOTS-1:0] slot_mem_addr;
|
// N_SLOTS is the shared activation_cache's own backend port. Each
|
||||||
wire [16*N_SLOTS-1:0] slot_mem_wdata, slot_mem_rdata;
|
// index still gets its OWN independent behavioral memory (matches
|
||||||
wire [N_SLOTS-1:0] slot_mem_lb_n, slot_mem_ub_n;
|
// this testbench's own pre-existing scope: real shared-PSRAM
|
||||||
wire [N_SLOTS-1:0] slot_mem_ready;
|
// arbitration across slots is M8's job, not exercised here) --
|
||||||
|
// X data is poked ONCE into memory index N_SLOTS (the cache's own,
|
||||||
|
// single shared backing store) rather than duplicated per-slot,
|
||||||
|
// since X now genuinely flows through ONE shared path regardless
|
||||||
|
// of which slot a job lands on; W data is still poked into every
|
||||||
|
// slot's own memory (unchanged), since W is not shared.
|
||||||
|
wire [N_SLOTS:0] slot_mem_req, slot_mem_wr;
|
||||||
|
wire [ADDR_WIDTH*(N_SLOTS+1)-1:0] slot_mem_addr;
|
||||||
|
wire [16*(N_SLOTS+1)-1:0] slot_mem_wdata, slot_mem_rdata;
|
||||||
|
wire [N_SLOTS:0] slot_mem_lb_n, slot_mem_ub_n;
|
||||||
|
wire [N_SLOTS:0] slot_mem_ready;
|
||||||
|
|
||||||
dataflow_core #(
|
dataflow_core #(
|
||||||
.DATA_WIDTH(DATA_WIDTH), .P_IN(P_IN), .ACC_WIDTH(ACC_WIDTH), .ADDR_WIDTH(ADDR_WIDTH),
|
.DATA_WIDTH(DATA_WIDTH), .P_IN(P_IN), .ACC_WIDTH(ACC_WIDTH), .ADDR_WIDTH(ADDR_WIDTH),
|
||||||
@@ -118,7 +128,7 @@ module tb;
|
|||||||
|
|
||||||
genvar g;
|
genvar g;
|
||||||
generate
|
generate
|
||||||
for (g = 0; g < N_SLOTS; g = g + 1) begin : GEN_MEM
|
for (g = 0; g < N_SLOTS+1; g = g + 1) begin : GEN_MEM
|
||||||
sim_word_mem #(.ADDR_WIDTH(ADDR_WIDTH), .DEPTH(4096)) u_mem (
|
sim_word_mem #(.ADDR_WIDTH(ADDR_WIDTH), .DEPTH(4096)) u_mem (
|
||||||
.clk(clk), .rst(rst),
|
.clk(clk), .rst(rst),
|
||||||
.req(slot_mem_req[g]), .wr(slot_mem_wr[g]),
|
.req(slot_mem_req[g]), .wr(slot_mem_wr[g]),
|
||||||
@@ -143,6 +153,8 @@ module tb;
|
|||||||
else tb.GEN_MEM[0].u_mem.mem[word_addr][15:8] = val;
|
else tb.GEN_MEM[0].u_mem.mem[word_addr][15:8] = val;
|
||||||
1: if (byte_addr[0]==1'b0) tb.GEN_MEM[1].u_mem.mem[word_addr][7:0] = val;
|
1: if (byte_addr[0]==1'b0) tb.GEN_MEM[1].u_mem.mem[word_addr][7:0] = val;
|
||||||
else tb.GEN_MEM[1].u_mem.mem[word_addr][15:8] = val;
|
else tb.GEN_MEM[1].u_mem.mem[word_addr][15:8] = val;
|
||||||
|
2: if (byte_addr[0]==1'b0) tb.GEN_MEM[2].u_mem.mem[word_addr][7:0] = val; // shared activation_cache backing store (N_SLOTS index)
|
||||||
|
else tb.GEN_MEM[2].u_mem.mem[word_addr][15:8] = val;
|
||||||
default: ;
|
default: ;
|
||||||
endcase
|
endcase
|
||||||
end
|
end
|
||||||
@@ -155,6 +167,7 @@ module tb;
|
|||||||
case (slot)
|
case (slot)
|
||||||
0: peek = (byte_addr[0]==1'b0) ? tb.GEN_MEM[0].u_mem.mem[word_addr][7:0] : tb.GEN_MEM[0].u_mem.mem[word_addr][15:8];
|
0: peek = (byte_addr[0]==1'b0) ? tb.GEN_MEM[0].u_mem.mem[word_addr][7:0] : tb.GEN_MEM[0].u_mem.mem[word_addr][15:8];
|
||||||
1: peek = (byte_addr[0]==1'b0) ? tb.GEN_MEM[1].u_mem.mem[word_addr][7:0] : tb.GEN_MEM[1].u_mem.mem[word_addr][15:8];
|
1: peek = (byte_addr[0]==1'b0) ? tb.GEN_MEM[1].u_mem.mem[word_addr][7:0] : tb.GEN_MEM[1].u_mem.mem[word_addr][15:8];
|
||||||
|
2: peek = (byte_addr[0]==1'b0) ? tb.GEN_MEM[2].u_mem.mem[word_addr][7:0] : tb.GEN_MEM[2].u_mem.mem[word_addr][15:8];
|
||||||
default: peek = 8'sdx;
|
default: peek = 8'sdx;
|
||||||
endcase
|
endcase
|
||||||
end
|
end
|
||||||
@@ -193,15 +206,15 @@ module tb;
|
|||||||
rst = 0;
|
rst = 0;
|
||||||
@(posedge clk);
|
@(posedge clk);
|
||||||
|
|
||||||
// Pre-load PSRAM-equivalent memory for both slots (a job could
|
// Pre-load PSRAM-equivalent memory. W (per-slot, not shared)
|
||||||
// land on either slot, first-free, so both need the data).
|
// still needs to land in EVERY slot's own memory (a job could
|
||||||
|
// land on either slot, first-free). X (post-DEC-0016) flows
|
||||||
|
// through the ONE shared activation_cache instead -- poked
|
||||||
|
// once into memory index N_SLOTS(=2)'s backing store.
|
||||||
for (i = 0; i < 8; i = i + 1) begin
|
for (i = 0; i < 8; i = i + 1) begin
|
||||||
poke(0, 23'h10+i, 8'sd2); poke(0, 23'h20+i, 8'sd3); // node0: x=2,w=3
|
poke(2, 23'h10+i, 8'sd2); poke(0, 23'h20+i, 8'sd3); poke(1, 23'h20+i, 8'sd3); // node0: x=2,w=3
|
||||||
poke(1, 23'h10+i, 8'sd2); poke(1, 23'h20+i, 8'sd3);
|
poke(2, 23'h30+i, 8'sd1); poke(0, 23'h40+i, 8'sd1); poke(1, 23'h40+i, 8'sd1); // node1: x=1,w=1
|
||||||
poke(0, 23'h30+i, 8'sd1); poke(0, 23'h40+i, 8'sd1); // node1: x=1,w=1
|
poke(2, 23'h50+i, 8'sd1); poke(0, 23'h60+i, 8'sd5); poke(1, 23'h60+i, 8'sd5); // node2: x=1,w=5
|
||||||
poke(1, 23'h30+i, 8'sd1); poke(1, 23'h40+i, 8'sd1);
|
|
||||||
poke(0, 23'h50+i, 8'sd1); poke(0, 23'h60+i, 8'sd5); // node2: x=1,w=5
|
|
||||||
poke(1, 23'h50+i, 8'sd1); poke(1, 23'h60+i, 8'sd5);
|
|
||||||
end
|
end
|
||||||
|
|
||||||
// node0, node1: no dependencies. node2: depends on BOTH.
|
// node0, node1: no dependencies. node2: depends on BOTH.
|
||||||
|
|||||||
@@ -50,15 +50,73 @@ module tb;
|
|||||||
wire mm_result_valid, mm_result_ready;
|
wire mm_result_valid, mm_result_ready;
|
||||||
wire signed [DATA_WIDTH-1:0] mm_result_data;
|
wire signed [DATA_WIDTH-1:0] mm_result_data;
|
||||||
|
|
||||||
// ---- memory_manager <-> memory_interface (word-level Memory
|
// ---- memory_manager's own WEIGHT backend port (word-level Memory
|
||||||
// Backend Interface, post-M10 DEC-0015 -- int8_memory_access is no
|
// Backend Interface, post-M10 DEC-0015) ----
|
||||||
// longer in this datapath, see memory_manager.v's own header) ----
|
wire mm_mem_req, mm_mem_wr;
|
||||||
wire mem_req, mem_wr;
|
wire [ADDR_WIDTH-1:0] mm_mem_addr; // WORD address
|
||||||
wire [ADDR_WIDTH-1:0] mem_addr; // WORD address
|
wire [15:0] mm_mem_wdata;
|
||||||
wire [15:0] mem_wdata;
|
wire mm_mem_lb_n, mm_mem_ub_n;
|
||||||
wire mem_lb_n, mem_ub_n;
|
wire [15:0] mm_mem_rdata;
|
||||||
wire [15:0] mem_rdata;
|
wire mm_mem_ready;
|
||||||
wire mem_ready;
|
|
||||||
|
// ---- shared activation_cache (M10+, DEC-0016) -- N_SLOTS=1 here
|
||||||
|
// (a single memory_manager instance), routed through a real 2-port
|
||||||
|
// arbiter (weight port + cache port) into the SAME real
|
||||||
|
// memory_interface, mirroring dataflow_core.v/neural_multiprocessor.v's
|
||||||
|
// own real structure exactly, just scoped down to one slot. ----
|
||||||
|
wire xc_req;
|
||||||
|
wire [ADDR_WIDTH-1:0] xc_x_base;
|
||||||
|
wire [15:0] xc_tile_idx;
|
||||||
|
wire xc_ack;
|
||||||
|
wire signed [DATA_WIDTH*P_IN-1:0] xc_tile_x;
|
||||||
|
|
||||||
|
wire xc_mem_req, xc_mem_wr;
|
||||||
|
wire [ADDR_WIDTH-1:0] xc_mem_addr;
|
||||||
|
wire [15:0] xc_mem_wdata;
|
||||||
|
wire xc_mem_lb_n, xc_mem_ub_n;
|
||||||
|
wire [15:0] xc_mem_rdata;
|
||||||
|
wire xc_mem_ready;
|
||||||
|
|
||||||
|
activation_cache #(
|
||||||
|
.DATA_WIDTH(DATA_WIDTH), .P_IN(P_IN), .ADDR_WIDTH(ADDR_WIDTH), .N_SLOTS(1)
|
||||||
|
) u_xcache (
|
||||||
|
.clk(clk), .rst(rst),
|
||||||
|
.req(xc_req), .req_x_base(xc_x_base), .req_tile_idx(xc_tile_idx),
|
||||||
|
.ack(xc_ack), .tile_x_out(xc_tile_x),
|
||||||
|
.mem_req(xc_mem_req), .mem_wr(xc_mem_wr), .mem_addr(xc_mem_addr), .mem_wdata(xc_mem_wdata),
|
||||||
|
.mem_lb_n(xc_mem_lb_n), .mem_ub_n(xc_mem_ub_n),
|
||||||
|
.mem_rdata(xc_mem_rdata), .mem_ready(xc_mem_ready)
|
||||||
|
);
|
||||||
|
|
||||||
|
wire [1:0] arb2_req = {xc_mem_req, mm_mem_req};
|
||||||
|
wire [1:0] arb2_wr = {xc_mem_wr, mm_mem_wr};
|
||||||
|
wire [ADDR_WIDTH*2-1:0] arb2_addr = {xc_mem_addr, mm_mem_addr};
|
||||||
|
wire [31:0] arb2_wdata = {xc_mem_wdata, mm_mem_wdata};
|
||||||
|
wire [1:0] arb2_lb_n = {xc_mem_lb_n, mm_mem_lb_n};
|
||||||
|
wire [1:0] arb2_ub_n = {xc_mem_ub_n, mm_mem_ub_n};
|
||||||
|
wire [31:0] arb2_rdata;
|
||||||
|
wire [1:0] arb2_ready;
|
||||||
|
assign mm_mem_rdata = arb2_rdata[15:0];
|
||||||
|
assign mm_mem_ready = arb2_ready[0];
|
||||||
|
assign xc_mem_rdata = arb2_rdata[31:16];
|
||||||
|
assign xc_mem_ready = arb2_ready[1];
|
||||||
|
|
||||||
|
wire arb_m_req, arb_m_wr;
|
||||||
|
wire [ADDR_WIDTH-1:0] arb_m_addr;
|
||||||
|
wire [15:0] arb_m_wdata;
|
||||||
|
wire arb_m_lb_n, arb_m_ub_n;
|
||||||
|
wire [15:0] arb_m_rdata;
|
||||||
|
wire arb_m_ready;
|
||||||
|
|
||||||
|
slot_mem_arbiter #(.ADDR_WIDTH(ADDR_WIDTH), .N_PORTS(2)) u_arb2 (
|
||||||
|
.clk(clk), .rst(rst),
|
||||||
|
.s_req(arb2_req), .s_wr(arb2_wr), .s_addr(arb2_addr),
|
||||||
|
.s_wdata(arb2_wdata), .s_lb_n(arb2_lb_n), .s_ub_n(arb2_ub_n),
|
||||||
|
.s_rdata(arb2_rdata), .s_ready(arb2_ready),
|
||||||
|
.m_req(arb_m_req), .m_wr(arb_m_wr), .m_addr(arb_m_addr), .m_wdata(arb_m_wdata),
|
||||||
|
.m_lb_n(arb_m_lb_n), .m_ub_n(arb_m_ub_n),
|
||||||
|
.m_rdata(arb_m_rdata), .m_ready(arb_m_ready)
|
||||||
|
);
|
||||||
|
|
||||||
memory_manager #(
|
memory_manager #(
|
||||||
.DATA_WIDTH(DATA_WIDTH), .P_IN(P_IN), .ADDR_WIDTH(ADDR_WIDTH)
|
.DATA_WIDTH(DATA_WIDTH), .P_IN(P_IN), .ADDR_WIDTH(ADDR_WIDTH)
|
||||||
@@ -69,9 +127,11 @@ module tb;
|
|||||||
.operand_valid(mm_operand_valid), .operand_ready(mm_operand_ready),
|
.operand_valid(mm_operand_valid), .operand_ready(mm_operand_ready),
|
||||||
.input_data(mm_input_data), .weight_data(mm_weight_data), .tile_last(mm_tile_last),
|
.input_data(mm_input_data), .weight_data(mm_weight_data), .tile_last(mm_tile_last),
|
||||||
.result_valid(mm_result_valid), .result_ready(mm_result_ready), .result_data(mm_result_data),
|
.result_valid(mm_result_valid), .result_ready(mm_result_ready), .result_data(mm_result_data),
|
||||||
.mem_req(mem_req), .mem_wr(mem_wr), .mem_addr(mem_addr), .mem_wdata(mem_wdata),
|
.xc_req(xc_req), .xc_x_base(xc_x_base), .xc_tile_idx(xc_tile_idx),
|
||||||
.mem_lb_n(mem_lb_n), .mem_ub_n(mem_ub_n),
|
.xc_ack(xc_ack), .xc_tile_x(xc_tile_x),
|
||||||
.mem_rdata(mem_rdata), .mem_ready(mem_ready)
|
.mem_req(mm_mem_req), .mem_wr(mm_mem_wr), .mem_addr(mm_mem_addr), .mem_wdata(mm_mem_wdata),
|
||||||
|
.mem_lb_n(mm_mem_lb_n), .mem_ub_n(mm_mem_ub_n),
|
||||||
|
.mem_rdata(mm_mem_rdata), .mem_ready(mm_mem_ready)
|
||||||
);
|
);
|
||||||
|
|
||||||
// ---- real Neural Processor (M1), driven entirely by memory_manager ----
|
// ---- real Neural Processor (M1), driven entirely by memory_manager ----
|
||||||
@@ -121,9 +181,9 @@ module tb;
|
|||||||
|
|
||||||
memory_interface #(.ADDR_WIDTH(ADDR_WIDTH), .DATA_WIDTH(PSRAM_DATA_WIDTH)) u_memif (
|
memory_interface #(.ADDR_WIDTH(ADDR_WIDTH), .DATA_WIDTH(PSRAM_DATA_WIDTH)) u_memif (
|
||||||
.clk(clk), .rst(rst),
|
.clk(clk), .rst(rst),
|
||||||
.req(mem_req), .wr(mem_wr), .addr(mem_addr), .wdata(mem_wdata),
|
.req(arb_m_req), .wr(arb_m_wr), .addr(arb_m_addr), .wdata(arb_m_wdata),
|
||||||
.lb_n(mem_lb_n), .ub_n(mem_ub_n),
|
.lb_n(arb_m_lb_n), .ub_n(arb_m_ub_n),
|
||||||
.rdata(mem_rdata), .ready(mem_ready),
|
.rdata(arb_m_rdata), .ready(arb_m_ready),
|
||||||
.mem_req(pc_mem_req), .mem_wr(pc_mem_wr), .mem_addr(pc_mem_addr), .mem_wdata(pc_mem_wdata),
|
.mem_req(pc_mem_req), .mem_wr(pc_mem_wr), .mem_addr(pc_mem_addr), .mem_wdata(pc_mem_wdata),
|
||||||
.mem_lb_n(pc_mem_lb_n), .mem_ub_n(pc_mem_ub_n),
|
.mem_lb_n(pc_mem_lb_n), .mem_ub_n(pc_mem_ub_n),
|
||||||
.mem_rdata(pc_mem_rdata), .mem_ready(pc_mem_ready)
|
.mem_rdata(pc_mem_rdata), .mem_ready(pc_mem_ready)
|
||||||
|
|||||||
Reference in New Issue
Block a user