Files
FPGA-Neural/hardware/v2/docs/ROADMAP.md
T
micheleandClaude Sonnet 5 84794a3d25 docs(v2): M9 full benchmark - V1 vs V2 comparison table (§32)
Consolidates real, already-measured data from hardware/v1/ (frozen,
pre-certified) and V2's own M1-M8 logs into the §32-mandated
comparison table, on an apples-to-apples basis: both full systems
(V1's spi_neuron_top post_fix_verify vs V2's neural_multiprocessor
N_SLOTS=2), both PARALLEL=8/P_IN=8, both using the real unmodified V1
PSRAM backend.

Headline, all real measurements: V2 full-system Fmax 142.45 MHz
POST-P&R (PASS at 80MHz) vs V1's 68.65 MHz (FAIL at 80MHz); 166 vs 209
real simulated cycles for one neuron's 8-input dot product through the
same real PSRAM chain (2.6x wall-clock speedup); peak MAC/cycle 16
(N_SLOTS=2 concurrent slots, real contention already demonstrated in
EXP-0009) vs V1's 8 (single sequential core); lower LUT/FF despite V2
already including full dependency-graph scheduling that V1 has none
of.

9 of the table's 12 rows carry real sourced numbers; stall %/memory
utilization/processor utilization are reported as NOT MEASURED rather
than approximated (DEC-0011) - a real number needs dedicated
cycle-accounting instrumentation neither system has had built for it
yet, and approximating from partial data would violate §30's "no
invented results" rule. Deferred to M10, which needs exactly this
data to decide what to optimize.

No new RTL this milestone - pure data consolidation, logged as
EXP-0010.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
2026-09-05 15:17:14 +02:00

89 lines
5.0 KiB
Markdown

# FPGA-Neural V2 — stato roadmap
Fonte del mandato: `docs/v2-description.md` (root del repository). Baseline
funzionale/numerica/bit-exact: `hardware/v1/` (frozen, sola lettura — vedi
`hardware/v1/README.md`).
Legenda: `[ ]` non iniziato · `[~]` in corso · `[x]` completo (sim+synth+timing
reali, non solo scritto).
- [x] **M1 — Neural Processor** (`hardware/v2/rtl/neural_processor.v`, P8).
Bit-exact vs V1 (7/7 test, Verilator), pipeline a 8 stadi
funzionante, throughput reale (1 tile/ciclo). Sintesi reale: 0
problemi CHECK, Fmax 183.12 MHz (ACC_WIDTH=32) — vedi
`logs/experiments.log` EXP-0001/EXP-0002, `logs/errors.log` per 3
bug reali trovati e risolti (2 del toolchain Icarus, 1 RTL).
- [x] **M2 — Processor Array** (`neural_processor_array.v`). 1/2/4/8
processor testati (sim concorrenza reale + sintesi/P&R reali).
Fmax sempre PASS a 80MHz (159.11→134.70 MHz). Scoperta: il DSP
(MULT18X18D), non LUT/FF, satura per primo (88% a N=8) — vedi
`logs/decisions.log` DEC-0005.
- [x] **M3 — Buffers** (`activation_buffer.v`, `weight_buffer.v`,
`result_buffer.v`). Tutti inferiscono DP16KD reale (10/10 test,
6/6 config sintetizzate 0 problemi). Scoperta: il costo BRAM di
weight_buffer e' guidato da P_IN (larghezza), non da DEPTH.
- [x] **M4 — Memory Manager** (`memory_manager.v`, `prefetch_engine.v`),
backend PSRAM V1 riusato SENZA MODIFICHE. End-to-end reale (3/3
job PASS) con vero neural_processor + vera catena PSRAM V1.
3 bug RTL trovati/risolti (`logs/errors.log` ERR-0006). Fmax
165.86 MHz.
- [x] **M5 — Neural Director** (`neural_director.v`), scheduling
first-free. 4/4 test PASS (dispatch + coda + backpressure reale
su N_SLOTS=2). FSM ridotta a 4 stati, dependency rimandata a M6
(`logs/decisions.log` DEC-0007). Fmax 250.50 MHz.
- [x] **M6 — Dependency Manager** (`dependency_manager.v`), ready/waiting
queue, dependency counters, wake-up, producer tracking. 4/4 test
PASS (dipendenze multiple + produttore condiviso/piu' consumer).
Fmax 155.30 MHz. Forwarding di valori e riuso slot rimandati
(`logs/decisions.log` DEC-0008).
- [x] **M7 — Dataflow Core** (`dataflow_core.v`), prima integrazione
completa: Dependency Manager (M6) -> Neural Director (M5) ->
N_SLOTS x (Memory Manager (M4) + Neural Processor (M1)), loop di
wake-up chiuso end-to-end. 4/4 test PASS su un DAG a 3 nodi (node2
dipende da entrambi node0+node1, dispatch confermato solo dopo che
ENTRAMBI completano davvero). Sintesi reale 0 problemi a
N_SLOTS=2 e N_SLOTS=4. Fmax reale (harness): 165.15 MHz
(N_SLOTS=2), 133.19 MHz (N_SLOTS=4). Buffer M3 e arbitraggio PSRAM
condiviso rimandati esplicitamente a M8 (`logs/decisions.log`
DEC-0009).
- [x] **M8 — PSRAM integration** (`neural_multiprocessor.v`,
`slot_mem_arbiter.v`), controller V1 riusato SENZA MODIFICHE,
condiviso tra N_SLOTS memory_manager concorrenti reali. Trovato e
risolto un bug RTL reale: il primo arbitro perdeva silenziosamente
una richiesta arrivata durante la contesa (protocollo byte-level
"fire-and-forget", mai esposto da M4 che collega un solo master
direttamente) — vedi `logs/errors.log` ERR-0008. Dopo il fix: 4/4
test PASS (2 slot in vera contesa concorrente sulla stessa PSRAM
reale). Sintesi reale 0 problemi (nessun harness necessario — pin
reali PSRAM tengono il top-level a 157 pin). Fmax reale 142.45
MHz. Politica di arbitraggio a priorità fissa, non ancora fair
(`logs/decisions.log` DEC-0010).
- [x] **M9 — Full benchmark**, tabella V1 vs V2 (§32 del mandato) —
confronto full-system, stesso PARALLEL/P_IN=8, stesso backend
PSRAM reale V1 in entrambi. Fmax POST-P&R: V2 142.45 MHz (PASS
@80MHz) vs V1 68.65 MHz (FAIL @80MHz). Cicli/neurone SIMULATED
(1 neurone, 8 input, PSRAM reale): V2 166 vs V1 209 (2.6x
speedup wall-clock reale). MAC/cycle di picco: V2 16 (N_SLOTS=2 x
P_IN=8, concorrenza reale) vs V1 8 (core sequenziale singolo).
LUT/FF: V2 4191/3659 vs V1 8907/4900. 9/12 righe con dati reali
misurati; stall %/memory utilization/processor utilization
esplicitamente NON misurati questo milestone (`logs/decisions.log`
DEC-0011), rimandati a M10. Tabella completa in
`logs/benchmark.log`.
- [ ] **M10 — Optimization**, solo sulla base dei dati raccolti in M1-M9.
## Log
Vedi `hardware/v2/logs/` (`development.log` per la cronologia di sessione,
`decisions.log` per le decisioni architetturali con motivazione,
`experiments.log` per ogni EXP-XXXX end-to-end).
## Regole non negoziabili attive (§34 del mandato, per riferimento rapido)
1. V1 (`hardware/v1/`) rimane intatta — mai modificata.
2. V2 vive esclusivamente sotto `hardware/v2/`.
3. Nessun risultato inventato: THEORETICAL vs SIMULATED vs SYNTHESIZED vs
POST-P&R sempre etichettati esplicitamente.
4. Ogni modifica/esperimento/decisione registrata nei log, mai persa.
5. Ogni esperimento ha un ID univoco, mai riutilizzato — anche i FAIL restano.