test(v2): final benchmark campaign - real end-to-end characterization (EXP-0014)
Post-M10, user-requested final benchmark campaign: 6 realistic workloads (16-256 independent neurons in a shared-input dense-layer shape, a random-seeded 2-layer network with real cross-node PSRAM forwarding, and a 6-node 2-hop dependency diamond) x 4 concurrency levels (N_SLOTS=1/2/4/8) through the real, full neural_multiprocessor system (real V1 PSRAM chain, real slot_mem_arbiter). 24/24 runs PASS bit-exact against a software golden model (11,520 individual neuron/ node checks, zero mismatches). Three real bugs found and fixed during the campaign itself (ERR-0009): 1. neural_director.v (M5) had a real RTL bug at N_SLOTS=1 ($clog2(1)=0 makes a replication expression illegal) - never caught because M5-M10 only ever tested N_SLOTS=2/4/8. Fixed with a width-agnostic '0 literal; M5's own testbench re-verified unaffected. 2/3. Two testbench sizing bugs in tb_benchmark_suite.v itself (psram_model DEPTH too small for the Large workload's address range; N_NODES too small for the Stress workload's node-id range, causing a real deadlock via node-id wraparound colliding with an already-DISPATCHED node - a real, honest consequence of DEC-0008's own "no node-slot reclamation" design choice). Headline finding: real parallel scaling is essentially flat beyond N_SLOTS=2 - the single shared PSRAM port saturates at ~91% utilization regardless of slot count, so memory-bound workloads gain only 1.05-1.06x real speedup from N=1 to N=8. Once real POST-P&R Fmax degradation is also factored in, N_SLOTS=4 is measurably 21% SLOWER in real wall-clock time than N_SLOTS=1 for the largest workload tested. N_SLOTS=2 is recommended as the default (DEC-0014, superseding DEC-0012's resource-only "N_SLOTS=8 ceiling" framing for general use). Full 21-section report (every number classified THEORETICAL/ SIMULATED/POST-P&R MEASURED/DERIVED, per the user's own methodology requirements): hardware/v2/docs/benchmarks/ final-benchmark.md Logged: simulation/synthesis/timing/benchmark/decisions (DEC-0014)/ experiments (EXP-0014)/errors (ERR-0009)/development.log, ROADMAP.md updated. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
This commit is contained in:
@@ -284,3 +284,30 @@ next_action: nessuna prevista dal mandato (§33 termina a M10) -- la
|
||||
P_IN<8 per N_SLOTS ancora piu' alto (DEC-0012 alternativa #2), riuso
|
||||
dei buffer M3 come cache condivisa (DEC-0009), strumentazione
|
||||
stall%/utilization anche lato V1 (DEC-0011).
|
||||
|
||||
[2026-09-05] Final Benchmark Campaign (post-M10, richiesta diretta
|
||||
utente, non parte del roadmap §33)
|
||||
reason: caratterizzare le prestazioni reali end-to-end di V2 prima di
|
||||
decidere il numero finale di unita' parallele (N_SLOTS) e prima di
|
||||
scrivere il datasheet V2. Nessun test funzionale isolato, nessuna
|
||||
ottimizzazione prima della misura, come esplicitamente richiesto.
|
||||
result: 6 workload realistici (16-256 neuroni indipendenti + un layer
|
||||
multilivello con dati random seed loggato + un DAG a diamante a 2
|
||||
hop) verificati bit-exact su 4 configurazioni (N_SLOTS=1/2/4/8) --
|
||||
24/24 PASS dopo aver risolto 3 problemi reali (errors.log ERR-0009).
|
||||
Scoperta principale: lo scaling parallelo reale e' sostanzialmente
|
||||
PIATTO per i workload memory-bound (1.05-1.06x da N=1 a N=8) -- il
|
||||
vero collo di bottiglia e' la porta PSRAM condivisa (~91% utilizzo
|
||||
indipendentemente da N_SLOTS>=2), non il numero di processori.
|
||||
Considerando anche il vero Fmax POST-P&R (152.46/142.45/113.38 MHz
|
||||
per N=1/2/4), N_SLOTS=4 e' realmente PIU' LENTO del 21% in
|
||||
wall-clock rispetto a N_SLOTS=1 per il workload Stress. Trovato
|
||||
squilibrio reale di scheduling (slot a indice basso fanno quasi
|
||||
tutto il lavoro).
|
||||
errors: vedi errors.log ERR-0009 (1 bug RTL reale in neural_director.v,
|
||||
mai testato prima a N_SLOTS=1; 2 bug nel testbench stesso).
|
||||
decision: vedi decisions.log DEC-0014 -- N_SLOTS=2 raccomandato come
|
||||
configurazione di default, non N_SLOTS=8 (DEC-0012 resta valido come
|
||||
tetto DSP mafisico, non come raccomandazione d'uso generale).
|
||||
next_action: datasheet V2 in stile professionale (richiesta utente),
|
||||
ora sbloccato dalla decisione su N_SLOTS.
|
||||
|
||||
Reference in New Issue
Block a user