Consolidates real, already-measured data from hardware/v1/ (frozen, pre-certified) and V2's own M1-M8 logs into the §32-mandated comparison table, on an apples-to-apples basis: both full systems (V1's spi_neuron_top post_fix_verify vs V2's neural_multiprocessor N_SLOTS=2), both PARALLEL=8/P_IN=8, both using the real unmodified V1 PSRAM backend. Headline, all real measurements: V2 full-system Fmax 142.45 MHz POST-P&R (PASS at 80MHz) vs V1's 68.65 MHz (FAIL at 80MHz); 166 vs 209 real simulated cycles for one neuron's 8-input dot product through the same real PSRAM chain (2.6x wall-clock speedup); peak MAC/cycle 16 (N_SLOTS=2 concurrent slots, real contention already demonstrated in EXP-0009) vs V1's 8 (single sequential core); lower LUT/FF despite V2 already including full dependency-graph scheduling that V1 has none of. 9 of the table's 12 rows carry real sourced numbers; stall %/memory utilization/processor utilization are reported as NOT MEASURED rather than approximated (DEC-0011) - a real number needs dedicated cycle-accounting instrumentation neither system has had built for it yet, and approximating from partial data would violate §30's "no invented results" rule. Deferred to M10, which needs exactly this data to decide what to optimize. No new RTL this milestone - pure data consolidation, logged as EXP-0010. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
169 lines
10 KiB
Plaintext
169 lines
10 KiB
Plaintext
# V2 benchmark log -- solo append, mai troncato/sovrascritto (vedi README.md)
|
|
# Nessuna entry ancora -- popolato incrementalmente man mano che avanza lo sviluppo V2.
|
|
|
|
[2026-09-05] M1 single Neural Processor, isolated (no array/director/
|
|
memory manager yet -- system-level numbers deferred to M9)
|
|
|
|
| Config | Fmax (POST-P&R) | LUT | FF | DSP | BRAM |
|
|
|----------------------|------------------|-----|-----|-----|------|
|
|
| P_IN=8, ACC_WIDTH=32 | 183.12 MHz | 55 | 533 | 8 | 0 |
|
|
| P_IN=8, ACC_WIDTH=24 | 176.21 MHz | 49 | 509 | 8 | 0 |
|
|
|
|
Reference (V1, hardware/v1/synthesis/p8/, isolated neuron_parallel,
|
|
PARALLEL=8, N_INPUTS=256): 61.71 MHz POST-P&R.
|
|
|
|
All Fmax figures above are POST-P&R (real nextpnr-ecp5), not
|
|
theoretical or simulated-only. MAC/cycle, cycles/neuron, neurons/s,
|
|
stall %, effective MAC/s: not yet meaningful at this milestone (single
|
|
isolated processor, no streaming benchmark harness yet -- deferred to
|
|
M2 once neural_processor_array.v exists and a real workload can be
|
|
timed end-to-end).
|
|
|
|
[2026-09-05] M2 Neural Processor Array, N_PROCESSORS sweep (P_IN=8,
|
|
ACC_WIDTH=32 each; synthesized via the timing harness, see errors.log
|
|
ERR-0005 for why)
|
|
|
|
| N_PROCESSORS | Fmax (POST-P&R) | LUT | FF | DSP (MULT18X18D) | DSP % of 72 |
|
|
|--------------|------------------|-----|------|-------------------|-------------|
|
|
| 1 | 159.11 MHz | 59 | 409 | 8 | 11% |
|
|
| 2 | 149.59 MHz | 106 | 786 | 16 | 22% |
|
|
| 4 | 151.01 MHz | 207 | 1540 | 32 | 44% |
|
|
| 8 | 134.70 MHz | 374 | 3048 | 64 | 88% |
|
|
|
|
All figures POST-P&R (real nextpnr-ecp5), all PASS at the 80MHz
|
|
target. LUT4 utilization stays under 6% of the device even at N=8;
|
|
DSP is the binding resource (see decisions.log DEC-0005), reaching 88%
|
|
at N=8 -- N_PROCESSORS=9 would already exceed the LFE5U-45F's 72
|
|
MULT18X18D budget at P_IN=8. Theoretical MAC/cycle (THEORETICAL, not
|
|
yet measured end-to-end -- no real workload/benchmark harness exists
|
|
until M9): N_PROCESSORS * P_IN MACs/cycle when all processors are
|
|
simultaneously streaming tiles (8, 16, 32, 64 for N=1/2/4/8 -- verified
|
|
achievable in principle by EXP-0003's concurrent/staggered simulation,
|
|
not yet measured as a sustained throughput number).
|
|
|
|
[2026-09-05] M3 buffers -- BRAM cost vs DEPTH (real Yosys synth_ecp5)
|
|
|
|
| Module | DEPTH | DP16KD | LUT4 | FF |
|
|
|--------------------|-------|--------|------|-----|
|
|
| activation_buffer | 4096 | 2 | 37 | 30 |
|
|
| activation_buffer | 256 | 1 | 21 | 26 |
|
|
| weight_buffer | 512 | 2 | 88 | 139 |
|
|
| weight_buffer | 64 | 2 | 73 | 136 |
|
|
| result_buffer | 4096 | 2 | 37 | 30 |
|
|
| result_buffer | 256 | 1 | 21 | 26 |
|
|
|
|
weight_buffer's DP16KD count is flat across an 8x depth reduction --
|
|
its 64-bit TILE_WIDTH (P_IN=8 * DATA_WIDTH=8), not DEPTH, determines
|
|
BRAM count for this module. activation_buffer/result_buffer (byte-
|
|
wide) scale as expected with depth. Confirms §14's warning literally:
|
|
"non assumere che buffer piu' grandi siano automaticamente migliori"
|
|
-- here, smaller was not cheaper either, because depth was the wrong
|
|
lever for this specific buffer's cost.
|
|
|
|
[2026-09-05] M4 Memory Manager + Prefetch Engine (standalone resource
|
|
count; Fmax via timing harness -- see errors.log ERR-0005)
|
|
|
|
| Module | Fmax (POST-P&R) | LUT | FF | DSP | CCU2C |
|
|
|----------------------------------|------------------|-----|-----|-----|-------|
|
|
| memory_manager + prefetch_engine | 165.86 MHz | 851 | 789 | 0 | 108 |
|
|
|
|
End-to-end (real PSRAM + real neural_processor, SIMULATED only, no
|
|
system-level P&R yet -- deferred to M7/M9): 3-tile job = 446 cycles,
|
|
1-tile job = 166 cycles, 5-tile job = 728 cycles. ~140-150 cycles/tile,
|
|
dominated by psram_model.v's real ~70ns TAA access latency, not by
|
|
memory_manager's own control overhead (its bank-swap turnaround is
|
|
documented as a fixed +1 cycle/tile in decisions.log DEC-0006, a small
|
|
fraction of the ~140-cycle PSRAM-dominated total).
|
|
|
|
[2026-09-05] M5 Neural Director (standalone resource count; Fmax via
|
|
timing harness -- see errors.log ERR-0005)
|
|
|
|
| Module | Fmax (POST-P&R) | LUT | FF | DSP | CCU2C |
|
|
|------------------|------------------|-----|-----|-----|-------|
|
|
| neural_director (N_SLOTS=4) | 250.50 MHz | 382 | 366 | 0 | 4 |
|
|
|
|
[2026-09-05] M6 Dependency Manager (real standalone synthesis + P&R,
|
|
no harness needed)
|
|
|
|
| Module | Fmax (POST-P&R) | LUT | FF | DSP | CCU2C |
|
|
|----------------------|------------------|-----|-----|-----|-------|
|
|
| dependency_manager (N_NODES=16) | 155.30 MHz | 763 | 474 | 0 | 0 |
|
|
|
|
[2026-09-05] M7 Dataflow Core (full M1-M6 integration; resources via
|
|
real standalone synthesis, Fmax via timing harness -- see errors.log
|
|
ERR-0005)
|
|
|
|
| Module (config) | Fmax (POST-P&R) | LUT4 | CCU2C | FF | DSP | BRAM |
|
|
|-------------------------------|------------------|------|-------|------|-----|------|
|
|
| dataflow_core (N_SLOTS=2) | 165.15 MHz | 2127 | 248 | 2505 | 16 | 0 |
|
|
| dataflow_core (N_SLOTS=4) | 133.19 MHz | 3953 | 500 | 4688 | 32 | 0 |
|
|
|
|
DSP budget on the LFE5U-45F is 72 MULT18X18D total: N_SLOTS=4 already
|
|
uses 32/72 (44%), consistent with DEC-0005's finding that DSP, not
|
|
LUT/FF, is the first resource to saturate as concurrency grows (M2's
|
|
own N_PROCESSORS=8 measurement: 88%). BRAM=0 on both is expected --
|
|
M3's buffers are not wired into dataflow_core yet (DEC-0009).
|
|
|
|
[2026-09-05] M8 Neural Multiprocessor top (real standalone synthesis +
|
|
P&R, no harness needed -- real PSRAM pins keep the bare top-level
|
|
pin count at 157, under the TRELLIS_IO budget)
|
|
|
|
| Module (config) | Fmax (POST-P&R) | LUT4 | CCU2C | FF | DSP | BRAM |
|
|
|--------------------------------------|------------------|------|-------|------|-----|------|
|
|
| neural_multiprocessor (N_SLOTS=2) | 142.45 MHz | 3145 | 388 | 3659 | 16 | 0 |
|
|
|
|
Compare to M7's dataflow_core alone (N_SLOTS=2): 165.15 MHz / LUT4=2127
|
|
/ CCU2C=248 / FF=2505 / DSP=16. Adding the real V1 PSRAM chain +
|
|
slot_mem_arbiter costs ~1000 LUT4/140 CCU2C/1150 FF and drops Fmax by
|
|
~23 MHz (165.15 -> 142.45) -- both real, measured costs of real PSRAM
|
|
integration, not assumed.
|
|
|
|
[2026-09-05] M9 -- Confronto finale V1 vs V2 (docs/v2-description.md
|
|
§32). Basi di confronto: entrambi i sistemi COMPLETI (full-system,
|
|
non moduli isolati), stessa larghezza di dot-product per neurone
|
|
(PARALLEL=8 / P_IN=8), stesso backend PSRAM reale V1 non modificato
|
|
in entrambi i casi.
|
|
V1 = hardware/v1's spi_neuron_top (post_fix_verify synthesis,
|
|
PARALLEL=8) + neuron_memory_tb.v TEST 5 (1 neuron, N_INPUTS
|
|
ridotto a 8, catena PSRAM reale) -- entrambi dati gia'
|
|
certificati/congelati in hardware/v1/.
|
|
V2 = hardware/v2/rtl/neural_multiprocessor.v (N_SLOTS=2, P_IN=8,
|
|
M8) -- EXP-0009 (1 tile/8 input via memory_manager standalone,
|
|
EXP-0005) + EXP-0009 stesso (catena PSRAM reale condivisa).
|
|
|
|
| Metrica | V1 | V2 | Classificazione |
|
|
|------------------------|---------------------------|------------------------------|-------------------|
|
|
| Fmax | 68.65 MHz (FAIL @80MHz) | 142.45 MHz (PASS @80MHz) | POST-P&R (reale, nextpnr-ecp5, entrambi full-system) |
|
|
| LUT (Total LUT4s) | 8907 | 4191 | POST-P&R (nextpnr "Total LUT4s", stessa metrica per entrambi) |
|
|
| FF (Total DFFs) | 4900 | 3659 | POST-P&R |
|
|
| DSP (MULT18X18D) | 16 | 16 | POST-P&R |
|
|
| BRAM (DP16KD) | 2 | 0 (M3 non ancora collegato, DEC-0009) | POST-P&R |
|
|
| MAC/cycle (picco) | 8 (PARALLEL=8, core singolo sequenziale) | 16 (N_SLOTS=2 x P_IN=8, concorrenti) | THEORETICAL (parametri architetturali noti, non ancora un picco sostenuto misurato con contesa reale su entrambi gli slot) |
|
|
| cycles/neuron | 209 (1 neurone, 8 input reali, PSRAM reale -- hardware/v1/docs/FPGA-NeuralNetwork-Engine.md, TEST 5 neuron_memory_tb.v, gia' certificato) | 166 (1 tile/8 input, PSRAM reale -- EXP-0005) | SIMULATED (Verilator per V2; V1 dato gia' certificato in hardware/v1/, non ri-simulato in questa sessione) |
|
|
| neurons/s | 328,469 (Fmax POST-P&R / cycles SIMULATED) | 858,133 (idem) | DERIVATO (Fmax POST-P&R reale x cycles/neuron SIMULATED reale -- non un numero misurato direttamente in un'unica prova, ma calcolato da due misure reali indipendenti, entrambe citate) |
|
|
| stall % | NON MISURATO questo milestone | NON MISURATO questo milestone | -- vedi decisions.log DEC-0011 |
|
|
| memory utilization | NON MISURATO questo milestone | NON MISURATO questo milestone | -- vedi decisions.log DEC-0011 |
|
|
| processor utilization | NON MISURATO questo milestone | NON MISURATO questo milestone | -- vedi decisions.log DEC-0011 |
|
|
| effective MAC/s | 2.63 MMAC/s (8 MAC/neurone / 209 cicli x 68.65MHz) | 6.87 MMAC/s (8 MAC/neurone / 166 cicli x 142.45MHz) | DERIVATO (stessa base di neurons/s) |
|
|
| effective MAC/s (picco teorico) | 549.2 MMAC/s (8 x 68.65MHz) | 2279.2 MMAC/s (16 x 142.45MHz) | THEORETICAL |
|
|
|
|
Osservazioni (tutte da dati reali sopra, nessun numero inventato):
|
|
- V2 e' 2.6x piu' veloce in wall-clock per un singolo neurone/8 input
|
|
attraverso la catena PSRAM reale (1.165us vs 3.044us), pur usando lo
|
|
STESSO backend PSRAM V1 non modificato -- il guadagno viene sia da
|
|
meno cicli (166 vs 209, pipeline M1 piu' efficiente) sia da un Fmax
|
|
POST-P&R piu' che doppio (142.45 vs 68.65 MHz).
|
|
- V1's full system FALLISCE il target 80MHz (68.65 MHz); V2's full
|
|
system lo supera con margine (142.45 MHz) -- confermato da vero
|
|
nextpnr-ecp5 place&route su entrambi.
|
|
- Il vantaggio "MAC/cycle di picco" di V2 (16 vs 8) viene dalla
|
|
concorrenza reale a livello di sistema (N_SLOTS=2, EXP-0009 lo
|
|
dimostra con vera contesa PSRAM tra 2 slot), non da un core piu'
|
|
largo -- V1 e' strutturalmente un acceleratore sequenziale (un solo
|
|
neuron_parallel attivo alla volta), esattamente il limite che
|
|
l'intero mandato V2 (docs/v2-description.md, titolo) si propone di
|
|
superare.
|
|
- BRAM=0 per V2 riflette una scelta di scope esplicita (M3 non ancora
|
|
collegato, DEC-0009), non un vantaggio architetturale reale -- un
|
|
confronto onesto lo nota piuttosto che nasconderlo.
|