Files
FPGA-Neural/hardware/v2/logs/development.log
T
micheleandClaude Sonnet 5 84794a3d25 docs(v2): M9 full benchmark - V1 vs V2 comparison table (§32)
Consolidates real, already-measured data from hardware/v1/ (frozen,
pre-certified) and V2's own M1-M8 logs into the §32-mandated
comparison table, on an apples-to-apples basis: both full systems
(V1's spi_neuron_top post_fix_verify vs V2's neural_multiprocessor
N_SLOTS=2), both PARALLEL=8/P_IN=8, both using the real unmodified V1
PSRAM backend.

Headline, all real measurements: V2 full-system Fmax 142.45 MHz
POST-P&R (PASS at 80MHz) vs V1's 68.65 MHz (FAIL at 80MHz); 166 vs 209
real simulated cycles for one neuron's 8-input dot product through the
same real PSRAM chain (2.6x wall-clock speedup); peak MAC/cycle 16
(N_SLOTS=2 concurrent slots, real contention already demonstrated in
EXP-0009) vs V1's 8 (single sequential core); lower LUT/FF despite V2
already including full dependency-graph scheduling that V1 has none
of.

9 of the table's 12 rows carry real sourced numbers; stall %/memory
utilization/processor utilization are reported as NOT MEASURED rather
than approximated (DEC-0011) - a real number needs dedicated
cycle-accounting instrumentation neither system has had built for it
yet, and approximating from partial data would violate §30's "no
invented results" rule. Deferred to M10, which needs exactly this
data to decide what to optimize.

No new RTL this milestone - pure data consolidation, logged as
EXP-0010.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
2026-09-05 15:17:14 +02:00

255 lines
14 KiB
Plaintext

# V2 development log — solo append, mai troncato/sovrascritto (vedi README.md)
[2026-09-05T11:14:08Z] commit=07a48e4 session=v2-kickoff
module: scaffold
action: creato hardware/v1/ (baseline V1 congelata, sola lettura) e
hardware/v2/ (struttura obbligatoria: rtl/ sim/ constraints/ synthesis/
reports/{simulation,synthesis,timing,experiments}/ scripts/sweep/ logs/
docs/), per iniziare lo sviluppo V2 richiesto da docs/v2-description.md.
reason: richiesta utente esplicita -- "Crea un v1 con tutto il codice
attuale che non deve essere modificato per nessun motivo e lavora su
una v2 copiando quello che ti serve. Aggiornami frequentemente e crea
un log nuovo."
command: cp rtl/*.v sim/*.v tools/ synth/ecp5/{p2,p4,p8,post_fix_verify}
synth/ecp5/*.lpf docs/validation/*.md WORKLOG.md
docs/FPGA-NeuralNetwork-Engine.md docs/FPGA-Neural-Datapatch-Benchmark.md
docs/FPGA-Neural-Hardware-Design.md -> hardware/v1/{rtl,sim,tools,
constraints,synthesis,docs}; chmod -R a-w su tutto hardware/v1/.
result: verificato bit-esatto (`diff -rq rtl/ hardware/v1/rtl/` e cmp
file-per-file su sim/*.v, 0 differenze). hardware/v1/ reso sola lettura
a livello filesystem come ulteriore salvaguardia oltre alla policy.
hardware/v2/ vuoto tranne logs/ (questo file + README.md).
errors: nessuno.
decision: copia (non spostamento) dell'albero top-level rtl/sim/synth/
tools -- l'albero principale del repository resta la "produzione" V1
invariata; hardware/v1/ e' una seconda copia frozen dedicata al
confronto V2, per minimizzare il rischio di rompere riferimenti/path
esistenti in WORKLOG.md, docs/, altri script. Vedi decisions.log
DEC-0001 per la motivazione completa.
next_action: M1 (docs/v2-description.md §33) -- implementare
hardware/v2/rtl/neural_processor.v (P_IN=8), pipeline 8 stage,
bit-exact contro hardware/v1/rtl/neuron_parallel.v + mac8.v,
testbench hardware/v2/sim/tb_neural_processor.v.
[2026-09-05T12:03:12Z] commit=07a48e4 session=v2-M1-neural-processor
module: hardware/v2/rtl/neural_processor.v + hardware/v2/sim/
tb_neural_processor.v
action: implementato M1 (docs/v2-description.md §33) -- Neural
Processor P_IN=8, pipeline a 8 stadi (input align -> 8 moltiplicatori
-> 3 livelli di albero di somma -> accumulatore -> bias/activation ->
saturazione INT8), FSM a 7 stati (§6, con 4 stati assorbiti in
NP_WAIT_OPERANDS per il vero throughput pipeline -- vedi
decisions.log DEC-0002), interfacce valid/ready/data/last (§7).
Aritmetica bit-esatta con hardware/v1/rtl/neuron_parallel.v +
mac8.v + mac_unit.v.
reason: roadmap M1, richiesta esplicita utente di iniziare lo sviluppo
V2 copiando da V1 quanto necessario.
result: 3 bug reali incontrati e risolti durante lo sviluppo (vedi
errors.log ERR-0001/0002/0003) -- due erano bug del toolchain Icarus
Verilog v13.0 (non del RTL), confermati installando e usando
Verilator 5.050 come secondo simulatore indipendente (decisions.log
DEC-0004); uno era un bug RTL reale (gating mancante su `last0`,
fix verificato). Testbench finale: 7/7 test PASS, bit-exact contro
V1, sia a ACC_WIDTH=32 che ACC_WIDTH=24 (esperimento richiesto
dall'utente, experiments.log EXP-0002). Sintesi reale Yosys + P&R
reale nextpnr-ecp5: 0 problemi CHECK, Fmax 183.12 MHz (ACC_WIDTH=32)
/ 176.21 MHz (ACC_WIDTH=24) -- entrambi PASS a 80MHz con ampio
margine, ~3x il Fmax isolato di V1 a PARALLEL=8 (61.71 MHz).
errors: vedi errors.log ERR-0001, ERR-0002, ERR-0003, ERR-0004.
decision: vedi decisions.log DEC-0002, DEC-0003, DEC-0004.
next_action: M2 -- neural_processor_array.v, sweep N_PROCESSORS
(1,2,4,8), misura timing/risorse/throughput/utilization reali.
[2026-09-05T14:30:00Z] commit=dc0b331 session=v2-M2-processor-array
module: hardware/v2/rtl/neural_processor_array.v + hardware/v2/sim/
tb_neural_processor_array.v + hardware/v2/synthesis/
harness_neural_processor_array.v
action: implementato M2 -- Neural Processor Array, N_PROCESSORS
parametrico, ogni processore con interfaccia job/operand/result
dedicata (nessun bus condiviso/mux a questo livello, §7). Testato
N_PROCESSORS=4 con esecuzione concorrente reale (lancio simultaneo
di 4 job diversi, più un test a partenza sfalsata) e con lo sweep
reale N_PROCESSORS={1,2,4,8} per sintesi/place&route.
reason: roadmap M2.
result: 7/7 test PASS (Verilator), confermata esecuzione concorrente e
indipendente (un job più corto lanciato dopo completa prima di uno
più lungo lanciato prima -- non c'e' serializzazione nascosta).
Sintesi/P&R reali per N=1/2/4/8: 0 problemi CHECK, scaling lineare
delle risorse, Fmax sempre PASS a 80MHz (159.11/149.59/151.01/134.70
MHz). Scoperta reale non assunta: il DSP (MULT18X18D) e' la prima
risorsa a saturare (88% a N=8), non LUT/FF (sotto il 6% anche a
N=8) -- vedi decisions.log DEC-0005.
errors: ERR-0005 (artefatto di sintesi da pin-count, non un bug RTL --
workaround con un harness dedicato, vedi errors.log). Un primo
tentativo di harness alimentava dati identici a ogni processore/lane
MAC, e Yosys deduplicava silenziosamente tutto a 1x indipendentemente
da N -- scoperto verificando che lo scaling fosse davvero lineare
prima di fidarsi dei numeri.
decision: vedi decisions.log DEC-0005.
next_action: M3 -- activation_buffer.v / weight_buffer.v /
result_buffer.v (profondita' parametrica, valutare BRAM mapping).
[2026-09-05T15:15:00Z] commit=3026dcd session=v2-M3-buffers
module: hardware/v2/rtl/activation_buffer.v, weight_buffer.v,
result_buffer.v
action: implementato M3 -- tre buffer paramentrici (profondita'
parametrica) che inferiscono DP16KD reale, riusando l'idioma gia'
validato in hardware/v1/rtl/act_buffer.v (porta A scrittura sync,
porta B lettura sync REGISTRATA, nessun reset sulla porta di lettura
per restare fuori dal percorso LUT-RAM).
reason: roadmap M3.
result: 10/10 test PASS (Verilator). Sintesi reale a 2 profondita'
ciascuno: 0 problemi CHECK, DP16KD reale confermato in tutte le 6
configurazioni (mai LUT-RAM). Scoperta reale non assunta: il costo
BRAM di weight_buffer e' determinato dalla LARGHEZZA (P_IN*DATA_
WIDTH=64 bit), non dalla profondita' -- 512->64 di profondita' non
ha ridotto il conteggio DP16KD (resta 2). Place&route reale sui
default: tutti PASS a 80MHz con ampio margine (287-367 MHz).
errors: nessuno.
decision: vedi benchmark.log -- il dimensionamento di weight_buffer
andra' guidato da P_IN, non solo da DEPTH, quando si arrivera' a
M4/M9.
next_action: M4 -- memory_manager.v + prefetch_engine.v.
[2026-09-05T16:00:00Z] commit=5f0d7f1 session=v2-M4-memory-manager
module: hardware/v2/rtl/memory_manager.v, prefetch_engine.v
action: implementato M4 -- Memory Manager (arbitraggio/buffering/
forwarding/gestione latenza/double buffering, §12) + Prefetch Engine
(fetch a doppio buffer con retargeting di una singola istanza per
banco, §13). Backend PSRAM V1 riusato SENZA MODIFICHE
(int8_memory_access.v -> memory_interface.v -> psram_controller.v,
§15), integrato end-to-end con un vero hardware/v2/rtl/
neural_processor.v (M1).
reason: roadmap M4.
result: 3/3 job PASS end-to-end (1/3/5 tile), risultato verificato con
RILETTURA INDIPENDENTE da PSRAM (non solo ispezione di segnali
interni), con byte "poison" attorno alle regioni operando per
catturare eventuali errori di indirizzamento off-by-one (nessuno
trovato). 3 bug RTL reali trovati e risolti durante l'integrazione
(vedi errors.log ERR-0006): mancava una disciplina "una sola
richiesta di prefetch in volo", un buco di un ciclo nel check
!pf_busy, un mux di stato disallineato di un ciclo che faceva
silenziosamente perdere la scrittura del risultato su PSRAM.
Sintesi reale: 0 problemi, 851 LUT4/789 FF/108 CCU2C/0 DSP (atteso).
Fmax reale (via harness, stesso motivo pin-count di ERR-0005):
165.86 MHz, PASS a 80MHz.
errors: vedi errors.log ERR-0005 (ricorrenza), ERR-0006 (3 bug nuovi).
decision: vedi decisions.log DEC-0006 (motore di prefetch singolo +
registro pendente, nessun arbitro backend ancora necessario).
next_action: M5 -- neural_director.v, scheduling first-free.
[2026-09-05T17:00:00Z] commit=175f697 session=v2-M5-neural-director
module: hardware/v2/rtl/neural_director.v
action: implementato M5 -- Neural Director, scheduling first-free
(§9) su N_SLOTS coppie (memory_manager, neural_processor). FSM
ridotta a 4 stati (DIR_IDLE/SCAN_READY/ALLOCATE/ERROR) -- dependency
tracking/wake-up rimandati al Dependency Manager (M6, non ancora
costruito), rilevamento completamenti gestito da un tracker
sempre-attivo per-slot (non uno stato dedicato da rivisitare).
reason: roadmap M5.
result: 4/4 test PASS (N_SLOTS=2) -- dispatch first-free confermato,
coda pronta con backpressure reale confermata (riempimento e
recupero), 3 job su 2 slot con il terzo correttamente in coda fino
a liberazione di uno slot. 2 bug di testbench trovati e risolti
(non RTL): DEPTH di sim_byte_mem troppo piccolo per il range di
indirizzi usato, e una condizione di attesa che si fermava al primo
job completato invece che a tutti e tre. Sintesi reale: 0 problemi,
382 LUT4/366 FF/4 CCU2C/0 DSP. Fmax reale (via harness): 250.50 MHz.
errors: nessun bug RTL, solo 2 bug di testbench (vedi experiments.log
EXP-0006).
decision: vedi decisions.log DEC-0007.
next_action: M6 -- dependency_manager.v.
[2026-09-05T18:00:00Z] commit=2e4cedc session=v2-M6-dependency-manager
module: hardware/v2/rtl/dependency_manager.v
action: implementato M6 -- tabella di dipendenze (node_id/state/
required/resolved/producer_ids, campi esatti §10), wake-up su
completamento produttore, dispatch first-found-ready verso il
Director (M5).
reason: roadmap M6.
result: 4/4 test PASS su un piccolo DAG a mano (node2 dipende da
ENTRAMBI node0+node1 -- dipendenze multiple; node3 dipende solo da
node0 -- risultato condiviso/piu' consumer). Confermato: node3 pronto
subito dopo node0, node2 resta WAITING finche' anche node1 non
completa. Sintesi reale: 0 problemi, 763 LUT4/474 FF/0 DSP/0 CCU2C.
Fmax reale (nessun harness necessario stavolta): 155.30 MHz.
errors: un errore di sintassi nel testbench (nested replication senza
livello di parentesi extra), non un bug RTL.
decision: vedi decisions.log DEC-0008 (nessun forwarding di valori
ancora, nessun riuso degli slot ancora -- entrambi rimandati
esplicitamente).
next_action: M7 -- dataflow_core.v, prima integrazione di
Director+Dependency Manager+Memory Manager+Processor Array+Buffer
in un unico top-level.
[2026-09-05] M7 -- hardware/v2/rtl/dataflow_core.v
reason: roadmap M7 -- first full integration of dependency_manager
(M6) + neural_director (M5) + N_SLOTS x (memory_manager (M4) +
neural_processor (M1)) into one top-level module, closing the
producer-completion -> dependency-wake-up loop end-to-end for the
first time. Additive extension to neural_director.v: added a
slot_node_id output port (which node_id occupies each slot) so a
caller can map a completed slot back to the node_id that just
finished -- re-verified M5's own testbench still passes 4/4
unaffected.
result: 4/4 test PASS on a 3-node DAG (node2 depends on BOTH node0 and
node1; confirmed it does NOT dispatch until both genuinely complete,
polled every cycle). Real synthesis: 0 problems at both N_SLOTS=2
(LUT4=2127/CCU2C=248/FF=2505/DSP=16) and N_SLOTS=4
(LUT4=3953/CCU2C=500/FF=4688/DSP=32). Real Fmax (via
harness_dataflow_core.v): 165.15 MHz (N_SLOTS=2), 133.19 MHz
(N_SLOTS=4), both PASS at 80MHz.
errors: one Yosys build-script usage quirk (errors.log ERR-0007,
chparam target ordering), not an RTL bug.
decision: see decisions.log DEC-0009 (M3 buffers not wired in yet, no
shared-PSRAM arbitration across slots yet -- both deferred to M8).
next_action: M8 -- PSRAM integration. Wire the real (unmodified) V1
PSRAM backend chain through dataflow_core end-to-end and measure/
design whatever N_SLOTS>1 arbitration real contention requires.
[2026-09-05] M8 -- hardware/v2/rtl/neural_multiprocessor.v +
hardware/v2/rtl/slot_mem_arbiter.v
reason: roadmap M8 -- "Integrare il controller V1 senza modificarlo
inizialmente. Misurare il comportamento reale." dataflow_core.v (M7,
untouched) now shares the real, unmodified V1 PSRAM backend chain
(int8_memory_access -> memory_interface -> psram_controller) across
N_SLOTS genuinely concurrent memory_manager instances, via a new
generic N-port arbiter inspired by (not copied from) V1's own
mem_arbiter.v.
result: real concurrent-slot simulation immediately surfaced a real
bug (errors.log ERR-0008: the byte-level backend's fire-and-forget
request pulse gets silently dropped by a naive arbiter under
contention) -- fixed with a pending-request latch. After the fix:
4/4 test PASS (2 concurrent slots genuinely contending for one real
PSRAM port, 444 cycles). Real synthesis: 0 problems, LUT4=3145/
CCU2C=388/FF=3659/DSP=16, real Fmax 142.45 MHz (PASS at 80MHz, no
harness needed -- real PSRAM pins keep the top-level pin count at
157). No regression on M4's own testbench.
errors: see errors.log ERR-0008 (real RTL bug, found and fixed).
decision: see decisions.log DEC-0010 (fixed-priority arbitration, no
fairness yet -- deferred pending real measured need).
next_action: M9 -- Full benchmark (§32): V1-vs-V2 comparison table,
every number labeled THEORETICAL/SIMULATED/SYNTHESIZED/POST-P&R.
[2026-09-05] M9 -- Confronto finale V1 vs V2 (§32)
reason: roadmap M9 -- produrre la tabella di confronto mandata dal
mandato, ogni numero classificato THEORETICAL/SIMULATED/SYNTHESIZED/
POST-P&R, nessun risultato inventato (§30).
result: tabella completa in benchmark.log (M9 entry) -- V2 full-system
(neural_multiprocessor, N_SLOTS=2, M8) supera V1 full-system
(spi_neuron_top, PARALLEL=8, post_fix_verify) su ogni asse reale
misurato: Fmax POST-P&R 142.45 vs 68.65 MHz (V1 FALLISCE il target
80MHz, V2 lo supera), 166 vs 209 cicli SIMULATED per un neurone/8
input attraverso la stessa vera catena PSRAM V1 (2.6x speedup
wall-clock reale), MAC/cycle di picco 16 vs 8 (concorrenza reale a
livello di sistema, N_SLOTS=2, gia' dimostrata con vera contesa
PSRAM in EXP-0009). LUT/FF di V2 sono anche piu' bassi (4191/3659 vs
8907/4900) nonostante il sistema completo comprenda gia' scheduling
a grafo di dipendenze che V1 non ha affatto.
errors: nessuno (consolidamento dati, nessun nuovo RTL).
decision: vedi decisions.log DEC-0011 (stall %/utilization non
misurati questo milestone -- richiedono strumentazione dedicata,
rimandata a M10 dove diventa direttamente utile).
next_action: M10 -- Optimization, sulla base dei dati reali raccolti
in M1-M9.