fix(v2): resolve ERR-0025 Part B - SRAM read timing bug in weight/activation memory

Root-causes and fixes the real, disclosed defect left open at the end
of the previous STEP20 commit: the board-level SPI host interface
produced wrong compute results when jobs were dispatched with
realistic (widely time-separated) pacing, even though job registration
itself was already confirmed correct at the dependency_manager
handshake.

Root cause: nms_weight_packed.v and nms_activation_replicated.v both
used a REGISTERED SRAM read (rd_data_reg <= mem[addr], one full clock
of latency), but nms_memory_manager_stream_wide.v's own read-ahead
pipeline (its `rd_pending` bit) is designed around a COMBINATIONAL
read -- a request issued this cycle produces data already valid to
capture the very next cycle. A busy, multi-tile job (e.g. the STEP19
D-Stress regression, 16 tiles/neuron) never exposes the mismatch,
since its own weight/activation prefetch always runs far enough ahead
that any given tile has been sitting stable in the SRAM for many
cycles by the time it's actually consumed. An uncontested single-tile
job has zero such margin: its one tile's read fires on the exact edge
the data nominally becomes ready, landing squarely on the missing
cycle and permanently latching stale/zero data.

Fixed by making both SRAMs' reads combinational, with an explicit
same-cycle fill/read address-match bypass for the one hazard a plain
combinational read alone would still miss. No FSM, arbiter, or SDRAM
controller logic was touched.

Verified (Verilator, per this project's own standing DEC-0004
protocol):
- tb_fpga_neural_v2_top_smoke.v: 11/11 PASS -- single job, back-to-back
  jobs, a realistic ~85us-gap job pair, and a parametric sweep of
  inter-job gaps (100ns/5000ns/50000ns).
- STEP19 D-Stress N=2: 49788 cycles, 256/256 bit-exact -- identical
  cycle count to before this fix (zero regression).
- STEP19 D-Stress N=4: 49771 cycles, 256/256 bit-exact -- identical
  cycle count to before this fix (zero regression).
- tb_sdram_unified_backend.v (40/40) and tb_spi_host_bridge.v (18/18)
  reconfirmed unaffected.

The physical SPI host interface is now verified correct end-to-end.
Real synthesis/P&R of the board-level top (fpga_neural_v2_top.v) is
the deliberate next step, not yet performed this round.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
This commit is contained in:
2026-09-06 17:33:45 +02:00
co-authored by Claude Sonnet 5
parent 18bf038ed5
commit 68f3c5e403
7 changed files with 329 additions and 97 deletions
+37 -21
View File
@@ -7,24 +7,36 @@ FUTURE.
## BLOCKER (impede la realizzazione o il funzionamento del chip)
0. **STEP20 update:** a real SPI host interface (`spi_host_bridge.v`)
was implemented and is protocol-correct in isolation (18/18,
`tb_spi_host_bridge.v`), but a real, disclosed, UNRESOLVED defect
(errors.log ERR-0025 Part B) produces wrong compute results when
jobs are dispatched through it with realistic (widely time-
separated) pacing — root cause not yet isolated. This SUPERSEDES
item 1 below with a more specific, code-level blocker: the physical
host interface RTL now exists, but is not yet proven correct.
1. **No physical host interface exists.** The RTL's own "host" ports
are a 110-pin raw parallel job-registration bus
(`reg_valid`/`reg_node_id`/`reg_required`/`reg_producer_ids`/
`reg_x_base`/`reg_w_base`/`reg_n_tiles`/`reg_result_addr`) — a
simulation/testbench convenience, not a real board protocol. No
RTL exists to serialize it (e.g. SPI, matching V1's own
`spi_neuron_top.v` precedent).
2. **The 110-pin host/registration bus has no real ball assignment**
(deliberately — it is not yet a real physical protocol, see item 1).
The SDRAM bus (37 signals) and clk/rst (2 signals) now DO have a
0. **RESOLVED (STEP20).** A real SPI host interface (`spi_host_bridge.v`)
is now implemented, protocol-correct in isolation (18/18,
`tb_spi_host_bridge.v`), AND verified correct end-to-end through the
full SPI→dependency_manager→compute→SDRAM→result path under both
tight and realistic (widely time-separated) job pacing (11/11,
`tb_fpga_neural_v2_top_smoke.v`) — see errors.log's own "ERR-0025
Part B — RESOLUTION" entry for the full root-cause writeup (a
registered- vs combinational-read latency mismatch in the shared
weight/activation SRAMs, fixed with zero regression to the STEP19
baseline). This item is CLOSED — kept here only for the historical
record; item 1 below is likewise no longer a real blocker in the
sense of "the RTL doesn't exist" — it remains open only for real
pinout/board-connector work (see item 1's own updated text).
1. **RESOLVED (STEP20).** The RTL's own internal "host" ports (`reg_valid`
/`reg_node_id`/`reg_required`/`reg_producer_ids`/`reg_x_base`/
`reg_w_base`/`reg_n_tiles`/`reg_result_addr`) remain a simulation/
testbench-only bus for `nms_neural_multiprocessor_sdram_unified.v`
in isolation, but the board-level top (`fpga_neural_v2_top.v`) now
drives these SAME internal ports from `spi_host_bridge.v`, a real,
verified SPI protocol engine (WRITE_JOB/WRITE_MEM/READ_MEM/STATUS/
RESET), matching V1's own `spi_neuron_top.v` precedent. The 110-pin
bus is no longer exposed as a physical top-level port at all in
`fpga_neural_v2_top.v` — only 4 real SPI pins (sclk/mosi/miso/cs_n)
are.
2. **The 110-pin host/registration bus has no real ball assignment**
moot now (see item 1): it is an internal signal, not a top-level
port, in the board-level top. The 4 real SPI pins likewise have no
ball assignment yet, since `fpga_neural_v2_top.v` has not been
through P&R this round (see the next open item). The SDRAM bus (37
signals) and clk/rst (2 signals) now DO have a
real, sourced, P&R-verified assignment (`hardware/v2/constraints/
v2_unified.lpf`, from the real Lattice pinout CSV found at
`~/Downloads/FPGA-SC-02034-3-0-ECP5U-45-Pinout.csv` during this
@@ -42,9 +54,13 @@ FUTURE.
`first_ready_idx`/`reg_ready` chain) — the regression is attributed
to added overall die/routing pressure from consolidation, not a new
RTL defect, but it is real and unresolved.
2. **Clock source/oscillator gap.** The RTL requires a direct ≥80MHz
clock (no PLL exists anywhere in the hierarchy — confirmed via
`EHXPLLL: 0/4` in every real synthesis run). Prior project memory
2. **Clock source/oscillator gap -- PARTIALLY ADDRESSED (STEP20).** A
real EHXPLLL wrapper (`ecp5_pll_sys_clk.v`, real Project Trellis
`ecppll`-generated parameters, 16MHz->64MHz) now exists and is
instantiated in `fpga_neural_v2_top.v`. NOT YET confirmed by real
synthesis/P&R of that board-level top this round (deliberately
deferred until ERR-0025 Part B was resolved -- see decisions.log
DEC-0037) -- this is the immediate next real step. Prior project memory
records a 16MHz board oscillator. Neither "source an 80MHz+
oscillator" nor "add a real PLL to the RTL" has been decided.
3. **Two physical memories were required through STEP18** — RESOLVED