Files
FPGA-Neural/hardware/v2/docs/OPEN_ITEMS.md
T
micheleandClaude Sonnet 5 68f3c5e403 fix(v2): resolve ERR-0025 Part B - SRAM read timing bug in weight/activation memory
Root-causes and fixes the real, disclosed defect left open at the end
of the previous STEP20 commit: the board-level SPI host interface
produced wrong compute results when jobs were dispatched with
realistic (widely time-separated) pacing, even though job registration
itself was already confirmed correct at the dependency_manager
handshake.

Root cause: nms_weight_packed.v and nms_activation_replicated.v both
used a REGISTERED SRAM read (rd_data_reg <= mem[addr], one full clock
of latency), but nms_memory_manager_stream_wide.v's own read-ahead
pipeline (its `rd_pending` bit) is designed around a COMBINATIONAL
read -- a request issued this cycle produces data already valid to
capture the very next cycle. A busy, multi-tile job (e.g. the STEP19
D-Stress regression, 16 tiles/neuron) never exposes the mismatch,
since its own weight/activation prefetch always runs far enough ahead
that any given tile has been sitting stable in the SRAM for many
cycles by the time it's actually consumed. An uncontested single-tile
job has zero such margin: its one tile's read fires on the exact edge
the data nominally becomes ready, landing squarely on the missing
cycle and permanently latching stale/zero data.

Fixed by making both SRAMs' reads combinational, with an explicit
same-cycle fill/read address-match bypass for the one hazard a plain
combinational read alone would still miss. No FSM, arbiter, or SDRAM
controller logic was touched.

Verified (Verilator, per this project's own standing DEC-0004
protocol):
- tb_fpga_neural_v2_top_smoke.v: 11/11 PASS -- single job, back-to-back
  jobs, a realistic ~85us-gap job pair, and a parametric sweep of
  inter-job gaps (100ns/5000ns/50000ns).
- STEP19 D-Stress N=2: 49788 cycles, 256/256 bit-exact -- identical
  cycle count to before this fix (zero regression).
- STEP19 D-Stress N=4: 49771 cycles, 256/256 bit-exact -- identical
  cycle count to before this fix (zero regression).
- tb_sdram_unified_backend.v (40/40) and tb_spi_host_bridge.v (18/18)
  reconfirmed unaffected.

The physical SPI host interface is now verified correct end-to-end.
Real synthesis/P&R of the board-level top (fpga_neural_v2_top.v) is
the deliberate next step, not yet performed this round.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
2026-09-06 17:33:45 +02:00

6.1 KiB
Raw Blame History

FPGA-Neural V2 — OPEN ITEMS

Consolidated from HARDWARE_FREEZE.md, PINOUT.md, CLOCK_ARCHITECTURE.md, POWER_ARCHITECTURE.md, SCHEMATIC_READINESS.md. Classified per the governing spec's own rule: BLOCKER / CRITICAL / WARNING / OPEN / FUTURE.

BLOCKER (impede la realizzazione o il funzionamento del chip)

  1. RESOLVED (STEP20). A real SPI host interface (spi_host_bridge.v) is now implemented, protocol-correct in isolation (18/18, tb_spi_host_bridge.v), AND verified correct end-to-end through the full SPI→dependency_manager→compute→SDRAM→result path under both tight and realistic (widely time-separated) job pacing (11/11, tb_fpga_neural_v2_top_smoke.v) — see errors.log's own "ERR-0025 Part B — RESOLUTION" entry for the full root-cause writeup (a registered- vs combinational-read latency mismatch in the shared weight/activation SRAMs, fixed with zero regression to the STEP19 baseline). This item is CLOSED — kept here only for the historical record; item 1 below is likewise no longer a real blocker in the sense of "the RTL doesn't exist" — it remains open only for real pinout/board-connector work (see item 1's own updated text).
  2. RESOLVED (STEP20). The RTL's own internal "host" ports (reg_valid /reg_node_id/reg_required/reg_producer_ids/reg_x_base/ reg_w_base/reg_n_tiles/reg_result_addr) remain a simulation/ testbench-only bus for nms_neural_multiprocessor_sdram_unified.v in isolation, but the board-level top (fpga_neural_v2_top.v) now drives these SAME internal ports from spi_host_bridge.v, a real, verified SPI protocol engine (WRITE_JOB/WRITE_MEM/READ_MEM/STATUS/ RESET), matching V1's own spi_neuron_top.v precedent. The 110-pin bus is no longer exposed as a physical top-level port at all in fpga_neural_v2_top.v — only 4 real SPI pins (sclk/mosi/miso/cs_n) are.
  3. The 110-pin host/registration bus has no real ball assignment — moot now (see item 1): it is an internal signal, not a top-level port, in the board-level top. The 4 real SPI pins likewise have no ball assignment yet, since fpga_neural_v2_top.v has not been through P&R this round (see the next open item). The SDRAM bus (37 signals) and clk/rst (2 signals) now DO have a real, sourced, P&R-verified assignment (hardware/v2/constraints/ v2_unified.lpf, from the real Lattice pinout CSV found at ~/Downloads/FPGA-SC-02034-3-0-ECP5U-45-Pinout.csv during this step's own pre-commit review) — this item is narrower than originally scoped.
  4. No schematic exists; no PCB has been started.

CRITICAL (rischio elevato, deve essere risolto prima del freeze)

  1. Timing closure is MARGINAL, with an unfavorable pass rate. 8 real P&R seeds for the frozen N=4 single-SDRAM design: only 1/8 reach ≥80MHz (66.9781.84MHz range). This is WORSE than STEP18's own dual-memory design (5/8 pass). The critical path itself is unchanged (still dependency_manager.v's own pre-existing first_ready_idx/reg_ready chain) — the regression is attributed to added overall die/routing pressure from consolidation, not a new RTL defect, but it is real and unresolved.
  2. Clock source/oscillator gap -- PARTIALLY ADDRESSED (STEP20). A real EHXPLLL wrapper (ecp5_pll_sys_clk.v, real Project Trellis ecppll-generated parameters, 16MHz->64MHz) now exists and is instantiated in fpga_neural_v2_top.v. NOT YET confirmed by real synthesis/P&R of that board-level top this round (deliberately deferred until ERR-0025 Part B was resolved -- see decisions.log DEC-0037) -- this is the immediate next real step. Prior project memory records a 16MHz board oscillator. Neither "source an 80MHz+ oscillator" nor "add a real PLL to the RTL" has been decided.
  3. Two physical memories were required through STEP18 — RESOLVED this round (DEC-0034): the V2 physical path no longer instantiates hardware/v1/rtl/psram_controller.v at all. Kept here only as a closed CRITICAL item for the historical record.

WARNING (non blocca il prototipo ma deve essere documentato)

  1. N=2's real Fmax (86.04MHz in STEP18's own dual-memory design) and the STEP19 single-SDRAM N=2 config were not both measured with the same best-of-N-seed rigor as N=4 — a real, disclosed gap in measurement thoroughness, not a functional issue.
  2. W_ENTRIES/cache sizing in the weight-fetch path was set to match N_SLOTS (4) by construction reasoning, not swept for optimality.
  3. I/O standard (LVCMOS33 assumed for all 149 signals) has not been verified per real VCCIO bank once ball assignment becomes possible.

OPEN (decisione ancora da prendere)

  1. Configuration-flash part number / SPI-flash-boot vs JTAG-only bring-up.
  2. Power regulator topology and part numbers (the previously-recorded ../basic-ecp5-pcb reference design is not accessible this session to confirm as a concrete plan).
  3. Real current budget (requires running a real power-estimation tool against the actual synthesized netlist — not done this round).
  4. Decoupling/bulk capacitance values (depend on regulator selection).
  5. Reset synchronization to a real external POR/supervisor source.
  6. Real per-bank VCCIO/I-O-standard verification once ball data is available.

FUTURE EVOLUTION (miglioramento post-freeze — explicitly deferred)

  1. N=8 evaluation.
  2. A smarter W/AR priority scheme in sdram_unified_backend.v to recover some of the +10.8% (N=4) / +5.1% (N=2) cycle-count cost of single-SDRAM unification (STEP18 EXP-0046's own packing win is still present — this is about the NEW W-vs-AR contention specifically). generic
  3. Page-mode / keep-row-open SDRAM controller redesign (STEP18's own identified next bottleneck for raw memory bandwidth, independent of the single-vs-dual-memory question).
  4. True multi-outstanding SDRAM request pipelining (STEP18 Part E's own documented, deliberately out-of-scope boundary).
  5. A real physical host-interface RTL bridge (SPI or similar), resolving BLOCKER #1 above.
  6. Floorplanning / seed-pinning work to convert the current MARGINAL timing result into a reliable PASS.