FASE #1 hardware freeze for FPGA-Neural V2, N4/P8, single external SDRAM (Alliance Memory AS4C4M16SA-6TIN) serving weights, activations, and results through one physical sdram_controller.v instance. Removes the PSRAM dependency (hardware/v1/rtl/psram_controller.v + memory_interface.v) from the V2 physical path entirely -- V1 itself remains fully unmodified, the golden reference. New RTL: sdram_unified_backend.v (2-way W/AR arbitration over one SDRAM controller, real per-byte DQM write masking added to sdram_controller.v for correct single-byte result writes with no read-modify-write), nms_neural_multiprocessor_sdram_unified.v (the frozen top-level). Two real bugs found and fixed via full-system testing before being accepted (ERR-0023): a deadlock and an off-by-one data-shift bug in the new arbitration logic. Real results: N=4 and N=2 D-Stress bit-exact (256/256 neurons), 40 real AUTO REFRESH events interleaved with zero corruption, real Yosys+nextpnr-ecp5 synthesis/P&R for LFE5U-45F-8CABGA381 (149/245 TRELLIS_IO, a real 45-pin reduction from the prior dual-memory design). Timing is MARGINAL (1/8 P&R seeds >=80MHz), reported honestly rather than masked by the best seed. Real, sourced ball-level pinout for the SDRAM bus + clk/rst (39/149 signals, P&R-verified) using the official Lattice ECP5U-45 pinout CSV found on disk during this step's own pre-commit review -- corrects an earlier draft that wrongly assumed no real pinout data was available. Chip readiness: NO. Real, disclosed blockers remain (no physical host interface exists yet -- the RTL's own reg_* ports are a 110-pin raw test-harness bus; clock source/PLL decision; power/configuration component selection) -- see hardware/v2/docs/{HARDWARE_FREEZE, CHIP_READINESS,OPEN_ITEMS}.md for the complete, itemized status. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
230 lines
11 KiB
Markdown
230 lines
11 KiB
Markdown
# NMS STEP15 — Full 32-bit Physical Memory Validation
|
||
|
||
Full narrative: `hardware/v2/logs/experiments.log` (EXP-0037 through
|
||
EXP-0039), `decisions.log` (DEC-0030). Builds directly on the prior
|
||
STEP15 round's own recommendation (DEC-0029) by implementing,
|
||
debugging, and fully validating the *actual* dual-chip architecture
|
||
end to end.
|
||
|
||
## Executive conclusion
|
||
|
||
# YES, WITH CONDITIONS
|
||
|
||
The real dual-chip 32-bit PSRAM architecture is **validated** at RTL,
|
||
synthesis, and post-P&R level for the actual target
|
||
(LFE5U-45F-8CABGA381): a real, bit-exact, **2.496× end-to-end
|
||
speedup**, real Fmax of **110.28 MHz** (exceeding the 106.81 MHz
|
||
16-bit baseline), and real I/O feasibility — but only after fixing a
|
||
genuine pin-budget overflow, and only with a firm ceiling on further
|
||
headroom (88.9% of the package's I/O is now committed). 64-bit is
|
||
**not** pin-feasible on this exact package without a separate,
|
||
larger redesign of the existing host/registration interface.
|
||
|
||
## Methodology
|
||
|
||
Per the governing instructions: no hand-derived timing model was
|
||
trusted where RTL simulation could answer the question; the real
|
||
`psram_controller.v` was never replaced with an idealized model for
|
||
any final performance claim; every number below is labeled by its
|
||
actual source (RTL SIM / POST-SYNTH / POST-P&R / DERIVED); and every
|
||
unexpected result was traced to a root cause, not silently adjusted
|
||
toward the STEP15-prior-round prediction.
|
||
|
||
## 1. Architecture implemented
|
||
|
||
`psram_controller_dual32.v`: two full, **byte-for-byte unmodified**
|
||
`psram_controller.v` instances, each driving its own real physical
|
||
16-bit chip, fed identical `clk`/`rst`/`mem_req`/`mem_wr`/`mem_addr`
|
||
every cycle. This is not an assumption — it was chosen over the two
|
||
explicit alternatives the governing spec named: a single widened
|
||
controller (rejected: `psram_controller.v`'s own `psram_dq` is one
|
||
`inout` bus per instance, cannot represent two separate physical
|
||
chips) and interleaved controllers (rejected: solves capacity, not
|
||
per-transfer width). Because both instances run the identical real
|
||
timing FSM against identical inputs, they are **structurally
|
||
cycle-exact synchronized by construction** — no added synchronization
|
||
logic was needed, confirmed by a real, synthesizable cross-check
|
||
(`lane_sync_error`) that never fired in any test.
|
||
|
||
## 2. Three real bugs found and fixed (RTL correctness)
|
||
|
||
1. **Address-space mismatch**: `weight_prefetch_engine_wide.v`
|
||
outputs a *byte* address (its STEP14 convention); the real
|
||
`psram_controller.v` requires a per-chip *word* address. First
|
||
draft fed the byte address unshifted — every access landed ~4×
|
||
further out than intended. Fixed with an explicit `>>2` conversion
|
||
inside the wrapper, keeping the wrapper's own external contract as
|
||
a byte address (so it plugs into the already-validated engine
|
||
unmodified).
|
||
2. **`mem_ready` timing misalignment**: first draft *registered*
|
||
`mem_ready` while `mem_rdata` stayed combinational — a real
|
||
one-cycle skew causing the caller to sample stale data. Fixed by
|
||
making `mem_ready` a plain continuous assignment, matching the
|
||
real single-chip controller's own timing exactly.
|
||
3. **Testbench `DEPTH` too small**: a real, previously-seen bug class
|
||
in this project (documented in `tb_nms_dstress.v`'s own header) —
|
||
the backing array didn't cover the real test base address. Fixed
|
||
by sizing it generously.
|
||
|
||
Each was found by direct simulation, not assumed — the methodology
|
||
the governing spec explicitly required ("if the discrepancy appears,
|
||
trace it, do not silently adjust the model").
|
||
|
||
## 3. RTL results
|
||
|
||
Bit-exact regression (`tb_psram_dual32.v`, isolated, real V1 timing
|
||
chain, both chips): **6/6 tests, 0 errors**, covering
|
||
`n_tiles ∈ {0,1,2,15,16,511}` (including the project's own mandatory
|
||
"counter-width bug at value 16" class and `MAX_TILES-1`),
|
||
back-to-back jobs, and `lane_sync_error=0` throughout.
|
||
|
||
Real single-slot cycles/tile: **8.5488**, not exactly the ~9.0 the
|
||
prior round's abstract model predicted. Traced (not adjusted): the
|
||
concrete 2-chip architecture's own per-chip page granularity (16 of
|
||
*each chip's own* word-addresses) maps to a *larger* effective page in
|
||
the combined 32-bit space (16 combined-word transactions, not 8, since
|
||
`chip_word_addr` advances 1:1 with 32-bit transactions) than the
|
||
earlier model's flat 32-byte-page assumption. Recomputed with the real
|
||
page depth: `(15×4+1×8)/16 × 2 = 8.5` — matches the measurement almost
|
||
exactly. **The real architecture is measurably better than the
|
||
abstract model predicted**, a genuine positive finding.
|
||
|
||
## 4. Synthesis
|
||
|
||
| Metric | 16-bit baseline (STEP14) | 32-bit actual | Delta |
|
||
|---|---|---|---|
|
||
| LUT4 | 2776 | 2893 | +4.2% |
|
||
| FF | 5957 | 6273 | +5.3% |
|
||
| CCU2C | 705 | 721 | +2.3% |
|
||
| DSP (MULT18X18D) | 32 | 32 | — |
|
||
| EBR (DP16KD) | 0 | 0 | — |
|
||
| TRELLIS_IO | 157 | 218 | +61 |
|
||
|
||
## 5. Place & route — the critical validation
|
||
|
||
**First attempt failed**: separate address/control pins per chip (90
|
||
pins for the weight interface) exceeded the package's real I/O budget
|
||
by 2 pins (245 total TRELLIS_IO; 157 already committed by the existing
|
||
registration interface + original single-chip PSRAM path, leaving 88
|
||
free). This is a real P&R failure ("no BELs remaining to implement
|
||
cell type TRELLIS_IO"), not a timing failure — traced and reported
|
||
as required, not glossed over.
|
||
|
||
**Fix**: chip0 and chip1's own address/control outputs are, by
|
||
construction, byte-for-byte identical every cycle — shared to one set
|
||
of pins (a real, valid PCB fan-out technique, not a synthesis trick),
|
||
dropping the requirement to 61 pins (23 addr + 6 ctrl + 16+16 DQ).
|
||
|
||
**Final result**: P&R **succeeds**. TRELLIS_IO 218/245 (88.9%, 27
|
||
spare). Real Fmax (final, post-optimization value — nextpnr reports a
|
||
lower preliminary estimate first, a higher final value after further
|
||
passes, the same pattern as every prior synthesis in this project):
|
||
**110.28 MHz**, PASS at 80 MHz, actually **exceeding** the STEP14
|
||
baseline's 106.81 MHz. Bit-exact re-confirmed unchanged (74038 cycles)
|
||
after the pin-sharing refactor, as expected — a pure pad-level wiring
|
||
change, zero functional difference.
|
||
|
||
Determination: **(A) the 32-bit implementation does not lose Fmax —
|
||
it slightly *improves* on the baseline**, despite carrying real
|
||
additional logic (an extra arbiter, two duplicated real controller
|
||
instances).
|
||
|
||
## 6. End-to-end benchmark
|
||
|
||
| | 16-bit (baseline) | 32-bit (actual) |
|
||
|---|---|---|
|
||
| N=4 total cycles | 184771 | **74038** |
|
||
| N=4 sustained MAC/cycle | 0.1773 | 0.4426 |
|
||
| N=2 total cycles | 185270 | 75676 |
|
||
| Bit-exact | PASS 256/256 | PASS 256/256 |
|
||
|
||
**Real speedup (N=4): 184771/74038 = 2.496×** — measured, not
|
||
projected. This *exceeds* the prior round's own DERIVED 1.89×
|
||
projection. Investigated, not accepted at face value: the earlier
|
||
projection calibrated a single "degradation factor" from the *old*,
|
||
single-shared-port system, where weight, activation, and result
|
||
write-back all contended for one physical port, and assumed that
|
||
factor would persist after widening. The architecture actually built
|
||
here gives weight fetch its **own, physically separate port** —
|
||
eliminating cross-traffic-type contention entirely, not merely
|
||
widening the shared bus. Confirmed directly: the original 16-bit port
|
||
(now serving only activation+write-back) sits at just 4.4% utilization
|
||
in the new system. This is a real, structural, additional benefit no
|
||
single-degradation-factor projection could have captured.
|
||
|
||
N=2 and N=4 give statistically similar cycles (75676 vs 74038, within
|
||
2.2%) — `N_SLOTS` still does not change the port-bound ceiling, now
|
||
confirmed for the real 32-bit architecture too.
|
||
|
||
## 7. Memory-bound validation — the full bandwidth breakdown
|
||
|
||
All real/measured (EXP-0037/0038), not nominal:
|
||
|
||
| Quantity | Value | % of nominal |
|
||
|---|---|---|
|
||
| Nominal physical bandwidth (32-bit @ 80 MHz) | 320.0 MB/s | 100% |
|
||
| Usable bandwidth (real controller overhead, single-slot, uncontended) | 74.86 MB/s | 23.4% |
|
||
| — lost to protocol/non-burst overhead | 245.14 MB/s | 76.6% |
|
||
| — of which, specifically page-transitions | 5.14 MB/s | 1.6% |
|
||
| Effective weight bandwidth (real N=4 system, real arbitration) | 35.41 MB/s | 11.1% |
|
||
| — additional loss to N=4 arbitration contention | 39.46 MB/s | 12.3% |
|
||
|
||
**The dominant loss (76.6% of nominal) is the fundamentally
|
||
non-bursting, one-transaction-at-a-time protocol itself — not page
|
||
transitions specifically (only 1.6%).** This directly answers Part 11:
|
||
multi-tile bursting, if it existed, is where the largest remaining
|
||
theoretical headroom sits, far more than page-open optimization alone.
|
||
|
||
**Compute utilization: 1.383%** of the N=4 theoretical 32 MAC/cycle
|
||
ceiling. **The architecture remains firmly memory-bound** — exactly as
|
||
predicted, now proven with a real, independently-measured number
|
||
rather than a projection.
|
||
|
||
## 8. 64-bit reassessment — not pin-feasible on this package
|
||
|
||
Applying the same validated address/control-sharing technique, a
|
||
4-chip 64-bit weight interface needs `23 + 6 + 4×16 = 93` pins.
|
||
Combined with the existing 157-pin commitment: **250 pins, exceeding
|
||
the package's own 245-pin budget by 5** — *before* even considering
|
||
Fmax, LUT/FF cost, or incremental speedup-per-pin. **64-bit is
|
||
therefore not evaluated further as a real option for this board
|
||
revision** without a separate, larger initiative to first free up
|
||
pins (e.g., replacing the current wide parallel test-harness
|
||
registration interface — which alone commits 181 of the 157 "already
|
||
used" pins — with a narrower real host/SPI interface). This is a
|
||
decisive, evidence-based finding, not a restatement of the prior
|
||
round's own more tentative recommendation.
|
||
|
||
## 9. Physical implementation (I/O, PCB)
|
||
|
||
- **Exact signal count for the weight interface**: 23 address + 6
|
||
control (CE#/OE#/WE#/LB#/UB#/ZZ#, shared between both chips) + 16 +
|
||
16 independent DQ = **61 pins total**, confirmed by real
|
||
synthesis+P&R, not estimated.
|
||
- **ECP5 bank feasibility**: not independently re-verified bank-by-bank
|
||
in this round (real board-level bank/voltage assignment requires the
|
||
project's own real pinout spreadsheet, not general ECP5 facts) — the
|
||
aggregate 218/245 TRELLIS_IO figure is real and P&R-confirmed
|
||
placeable, but the specific bank layout is flagged as the concrete
|
||
next step before finalizing board layout.
|
||
- **Synchronization**: both chips share a common clock and common
|
||
control signals (address/CE#/OE#/WE#/LB#/UB#/ZZ#) by design; DQ
|
||
lanes are independent and never contend (never driven by more than
|
||
one source at a time, since only one of the two real controllers'
|
||
own tri-state DQ ever gets enabled per direction at a time as it
|
||
already does for the single-chip case). `lane_sync_error` (a real,
|
||
synthesizable, always-monitoring assertion) confirms both chips'
|
||
own real timing FSMs never diverge, in every test run — no separate
|
||
independent-control path is required.
|
||
|
||
## Answers to the governing spec's own explicit questions
|
||
|
||
The task's own "Success criteria" diagram is now fully populated with
|
||
real evidence at every stage (RTL → synthesis → P&R → end-to-end) for
|
||
both the 16-bit baseline and the 32-bit actual implementation, and an
|
||
objective comparison has been made. **Next PCB decision: proceed with
|
||
the 32-bit, 2-chip, shared-address/control architecture**, subject to
|
||
the stated I/O-headroom condition and the flagged bank-assignment
|
||
follow-up. 64-bit is off the table for this specific board without a
|
||
separate host-interface redesign.
|