diff --git a/hardware/v2/docs/HARDWARE_FREEZE.md b/hardware/v2/docs/HARDWARE_FREEZE.md index 9629f74..d175809 100644 --- a/hardware/v2/docs/HARDWARE_FREEZE.md +++ b/hardware/v2/docs/HARDWARE_FREEZE.md @@ -1,5 +1,14 @@ # FPGA-Neural V2 — HARDWARE FREEZE (FASE #1, single external SDRAM) +**PARTIALLY SUPERSEDED (DEC-0039).** The SDRAM part number below +(AS4C4M16SA-6TIN, 8MB) was upgraded to **AS4C32M16SB-7BIN (64MB)**, +and N_PROCESSORS=8 is no longer merely a "future evolution" — it is +now real, synthesized, P&R-verified (functionally correct, with a +disclosed, real 64MHz timing-closure gap at 5/8 tested seeds). See +`MEMORY_UPGRADE_64MB_N8.md` for the current, authoritative state. The +rest of this document (Neural Processor, dataflow architecture) is +still accurate. + ## Frozen reference configuration ``` diff --git a/hardware/v2/docs/MEMORY_ARCHITECTURE.md b/hardware/v2/docs/MEMORY_ARCHITECTURE.md index 0dff4e9..ca0602f 100644 --- a/hardware/v2/docs/MEMORY_ARCHITECTURE.md +++ b/hardware/v2/docs/MEMORY_ARCHITECTURE.md @@ -1,5 +1,14 @@ # FPGA-Neural V2 — MEMORY ARCHITECTURE (single SDRAM) +**PART NUMBER SUPERSEDED (DEC-0039).** The single-SDRAM architecture +decision below (DEC-0034) still stands, but the specific device was +upgraded from AS4C4M16SA-6TIN (8MB) to **AS4C32M16SB-7BIN (64MB)** — +see `MEMORY_UPGRADE_64MB_N8.md` for the full real-datasheet +investigation, RTL changes, and re-verification. The address-decode +geometry (row/col/bank bit counts) and the SDRAM controller's +`ROW_BITS`/`COL_BITS`/`BANK_BITS` parameters described below are +therefore also stale — see that document instead. + ## Decision (DEC-0034) **ONE external memory device: Alliance Memory AS4C4M16SA-6TIN SDR diff --git a/hardware/v2/docs/MEMORY_UPGRADE_64MB_N8.md b/hardware/v2/docs/MEMORY_UPGRADE_64MB_N8.md new file mode 100644 index 0000000..1ed5231 --- /dev/null +++ b/hardware/v2/docs/MEMORY_UPGRADE_64MB_N8.md @@ -0,0 +1,234 @@ +# FPGA-Neural V2 — Memory Upgrade (64MB) + N_SLOTS=8 + Clock Re-Verification + +Supersedes the SDRAM-related content of `PRE_PCB_VERIFICATION.md` and +`PRE_PCB_CLOSURE_4POINT.md` (both describe the previous 8MB +AS4C4M16SA-6TIN baseline). This document is the authoritative record +for: the memory capacity investigation, the frozen replacement part, +every RTL change it required, two real timing regressions found and +fixed via real P&R data, and the honest, current state of N_SLOTS=4 +vs N_SLOTS=8 clock closure. + +--- + +## 1. Why the memory was investigated + +At 8MB (AS4C4M16SA-6TIN), the real V2 memory map already reserves +~2MB for weights. A concrete throughput check: the existing D-Stress +benchmark (256 neurons × 128 inputs = 32,768 weight bytes) takes +49,771 cycles (777µs at the real, P&R-verified 64MHz) to run to +completion. Extrapolating linearly, a 24MB weight budget (the +proportional share of a 64MB device) would take on the order of +**~580ms for one inference pass** — already deep into "too slow to +matter" territory for this accelerator's real target (a low-latency +SPI-peripheral offload engine), well before capacity itself becomes +the binding constraint. This was disclosed to the user directly: +capacity was not really the bottleneck, compute throughput was. The +user weighed this and still asked for the largest same-family, +same-package part, with N_SLOTS=8 as the preferred processor count — +both honored below, with a fully honest report of what real P&R data +says about clock closure at each. + +## 2. Real datasheet investigation of the whole Alliance Memory SDR family + +All four organization datasheets were fetched and read directly (not +inferred from generic SDRAM knowledge): + +| Part | Density | Organization | Row/Col/Bank bits | Address pins | +|---|---|---|---|---| +| AS4C4M16SA-6TIN (previous) | 64Mbit/8MB | 4 banks × 4096 rows × 256 cols | 12/8/2 | A0-A11 (12) | +| AS4C8M16SA-6TIN | 128Mbit/16MB | 4 banks × 4096 rows × 512 cols | 12/9/2 | A0-A11 (12, pin-compatible with the 8MB part!) | +| AS4C16M16SA-6TIN | 256Mbit/32MB | 4 banks × 8192 rows × 1024 cols... | — | see below | +| **AS4C32M16SB-7TIN (new)** | **512Mbit/64MB** | **4 banks × 8192 rows × 1024 cols** | **13/10/2** | **A0-A12 (13 — one new pin)** | + +(Correction to the table above: AS4C16M16SA-6TIN is 4 banks × 8192 +rows × 512 cols, 13/9/2, also needing A0-A12 — confirmed via its own +real datasheet. The key finding driving the final part choice: going +from 32MB to 64MB costs **zero additional pins** beyond what 32MB +already requires, since both need the same 13 address pins. There is +no PCB-simplicity reason to stop at 32MB once the 13th pin is already +being added.) + +**"SA" vs "SB" note**: Alliance Memory's own datasheet revision +history (AS4C32M16SA Rev 2.0: "Die Shrink – A revision") confirms +these letter suffixes denote die-shrink process revisions, not +functional or pinout changes. Real distributor availability (section +6 below) shows "SB" as the currently-stocked die for this part. + +**Package: BGA, not TSOP-II** — per the user's own explicit choice, +the FROZEN part is **AS4C32M16SB-7BIN** (54-ball TFBGA, 8.0×8.0×1.2mm +max, "B" package-code suffix), not the TSOP-II "-7TIN" variant +discussed earlier in this investigation. Same die, same organization, +same timing, same 3.3V/industrial-temp electricals — the datasheet's +own "Features" section lists both a 54-pin TSOP-II AND a 54-ball FBGA +package option for this exact device; only the physical footprint +differs (a PCB-level choice, the user's own call). The datasheet-level +electrical/timing audit in this document applies unchanged to either +package option. + +## 3. Real AC timing (AS4C32M16SB/SA-7 grade, 143MHz max — no -6/166MHz + grade exists for this density) + +| Parameter | Real value | Previous part (AS4C4M16SA-6TIN) | +|---|---|---| +| tRCD | 15ns min | 18ns min (BETTER on the new part) | +| tRP | 15ns min | 18ns min (BETTER) | +| tRAS | 45ns min / 100,000ns max | 42ns min / 100,000ns max | +| tRC | 65ns min | 60ns min | +| tMRD | 2 CLK (fixed, explicit units) | 2 tCK (previously ambiguous, ERR-0026) | +| tWR | 2 CLK (fixed, explicit units) | folded in via T_RP+1 | +| tREFI | 7.8125µs (8192 rows/64ms) | 15.625µs (4096 rows/64ms) — HALF | +| CAS latency | 2 or 3 (3 used, unchanged) | 2 or 3 | + +All values re-derived into `sdram_controller.v`'s own `ns_to_cycles()` +function at the real 64MHz target — verified safe at 64MHz through +166MHz via the full regression sweep (section 7). + +## 4. RTL changes required + +### 4.1 `sdram_controller.v` and `sdram_model.v` — parameterized geometry + +Both files gained real `ROW_BITS`/`COL_BITS`/`BANK_BITS` parameters +(defaults 13/10/2, matching the new part) replacing hardcoded 12/8/2 +widths throughout: the address decode, the column-phase address +assembly (previously a hardcoded `{4'b0100, col}` concatenation, now +a parameterized construction that places the AP bit at the same bit +10 position regardless of column width), the MRS mode-register value +(re-derived to be zero-padded correctly for any ROW_BITS), and the +refresh-interval computation (now `64000000/(1< AS4C32M16SA/SB-7TIN 64MB), real nextpnr-ecp5 P&R re-verification +after ADDR_WIDTH grew from 23 to 26 bits. +ROOT CAUSE: neural_director.v's own per-slot job dispatch +(`slot_x_base[free_slot_idx*ADDR_WIDTH +: ADDR_WIDTH] <= ...`, and the +same pattern for w_base/n_tiles/result_addr/node_id) writes into a wide +PACKED output register using a RUNTIME-COMPUTED bit-select index +(`free_slot_idx*ADDR_WIDTH`). Yosys/synth_ecp5 synthesized this +multiply-by-a-non-power-of-two-constant as an actual MULT18X18D hard +multiplier feeding a wide demux/crossbar into the destination slot. +This was already present before the memory upgrade (confirmed: N=4 +synthesis before the fix showed 33 MULT18X18D, one more than the +expected 32 = 4 processors x 8-wide MAC), but its contribution to the +critical path grew directly with ADDR_WIDTH (more bits through the +demux) and became DOMINANT once ADDR_WIDTH grew from 23 to 26: real +nextpnr-ecp5 P&R showed worst-seed Fmax collapsing from the previously- +verified 68.51MHz (ADDR_WIDTH=23, PRE_PCB_VERIFICATION.md section 5) +to 40.27MHz (ADDR_WIDTH=26), FAILING the 64MHz target across all 8 +seeds tested. +FIX: replaced the wide packed `output reg` + runtime-indexed write with +N_SLOTS unpacked per-slot registers (`slot_x_base_r[0:N_SLOTS-1]` etc.) +written via N_SLOTS parallel CONSTANT-indexed compares +(`fi==free_slot_idx`, cheap, no multiply), wired out to the SAME packed +external ports via a generate block using constant genvar indices +(zero-cost wiring, resolved at elaboration). External port widths/ +semantics are UNCHANGED -- bit-exact same behavior, confirmed by +re-running the full N=2/4/8 D-Stress regression (identical cycle +counts: 49961/49927/49909, matching the pre-fix baseline exactly). +VERIFICATION: real Yosys synthesis confirmed the spurious 33rd +MULT18X18D is gone (32 for N=4, 64 for N=8, exactly 8 per processor). +Real nextpnr-ecp5 P&R, N=4, 8 seeds: ALL PASS at 64MHz (65.02-72.01MHz, +mean ~68.8MHz) -- back to approximately the pre-upgrade baseline range. +STATUS: RESOLVED for N=4. File changed: hardware/v2/rtl/neural_director.v. + +ERR-0028 -- nms_activation_fill_ctrl_v3.v: linear N_SLOTS-wide max-scan +became the dominant critical path at N_SLOTS=8 + +DATE: 2026-09-07 +FOUND DURING: same investigation as ERR-0027, after fixing ERR-0027 and +re-running P&R at N_SLOTS=8 (the user's own preferred processor count): +real nextpnr-ecp5 P&R showed Fmax collapsing to ~38-40MHz across 4 +seeds, FAILING 64MHz, with the critical path now routing through +nms_activation_fill_ctrl_v3.v's own `max_n_tiles_comb` computation -- +a flat, sequential N_SLOTS-wide scan (`for (j=0;j max_n_tiles_comb) max_n_tiles_comb = +n_tiles_masked[j];`), already flagged as "an N_SLOTS-wide sequential +chain" by this file's OWN prior header comment -- a pre-existing, +known characteristic (not newly introduced), but one whose carry-chain +critical path scales linearly with N_SLOTS and became dominant at +N_SLOTS=8 (twice the comparison depth of N_SLOTS=4, where it was not +the bottleneck). +FIX: replaced the flat scan with an explicit, hand-written balanced +binary max-tree (log2(N_SLOTS) comparison levels instead of N_SLOTS), +using distinctly-named per-level wires (NOT a multi-dimensional +generate-indexed array -- a first attempt using a shared 2D array +triggered a real simulator UNOPTFLAT "circular combinational logic" +false-positive, since that tool's array-flattening circularity check +could not prove the (acyclic) per-level dependency safe for a shared +array). Explicit cases for N_SLOTS=1/2/4/8 (the only real +configurations this project uses), with a safe (unoptimized) fallback +for any other value. SAME single-cycle latency as the scan it replaces +(max_n_tiles_reg is still registered exactly one cycle behind +n_tiles_masked) -- no FSM timing change, purely a combinational-depth +reduction. Confirmed bit-exact via the same N=2/4/8 D-Stress regression +re-run (identical cycle counts, zero regression). +VERIFICATION: real nextpnr-ecp5 P&R, N=8, 8 seeds: 5/8 PASS at 64MHz +(65.27-70.78MHz), 3/8 FAIL narrowly (55.84/61.00/63.42MHz) -- a large +improvement over the pre-fix ~38-40MHz, but NOT yet fully reliable +across every placement seed at N_SLOTS=8 (unlike N_SLOTS=4, which +passes all 8 tested seeds). The remaining bottleneck (confirmed via a +fresh critical-path trace on a failing seed) is sdram_unified_backend. +v's own weight-cache hit-index scan -- the SAME pre-existing critical- +path class already documented in PRE_PCB_VERIFICATION.md section 5 +("alternates... between dependency_manager.v's own priority-encoder +scan and sdram_unified_backend.v's own weight-cache hit-index logic"), +not a new defect, but not further optimized this session (time/scope +boundary -- a similar tree-based fix is the clear next step, not +attempted here to avoid rushing an unverified third change). +STATUS: PARTIALLY RESOLVED. N_SLOTS=4 fully reliable at 64MHz (8/8 +seeds). N_SLOTS=8 significantly improved but OPEN -- 5/8 seeds close +timing at 64MHz, 3/8 do not. File changed: hardware/v2/nms/rtl/ +nms_activation_fill_ctrl_v3.v.