Files
FPGA-Neural/hardware/v2/reports/step18_sdram_transaction_efficiency.md
T
micheleandClaude Sonnet 5 8e014d8d49 V2.0.0 hardware freeze - single SDRAM
FASE #1 hardware freeze for FPGA-Neural V2, N4/P8, single external
SDRAM (Alliance Memory AS4C4M16SA-6TIN) serving weights, activations,
and results through one physical sdram_controller.v instance. Removes
the PSRAM dependency (hardware/v1/rtl/psram_controller.v +
memory_interface.v) from the V2 physical path entirely -- V1 itself
remains fully unmodified, the golden reference.

New RTL: sdram_unified_backend.v (2-way W/AR arbitration over one
SDRAM controller, real per-byte DQM write masking added to
sdram_controller.v for correct single-byte result writes with no
read-modify-write), nms_neural_multiprocessor_sdram_unified.v (the
frozen top-level). Two real bugs found and fixed via full-system
testing before being accepted (ERR-0023): a deadlock and an off-by-one
data-shift bug in the new arbitration logic.

Real results: N=4 and N=2 D-Stress bit-exact (256/256 neurons), 40
real AUTO REFRESH events interleaved with zero corruption, real
Yosys+nextpnr-ecp5 synthesis/P&R for LFE5U-45F-8CABGA381 (149/245
TRELLIS_IO, a real 45-pin reduction from the prior dual-memory
design). Timing is MARGINAL (1/8 P&R seeds >=80MHz), reported honestly
rather than masked by the best seed.

Real, sourced ball-level pinout for the SDRAM bus + clk/rst (39/149
signals, P&R-verified) using the official Lattice ECP5U-45 pinout CSV
found on disk during this step's own pre-commit review -- corrects an
earlier draft that wrongly assumed no real pinout data was available.

Chip readiness: NO. Real, disclosed blockers remain (no physical host
interface exists yet -- the RTL's own reg_* ports are a 110-pin raw
test-harness bus; clock source/PLL decision; power/configuration
component selection) -- see hardware/v2/docs/{HARDWARE_FREEZE,
CHIP_READINESS,OPEN_ITEMS}.md for the complete, itemized status.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
2026-09-06 13:39:55 +02:00

19 KiB
Raw Blame History

STEP18 — SDRAM Transaction Efficiency & Weight-Path Scaling

Governing spec: the user's STEP18 message, in full. SDRAM remains the closed, unreopened V2 external memory (DEC-0031/DEC-0032). All evidence below is classified per the spec's own discipline. Raw data: step18_bandwidth_ladder.csv, step18_configs.csv, step18_resources.csv. Full narrative: experiments.log (EXP-0046/ 0047), errors.log (ERR-0022), decisions.log (DEC-0033).

1. Executive Summary

The working hypothesis — "the 16-bit interface requires multiple transactions to deliver one P8 tile" — is FALSE for the STEP16 baseline: the existing architecture already delivers exactly one SDRAM transaction per tile (MEM_DATA_WIDTH=64 was chosen in STEP16 specifically so BURST_LEN=4 = 8 bytes = 1 tile). The real inefficiency is the opposite direction: each transaction pays a large fixed per-transaction overhead (always-precharge, no page-mode), so 33.2% bandwidth utilization comes from paying that fixed cost once per tile rather than amortizing it over more data.

Built and validated a new memory-side-only module (sdram_weight_backend_pack128.v) that packs 2 tiles per real SDRAM transaction (BURST_LEN=8, 128 bits) using a small address-tagged cache — weight_prefetch_engine_wide.v and neural_processor.v are completely unmodified. A first (1-entry cache) draft regressed throughput by 50% when tested end-to-end (caught and documented, ERR-0022) because N=4's interleaved requests thrash a single cache entry; fixing the cache to N_ENTRIES=4 (sized to N_SLOTS) turned this into a real, validated 9.1% cycle reduction at both N=2 and N=4, bit-exact, with Fmax essentially unchanged at N=4 (81.47 vs 81.55MHz) and a smaller Fmax reduction at N=2 (86.04 vs 103.99MHz — still comfortably >80MHz).

The system remains memory-transaction-bound after this optimization — cycles/tile improved from 12.07 to 10.97 at N=4, but compute utilization is still only ~2.3%. The next bottleneck, after transaction packing, is the SDRAM controller's own fixed always-precharge overhead itself (not bus width, not burst length, not N-scaling) — recoverable only via a page-mode/keep-row-open controller redesign, explicitly a larger change than this round's "smallest possible" scope.

2. STEP17 baseline verification

Cross-checked against step17_n4_timing_throughput.md and its CSVs: D-Stress N4=49,430 cycles ✓, Fmax=81.55MHz best-of-3-seeds ✓, I/O= 194/245 ✓, measured bandwidth 53.04MB/s of 160MB/s nominal (33.2%) ✓, controller busy 99.92% ✓, N2=52,161 cycles ✓, N2→N4 improvement (52161-49430)/52161=5.2% ✓ (matches the spec's own cited "≈5.2%" exactly). All verified consistent with the source reports — no discrepancy found.

3. Raw SDRAM controller ceiling (Part A)

Rather than re-deriving numbers STEP16/17 already measured, this section reframes the existing REAL, real measured data into the required ladder, plus targeted new checks for patterns STEP16/17 did not explicitly frame:

Layer Value Classification
THEORETICAL SDRAM bandwidth (nominal) 160.0 MB/s (2B×16bit×80MHz) THEORETICAL
CONTROLLER MAXIMUM (isolated, BURST_LEN=4, back-to-back single requester) 64.0 MB/s (40.0%) RTL SIMULATION (STEP16 EXP-0041: 10 cyc/txn @80MHz)
CONTROLLER MAXIMUM (isolated, BURST_LEN=8 packed, back-to-back single requester) ~114.3 MB/s (~71.4%) RTL SIMULATION (STEP18: 17 cyc for 16 bytes = real fetch 16cyc + cache-hit 1cyc, single requester)
REALISTIC SUSTAINABLE (N=4, real arbitrated, BURST_LEN=4 baseline) 53.04 MB/s (33.2%) INTEGRATED BENCHMARK
REALISTIC SUSTAINABLE (N=4, real arbitrated, BURST_LEN=8 packed) 58.34 MB/s (36.5%) INTEGRATED BENCHMARK

Patterns A-D (sequential R/W, bursts, R/W turnaround): already covered by STEP16 Phase 3/4 (tb_sdram_controller.v, tests A/B/C/D) and STEP17's own controller-port instrumentation — reused, not re-measured. Writes are not part of the weight-fetch path at all (0 writes observed in every real D-Stress run, STEP17 EXP-0045), so Pattern D (R/W turnaround) has no real-workload relevance for this system's weight path; it was already validated for protocol correctness in STEP16 Phase 3 but does not affect the throughput analysis below.

Patterns E/F (row locality): the controller's own design (STEP16) always issues auto-precharge on every transaction — no row is ever kept open across transactions, by explicit, documented design choice (correctness-first, no per-row state to track). This means row locality provides zero throughput benefit in this controller, BY CONSTRUCTION — same-row and different-row accesses cost identically (confirmed structurally from the RTL, not re-benchmarked, since the FSM has no conditional path that could make them differ). This is the single largest lever available for a FUTURE controller redesign (not pursued here — see §15/§18).

Pattern G (refresh): STEP17 already measured this directly — 40 real AUTO REFRESH events during the N=4 D-Stress run, contributing negligibly to total cycles (each refresh costs ~10 cycles at 80MHz, 40×10=400 cycles of 49,430 total = 0.8%). Refresh is not a meaningful throughput factor.

4. Current transaction analysis (Part B)

Verified directly against the RTL and STEP17's own instrumentation: one P8 tile (8×INT8=64 bits=8 bytes) already costs exactly ONE physical SDRAM transaction in the STEP16/17 baseline — confirmed by nms_dataflow_core_sdram.v's own MEM_DATA_WIDTH=64 parameter (chosen in STEP16 specifically so WORDS_PER_TILE in weight_prefetch_engine_wide.v equals exactly 1), and by STEP17's own measured sdram_req_count=4096 exactly matching tiles_delivered= 4096. The working hypothesis that multiple transactions were needed per tile is refuted by direct inspection of the existing implementation, not assumed.

Why 8 bytes costs 10 cycles (not fewer): the 10 cycles decompose as tRCD(2)+CAS latency(3)+burst data(4)+tRP(2), with 1 cycle of real pipeline overlap (measured, not the naive 11-cycle sum) — 6 of the 10 cycles (60%) are fixed row-open/row-close overhead, independent of burst length. This is a controller-architecture fact (always- precharge), not a consequence of the bus being 16 bits wide — a 32-bit or 64-bit physical bus with the SAME always-precharge design would show the identical %-overhead ratio, just at a higher absolute byte count per transaction.

5. Physical vs logical bandwidth

Physical SDRAM bus width (16 bits) and internal logical delivery width (MEM_DATA_WIDTH, currently 64 bits in the baseline) are correctly already DECOUPLED in this architecture — weight_prefetch_engine_ wide.v was built in STEP14 specifically to support this distinction, and STEP16 exploited it (64-bit internal width over a 16-bit physical bus, 4 physical 16-bit beats per logical fetch via BURST_LEN=4). The STEP18 experiment tests whether widening the INTERNAL delivery further (to 128 bits, 2 tiles/fetch) — while the physical bus stays 16 bits — helps. See §6.

6. Weight packing results (Part C)

Implemented sdram_weight_backend_pack128.v: internally BURST_LEN=8 (128 bits/16 bytes per real transaction = 2 P8 tiles), externally still presents the exact same 64-bit mem_req/mem_addr/mem_rdata/ mem_ready contract weight_prefetch_engine_wide.v already uses — that engine and neural_processor.v are byte-for-byte unchanged. An address-tagged cache holds the "other half" of each real 128-bit fetch for the next sequential request.

First draft (1 cache entry): REJECTED. Isolated single-requester unit test passed 20/20 (tb_sdram_weight_backend_pack128.v), but the full N=4 integration benchmark REGRESSED to 74,004 cycles (+49.7% vs baseline) — N=4's interleaved multi-slot requests thrash a single cache entry before the natural pair completes (ERR-0022, full root-cause below).

Fixed (N_ENTRIES=4, sized to N_SLOTS): ACCEPTED.

Baseline (BURST_LEN=4) Packed (BURST_LEN=8, N_ENTRIES=4) Δ
SDRAM transactions/tile (N=4) 1.0 0.5 (2 tiles/real txn) -50%
Cycles/tile (N=4) 12.07 10.97 -9.1%
D-Stress cycles (N=4) 49,430 44,935 -9.1%
D-Stress cycles (N=2) 52,161 47,399 -9.1%
Sustained weight bandwidth (N=4) 53.04 MB/s 58.34 MB/s +10.0%
Sustained MAC/cycle (N=4) 0.6629 0.7292 +10.0%
Bit-exact PASS PASS

7. Burst results (Part D)

BURST_LEN=8 was already protocol-validated in STEP16 Phase 3 (460/460 tests, all frequencies including 80MHz) — reused directly, not re-verified from scratch. Confirmed here: address alignment (16-byte blocks), data ordering (rdata[63:0]=lower address half, rdata[127:64]= upper half, matching the controller's own word0-first convention exactly, no byte-order surprises), and real cycles/tile (10.97, measured, not estimated). Larger burst is NOT automatically better — confirmed directly by the rejected 1-entry-cache draft, where BURST_LEN=8's own real 16-cycle cost, applied to nearly every access (cache thrashed), made things WORSE than BURST_LEN=4's 10-cycle cost. The benefit only appears once the consuming logic (the cache) actually captures the 2x data-per-transaction ratio in real traffic, not merely in isolation.

8. Outstanding request results (Part E)

The controller is inherently single-transaction: confirmed, not assumed. sdram_controller.v's own FSM has exactly one busy state machine and cannot begin a new ACTIVATE while servicing a prior transaction (STEP16 architecture, unmodified). A true multi- outstanding-request redesign (overlapping ACTIVATE of transaction N+1 with the CAS/burst of transaction N) was not attempted this round — it would require re-architecting the controller's own FSM to track multiple in-flight bank states, response ordering, and address association simultaneously, a materially larger change than "smallest possible," and the packing experiment (§6) already recovers a comparable practical benefit (fewer, larger transactions) at much lower risk. This is documented as an explicit, deliberate scope boundary, not an oversight.

9. Weight buffer results (Part F)

The packing cache (§6) IS the weight tile buffer this Part asks to evaluate — a small (N_ENTRIES=4), non-blocking, per-address buffer between the real SDRAM burst and the 64-bit interface weight_ prefetch_engine_wide.v consumes. Its own PREFETCH_DISTANCE-based read-ahead mechanism (STEP11, unchanged) was ALREADY confirmed sufficient in STEP17 (mm.state busy ~97-99% of cycles despite only ~2% useful — the bottleneck was never insufficient buffering, it was SDRAM's own service rate). No separate double-buffer/FIFO experiment was built, since the packing cache already demonstrates the intended "eliminate bubbles between SDRAM delivery and P8 consumption" effect directly (measured: cycles/tile dropped, not merely resource usage changed).

10. Activation traffic (Part G)

Reused directly from STEP15/16/17's own consistent measurement: the shared 16-bit activation/result-writeback PSRAM port utilization is 6.5% (baseline) / 7.2% (packed, N=4) of total cycles — a small, stable fraction, confirming weight traffic dominates external memory traffic by a wide margin (>90% of all external-memory activity is weight fetch, not activation or result writeback). This was not re-instrumented this round (STEP17's own measurement already answers the question directly and the packing change does not touch the activation path at all — its port utilization moving from 6.5%→ 7.2% is a pure DERIVED consequence of total_cycles shrinking, not a change in absolute activation traffic).

11. N2/N4 comparison (Part I, addressed with §6's data)

N=2 baseline N=2 packed N=4 baseline N=4 packed
Cycles 52,161 47,399 49,430 44,935
Improvement vs own baseline -9.1% -9.1%
Fmax (best-of-seeds) 103.99 86.04 (1 seed) 81.55 81.47

Does improved transaction efficiency let N=4 extract more useful throughput than N=2? Both configurations improve by an IDENTICAL 9.1% — packing is a pure memory-side win that benefits N=2 and N=4 equally, because it reduces the FIXED per-transaction overhead regardless of how many requesters share the port. It does not change the fundamental N=2-vs-N=4 story: N=4 remains only modestly faster than N=2 (44935 vs 47399, 5.2% — essentially the SAME relative gap as the baseline's own 49430-vs-52161, 5.2%), because the shared SDRAM port is still the binding resource in both cases; packing raises the ceiling for BOTH equally without changing which layer is the bottleneck. N=8 was not run (explicitly optional/exploratory per the governing spec, and the N=2/N=4 result already answers the scaling question the spec asks).

12. Post-synthesis / P&R (Part J)

Real Yosys 0.68+/nextpnr-ecp5 0.11.1 results, --45k --package CABGA381 --lpf-allow-unconstrained, identical methodology to STEP16/17:

Config TRELLIS_IO TRELLIS_FF TRELLIS_COMB MULT18X18D DP16KD Fmax (best-of-3)
Baseline N=4 194/245 6215 5516 32 0 81.55 MHz
Packed N=4 194/245 6483 (+4.3%) 6106 (+10.7%) 32 0 81.47 MHz (PASS)
Packed N=2 194/245 not captured not captured 16 0 86.04 MHz (1 seed, PASS)

I/O is unchanged (packing is purely internal, no new pins). Modest FF/ COMB increase for the multi-entry cache logic. Fmax at N=4 is essentially unchanged (81.47 vs 81.55MHz, within normal seed variance) — the decision criterion (N4 Fmax ≥80MHz) is met with the SAME margin as the STEP17 baseline. N=2's Fmax drop (103.99→86.04) is real but only single-seed-measured here (not best-of-3) and still comfortably clears 80MHz; N=4 is the primary target per the governing spec and shows no meaningful Fmax cost.

13. Throughput roofline (Part K)

Ceiling N=4 baseline N=4 packed Classification
Compute ceiling (N×P_IN) 32 MAC/cycle 32 MAC/cycle THEORETICAL
External SDRAM ceiling (measured sustained) 53.04 MB/s 58.34 MB/s INTEGRATED BENCHMARK
System ceiling (actual) 0.6629 MAC/cycle 0.7292 MAC/cycle INTEGRATED BENCHMARK
Compute utilization 2.07% 2.28% DERIVED

The external SDRAM ceiling moved up (transaction efficiency improved), and the system ceiling moved up proportionally with it — confirming the system is STILL memory-bound (compute utilization barely changed, 2.07%→2.28%), just against a slightly higher memory ceiling than before.

14. Bottleneck analysis (Part L)

transaction packing         -> IMPROVED (this round, -9.1% cycles)
        |
controller fixed overhead   -> STILL DOMINANT (always-precharge pays
        |                       the same 6-cycle row-open/close cost
        |                       per transaction regardless of packing)
        v
N-way arbitration overhead  -> small, ~2 cycles/tile, unchanged
        |
        v
SDRAM physical bandwidth    -> not yet the limit (160MB/s nominal vs
                                58.34MB/s sustained = 36.5% used)
        |
        v
compute                     -> far from the limit (2.28% utilization)

The NEXT bottleneck after this round's packing optimization is the SDRAM controller's own fixed always-precharge overhead — not bus width, not burst organization (already exploited), not N-scaling (unaffected by this change), not activation traffic (confirmed minor), and not compute (nowhere near saturated).

Adopt sdram_weight_backend_pack128.v (BURST_LEN=8, N_ENTRIES=4 address-tagged cache) as the new weight-fetch backend for the N=4 V2 baseline, replacing STEP16's sdram_weight_backend.v (BURST_LEN=4, no cache). All STEP18 decision criteria are met: bit-exact (✓), no deadlock/timeout/dropped-or-duplicated-jobs (✓, full D-Stress PASS at both N=2/N=4), SDRAM protocol correct (✓, reuses the already-validated sdram_controller.v unchanged, just at BURST_LEN=8), N=4 Fmax ≥80MHz (✓, 81.47MHz), D-Stress cycles improve (✓, -9.1%), sustained MAC/cycle improves (✓, +10.0%), memory efficiency improves (✓, 33.2%→36.5%), no hidden processor serialization (✓, neural_processor.v and the STEP13 streaming architecture are completely untouched).

16. Rejected alternatives

  • 1-entry cache (first draft): rejected — real, measured 49.7% throughput REGRESSION under N=4 interleaving (ERR-0022).
  • True multi-outstanding-request controller: not attempted — materially larger redesign risk for an uncertain additional gain once packing already captures the "amortize fixed overhead" benefit; deferred as explicit future work (§8).
  • Page-mode / keep-row-open controller redesign: not attempted — correctly identified (§3, §14) as the actual next bottleneck, but a genuinely large controller rewrite, explicitly out of this round's "smallest possible change" scope.
  • N=8: not run — explicitly optional/exploratory per the governing spec, and N=2/N=4 already answers the scaling question asked.

17. Risks

  • N=2's Fmax (86.04MHz) was measured with only 1 seed (not best-of-3) — a real, disclosed gap in rigor relative to the N=4 measurement; N=4 is the primary target and was measured properly.
  • The packing cache's address-alignment assumption (natural 16-byte pairing from w_base/TILE_BYTES strides) held for the real D-Stress workload but is not universally guaranteed for arbitrary future workloads — correctness is guaranteed regardless (a cache MISS always falls back to a real, address-exact fetch), but the 9.1% benefit is workload-pattern-dependent and could be smaller for a workload with misaligned or non-sequential weight access.
  • N_ENTRIES=4 was sized to match N_SLOTS=4 by construction reasoning, not swept (e.g. N_ENTRIES=2 or 8 were not separately measured) — the chosen size is justified analytically (§6/ERR-0022) but not proven optimal.
  • No new gate-level/post-P&R re-simulation was performed (same methodology limitation as STEP16/17).

18. Final decision

With the existing 16-bit SDRAM hardware, FPGA-Neural can now sustain ~36.5% of nominal physical bandwidth (58.34 of 160 MB/s) at N=4, up from 33.2% before this step — recovered entirely through transaction packing (2 tiles/real SDRAM transaction via BURST_LEN=8 + a small N_SLOTS-sized cache), with zero change to the SDRAM device, the physical bus width, the neural processor, or the STEP13 streaming architecture. The minimum memory-side architecture required to feed N=4/P8 efficiently, given the current controller's always-precharge design, is exactly this: pack the natural tile-pair granularity into one larger burst, cached per-outstanding-slot to survive arbitration interleaving — no further internal-width widening (256-bit etc.) is justified without ALSO addressing the controller's own fixed overhead first, since a wider pack alone cannot beat the row-open/row-close cost ratio without a page-mode redesign.

Next bottleneck: SDRAM controller efficiency (transaction overhead) — specifically, the always-precharge, no-page-mode design. It is NOT: SDRAM physical bandwidth (36.5% of 160MB/s used, headroom remains), burst organization (already exploited this round), Memory Manager (STEP17 confirmed adequate buffering pre-existing), activation traffic (confirmed minor, §10), internal delivery width in isolation (the 1-entry-cache failure proved width alone doesn't help without correct multi-slot handling), or compute (2.28% utilization, far from saturated). The memory device choice remains closed and unchanged, per the governing spec's own instruction.