Files
FPGA-Neural/docs/ARCHITECTURE_ANALYSIS.md
T
micheleandClaude Sonnet 5 376ccb6ee2 docs: capture two more exploratory scaling directions (BRAM cache, ESP32-side scheduling)
S5.6.1: opportunistic BRAM cache for activation tiles - exploits real,
currently 0%-utilized Block RAM to catch whatever locality the workload
happens to have, without committing to a specific reuse pattern the way
the systolic direction does. No cache-invalidation problem given the
current write-once-before-job protocol.

S5.6.2: host-side (ESP32) job-queue reordering - a software-only
"DDRManager" upstream of ddr_prefetch_mgr.v, grouping jobs with nearby
DDR3 addresses before submission to reduce row-switch cost with zero
RTL and zero timing-margin risk. Both marked exploratory, not decided,
not built - same as S5.6's systolic direction.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MUG92aM9m68TRc4rG55BcC
2026-09-20 13:22:17 +02:00

38 KiB
Raw Blame History

FPGA-Neural V3 — Architecture Analysis: Timing, Bottlenecks, and Recommended Interventions

Scope: the current, real, P&R-verified V3 design (hardware/v3/, branch v3-artix7), updated through EXP-0083 (DDRManager phase 1, real P&R: WNS +0.073ns). Every number in this document is either directly measured (real simulation trace, real P&R report) or a calculation built from directly- measured building blocks — the two are labeled explicitly throughout. Nothing here is guessed.

Status note (post EXP-0083): §5.1 (denser activation packing) described below as a recommendation is DONE and real-P&R-verified (EXP-0081/ 0082), and §5.2 (DDRManager) phase 1 is also DONE and real-measured (EXP-0083, a genuinely modest ~2.9% real benefit — see that section for the honest number and why the original hypothesis overstated it). See the "DONE" markers in those sections and the updated bandwidth numbers in §3. The document originally analyzed the pre-fix state; it's kept below (marked historical) because the comparison is itself informative, then updated with the real post-fix numbers throughout.


1. Executive summary

The single most important finding of this analysis: the system is DDR3 memory-bandwidth-bound, not DSP-bound, already at N=1 core — and this is true before considering any core-count scaling. Adding more compute cores (N=4/8/16) without first addressing memory bandwidth would not increase real throughput; it would only add more cores contending for the same saturated DDR3 channel.

Metric Value Source
Real DDR3 back-to-back burst bandwidth (physical channel) 1.24 GB/s (9.92 Gbps) measured, real JEDEC trace (§3.1) — unchanged by packing, this is a physical-channel limit
Real DDR3 bandwidth needed for ONE core at peak DSP throughput 4.96 GB/s calculated from measured DSP rate + memory layout (§3.2)
→ DDR3 can sustain, pre-EXP-0081 packing (1 tile/burst) ~25% of one core's peak compute throughput §3.2, historical
→ DDR3 can sustain, post-EXP-0081/0082 packing (2 tiles/burst, DONE) ~50% of one core's peak compute throughput §3.2, current, real
Real P&R timing margin (WNS) +0.073 ns measured, EXP-0083 real P&R (improved further from EXP-0082's +0.068ns, and EXP-0079's +0.030ns before that)
DDRManager phase 1 (single-slot look-ahead prefetch) real benefit 2.86% reduction in total real simulated time measured, real xsim A/B on tb_n2_system_ddr3.v (§5.2, EXP-0083) — modest, honestly reported, not oversold
DSP48E1 headroom for scaling 224/240 free (93%) measured, real P&R utilization

The DSP headroom is real and large. The memory-bandwidth ceiling is real, and denser activation packing (§5.1) has already doubled the real achievable fraction of it — from ~25% to ~50% of one core's peak DSP throughput, with the real P&R margin improving, not degrading, as a side effect. This was free leverage and it's now banked.

Even at ~50%, DDR3 is still the limiting resource, not DSP count — the next real interventions, per the user's own explicit direction, are: (a) widening the physical DDR3 channel from 16-bit to 32-bit (§5.5 — doubles the physical 1.24 GB/s ceiling itself, unlike §5.1 which only reduced waste against a fixed ceiling), and (b) an intelligent DDRManager (§5.2) to hide latency via orchestrator-driven prefetch. Both are required together — a wider channel without a smarter prefetcher still stalls on latency; a smarter prefetcher against a 16-bit channel still hits the same physical bandwidth wall.


2. Real signoff history (all real P&R runs to date)

EXP What changed WNS (ns) LUTs DSP48E1 Notes
0059 isolated packed core, out-of-context n/a (isolated) 507 8 first real P&R, out-of-context only
0074 first real in-context P&R: DDR3 + pins + register file not yet added +0.040 5140 16 first trustworthy board-accurate number
0076 + register file, + pin constraints, + SPI physical-layer fix +0.056 5173 16 margin improved slightly (P&R is not perfectly monotonic run to run)
0078 + config-flash bridge (real STARTUPE2 placement) +0.013 5213 16 margin dropped — real added logic
0079 + real activation-fetch engine (act_tile_fetch.v) +0.030 5379 16 pre-packing baseline
0082 + denser activation packing (2 tiles/burst, EXP-0081) +0.068 5437 16 pre-DDRManager baseline
0083 + DDRManager phase 1 (ddr_prefetch_mgr.v, single-slot look-ahead prefetch) +0.073 5644 16 current, final, trustworthy number — margin IMPROVED again despite +207 LUTs

Observation: WNS does not move monotonically with LUT count (0.056 → 0.013 → 0.030 → 0.068 → 0.073 while LUTs only ever grow) — this is normal P&R behavior (placer/router heuristics find different solutions each run, small logic changes can shift which path is critical). Do not extrapolate a trend line from 3-4 data points — the only safe practice is a fresh real P&R after every real change, which this project already does.

DSP48E1 has stayed at 16 across every real change since EXP-0059's core design was fixed — confirms the packed-DSP MAC design (§4.1) is the efficient, stable part of this architecture; all the margin pressure has come from control/glue logic (arbitration, the flash bridge, the activation engine's FSM), not from the compute datapath itself.


3. The real memory-bandwidth bottleneck (the analysis's central finding)

3.1 Real measured DDR3 throughput

From the actual ddr3_model.sv JEDEC command trace captured during the EXP-0079 real simulation run (tb_n2_system_ddr3.v via xsim), two back-to-back Read commands to the same open row:

Read  bank 0 col 000 @ 7090615 ps
Read  bank 0 col 000 @ 7103515 ps   (delta = 12900 ps = 12.9 ns)

One BURST_LEN=8 transaction moves 8 × 16 bits = 128 bits. Real measured throughput for same-row back-to-back bursts:

128 bits / 12.9 ns = 9.92 Gbps = 1.24 GB/s

This exactly matches the theoretical peak for a 16-bit DDR3 interface at 310.078 MHz (16 bits × 2 (DDR) × 310.078 MHz = 9.92 Gbps) — confirming the real controller achieves its theoretical ceiling for the best case (same-row, no row switches). This is the best-case number; anything requiring a row change (Activate/Precharge) is real-measured to cost more (§3.3).

3.2 Real compute-side bandwidth requirement

Each packed core (EXP-0059's design, unchanged since) uses 8 DSP48E1, each computing 2 packed INT8 MACs per cycle (lane A + lane B sharing one resident weight) = 16 MACs/cycle/core. At the real measured 155.039 MHz compute clock:

16 MACs/cycle × 155.039 × 10^6 cycles/s = 2.48 GMAC/s per core (real, from measured Fmax)

Each compute cycle consumes 1 activation byte per MAC lane (16 bytes total: 8 for lane A, 8 for lane B).

Historical (EXP-0079, "1 tile = 1 burst"): each tile fetch moved a full 16-byte burst for only 8 useful bytes:

2 tiles (A+B) × 16 bytes/burst = 32 bytes moved DDR3 traffic → 16 MACs
= 2 real DDR3 bytes moved per MAC operation
2.48 GMAC/s × 2 bytes/MAC = 4.96 GB/s needed per core

Current (EXP-0081/0082, "2 tiles = 1 burst", DONE and real-P&R-verified): two tiles now share one 16-byte burst, halving the moved-bytes-per-useful-byte ratio:

2 tiles (A+B) × 16 bytes/burst ÷ 2 tiles-per-burst = 16 bytes moved DDR3 traffic → 16 MACs
= 1 real DDR3 byte moved per MAC operation (halved)
2.48 GMAC/s × 1 byte/MAC = 2.48 GB/s needed per core (halved)

3.3 The real gap

Historical:  1.24 GB/s available  vs  4.96 GB/s needed per core → ~25% sustainable
Current:     1.24 GB/s available  vs  2.48 GB/s needed per core → ~50% sustainable

The physical channel ceiling (1.24 GB/s, §3.1) did not change — §5.1's fix reduced waste against a fixed ceiling, it did not raise the ceiling itself. Even at ~50%, DDR3 remains the binding constraint, not DSP count, even in the BEST case (zero row-switch overhead, one core, nothing else sharing the bus). Raising the physical ceiling requires a wider channel (§5.5) or a second channel — both analyzed below.

This ceiling gets worse, not better, with:

  • Row switches: real measured Activate→Read latency is 16.125 ns (5 DDR3 clock cycles at 3.225 ns = real CAS-latency-5 timing, matches the MIG's own configured CL=5). A tile fetch that requires a fresh row activation costs ~16-30 ns instead of the 12.9 ns same-row case — a real, measured 25-130% penalty per row switch.
  • More cores sharing the one DDR3 channel (N=2 today, N=4/8/16 proposed): the 1.24 GB/s ceiling is shared across ALL active requesters, not per-core. Adding cores divides an already-insufficient budget further.
  • Weight fetching (amortized, but not free): each job pair's weight prefetch (8 bursts for a 128-byte layer) adds real DDR3 traffic on top of activation fetching, though this cost is shared across the M reuse positions and becomes negligible for large M.

Conclusion: this was a genuine architectural ceiling, not a tuning problem, and half of it has now been recovered for free. The 2× byte- overhead from the original "1 tile = 1 full burst" convention (chosen in EXP-0079 specifically to avoid a runtime-indexed part-select, given the then-already-thin timing margin) was confirmed, with real numbers, to be the single most expensive design decision in the memory path — and has since been fixed (§5.1, DONE, EXP-0081/0082) by registering the select bit at request time instead of avoiding the select entirely. The remaining gap (~50% sustainable, not 100%) is now a physical channel width problem, not a packing-waste problem — see §5.5.


4. Module-by-module review

4.1 Compute core (mac2_dsp_packed.v, neural_processor_packed.v)

Real, stable, efficient. 8 DSP48E1/core unchanged since EXP-0059. The 2-INT8-MAC-per-DSP48E1 packing technique is the correct choice for this INT8 workload — real utilization (16/240 DSP = 6.67% at N=2) confirms there is no DSP-side pressure at all; all scaling headroom is here, and all scaling risk is elsewhere (§3, §4.5).

No real change recommended here. This is the part of the design that is not the bottleneck.

4.2 Weight-reuse path (layer_prefetch_ctrl.v, layer_weight_buffer.v,

weight_tile_gather.v)

Reused unmodified from V2 (ECP5 era), real and well-verified across many EXPs (0057/0058/0061/0062 and every integration test since). Correctly amortizes DDR3 traffic across M reuse positions — this part of the design already does the "fetch once, use many times" optimization the activation path currently lacks (§5.1's recommendation follows the SAME philosophy).

No real change recommended; this module is a good template for how the activation path should evolve.

4.3 Activation-fetch path (act_tile_fetch.v, EXP-0079, updated EXP-0081/0082)

Real, correct (verified 3 levels deep, §2 of EXP-0079's own log entry; the EXP-0081 packing change re-verified at all 3 levels again — tb_act_tile_fetch.v 8/8, tb_packed_slot.v 9/9, tb_n2_system_ddr3.v 8/8, plus real P&R). Two real design choices identified in this analysis:

  1. 1 tile = 1 full burst (2× byte overhead)FIXED (EXP-0081/0082, DONE): now 2 tiles share 1 burst via a request-time-registered select bit (sel_lat), avoiding the runtime-indexed-part-select Fmax risk while still halving DDR3 waste. Real P&R confirms margin improved, not degraded (+0.030ns → +0.068ns). See §5.1.
  2. Lane A then lane B, sequential, per tile: still doubles the real number of DDR3 transactions (and row-switch risk) versus a design that could fetch both lanes in a single wider transaction when they happen to be adjacent in memory. Not changed — flagged for future work; the DDRManager (§5.2) and 32-bit channel widening (§5.5) are higher-leverage and were prioritized first per the user's explicit direction.

4.4 Scheduling (neural_director_packed.v) and arbitration

(sdram_arbiter_n.v)

Real, correct, and — importantly for §5.2 — neural_director_packed.v already maintains a real queue of pending jobs (QUEUE_DEPTH=8, q_x_base/q_w_base/etc arrays). This is directly relevant to the DDRManager/prefetch proposal (§5.2): the information needed to "know what's coming next" already exists in this module, it just isn't currently used for anything beyond pairing/dispatch decisions.

sdram_arbiter_n.v's combinational-first-grant design (EXP-0066) is real, proven, and already generalized to N-way (verified at N=3, EXP-0069) — no real change needed to scale its own requester count for a DDRManager addition or for N=4/8/16 scaling, though the arbiter's own fairness policy (lowest-index-wins, a deliberate simplicity choice, EXP-0066) would need re-examination if a DDRManager starts issuing speculative/anticipatory requests that could starve a real, urgent request — see §5.2's own caveat.

4.5 Host interface (spi_host_bridge_v3.v, flash_spi_master.v,

host_mem_bridge.v)

Real, verified, low resource cost, not on any critical performance path (host commands are inherently much slower than the internal compute/memory loop). No bottleneck here. Not a scaling concern.

4.6 Missing: result-writeback engine

Still genuinely absent (disclosed since packed_slot.v's own original header, unchanged through EXP-0079). Currently result_data_a/b are literal top-level pins — functional at N=2 (32 pins), but this is the exact same class of mistake already caught once for activation data (EXP-0074: ~360 pins nearly exceeded the whole package's I/O budget). At N=16 this port alone would need 8 bits × 2 lanes × 16 cores = 256 pins — a real, hard blocker for any scaling beyond a handful of cores, independent of the memory-bandwidth ceiling in §3. Recommended fix in §5.3.


5.1 [Highest leverage] Denser activation packing — attack the real bandwidth ceiling directly — DONE (EXP-0081/0082)

What was built: 2 tiles (even/odd tile index) now share one BURST_LEN=8 burst instead of one tile per burst, halving real DDR3 bytes-per-MAC from 2 to 1. Real, measured effect: the achievable fraction of one core's peak DSP throughput rose from ~25% to ~50% (§3.2/§3.3, current numbers).

Why this was avoided in EXP-0079: doing so naively requires a runtime-indexed part-select (which half of the burst response to use, selected by a runtime tile-index bit) — the same anti-pattern weight_tile_gather.v (EXP-0061) already flagged as a real Fmax risk, and the P&R margin was already thin (+0.013ns) when that original design decision was made.

The safe path that was actually built: the tile-index LSB is registered into sel_lat at request time (S_IDLE, the same cycle tcnt is latched) — many ui_clk cycles before the real DDR3 round-trip completes and ctrl_rdata becomes valid. The eventual data-select mux therefore selects on an already-long-stable registered bit, never one racing the read data. Real P&R confirms this is genuinely timing-safe, not just functionally correct: margin improved from +0.030ns to +0.068ns (EXP-0082), despite the added mux logic (LUTs 5379→5437). Full detail: hardware/v3/rtl/act_tile_fetch.v header, docs/PHYSICAL_REALIZATION.md §4, hardware/v2/logs/experiments.log EXP-0081/EXP-0082.

Verification: tb_act_tile_fetch.v (8/8 PASS, covers even/odd-in-same- burst, new-burst crossing, back-to-back alternation), tb_packed_slot.v (9/9 PASS, bit-identical numeric results to the pre-change run), tb_n2_system_ddr3.v (8/8 PASS, real xsim against real ddr3_model.sv, JEDEC trace confirmed to show no more half-burst zero-padding).

5.2 [Complementary, addresses latency not bandwidth] DDRManager with orchestrator-driven prefetch (user's proposal) — phase 1 DONE (EXP-0083), real benefit smaller than the original hypothesis below predicted

The original hypothesis (written before building anything, now corrected by real measurement — kept here so the correction is visible, not silently edited away): the CURRENT design only ever requests a tile the moment packed_slot.v's own FSM reaches S_TILEREQ for it, so a DDRManager that issues tile N+1's fetch WHILE tile N is still being consumed should hide "today's design is very likely stalling ... for the majority of real time".

What the real EXP-0083 measurement actually found: this hypothesis overstated the achievable benefit, for a specific, now-confirmed reason — neural_processor_packed.v's own pipeline accepts one operand per cycle whenever it's in NP_WAIT_OPERANDS (operand_ready is state-only, not gated on any internal pipeline stall). The old design's real "dead time" between one tile's fetch completing and the next one's fetch being issued was therefore only the ~2-cycle request/consume handshake overhead (S_TILEREQ + S_OPERAND), not a large compute-bound stall — and that small overhead is what look-ahead prefetch can actually remove, not the DDR3 fetch latency itself (which is dominated by row activation/precharge, §3.3, and look-ahead cannot make a single fetch faster, only start it earlier).

Real, measured result (ddr_prefetch_mgr.v, real P&R WNS +0.073ns, up from EXP-0082's +0.068ns, LUTs 5644, DSP48E1 16 unchanged):

  • Real apples-to-apples comparison on the real DDR3 backend (tb_n2_system_ddr3.v via real xsim, same N=2/8-position workload, before vs after, same ddr3_model.sv): 2.86% reduction in total real simulated time (108370.88ns → 105268.43ns). This is the trustworthy headline number.
  • On the fast SDR placeholder backend (used for isolated glue-logic testing, tb_ddr_prefetch_mgr.v): 0.9% reduction in a row-switch-heavy scenario, and -1.4% (i.e. not faster) in an isolated same-row best case — that placeholder model's own per-fetch cost turned out to be dominated by a near-fixed protocol cost regardless of address locality, so it doesn't cleanly isolate the mechanism the real DDR3 backend's own row/bank timing does. Full detail: EXP-0083 in hardware/v2/logs/experiments.log.

What it does NOT solve (this part of the original reasoning holds): §3.2's bandwidth ceiling is a hard physical limit (bytes/second the DDR3 channel can physically move) — prefetching earlier doesn't move more bytes per second, it only avoids idle gaps. §5.1 (done) reduced bytes needed per MAC; §5.4 (32-bit widening, decided, pending) raises the physical ceiling itself; this phase-1 DDRManager only removes a small, now-quantified, per-tile dead-time — real, free (zero timing cost, margin still improving), but genuinely modest, not the larger win a first-principles estimate suggested before it was actually built and measured.

What was built (ddr_prefetch_mgr.v, real RTL, not a sketch): wraps act_tile_fetch.v (unmodified) with a depth-2 ping-pong buffer scoped to ONE slot's own activation-tile look-ahead, exactly the validated, scoped-first approach recommended below before this experiment ran. Bank selection uses a registered index bit at both fill and read time, same "known long before the data it gates" discipline as act_tile_fetch.v's own EXP-0081 layout — confirmed timing-safe by real P&R, not asserted.

Full multi-slot / whole-Director-queue scheduler — still NOT built, and now a more deliberate call, not just deferred: given phase 1's real measured benefit was modest, the cost/benefit case for the larger design below should be re-examined against the 32-bit-widened channel's real numbers (§5.4) before committing more engineering time to it — building it now, on the still-16-bit channel, risks the same gap between hypothesis and measurement this phase-1 experiment just corrected.

Concrete design sketch for the full version (informed by what already exists in this codebase; kept for when it's revisited):

  • neural_director_packed.v already queues up to QUEUE_DEPTH=8 pending jobs, each with a known x_base/w_base/n_tiles — this is exactly the "reservation" information a full DDRManager needs. No new bookkeeping is required at the Director level; it would READ this existing queue, not need the Director to change its own job-acceptance logic.
  • A cross-slot manager would sit between sdram_arbiter_n.v and each slot's own ddr_prefetch_mgr.v/layer_prefetch_ctrl.v instances, scheduling across slots (not just within one slot's own tile loop as phase 1 does) — e.g. prioritizing requests that share an already-open DDR3 row across DIFFERENT slots, which phase 1 cannot see or exploit.
  • Real caveat, not glossed over: this adds real arbitration complexity — a prefetched-but-not-yet-consumed request competing with another slot's genuinely urgent request needs a real priority policy, not just sdram_arbiter_n.v's current lowest-index-wins simplicity (§4.4). A speculative prefetch that turns out to be wrong (e.g., the Director reorders/never dispatches that queued job) also wastes real bandwidth — needs a real cancellation/staleness mechanism, not assumed away.

5.3 [Blocking for any real scaling] Result-writeback engine

Must exist before N>2 is even attemptable (§4.6) — result data needs to go into DDR3 (or through the SPI status/register path for small result sets), never as N-scaled literal top-level pins again. Same architectural shape as the weight-fetch path, in reverse (write instead of read) — a reasonable, bounded scope, and a real prerequisite, not optional polish.

5.4 [Decided] Widening the physical DDR3 channel: 32-bit single channel vs. a second independent 16-bit channel

The real question: §5.1 halved waste against a fixed 1.24 GB/s physical ceiling; it did not raise the ceiling itself. Getting past ~50% sustained DSP utilization requires more physical bytes/second, which means either (a) widening the existing channel from 16-bit to 32-bit data width, or (b) adding a second, independent 16-bit DDR3 channel. Both roughly double the real 1.24 GB/s ceiling to ~2.48 GB/s. The user asked for an honest comparison, not a diplomatically-balanced non-answer — here it is, based on real device data, not guessed.

Real device data (queried directly from the actual Vivado part database for this exact part/package, XC7A100T-CSG324): this package has only 5 total I/O banks — 14 (56 pins), 15 (56 pins), 16 (11 pins), 34 (56 pins), 35 (56 pins). All report BANK_TYPE=BT_HIGH_RANGE (Artix-7 has no separate "HP" bank class the way some other families do). Banks 14 and 15 each expose 8 DQS-capable pin pairs — the same memory-PHY signature already used by the real, placed DDR3 controller on banks 34/35. This means a second, independent DDR3 channel is physically plausible on this package (the DQS-capable pins exist), but:

Factor 32-bit single channel Second independent 16-bit channel
Real ceiling gain ~2× (1.24 → ~2.48 GB/s) ~2× (aggregate, same total)
Pin cost Reuses/extends the existing MIG's own bank(s); no new bank claimed Would claim banks 14 and/or 15 (the only banks with free DQS-capable pins)
Conflict with already-placed I/O None Real, direct: the management-SPI bus (bank 15: A15/B16/B17/A16) and the config-flash bus (bank 14: K17/K18/L13) are already placed in exactly the banks that would need to host a second channel. Only bank 16 (11 pins) would remain free — not enough margin for either SPI bus, let alone both.
Controller logic cost One MIG instance, wider data path (mostly automatic — the MIG wizard regenerates CAS/CWL/MMCM ratios for the new width) A full second MIG instance: second calibration sequence, second ui_clk domain, and critically the DDRManager/arbiter would need to become channel-aware, not just requester-aware — real added complexity on top of §5.2's own design, not a simplification
Real risk given current thin margin (+0.073ns) Lower — one controller, one clock domain, incremental change to an already-proven design Higher — two independent PHYs, two calibration state machines, cross-channel coordination logic, all new
PCB impact None (same pins, same DDR3 part, different bus width usage — the physical board the user is designing does not need to change for this) Would require re-routing/relocating whichever board-level bus (SPI mgmt or flash) currently occupies bank 14/15 pins — real PCB-level rework, not just RTL

Recommendation (honest, not deferential, as requested): 32-bit single- channel widening, not a second independent channel. The bandwidth gain is identical, but the 32-bit path has zero pin conflicts with already-placed, already-verified I/O (SPI management bus, config-flash bus), a much smaller real risk profile against the current thin timing margin, doesn't require a second full MIG/calibration instance, and — most importantly for the user's own board — needs no PCB changes, since it reuses the DDR3 part's own existing data pins at a wider access width rather than claiming new banks. A second channel's only real advantage (aggregate bandwidth could in principle scale further with a 3rd/4th channel later) doesn't apply here — this package genuinely has no more free DQS-capable banks to grow into after banks 14/15/34/35, so there's no future-proofing benefit being given up. User confirmed this recommendation and it is the decided path forward.

What this requires (not yet done, real, disclosed): the real Xilinx MIG "Customize IP" wizard must be re-run interactively (Data Width 16→32 AND Input Clock Period both changed in the same wizard session, since both require the wizard's own JEDEC/PLL calculator to recompute CAS Latency/CWL/ MMCM ratios correctly — this is not safe to hand-edit in mig_a.prj the way the earlier TargetFPGA speed-grade field was, per this project's own established discipline). This is a real, outstanding, user-gated prerequisite before §5.2's DDRManager and any N>2 scaling test can use the wider channel.

5.5 Scaling path recommendation (real numbers, not a guess) — updated per user's final directive

Given §3's real bandwidth ceiling: scaling core count alone, without §5.2/§5.4, provides no real additional throughput past whatever N already saturates the (post-§5.1) ~2.48 GB/s-equivalent demand ceiling — back-of-envelope, using §3.2's current numbers (2.48 GB/s needed per core at peak, 1.24 GB/s physically available today pre-widening), that's already close to N≈1 in the worst case and at most N≈2 in the best (zero-row-switch) case on the current 16-bit channel. Building N=4/8/16 before §5.2/§5.4 are in place would very likely show near-IDENTICAL real throughput to N=2 — a real, wasted engineering cycle the analysis recommends avoiding.

Decided real order of work (§5.1 and §5.2-phase-1 already done; this reflects the user's own explicit final direction — 32-bit widening, then the DDRManager, then N=2/4/8/16 tests, real target N=8, N=16 built specifically to document where/how it breaks rather than to succeed):

  1. §5.3 (result-writeback) / §5.1 (denser activation packing) — §5.1 DONE (EXP-0081/0082). §5.3 remains a genuine blocker for N>2 and must land before any scaling test that needs real result data out of more than 2 cores' worth of pins.
  2. §5.4 (32-bit channel widening) — decided, user-gated on a real interactive MIG wizard session (Data Width + Input Clock Period changed together). Doubles the physical ceiling itself, which §5.1 alone could not do.
  3. §5.2 (DDRManager) — phase 1 (single-slot look-ahead prefetch) is DONE (EXP-0083, real but modest ~2.9% benefit on the still-16-bit channel). Re-measure this SAME real A/B once §5.4 lands, since a wider channel may change how much idle-channel time there is left to fill — only build the larger multi-slot scheduler version if that re-measurement justifies it, not on the original (now-corrected) hypothesis alone.
  4. Real N=2/4/8/16 tests, each with its own real P&R signoff (margin is thin, §2 — do not assume a prior N's timing closure predicts the next). N=8 is the real target configuration; N=16 is expected to expose real bus/arbitration/timing limits and is built specifically to document that breakdown, not to be a viable production configuration.

5.6 [EXPLORATORY — captured, not decided, not built] Hybrid systolic scaling: 4 groups × 4-PE weight-stationary chains

Captured from a 2026-09-20 brainstorming session (same continuation), before any N=4/8/16 scaling work starts, so the direction isn't lost. Nothing in this subsection is implemented or committed to — it's a working hypothesis for a future architecture, explicitly not yet an RTL task.

The problem it targets: plain N=16 independent cores (§5.5's own "documentary, expected to break" framing) means 16 independent DDR3 requesters contending for one arbitrated channel — real congestion that neither §5.1 (packing) nor §5.4 (32-bit widening) alone removes, since both attack bytes-per-MAC or raw bandwidth, not the number of independent consumers.

The idea: instead of 16 flat, independent packed_slot.v instances, group them into 4 systolic chains of 4 PEs each. Within a chain: the weight tile stays resident (loaded once, same "weight-stationary" pattern layer_prefetch_ctrl.v/layer_weight_buffer.v already implement for the existing A/B lane reuse — this is a direct extension to 4 positions instead of 2, not a new mechanism), and activation data streams through the chain position by position, fetched from DDR3 once per chain rather than once per PE. Between the 4 chains (groups), full task-level parallelism is preserved — each group can run an independent job, same as today's model.

Why weight-stationary specifically (not activation-broadcast): chosen because it generalizes to any layer type (FC, conv-via-im2col, attention — anything reducible to "same weight matrix, many activation vectors") without assuming a specific model's channel count or convolution overlap pattern — important since this is a general-purpose accelerator, not built for one fixed network.

Real, quantifiable rationale (order-of-magnitude, not yet measured — flagged explicitly as a projection):

  • DSP48E1: 16 cores × 8 DSP/core = 128/240 (53%) — real, fits with margin.
  • DDR3 requesters: drops from 16 (flat) to 4 (one per chain) — a real 4× reduction in the number of independent contenders for the arbitrated channel, on top of (not instead of) §5.1's packing and §5.4's widening.
  • Combined with §5.4's 32-bit widening (2× physical ceiling), the available-bandwidth-to-demand ratio improves by roughly 8× versus the naive flat-N=16 baseline — a real, meaningful, but NOT a full solve by itself; still needs real measurement once anything is built.

Real open questions, not resolved yet (deliberately not designed further until the current in-flight work — §5.4's channel widening, N=2/4/8 real testing — lands first, per this project's own "one variable at a time" discipline):

  • Intra-chain dataflow RTL (result propagation between adjacent PEs, pipeline drain/fill at chain boundaries) is a real, new design, not a trivial extension — needs its own isolated verification before wiring into anything real, same as every other module in this project.
  • Result collection actually gets SIMPLER under this model versus flat N=16 (4 chain-output events instead of 16 independent ones) — relevant to §5.3's result-writeback engine, worth designing writeback with this in mind rather than for flat N=16 if this direction is pursued.
  • Arbitration simplifies too: 4 group-level requesters instead of 16, though sdram_arbiter_n.v's own NUM_REQ parameter already generalizes to either case without changes.

Decision: not decided. Revisit after §5.4 (32-bit widening) and the real N=2/4/8 flat-core scaling tests produce real numbers — those numbers will tell us whether flat scaling is "good enough" up to some N, making this restructuring unnecessary, or whether the real congestion at N=8/16 justifies it.

5.6.1 [EXPLORATORY] Opportunistic BRAM cache for activation tiles

Smaller and more incremental than §5.6's systolic restructuring — doesn't require knowing anything about the target network's structure in advance. Artix-7 100T's Block RAM is real and currently 0% utilized (§3, every real P&R signoff to date) — real, free, unused capacity.

The idea: a small direct-mapped or low-associativity cache, in BRAM, remembering the last few activation tiles fetched from DDR3 (address + data). Before act_tile_fetch.v (or ddr_prefetch_mgr.v, EXP-0083) issues a real DDR3 request, check the cache first — on a hit (e.g. two jobs, on the same or different slots, requesting overlapping/adjacent tiles, common in convolution with sliding-window overlap), skip the DDR3 round-trip entirely.

Why it's attractive: catches real reuse the design doesn't have to predict or assume in advance — unlike §5.6's systolic chains (which commit to a specific reuse pattern, weight-stationary), a cache opportunistically exploits WHATEVER locality the real workload happens to have, including patterns nobody designed for. Composable with everything else already built or decided (§5.1 packing, §5.4 widening, §5.6 systolic groups if that direction is taken) — it's a cache in front of the existing fetch path, not a replacement for it.

Real open questions: cache size vs. real hit rate is workload-dependent and NOT measured — would need a real trace-driven estimate (or a real simulation with representative test data) before sizing it, not guessed. Coherency is simple here (activation data in DDR3 is written once by the host before a job runs and never modified during compute, per the current protocol) — no cache-invalidation problem to solve, a real simplification versus a general-purpose cache design.

5.6.2 [EXPLORATORY] Host-side (ESP32) job-queue optimization — a "software DDRManager"

A different kind of lever than anything else in this section: instead of adding hardware intelligence inside the FPGA, exploit the fact that the ESP32 already has full visibility of the whole job queue before submitting itneural_director_packed.v's own QUEUE_DEPTH=8 hardware queue only sees jobs one at a time as they're written over SPI; the ESP32 firmware, upstream of that, could see and reorder ALL pending jobs at once.

The idea: the ESP32's own job-submission firmware groups/reorders jobs before writing them over the management SPI bus (§2.2 of docs/PHYSICAL_REALIZATION.md), so that jobs whose DDR3 addresses are close together (same or adjacent rows) are submitted close together in time — directly reducing the real row-switch cost (§3.3) that dominates per-tile latency, WITHOUT any new RTL at all. A real, software-only "DDRManager" living in ESP32 firmware, upstream of and complementary to ddr_prefetch_mgr.v (EXP-0083, which only looks ahead within one already- submitted job).

Why this is worth capturing seriously, not just as a curiosity: it's the cheapest possible lever in this whole list — zero RTL, zero real P&R risk, zero timing-margin cost (the project's real margin is thin, §2, and every RTL addition risks it; this doesn't touch RTL at all) — and it's far faster to iterate on than Verilog (the user's own established preference for where complexity is easiest to absorb). It doesn't compete with any other idea in this section — it can be built independently, at any time, by whoever writes the ESP32-side firmware, and composes with all of them.

Real open question: requires the host firmware to know DDR3 addresses well enough to group by row locality (ROW_BITS/COL_BITS/BANK_BITS convention, §2 of docs/PHYSICAL_REALIZATION.md) — a real firmware-side design task, not yet scoped, and out of this repository's own RTL scope (ESP32 firmware isn't part of hardware/v3/).


6. Summary table: what's real vs. what's a calculation

Claim Status
WNS/WHS/LUT/DSP numbers throughout Measured (real Vivado P&R reports, current: EXP-0083)
DDR3 back-to-back burst throughput (1.24 GB/s) Measured (real ddr3_model.sv JEDEC trace) — physical channel limit, unchanged by §5.1's packing fix
Real Activate→Read latency (16.125 ns) Measured (same trace)
Per-core compute throughput (2.48 GMAC/s) Calculated from measured Fmax (155.039MHz) + known, fixed DSP-packing factor
Bandwidth needed per core, pre-packing (4.96 GB/s) Calculated, historical (EXP-0079 layout)
Bandwidth needed per core, post-packing (2.48 GB/s) Calculated from the real, as-built EXP-0081/0082 memory layout — current
"~25% of peak sustainable" (pre-packing) / "~50%" (post-packing, current) Calculated ratios; post-packing figure re-verified against real P&R (EXP-0082) and real simulation (§5.1)
I/O bank/DQS pin counts for XC7A100T-CSG324 (banks 14/15/16/34/35) Measured — queried directly from the real Vivado part database for this exact part/package, used in §5.4's dual-channel-vs-widening analysis
Row-switch penalty as a fraction of real workloads Not measured — depends on host-chosen memory layout, flagged as an open question, not asserted
DDRManager phase-1 real stall-reduction benefit (2.86%) Measured — real xsim A/B on tb_n2_system_ddr3.v, real DDR3 backend, before vs after ddr_prefetch_mgr.v (EXP-0083). Modest, not the larger figure the original hypothesis (§5.2) suggested before it was built.
Full multi-slot DDRManager's real benefit Not measured — not built; §5.2 recommends re-measuring phase 1 against the 32-bit-widened channel before deciding whether to build it
32-bit widening's real post-change bandwidth/timing numbers Not measured — requires the user's own interactive MIG wizard session (§5.4); this document's ~2.48 GB/s figure is a doubling projection, not yet re-verified by real P&R/simulation