diff --git a/docs/ARCHITECTURE_ANALYSIS.md b/docs/ARCHITECTURE_ANALYSIS.md index 28fc3d0..fd9c746 100644 --- a/docs/ARCHITECTURE_ANALYSIS.md +++ b/docs/ARCHITECTURE_ANALYSIS.md @@ -659,7 +659,7 @@ specifically to document where/how it breaks rather than to succeed): --- -### 5.6 [Functionally DONE, TIMING SUBSTANTIALLY IMPROVED BUT NOT YET CLOSED — EXP-0089/0090/0091/0092/0093/0094] Hybrid systolic scaling: 4 groups × 4-PE weight-stationary chains +### 5.6 [Functionally DONE, WNS -0.913ns→-0.338ns via 2 safe fixes, small intrinsic gap remains — EXP-0089…0094] Hybrid systolic scaling: 4 groups × 4-PE weight-stationary chains Captured from a 2026-09-20 brainstorming session as a purely exploratory idea; the same day, per the user's own explicit reprioritization, the real @@ -777,17 +777,39 @@ NOT fully closed timing result (EXP-0094)** — honestly not yet ready for real hardware at the target clock. N=2 (EXP-0088) remains the real, trustworthy, deployable signoff. +- **Real P&R strategy tuning (EXP-0094, continued)**: same RTL, + `opt_design -directive Explore` + `place_design -directive + ExtraNetDelay_high` + a new `phys_opt_design -directive + AggressiveExplore` step + `route_design -directive + AggressiveExplore` — zero RTL risk. **Real result: WNS improved + further to −0.338ns, TNS to −5.299ns, failing endpoints to 60** — + another real, substantial, measured improvement (recovered roughly + half of the remaining gap). Still not fully met. The new worst path + is the SAME structural class (a different PE instance, confirming a + recurring per-PE issue) but now measured at 84% logic delay (up from + 79%) — real evidence the recoverable route-delay slack is largely + exhausted; what remains is dominated by the FPGA's own intrinsic + DSP48E1→CARRY4 interconnect delay, unlikely to shrink further via + more P&R strategy tuning alone. + +**Real, cumulative progress**: two safe, real improvements (the +hierarchical arbiter + P&R directive tuning), neither touching +`neural_processor_packed.v`, together close ~89% of the original TNS +gap (−690ns → −5.3ns) and ~63% of the original WNS gap (−0.913ns → +−0.338ns). **Deliberate stop here**: the remaining gap needs the +shared, load-bearing `neural_processor_packed.v` MAC datapath itself +touched to close via RTL — a meaningfully larger, more careful step +(must be verified against both N=2's own signoff and N=16) than +anything else this session, better started with explicit direction +than pushed further autonomously. + **Not yet done, real and disclosed, real options for closing the -remaining gap** (need a real decision on direction before committing -more RTL effort — `neural_processor_packed.v` is the compute core -shared by every real PE in this project): (1) real pipelining inside +remaining gap**: (1) real, careful pipelining inside `neural_processor_packed.v`'s own MAC datapath at the specific `GEN_MAC_PACKED`/`prodb1_reg` boundary (fork-before-promote, N=2's own -signoff must stay protected). (2) a real, measured lower target clock -for the N=16 variant specifically (unquantified throughput trade-off -against N=2). (3) real Vivado placement/timing directives (e.g. a -`PBLOCK` per group to reduce congestion) as a lower-RTL-risk first -attempt. See EXP-0094's own `next_action`. +signoff must stay protected, verify both). (2) a real, measured lower +target clock for the N=16 variant specifically (unquantified +throughput trade-off against N=2). See EXP-0094's own `next_action`. **The problem it targets**: plain N=16 independent cores (§5.5's own "documentary, expected to break" framing) means 16 independent DDR3 diff --git a/hardware/v2/logs/experiments.log b/hardware/v2/logs/experiments.log index 93301d4..11d2a0e 100644 --- a/hardware/v2/logs/experiments.log +++ b/hardware/v2/logs/experiments.log @@ -6587,3 +6587,62 @@ N=2, not yet done). (3) real Vivado placement/timing directives (e.g. `-directive` on `place_design`/`route_design`, or a `PBLOCK` constraint per group to reduce congestion) as a lower-RTL-risk first attempt before touching `neural_processor_packed.v` itself. + +EXP-0094 (continued) -- real P&R with aggressive Vivado directives: a +SECOND real, substantial improvement, timing STILL not fully closed, +remaining gap now clearly intrinsic (2026-09-21, "cerca di terminare +il lavoro") + +METHOD: same real `n16_system_ddr3_top.v`/`sdram_arbiter_hier.v` RTL +(no changes), real, full P&R re-run with more aggressive Vivado +strategy directives (zero RTL risk, the real "lower-RTL-risk first +attempt" from this experiment's own prior `next_action`): +`opt_design -directive Explore`, `place_design -directive +ExtraNetDelay_high`, a new real `phys_opt_design -directive +AggressiveExplore` step, `route_design -directive AggressiveExplore`. + +REAL RESULT: utilization essentially unchanged (19936 LUTs/31.44%, +35397 regs/27.92%, 128 DSP48E1/53.33%). **Real timing: WNS improved +further from -0.646ns to -0.338ns, TNS from -97.541ns to -5.299ns, +failing endpoints from 771 to 60** -- another real, substantial, +measured improvement (P&R strategy tuning alone recovered roughly half +of the REMAINING gap). Timing constraints are STILL NOT MET. + +REAL, HONEST FINDING: the new worst violated path is the SAME +structural class as before (a different physical instance -- +group2/PE3 this time, vs. group3/PE2 -- confirming this is a +recurring, PER-PE-INTRINSIC issue, not a one-off placement fluke), +`neural_processor_packed.v`'s own `GEN_MAC_PACKED`->`prodb1_reg` +DSP48E1-to-CARRY4 path, now measured at 84% LOGIC delay (up from 79% +before the directive tuning) -- i.e. the ADDITIONAL P&R effort already +squeezed out essentially all the recoverable ROUTE-delay slack; what +remains is dominated by the FPGA's own intrinsic DSP48E1->fabric +interconnect delay, a fixed real silicon characteristic that further +P&R strategy exploration is unlikely to meaningfully shrink further -- +real evidence, not a guess, since the logic-delay FRACTION grew as the +route-delay component shrank. + +DECISION: two real, safe, substantial improvements now stacked (EXP- +0094's own hierarchical arbiter + this real directive tuning), +together closing ~89% of EXP-0093's original TNS gap (-690ns -> +-5.3ns) and ~63% of the original WNS gap (-0.913ns -> -0.338ns), +neither touching `neural_processor_packed.v` (the compute core shared +by EVERY real PE in this project, including N=2's own currently +deployable EXP-0088 signoff). The REMAINING gap is small but requires +touching that shared, load-bearing module to close via RTL (real +pipelining of its own MAC datapath) -- real, deliberate STOP here per +this project's own "fork before promote" caution for centrally-shared +modules: a change there needs explicit real verification against BOTH +N=2's own real signoff (must not regress it) and N=16, a meaningfully +larger, more careful real step than anything attempted so far this +session, better started with the user's own explicit direction than +pushed further autonomously. + +next_action: (1) real, careful pipelining of `neural_processor_ +packed.v`'s own MAC datapath (the `GEN_MAC_PACKED`/`prodb1_reg` +boundary specifically) -- fork-before-promote, real isolated +verification against BOTH N=2 and N=16 before trusting any result. (2) +alternatively, a real, measured lower target clock for the N=16 +variant specifically, if the user prefers not to touch the shared +compute core. Both are real, larger next steps for deliberate pickup, +not attempted further without explicit direction.