docs: N=8 re-verified with MAC pipeline fix (improves margin), N=16 margin-hunt result (EXP-0097 cont.)
N=8 with the same pipelined neural_processor_packed.v: 16/16 functional PASS, real P&R WNS=+0.108ns (up from the original exact-zero 0.000ns), 0 failing endpoints, no regression -- the fix is a pure improvement for N=8 too. N=16 margin-hunt: worst path moved again (now inside neural_director_grouped.v's own queue update logic, still route- dominated) -- confirms the MAC fix resolved its own real bottleneck. A second real P&R attempt with alternate directives gave WNS=+0.168ns, worse than the first attempt's +0.269ns -- confirms real P&R variance, not further systematic gain available without touching Director RTL (not attempted, current margin already comfortable). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MUG92aM9m68TRc4rG55BcC
This commit is contained in:
@@ -6909,3 +6909,56 @@ branch's own real bug findings (the testbench latching fix, the
|
|||||||
Vivado incremental-synthesis generic-binding quirk) are worth folding
|
Vivado incremental-synthesis generic-binding quirk) are worth folding
|
||||||
into CLAUDE.md's own hard-won-lessons section regardless of the
|
into CLAUDE.md's own hard-won-lessons section regardless of the
|
||||||
promotion decision.
|
promotion decision.
|
||||||
|
|
||||||
|
EXP-0097 (continued) -- real N=8 re-verification with the same MAC
|
||||||
|
pipeline fix (improves, does not regress) + a real margin-hunt attempt
|
||||||
|
for N=16 (2026-09-21/22, user's own explicit direction: "verifica
|
||||||
|
anche su N=8 e verifica se possiamo guadagnare qualcosina ancora su
|
||||||
|
N16 perche' io implemento N16 se funziona")
|
||||||
|
|
||||||
|
REAL RESULT (1), N=8 with the same pipelined `neural_processor_
|
||||||
|
packed.v`: real functional xsim (`tb_n8_system_ddr3.v`, real DDR3
|
||||||
|
model) **16/16 PASS, 0 errors**, identical real completion time to the
|
||||||
|
pre-fix result (131145.8335ns) -- no functional regression. Real, full
|
||||||
|
P&R (`n8_system_ddr3_top.v`, `-generic N_GROUPS=2` explicit, confirmed
|
||||||
|
via a real post-synth 64 DSP48E1 count): **WNS=+0.108ns (up from the
|
||||||
|
pre-fix exact-zero 0.000ns), WHS=+0.036ns, 0 failing endpoints, 64
|
||||||
|
DSP48E1 (26.7%), 12536 LUTs (19.77%)**. The same fix that closes N=16
|
||||||
|
also gives N=8 real, comfortable margin instead of the exact-zero
|
||||||
|
margin its original (unmodified-core) signoff had -- a real, additive
|
||||||
|
improvement, not a tradeoff.
|
||||||
|
|
||||||
|
REAL RESULT (2), N=16 margin-hunt: the real worst path has moved AGAIN
|
||||||
|
(confirming the MAC-datapath fix genuinely resolved ITS OWN real
|
||||||
|
bottleneck) -- now inside `neural_director_grouped.v`'s own queue
|
||||||
|
update logic (`u_dir/q_head_reg[3]` -> `q_count_reg[0]/CE`), still
|
||||||
|
real route-dominated (73%), not logic-depth-dominated. A second real
|
||||||
|
P&R attempt with alternate directives (`opt_design -directive
|
||||||
|
ExploreWithRemap`, `place_design -directive Explore`, same
|
||||||
|
`phys_opt_design`/`route_design -directive AggressiveExplore`) gave
|
||||||
|
**WNS=+0.168ns -- WORSE than the first attempt's own real +0.269ns**,
|
||||||
|
confirming real run-to-run/directive-to-directive P&R variance, not a
|
||||||
|
systematic further improvement available from directive-tuning alone.
|
||||||
|
|
||||||
|
DECISION: the original directive stack (`Explore`/`ExtraNetDelay_
|
||||||
|
high`/`AggressiveExplore`) remains the best real N=16 result found
|
||||||
|
(WNS=+0.269ns). Further real margin would require touching `neural_
|
||||||
|
director_grouped.v`'s own queue RTL (a new, separate, not-yet-scoped
|
||||||
|
piece of real engineering) -- NOT attempted, given the current real
|
||||||
|
margin is already comfortably closed (better than N=2's own original
|
||||||
|
historical +0.0999962ns real signoff margin) and blind further P&R-
|
||||||
|
directive search already showed diminishing/negative real returns.
|
||||||
|
|
||||||
|
Real, cumulative state on this branch (`n16-timing-closure`), both
|
||||||
|
configurations using the SAME real pipelined `neural_processor_
|
||||||
|
packed.v`:
|
||||||
|
N=8 (`n8_system_ddr3_top.v`): WNS=+0.108ns, 64 DSP48E1, 16/16 functional PASS
|
||||||
|
N=16 (`n16_system_ddr3_top.v`): WNS=+0.269ns, 128 DSP48E1, 32/32 functional PASS
|
||||||
|
|
||||||
|
next_action: real, explicit decision with the user on whether to (1)
|
||||||
|
promote this fix to `v3-artix7` for a FUTURE board revision (the
|
||||||
|
current physical board is already in fabrication as the unmodified
|
||||||
|
N=8 design, unaffected), and/or (2) actually build the NEXT physical
|
||||||
|
board as N=16 instead of N=8, given N=16 is now real, functionally
|
||||||
|
verified, AND timing-closed with a real, comfortable margin -- a real,
|
||||||
|
consequential hardware decision, not an RTL one.
|
||||||
|
|||||||
Reference in New Issue
Block a user