4 Commits
Author SHA1 Message Date
micheleandClaude Sonnet 5 0303aacb9f docs: consolidate N=16 to N=8's own level of rigor (EXP-0098) -- whole family real and closed
N=2 re-verified with the same shared pipelined neural_processor_packed.v
(EXP-0097's fix): real functional xsim 8/8 PASS (identical completion
time to the pre-fix baseline), real full P&R WNS=+0.389ns (up from
+0.099962ns), 0 failing endpoints, 16 DSP48E1 -- no regression.

Real, consolidated family state, all sharing the same pipelined core:
  N=2:  WNS=+0.389ns, 16 DSP48E1,  8/8 functional PASS
  N=8:  WNS=+0.108ns, 64 DSP48E1,  16/16 functional PASS
  N=16: WNS=+0.269ns, 128 DSP48E1, 32/32 functional PASS

Updated the project's own primary real docs (docs/PHYSICAL_REALIZATION.md
S3, docs/ARCHITECTURE_ANALYSIS.md S5.6) to the same depth EXP-0096 gave
N=8 alone -- not just the docs/latex/ deliverables. Clear two-track
framing throughout: the physically fabricated board still uses N=8's
original, unmodified core (WNS=0.000ns exact); this consolidated,
improved family lives on the isolated n16-timing-closure branch,
pending a real hardware decision on whether/when to adopt it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MUG92aM9m68TRc4rG55BcC
2026-09-22 01:15:15 +02:00
micheleandClaude Sonnet 5 cc6cfe168e docs: update LaTeX docs for N16 real timing closure + honest ESP32-S3 projection (EXP-0097)
architecture.tex: document the real MAC-pipeline fix (Stage 1a/1b
split) and its real result on both N8 (WNS 0.000ns -> +0.108ns) and
N16 (now closed, WNS=+0.269ns) -- N16 is a real, verified candidate
for a future board revision, N8 remains the configuration in current
physical fabrication.

tests_timing.tex: extended P&R history table, real N8/N16 signoff
tables (post-fix), the new N16 critical path (Director queue logic,
still route-dominated) and the negative result of a second directive
attempt (confirms current result is the best found). Added a new,
clearly-labeled PROJECTION (not a measurement) for N16 vs ESP32-S3 --
30-68x range, built from real measured scaling factors (same core
count as the original ECP5 comparison, real clock ratio, real 2x INT8
packing factor) -- more grounded than the old, superseded ~55-85x
figure, with the same DDR3-bandwidth-bound honest caveat carried
forward.

datasheet.tex: N8/N16 side-by-side timing/utilization/performance
tables, open items and revision history updated to reflect physical
fabrication status and the N16 candidate decision.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MUG92aM9m68TRc4rG55BcC
2026-09-22 00:59:46 +02:00
micheleandClaude Sonnet 5 2ab2ff7769 docs: N=8 re-verified with MAC pipeline fix (improves margin), N=16 margin-hunt result (EXP-0097 cont.)
N=8 with the same pipelined neural_processor_packed.v: 16/16 functional
PASS, real P&R WNS=+0.108ns (up from the original exact-zero 0.000ns),
0 failing endpoints, no regression -- the fix is a pure improvement for
N=8 too.

N=16 margin-hunt: worst path moved again (now inside
neural_director_grouped.v's own queue update logic, still route-
dominated) -- confirms the MAC fix resolved its own real bottleneck.
A second real P&R attempt with alternate directives gave WNS=+0.168ns,
worse than the first attempt's +0.269ns -- confirms real P&R variance,
not further systematic gain available without touching Director RTL
(not attempted, current margin already comfortable).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MUG92aM9m68TRc4rG55BcC
2026-09-22 00:44:54 +02:00
micheleandClaude Sonnet 5 f9b366d747 feat: N=16 real timing CLOSED via extra MAC pipeline stage (EXP-0097, branch n16-timing-closure)
neural_processor_packed.v: split the original single "Stage 1" (packed
DSP48E1 multiply + INT8 unpack + register) into two real stages --
Stage 1a registers the raw DSP48E1 product with zero logic in between,
Stage 1b does the carry-heavy unpack (the real critical path EXP-0094
traced) from that already-registered value. Adds exactly one real
clock cycle of latency; throughput unaffected (real valid/ready
handshaking throughout, no fixed-latency assumption downstream).

Real verification: isolated bit-exact vs 2x real neural_processor.v
(18/18 PASS, testbench fixed to latch each core's result independently
since result_valid is a one-shot pulse and the DUT is now one cycle
deeper -- not an RTL bug). Full-system functional xsim on real DDR3:
32/32 PASS. Real, full P&R: WNS=+0.269ns, WHS=+0.026ns, 0 failing
setup or hold endpoints -- N=16 TIMING CLOSES.

Also root-caused (not an RTL bug, folded into CLAUDE.md): a real
Vivado incremental-synthesis quirk silently carried forward a
N_GROUPS=2 parameter binding from an earlier sweep run despite no
-generic override and an intervening reset_run -- fixed by always
passing -generic explicitly and confirming the real elaborated value
via a post-synth DSP48E1 count.

Isolated on this branch -- does not touch the physical board already
in fabrication on v3-artix7 (N=8, unmodified).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MUG92aM9m68TRc4rG55BcC
2026-09-22 00:17:18 +02:00
9 changed files with 785 additions and 206 deletions
+17
View File
@@ -81,6 +81,23 @@ unmodified by v3, e.g. `layer_prefetch_ctrl.v`/`layer_weight_buffer.v`).
Real fix: explicitly `set_property top <newmodule> [get_filesets Real fix: explicitly `set_property top <newmodule> [get_filesets
sources_1]` *before* calling `synth_design -top ...` — the `-top` sources_1]` *before* calling `synth_design -top ...` — the `-top`
command-line flag alone wasn't sufficient this time. command-line flag alone wasn't sufficient this time.
- **A top-level module's own default parameter value can silently NOT
apply**, even with no `-generic` override on the `synth_design`
command line, an empty real `GENERIC` property on the run, and no
stale imported RTL copy (EXP-0097) — a real elaboration bound
`N_GROUPS` to a value from a DIFFERENT, EARLIER `-generic` override
used against the SAME top module in the SAME Vivado session/project
(an N=8 sweep run before an N=16 run), despite an intervening
`reset_run`. Most likely Vivado's own "Incremental synthesis
strategy default" silently carrying forward a parameter binding.
Real fix: always pass every `-generic` value EXPLICITLY on every
real `synth_design` call for a parameterized top-level, never rely
on "no override = the RTL's own default" once that module has EVER
been synthesized with a different override earlier in the same
project — and confirm the real elaborated value afterward (e.g. a
post-synth DSP48E1/cell count check) before trusting anything
downstream, don't assume the log's own "Parameter ... bound to"
line will be checked in time otherwise.
- **Testbench stimulus must use nonblocking assignment (`<=`), not blocking - **Testbench stimulus must use nonblocking assignment (`<=`), not blocking
(`=`), when driving a DUT's inputs from a separate `always`/`initial` (`=`), when driving a DUT's inputs from a separate `always`/`initial`
block.** Blocking assignment races the DUT's own `posedge`-triggered block.** Blocking assignment races the DUT's own `posedge`-triggered
+84 -20
View File
@@ -1,15 +1,16 @@
# FPGA-Neural V3 — Architecture Analysis: Timing, Bottlenecks, and Recommended Interventions # FPGA-Neural V3 — Architecture Analysis: Timing, Bottlenecks, and Recommended Interventions
Scope: the current, real, P&R-verified V3 design (`hardware/v3/`, branch Scope: the current, real, P&R-verified V3 design (`hardware/v3/`, branch
`v3-artix7`), updated through EXP-0095/0096 (N=8 hybrid systolic, real, `v3-artix7` for the physically-fabricated board; branch
closed timing — the current, definitive real deployment target, chosen `n16-timing-closure` for the real, consolidated follow-up below),
by the user after a real, measured N=4/8/16 timing curve). The current updated through EXP-0097/0098 (a real MAC-pipeline fix that closes
real, trustworthy *timing* signoff is now EXP-0095/0096 timing for the WHOLE real family — N=2/N=8/N=16 — not just N=8). The
(`n8_system_ddr3_top.v`, 32-bit DDR3, 155.039MHz, WNS 0.000ns, 8 real board **currently in physical fabrication** uses N=8's own original,
parallel PEs), replacing EXP-0088's own N=2 pointer (which itself unmodified EXP-0095/0096 signoff (`n8_system_ddr3_top.v`, 32-bit DDR3,
replaced EXP-0086's, which replaced EXP-0083's 16-bit-era one). See §5.6 155.039MHz, WNS=0.000ns exact, 8 real parallel PEs). Separately, on an
for the full real story, including N=16's own real RTL (functionally isolated real branch, EXP-0097/0098 found that one shared-core fix
verified but NOT timing-closed, kept as documented future work). Every closes N=16 too (WNS=+0.269ns) and improves N=2 (+0.389ns) and N=8
(+0.108ns) as well — see §5.6 for the full real story. Every
number in this document is either number in this document is either
directly measured (real simulation trace, real P&R report) or a calculation directly measured (real simulation trace, real P&R report) or a calculation
built from directly-measured building blocks — the two are labeled built from directly-measured building blocks — the two are labeled
@@ -664,7 +665,7 @@ specifically to document where/how it breaks rather than to succeed):
--- ---
### 5.6 [N=8 IS THE REAL, DEFINITIVE DEPLOYMENT TARGET — `n8_system_ddr3_top.v`, EXP-0089…0096] Hybrid systolic scaling: up to 4 groups × 4-PE weight-stationary chains ### 5.6 [WHOLE FAMILY CLOSED — N=2/N=8/N=16, EXP-0089…0098] Hybrid systolic scaling: up to 4 groups × 4-PE weight-stationary chains
Captured from a 2026-09-20 brainstorming session as a purely exploratory Captured from a 2026-09-20 brainstorming session as a purely exploratory
idea; the same day, per the user's own explicit reprioritization, the real idea; the same day, per the user's own explicit reprioritization, the real
@@ -857,16 +858,79 @@ lanes exercised exactly once): **real functional xsim, 16/16 PASS, 0
errors** — closes the real, disclosed functional-verification gap errors** — closes the real, disclosed functional-verification gap
this section previously flagged for N=8 specifically. this section previously flagged for N=8 specifically.
**Current real status: N=8 (`n8_system_ddr3_top.v`) is BOTH **Current real status (physical board): N=8 (`n8_system_ddr3_top.v`)
functionally verified AND timing-closed under its own permanent real is BOTH functionally verified AND timing-closed under its own
name — the definitive deployment target.** N=2 (`n2_system_ddr3_top.v`, permanent real name — the configuration currently in physical
EXP-0088) remains documented as a real, valid, simpler fallback. N=16 fabrication.** N=2 (`n2_system_ddr3_top.v`, EXP-0088) remains
(`n16_system_ddr3_top.v`) remains real, functionally-verified RTL, not documented as a real, valid, simpler fallback. Real, disclosed caveat:
abandoned, just not timing-closed and not the current target. Real, N=8's original WNS=0.000ns is an exact-zero margin — any future RTL
disclosed caveat carried forward: N=8's WNS=0.000ns is an exact-zero change touching this top-level or its dependents needs a fresh real
margin — any future RTL change touching this top-level or its P&R before trusting timing again. **This exact-zero margin is what
dependents needs a fresh real P&R (same `Explore`/`ExtraNetDelay_high`/ motivated EXP-0097/0098 below.**
`AggressiveExplore` directive stack) before trusting timing again.
#### 5.6.1 [DONE, EXP-0097/0098] The whole family closes: a real MAC-pipeline fix, branch `n16-timing-closure`
**Context**: N=16's own real, traced remaining bottleneck (after the
hierarchical arbiter + P&R directive tuning already closed most of the
gap, WNS 0.913ns → 0.338ns, EXP-0094) was inside `neural_processor_
packed.v`'s own DSP48E1 MAC datapath — a pre-existing, N=2-era design
(unchanged since EXP-0059) with an already razor-thin real margin
(+0.099962ns) that N=16's own higher real die congestion eroded past
zero, and that N=8's own real signoff (§5.6, EXP-0095/0096) closed at
an exact zero with no slack to spare.
**Real root cause** (traced to the exact register-to-register hop, not
guessed): the critical path ran from a DSP48E1's own (Vivado-auto-
retimed) product register, straight through the real carry-heavy
INT8-unpack logic (`pb_comb`'s own shift + conditional +1
carry-propagate add, CARRY4-dominated), into `proda1`/`prodb1` in a
SINGLE real cycle.
**Real fix**: split the original single pipeline stage into two real
stages — Stage 1a registers the RAW DSP48E1 product with zero logic in
between (a real, explicit register boundary immediately after the
multiply); Stage 1b does the carry-heavy unpack FROM the already-
registered value (unchanged real math, now one real cycle later). Real,
deliberate consequence: end-to-end per-tile latency grows by exactly
ONE real clock cycle; throughput is unaffected (real valid/ready
handshaking throughout, no fixed-latency assumption anywhere
downstream).
**Real bug found and fixed in the testbench first** (not an RTL bug):
`tb_neural_processor_packed.v`'s own comparison logic required all
three cores (2 real reference `neural_processor.v` instances + the
DUT) to assert `result_valid` simultaneously — correct only when all
three share the exact same real pipeline depth, no longer true once
the DUT is deliberately one real cycle deeper. Fixed by latching each
core's own result independently the cycle its own `result_valid`
first pulses, then comparing the three LATCHED values.
**Real, consolidated result — EVERY real top-level this project has,
same shared pipelined core, real functional xsim + real, full P&R**:
| Target | Real WNS | Real DSP48E1 | Real functional test |
|---|---|---|---|
| N=2 (`n2_system_ddr3_top.v`) | +0.389ns (up from +0.099962ns) | 16 (6.67%) | 8/8 PASS |
| N=8 (`n8_system_ddr3_top.v`) | +0.108ns (up from 0.000ns exact) | 64 (26.7%) | 16/16 PASS |
| N=16 (`n16_system_ddr3_top.v`) | **+0.269ns** (up from 0.338ns FAILED) | 128 (53.3%) | 32/32 PASS |
A real margin-hunt attempt for even more N=16 headroom (alternate P&R
directives) gave a WORSE real result (+0.168ns vs. the original
+0.269ns) — confirms real P&R run-to-run variance, not a further
systematic gain available without touching more RTL (the new N=16
worst path, after this fix, moved to `neural_director_grouped.v`'s own
queue update logic, still real route-dominated — not pursued further,
current margin already comfortable).
**Decision**: this fix is a pure, unconditional real improvement
across the whole real family — no regression anywhere, real margin
improves at every N, N=16 goes from real FAILING to real CLOSED. It is
**not yet applied to the physically fabricated board** (which uses the
original, unmodified N=8 core) — adopting it (and/or moving to N=16)
for a future board revision is a real, open hardware decision, not an
RTL blocker. A real Vivado incremental-synthesis quirk was also found
and fixed along the way (a parameter binding silently carried forward
from an earlier `-generic` run despite `reset_run` — see CLAUDE.md).
**The problem it targets**: plain N=16 independent cores (§5.5's own **The problem it targets**: plain N=16 independent cores (§5.5's own
"documentary, expected to break" framing) means 16 independent DDR3 "documentary, expected to break" framing) means 16 independent DDR3
+65 -37
View File
@@ -165,25 +165,39 @@ these from an ESP32 doing real JTAG bit-banging (TAP state machine, IR/DR
shifting) rather than a bench programmer — that firmware is separate, shifting) rather than a bench programmer — that firmware is separate,
software-side work, not covered here. software-side work, not covered here.
## 3. Real timing signoff (EXP-0095/0096, the current, definitive target — N=8 hybrid systolic, `n8_system_ddr3_top.v`, CLOSED) ## 3. Real timing signoff — N=8 in physical fabrication; N=2/N=8/N=16 all real, closed on branch `n16-timing-closure` (EXP-0095…0098)
**This is the real, definitive deployment target as of 2026-09-21** **Real, current, two-track status (2026-09-22)**: the **physical board
(the user's own explicit decision, after a real, measured N=4/8/16 currently being fabricated** uses **N=8** (`hardware/v3/rtl/
timing curve found N=8 to be the largest real configuration that n8_system_ddr3_top.v`, unmodified, EXP-0095/0096 signoff below) — this
closes timing cleanly — see `docs/ARCHITECTURE_ANALYSIS.md` §5.6 for was the real, definitive decision made on 2026-09-21 after a real,
the full real story). Top-level RTL: `hardware/v3/rtl/ measured N=4/8/16 timing curve found N=8 the largest configuration
n8_system_ddr3_top.v` — 2 groups × 4 PEs, shared-weight-broadcast that closed cleanly at the time. Separately, on an isolated real
systolic architecture (`systolic_group.v`/`packed_pe.v`, development branch (`n16-timing-closure`, does NOT affect the board
`neural_director_grouped.v`, `sdram_arbiter_hier.v`'s own real already in fabrication), a real fix (EXP-0097 — an extra real pipeline
hierarchical 2-level arbiter), same real DDR3/SPI/flash bridge as stage in `neural_processor_packed.v`'s own MAC datapath, see
every other real V3 top-level. Real in-context Vivado place-and-route `docs/ARCHITECTURE_ANALYSIS.md` §5.6) has since **closed real timing
(not out-of-context, not estimated): for N=16 too, and improved N=2's and N=8's own margins** — the whole
real product family (N=2/N=8/N=16) now closes with positive margin
under that fix. Whether/when to adopt this fix for a future board
revision (potentially N=16) is a real, open hardware decision, not an
RTL blocker — see `docs/ARCHITECTURE_ANALYSIS.md` §5.6 for the full
story.
### 3.0 N=8, in physical fabrication (`n8_system_ddr3_top.v`, EXP-0095/0096, unmodified core)
Top-level RTL: `hardware/v3/rtl/n8_system_ddr3_top.v` — 2 groups × 4
PEs, shared-weight-broadcast systolic architecture
(`systolic_group.v`/`packed_pe.v`, `neural_director_grouped.v`,
`sdram_arbiter_hier.v`'s own real hierarchical 2-level arbiter), same
real DDR3/SPI/flash bridge as every other real V3 top-level. Real
in-context Vivado place-and-route (not out-of-context, not estimated):
| Metric | Value | | Metric | Value |
|---|---| |---|---|
| DDR3 PHY clock (sys_clk_p/n) | **310.078 MHz** (3.225ns period) | | DDR3 PHY clock (sys_clk_p/n) | **310.078 MHz** (3.225ns period) |
| Compute domain clock (ui_clk/clk_pll_i, PLL-derived 2:1 from sys_clk) | **155.039 MHz** | | Compute domain clock (ui_clk/clk_pll_i, PLL-derived 2:1 from sys_clk) | **155.039 MHz** |
| WNS (setup slack) | **0.000 ns** — real, closed, 0 failing setup endpoints (EXP-0095/0096). Re-verify with a fresh P&R after ANY further logic addition — this margin is exactly zero, no slack to spare. | | WNS (setup slack) | **0.000 ns** — real, closed, 0 failing setup endpoints (EXP-0095/0096). This is the exact RTL physically being fabricated. |
| WHS (hold slack) | +0.017 ns | | WHS (hold slack) | +0.017 ns |
| Failing endpoints | 0 of 63212 (setup), 0 of 63209 (hold) | | Failing endpoints | 0 of 63212 (setup), 0 of 63209 (hold) |
| LUTs used | 12535 / 63400 (19.77%) | | LUTs used | 12535 / 63400 (19.77%) |
@@ -200,20 +214,31 @@ real, closed, but with zero slack. Re-run a real P&R (same
ExtraNetDelay_high` / `phys_opt_design -directive AggressiveExplore` / ExtraNetDelay_high` / `phys_opt_design -directive AggressiveExplore` /
`route_design -directive AggressiveExplore` directive stack used to `route_design -directive AggressiveExplore` directive stack used to
reach this result, EXP-0094/0095) after ANY further RTL change before reach this result, EXP-0094/0095) after ANY further RTL change before
trusting timing again — do not assume it still holds. trusting timing again — do not assume it still holds. This zero-margin
result is exactly why EXP-0097's own pipeline fix (below) was pursued.
**Superseded, real predecessor signoff (N=2, flat architecture, ### 3.0.1 The real, consolidated family — EXP-0097/0098, branch `n16-timing-closure` (not yet on the fabricated board)
`n2_system_ddr3_top.v`, EXP-0088)** — kept below for reference; still
a real, valid, simpler fallback target if N=8's zero-margin closure
ever proves too fragile in practice:
| Metric | Value | A single real fix (`neural_processor_packed.v`, an extra real pipeline
|---|---| stage in the packed-MAC datapath — full real root-cause story in
| WNS (setup slack) | +0.099962 ns | `docs/ARCHITECTURE_ANALYSIS.md` §5.6) applies identically to every
| WHS (hold slack) | +0.036275 ns | real top-level in the family (it's the SAME shared compute core). Real
| LUTs used | 6642 / 63400 (10.48%) | functional xsim + real, full P&R, every real configuration this
| DSP48E1 used | 16 / 240 (6.67%) — 8 per compute core × 2 cores | project has a dedicated top-level for:
| Real parallelism | 2 real parallel processing elements |
| Target | Top-level | Real WNS | Real DSP48E1 | Real functional test |
|---|---|---|---|---|
| N=2 | `n2_system_ddr3_top.v` | **+0.389 ns** (up from +0.100ns) | 16/240 (6.67%) | 8/8 PASS |
| N=8 | `n8_system_ddr3_top.v` | **+0.108 ns** (up from 0.000ns exact) | 64/240 (26.7%) | 16/16 PASS |
| N=16 | `n16_system_ddr3_top.v` | **+0.269 ns** (up from 0.338ns FAILED) | 128/240 (53.3%) | 32/32 PASS |
**Real, honest takeaway**: this fix is a pure, unconditional
improvement across the whole real family — no regression found
anywhere, real margin improves at every N, and N=16 goes from real
FAILING to real CLOSED. It is not yet applied to the physically
fabricated board (which uses the original, unmodified N=8 core,
§3.0) — that board's own real signoff (WNS=0.000ns) remains accurate
for the hardware actually being built.
Signoff history (every real change, in-context P&R): Signoff history (every real change, in-context P&R):
@@ -230,7 +255,10 @@ Signoff history (every real change, in-context P&R):
| 0088 (+result-writeback engine, `result_writeback.v`) | N=2 | +0.100 (CLOSED) | 6642 | 16 | | 0088 (+result-writeback engine, `result_writeback.v`) | N=2 | +0.100 (CLOSED) | 6642 | 16 |
| 0093 (hybrid systolic, flat 21-way arbiter) | N=16 | -0.913 (FAILED) | 19751 | 128 | | 0093 (hybrid systolic, flat 21-way arbiter) | N=16 | -0.913 (FAILED) | 19751 | 128 |
| 0094 (+hierarchical 2-level arbiter, +P&R directive tuning) | N=16 | -0.338 (still FAILED) | 19936 | 128 | | 0094 (+hierarchical 2-level arbiter, +P&R directive tuning) | N=16 | -0.338 (still FAILED) | 19936 | 128 |
| 0095/0096 (`n8_system_ddr3_top.v`, hybrid systolic) | **N=8** | **0.000 (CLOSED)** | 12535 | 64 | | 0095/0096 (`n8_system_ddr3_top.v`, hybrid systolic**in physical fabrication**) | **N=8** | **0.000 (CLOSED)** | 12535 | 64 |
| 0097/0098 (+MAC pipeline fix, branch `n16-timing-closure`) | N=2 | +0.389 (CLOSED) | 6645 | 16 |
| 0097/0098 (+MAC pipeline fix, branch `n16-timing-closure`) | N=8 | +0.108 (CLOSED) | 12536 | 64 |
| 0097/0098 (+MAC pipeline fix, branch `n16-timing-closure`) | **N=16** | **+0.269 (CLOSED)** | 19903 | 128 |
### 3.1 Real 32-bit DDR3 widening — functionally verified AND timing CLOSED (EXP-0084 → EXP-0086) ### 3.1 Real 32-bit DDR3 widening — functionally verified AND timing CLOSED (EXP-0084 → EXP-0086)
@@ -397,17 +425,17 @@ when it actually fires, instead of polling every loop iteration.
## 7. Known-open items (honestly disclosed, not hidden) ## 7. Known-open items (honestly disclosed, not hidden)
- Scaling past N=2 compute cores: **RESOLVED, real N=8 is now the - Scaling past N=2 compute cores: **RESOLVED, N=8 is the configuration
definitive deployment target** (§3, EXP-0095/0096, `n8_system_ddr3_ in physical fabrication** (§3.0, EXP-0095/0096, `n8_system_ddr3_
top.v`, a real hybrid systolic shared-weight-broadcast architecture, top.v`, WNS=0.000ns CLOSED). **UPDATE (EXP-0097/0098)**: a real MAC-
2 groups × 4 PEs, WNS=0.000ns CLOSED). Real N=4/8/16 P&R curve pipeline fix, developed on an isolated branch (`n16-timing-closure`,
measured: N=4 0.005ns (2 failing endpoints, essentially at the does not affect the board in fabrication), closes real timing for
edge), N=8 **0.000ns (CLOSED)**, N=16 0.338ns (still failing even N=16 too (WNS=+0.269ns) and improves N=2's and N=8's own margins
after a real hierarchical-arbiter fix + P&R directive tuning, EXP- (+0.389ns, +0.108ns respectively) — see §3.0.1 and
0093/0094) — see `docs/ARCHITECTURE_ANALYSIS.md` §5.6 for the full `docs/ARCHITECTURE_ANALYSIS.md` §5.6 for the full real story. N=16
real story. N=16 remains real, functionally verified RTL is no longer blocked by a real timing problem; whether to adopt it
(`n16_system_ddr3_top.v`), just not yet timing-closed; not the for a future board revision is now a real hardware decision, not an
current real deployment target. RTL one.
- `sys_rst` has only a *tentative* real pin (G13, bank 15, EXP-0084 — - `sys_rst` has only a *tentative* real pin (G13, bank 15, EXP-0084 —
chosen just to unblock real P&R, not a final board decision) — assign chosen just to unblock real P&R, not a final board decision) — assign
its real, permanent location once the rest of the board layout (reset its real, permanent location once the rest of the board layout (reset
+72 -19
View File
@@ -6,12 +6,19 @@
Il design realizzato in FPGA \`e un acceleratore per reti neurali Il design realizzato in FPGA \`e un acceleratore per reti neurali
quantizzate INT8, organizzato come un insieme di \textbf{elementi di quantizzate INT8, organizzato come un insieme di \textbf{elementi di
elaborazione} (Processing Element, PE) paralleli che condividono un elaborazione} (Processing Element, PE) paralleli che condividono un
unico canale DDR3 reale. La configurazione attualmente definitiva unico canale DDR3 reale, secondo un'architettura sistolica a
(Capitolo~\ref{chap:tests}) \`e \textbf{N=8}: 8 PE reali, organizzati \emph{broadcast dei pesi condiviso} (\emph{shared-weight broadcast}).
in \textbf{2 gruppi da 4 PE ciascuno}, secondo un'architettura
sistolica a \emph{broadcast dei pesi condiviso} (\emph{shared-weight \textbf{Stato reale a due configurazioni (Capitolo~\ref{chap:tests})}:
broadcast}). Il modulo top-level reale \`e \textbf{N=8} (\texttt{n8\_system\_ddr3\_top.v}, 2 gruppi da 4 PE) \`e
\texttt{hardware/v3/rtl/n8\_system\_ddr3\_top.v}. la configurazione attualmente in fabbricazione fisica sulla scheda
reale. \textbf{N=16} (\texttt{n16\_system\_ddr3\_top.v}, 4 gruppi da 4
PE) \`e ora, a seguito di una correzione reale del datapath MAC
(\S\ref{sec:mac-pipeline}), \textbf{funzionalmente verificata E con
timing reale chiuso} su un branch di sviluppo dedicato
(\texttt{n16-timing-closure}), non ancora promossa alla scheda fisica
in produzione --- una decisione hardware reale, non RTL, ancora da
prendere.
L'intero design gira in un unico dominio di clock reale a L'intero design gira in un unico dominio di clock reale a
\textbf{155.039\,MHz} (\texttt{ui\_clk}, derivato dal MIG DDR3, \textbf{155.039\,MHz} (\texttt{ui\_clk}, derivato dal MIG DDR3,
@@ -38,9 +45,45 @@ Al clock reale di 155.039\,MHz, il picco teorico per singolo PE \`e:
Questo nucleo \`e rimasto \emph{invariato} (stesso numero di DSP, Questo nucleo \`e rimasto \emph{invariato} (stesso numero di DSP,
stessa struttura) dalla primissima sintesi reale del progetto fino stessa struttura) dalla primissima sintesi reale del progetto fino
alla configurazione N=8 attuale: \`e la parte pi\`u stabile ed efficiente alla configurazione N=8: \`e la parte pi\`u stabile ed efficiente
del design, e non \`e mai stato il collo di bottiglia delle prestazioni del design. A N=16 (\S\ref{sec:mac-pipeline}) \`e per\`o diventato,
(si veda \S\ref{sec:bottleneck}). per la prima volta, il vero collo di bottiglia di temporizzazione ---
non del throughput di calcolo (\S\ref{sec:bottleneck}), ma della
temporizzazione fisica del place-and-route.
\subsection{Correzione reale: pipeline aggiuntiva nel datapath MAC (per N=16)}
\label{sec:mac-pipeline}
A N=16, il margine di temporizzazione reale --- gi\`a estremamente
sottile a N=2 (+0.0999962\,ns) --- \`e stato eroso oltre lo zero dalla
maggiore congestione fisica complessiva del die (WNS reale misurato:
$-0.338$\,ns, dopo un arbitro gerarchico e un tuning delle direttive
di place-and-route, Capitolo~\ref{chap:tests}). Il percorso critico
reale, tracciato (non ipotizzato) fino al livello del singolo
registro, va dall'uscita del moltiplicatore DSP48E1 (gi\`a
ri-temporizzata automaticamente da Vivado) attraverso la logica di
``spacchettamento'' dei due prodotti INT8 impacchettati (uno shift e
una somma di riporto condizionale, dominata da primitive CARRY4) fino
al registro che cattura il risultato --- tutto in un solo ciclo di
clock.
\textbf{Correzione reale applicata}: lo stadio originale \`e stato
spezzato in due stadi di pipeline reali distinti --- il primo registra
il prodotto grezzo del DSP48E1 senza alcuna logica intermedia; il
secondo esegue lo spacchettamento (matematica identica, invariata bit
per bit) a partire dal valore gi\`a registrato. Il costo reale: un solo
ciclo di clock aggiuntivo di latenza per operazione, senza alcun
impatto sul throughput (l'interfaccia a maniglia valid/ready non
assume mai una latenza fissa). Verificato bit-esatto in isolamento
(18/18 PASS contro 2 core reali di riferimento) e funzionalmente
sull'intero sistema N=16 su DDR3 reale (32/32 PASS).
\textbf{Risultato reale}: con questa sola correzione, il timing di
N=16 \textbf{chiude realmente} (WNS $=+0.269$\,ns, WHS $=+0.026$\,ns,
0 endpoint falliti) --- si veda il Capitolo~\ref{chap:tests} per i
numeri completi. La stessa correzione, applicata anche a N=8, migliora
pure il suo margine (da 0.000\,ns esatto a +0.108\,ns) senza alcuna
regressione funzionale in nessuna delle due configurazioni.
\section{L'architettura sistolica a gruppi (N=8)} \section{L'architettura sistolica a gruppi (N=8)}
@@ -180,20 +223,30 @@ margine di scalabilit\`a residuo del chip \`e ampio sul lato
computazionale, ma \`e vincolato dal canale di memoria condiviso, non computazionale, ma \`e vincolato dal canale di memoria condiviso, non
dalla logica di calcolo. dalla logica di calcolo.
\section{Configurazioni alternative esplorate} \section{Configurazioni esplorate e stato reale attuale}
Lo stesso RTL sistolico, tramite il parametro reale \texttt{N\_GROUPS}, Lo stesso RTL sistolico, tramite il parametro reale \texttt{N\_GROUPS},
\`e stato realmente sintetizzato e verificato anche a: \`e stato realmente sintetizzato e verificato a pi\`u configurazioni:
\begin{itemize} \begin{itemize}
\item \textbf{N=4} (1 gruppo, 4 PE, 32 DSP48E1) --- funzionante, \item \textbf{N=4} (1 gruppo, 4 PE, 32 DSP48E1) --- funzionante,
margine di timing quasi nullo. margine di timing quasi nullo (non ancora rifinito con la
\item \textbf{N=16} (4 gruppi, 16 PE, 128 DSP48E1) --- funzionalmente correzione di \S\ref{sec:mac-pipeline}).
verificato su DDR3 reale, ma con temporizzazione non ancora chiusa \item \textbf{N=8} (\texttt{n8\_system\_ddr3\_top.v}) --- \textbf{la
(si veda \S~\ref{sec:pnr-history}) --- mantenuto come lavoro futuro configurazione fisicamente in fabbricazione sulla scheda reale
documentato, non abbandonato. attuale.} Timing chiuso, WNS $=+0.108$\,ns con la correzione MAC.
\item \textbf{N=16} (\texttt{n16\_system\_ddr3\_top.v}) --- dopo la
correzione di \S\ref{sec:mac-pipeline}, \textbf{funzionalmente
verificato E con timing reale chiuso} (WNS $=+0.269$\,ns), su un
branch di sviluppo reale (\texttt{n16-timing-closure}) separato
dalla scheda in produzione. \`E ora un candidato reale, verificato,
per una futura revisione della scheda --- non ancora promosso alla
produzione fisica corrente, una decisione hardware reale ancora da
prendere con l'utente.
\end{itemize} \end{itemize}
N=8 \`e stata scelta come configurazione reale definitiva perch\'e N=8 resta, ad oggi, la configurazione realmente fabbricata. N=16 non
\`e la pi\`u grande che chiude realmente il timing con margine \`e pi\`u un limite architetturale reale (come inizialmente sembrava),
positivo (Capitolo~\ref{chap:tests}). ma una reale, verificata alternativa a parallelismo doppio, la cui
adozione fisica dipende ora da una scelta dell'utente, non da un
vincolo tecnico residuo.
+73 -53
View File
@@ -106,12 +106,17 @@ shared-weight-broadcast systolic architecture), reducing redundant
DDR3 weight traffic without the complexity of a literal PE-to-PE DDR3 weight traffic without the complexity of a literal PE-to-PE
systolic pipeline. systolic pipeline.
N8 is the current, definitive, real deployment configuration of the N8 is the configuration currently in real physical board fabrication.
FPGA-Neural family: it is the largest core count in the family that A real MAC-datapath pipeline fix (developed on the
closes real static timing on the target part with a real, verified \texttt{n16-timing-closure} branch, isolated from the board already in
margin. Smaller (N=2, N=4) and larger (N=16) configurations of the production) has since ALSO closed real static timing for \textbf{N16}
same RTL exist and are documented (see \S\ref{sec:variants}), but N8 with a comfortable real margin (WNS=+0.269\,ns) and improved N8's own
is the one recommended for real board fabrication. margin further (WNS=+0.108\,ns, up from the original exact-zero
0.000\,ns) -- both configurations are now real, functionally verified,
timing-closed options (see \S\ref{sec:variants}). N8 remains the
configuration in current physical production; N16 is a real, verified
candidate for a future board revision, pending the user's own hardware
decision.
\clearpage \clearpage
@@ -132,14 +137,14 @@ for reference.
\toprule \toprule
\textbf{Variant} & \textbf{PEs} & \textbf{DSP48E1} & \textbf{Real WNS} & \textbf{Status} & \textbf{Top-level file} \\ \textbf{Variant} & \textbf{PEs} & \textbf{DSP48E1} & \textbf{Real WNS} & \textbf{Status} & \textbf{Top-level file} \\
\midrule \midrule
N2 & 2 & 16 & +0.100\,ns & \textbf{Closed, deployable fallback} & \texttt{n2\_system\_ddr3\_top.v} \\ N2 & 2 & 16 & +0.100\,ns & Closed, deployable fallback & \texttt{n2\_system\_ddr3\_top.v} \\
N4 & 4 & 32 & $-0.005$\,ns & Nearly closed (2 failing endpoints) & \emph{(generic override only)} \\ N4 & 4 & 32 & $-0.005$\,ns & Nearly closed (2 failing endpoints), pre-fix & \emph{(generic override only)} \\
\rowcolor{dslight} \rowcolor{dslight}
N8 & 8 & 64 & \textbf{0.000\,ns} & \textbf{RECOMMENDED --- closed} & \texttt{n8\_system\_ddr3\_top.v} \\ N8 & 8 & 64 & \textbf{+0.108\,ns} & \textbf{IN PHYSICAL PRODUCTION --- closed} & \texttt{n8\_system\_ddr3\_top.v} \\
N16 & 16 & 128 & $-0.338$\,ns & Functional, timing not closed & \texttt{n16\_system\_ddr3\_top.v} \\ N16 & 16 & 128 & \textbf{+0.269\,ns} & \textbf{Closed, real candidate for next revision} & \texttt{n16\_system\_ddr3\_top.v} \\
\bottomrule \bottomrule
\end{tabular} \end{tabular}
\caption{Real family configurations. All share identical DDR3/SPI/flash physical connections and identical host protocol.} \caption{Real family configurations, post MAC-pipeline fix (\texttt{n16-timing-closure} branch) for N8/N16. All share identical DDR3/SPI/flash physical connections and identical host protocol.}
\end{table} \end{table}
%====================================================================== %======================================================================
@@ -322,46 +327,51 @@ CFGBVS & P8 & & & \\
\clearpage \clearpage
\section{Timing Characteristics} \section{Timing Characteristics}
Real, in-context Vivado place-and-route on XC7A100T-CSG324-2, N8 Real, in-context Vivado place-and-route on XC7A100T-CSG324-2, both
configuration (\texttt{n8\_system\_ddr3\_top.v}), directive stack with the real MAC-pipeline fix (\texttt{n16-timing-closure} branch),
directive stack
\texttt{Explore}/\texttt{ExtraNetDelay\_high}/\texttt{AggressiveExplore}: \texttt{Explore}/\texttt{ExtraNetDelay\_high}/\texttt{AggressiveExplore}:
\begin{table}[h] \begin{table}[h]
\centering \centering
\renewcommand{\arraystretch}{1.2} \renewcommand{\arraystretch}{1.2}
\begin{tabular}{lll} \begin{tabular}{llll}
\toprule \toprule
\textbf{Parameter} & \textbf{Value} & \textbf{Condition} \\ \textbf{Parameter} & \textbf{N8 (in production)} & \textbf{N16 (verified)} & \textbf{Condition} \\
\midrule \midrule
Setup slack (WNS) & \textbf{0.000\,ns} & \texttt{clk\_pll\_i}, 155.039\,MHz, 0 of 63212 failing \\ Setup slack (WNS) & +0.108\,ns & +0.269\,ns & \texttt{clk\_pll\_i}, 155.039\,MHz, 0 failing \\
Hold slack (WHS) & +0.017\,ns & 0 of 63209 failing \\ Hold slack (WHS) & +0.036\,ns & +0.026\,ns & 0 failing \\
Compute-domain clock period & 6.450\,ns & 155.039\,MHz \\ Compute-domain clock period & 6.450\,ns & 6.450\,ns & 155.039\,MHz \\
DDR3 CAS latency (real, measured) & CL=5 (16.125\,ns) & matches MIG configuration \\ DDR3 CAS latency (real, measured) & \multicolumn{2}{c}{CL=5 (16.125\,ns)} & matches MIG configuration \\
Same-row read-to-read & 12.9\,ns & real JEDEC trace \\ Same-row read-to-read & \multicolumn{2}{c}{12.9\,ns} & real JEDEC trace \\
Row-switch (Activate$\to$Read) penalty & +25--130\% & vs.\ same-row access, real measured \\ Row-switch penalty & \multicolumn{2}{c}{+25--130\%} & vs.\ same-row access, real measured \\
\bottomrule \bottomrule
\end{tabular} \end{tabular}
\end{table} \end{table}
\noindent\textbf{Margin notice:} the real WNS is exactly zero --- closed, \noindent\textbf{Margin notice:} both real margins are positive and
but with no spare slack. Any RTL modification to this configuration closed. N8's own pre-fix signoff (\texttt{n8\_system\_ddr3\_top.v}
requires a fresh real place-and-route run (same directive stack) before without the MAC-pipeline fix, the version currently in physical
the design may be considered timing-clean again. fabrication) closed at an exact-zero WNS=0.000\,ns, no spare slack ---
the improved +0.108\,ns figure above applies once the pipeline fix
(\S\ref{sec:variants}) is adopted. Any RTL modification requires a
fresh real place-and-route run (same directive stack) before the
design may be considered timing-clean again.
%====================================================================== %======================================================================
\section{Resource Utilization (XC7A100T-CSG324-2, N8)} \section{Resource Utilization (XC7A100T-CSG324-2)}
\begin{table}[h] \begin{table}[h]
\centering \centering
\renewcommand{\arraystretch}{1.2} \renewcommand{\arraystretch}{1.2}
\begin{tabular}{lrrr} \begin{tabular}{lrrrr}
\toprule \toprule
\textbf{Resource} & \textbf{Used} & \textbf{Available} & \textbf{Utilization} \\ \textbf{Resource} & \textbf{N8 Used} & \textbf{N16 Used} & \textbf{Available} & \textbf{N8 / N16 Util.} \\
\midrule \midrule
Slice LUTs & 12535 & 63400 & 19.77\% \\ Slice LUTs & 12536 & 19903 & 63400 & 19.77\% / 31.39\% \\
Slice Registers & 19902 & 126800 & 15.70\% \\ Slice Registers & 19910 & 35409 & 126800 & 15.70\% / 27.93\% \\
DSP48E1 & 64 & 240 & 26.7\% \\ DSP48E1 & 64 & 128 & 240 & 26.7\% / 53.33\% \\
Block RAM & 0 & 135 & 0\% \\ Block RAM & 0 & 0 & 135 & 0\% / 0\% \\
\bottomrule \bottomrule
\end{tabular} \end{tabular}
\end{table} \end{table}
@@ -371,28 +381,29 @@ Block RAM & 0 & 135 & 0\% \\
\begin{table}[h] \begin{table}[h]
\centering \centering
\begin{tabular}{ll} \begin{tabular}{lll}
\toprule \toprule
\textbf{Parameter} & \textbf{Value} \\ \textbf{Parameter} & \textbf{N8} & \textbf{N16} \\
\midrule \midrule
MAC operations per DSP48E1 per cycle & 2 (packed INT8, lane A + lane B) \\ MAC operations per DSP48E1 per cycle & \multicolumn{2}{c}{2 (packed INT8, lane A + lane B)} \\
DSP48E1 per processing element & 8 \\ DSP48E1 per processing element & \multicolumn{2}{c}{8} \\
MAC/cycle per PE & 16 \\ MAC/cycle per PE & \multicolumn{2}{c}{16} \\
Real parallel PEs & 8 \\ Real parallel PEs & 8 & 16 \\
Aggregate MAC/cycle & 128 \\ Aggregate MAC/cycle & 128 & 256 \\
\textbf{Aggregate peak throughput} (calculated, 155.039\,MHz) & \textbf{$\sim$19.84\,GMAC/s} \\ \textbf{Aggregate peak throughput} (calc., 155.039\,MHz) & \textbf{$\sim$19.84\,GMAC/s} & \textbf{$\sim$39.69\,GMAC/s} \\
Real DDR3 physical bandwidth ceiling & $\sim$2.48\,GB/s (shared across all 8 PEs) \\ Real DDR3 bandwidth ceiling (shared) & \multicolumn{2}{c}{$\sim$2.48\,GB/s} \\
\bottomrule \bottomrule
\end{tabular} \end{tabular}
\end{table} \end{table}
\noindent The peak MAC figure is a real calculation from measured \noindent The peak MAC figures are a real calculation from measured
building blocks (DSP count $\times$ packing factor $\times$ measured building blocks (DSP count $\times$ packing factor $\times$ measured
Fmax) --- not itself an independently benchmarked sustained-throughput Fmax) --- not themselves independently benchmarked sustained-throughput
number. Sustained throughput on a real workload is bounded by the real numbers. Sustained throughput on a real workload is bounded by the real
DDR3 bandwidth ceiling above, shared across all active PEs; a direct DDR3 bandwidth ceiling above, shared across all active PEs (more
measurement of sustained aggregate throughput at N8 has not yet been tightly at N16, twice the PEs on the same physical channel); a direct
performed (see \S\ref{sec:openitems}). measurement of sustained aggregate throughput has not yet been
performed at either N (see \S\ref{sec:openitems}).
%====================================================================== %======================================================================
\section{Host Communication Protocol} \section{Host Communication Protocol}
@@ -470,8 +481,8 @@ Host MCU & ESP32-S3-WROOM-1-N16R8 & In stock, LCSC C2913202 \\
\label{sec:openitems} \label{sec:openitems}
\begin{itemize} \begin{itemize}
\item Physical board has not yet been fabricated; no real \item Physical board fabrication is underway (N8 configuration,
electrical/thermal measurement exists yet. unmodified); no real electrical/thermal measurement exists yet.
\item \texttt{sys\_clk} oscillator requires a custom-programmed order \item \texttt{sys\_clk} oscillator requires a custom-programmed order
(no off-the-shelf SKU at 310.077\,MHz). (no off-the-shelf SKU at 310.077\,MHz).
\item Power-tree regulators (VCCINT, VCCAUX, VTT, per-bank VCCO) not \item Power-tree regulators (VCCINT, VCCAUX, VTT, per-bank VCCO) not
@@ -481,10 +492,14 @@ Host MCU & ESP32-S3-WROOM-1-N16R8 & In stock, LCSC C2913202 \\
\item ESP32-side firmware (JTAG bit-banging, SPI host driver, \item ESP32-side firmware (JTAG bit-banging, SPI host driver,
result-readback address translation) not yet written --- software result-readback address translation) not yet written --- software
work outside this document's scope. work outside this document's scope.
\item Real, direct sustained-throughput benchmark at N8 (vs.\ a \item Real, direct sustained-throughput benchmark (vs.\ a real
real ESP32-S3 baseline) not yet performed --- see the project's ESP32-S3 baseline) not yet performed at either N8 or N16 --- see
Tests \& Timing chapter for the honest historical comparison the project's Tests \& Timing chapter for the honest historical
context. comparison context and a real, clearly-labeled projection for N16.
\item Whether to adopt N16 (real, functionally verified, timing-
closed on the \texttt{n16-timing-closure} branch) for a future
board revision is a real, open hardware decision, not an RTL
blocker.
\end{itemize} \end{itemize}
%====================================================================== %======================================================================
@@ -497,6 +512,11 @@ Host MCU & ESP32-S3-WROOM-1-N16R8 & In stock, LCSC C2913202 \\
\textbf{Rev.} & \textbf{Date} & \textbf{Description} \\ \textbf{Rev.} & \textbf{Date} & \textbf{Description} \\
\midrule \midrule
1.0 & 2026-09-21 & Initial preliminary datasheet, N8 configuration \\ 1.0 & 2026-09-21 & Initial preliminary datasheet, N8 configuration \\
1.1 & 2026-09-22 & N8 re-verified with a real MAC-pipeline fix
(WNS improved 0.000\,ns $\to$ +0.108\,ns); N16 now also real,
functionally verified, and timing-closed (WNS=+0.269\,ns) on a
separate development branch, documented as a real candidate for a
future board revision \\
\bottomrule \bottomrule
\end{tabular} \end{tabular}
\end{table} \end{table}
+118 -37
View File
@@ -75,12 +75,14 @@ di sintesi fuori contesto.
0094 & + arbitro gerarchico a 2 livelli & 16 & $-0.646$ (ancora fallito) & 19936 & 128 \\ 0094 & + arbitro gerarchico a 2 livelli & 16 & $-0.646$ (ancora fallito) & 19936 & 128 \\
0094 & + tuning delle direttive P\&R & 16 & $-0.338$ (ancora fallito) & 19936 & 128 \\ 0094 & + tuning delle direttive P\&R & 16 & $-0.338$ (ancora fallito) & 19936 & 128 \\
0095 & curva N=4 (1 gruppo) & 4 & $-0.005$ (2 endpoint falliti) & 8794 & 32 \\ 0095 & curva N=4 (1 gruppo) & 4 & $-0.005$ (2 endpoint falliti) & 8794 & 32 \\
\textbf{0095/0096} & \textbf{\texttt{n8\_system\_ddr3\_top.v}, definitivo} & \textbf{8} & \textbf{0.000 (CHIUSO)} & \textbf{12535} & \textbf{64} \\ 0095/0096 & \texttt{n8\_system\_ddr3\_top.v}, in fabbricazione fisica & 8 & 0.000 (CHIUSO) & 12535 & 64 \\
0097 & + pipeline MAC aggiuntiva (branch \texttt{n16-timing-closure}) & 8 & +0.108 (CHIUSO, migliorato) & 12536 & 64 \\
\textbf{0097} & \textbf{+ pipeline MAC aggiuntiva} & \textbf{16} & \textbf{+0.269 (CHIUSO)} & \textbf{19903} & \textbf{128} \\
\bottomrule \bottomrule
\caption{Cronologia reale dei segni di place-and-route, dal primo P\&R in contesto fino al target definitivo N=8.} \caption{Cronologia reale dei segni di place-and-route, dal primo P\&R in contesto fino ai risultati pi\`u recenti su entrambe le configurazioni N=8 e N=16.}
\end{longtable} \end{longtable}
\subsection{Il segno definitivo: N=8} \subsection{Il segno in fabbricazione fisica: N=8}
\begin{table}[h] \begin{table}[h]
\centering \centering
@@ -90,40 +92,72 @@ di sintesi fuori contesto.
\midrule \midrule
Clock PHY DDR3 (sys\_clk\_p/n) & 310.078\,MHz (periodo 3.225\,ns) \\ Clock PHY DDR3 (sys\_clk\_p/n) & 310.078\,MHz (periodo 3.225\,ns) \\
Clock di calcolo (ui\_clk/clk\_pll\_i) & 155.039\,MHz \\ Clock di calcolo (ui\_clk/clk\_pll\_i) & 155.039\,MHz \\
WNS (setup slack) & \textbf{0.000\,ns} --- chiuso, 0 endpoint falliti \\ WNS (setup slack) & \textbf{+0.108\,ns} --- chiuso, con la correzione pipeline MAC (EXP-0097) \\
WHS (hold slack) & +0.017\,ns \\ WHS (hold slack) & +0.036\,ns \\
Endpoint falliti & 0 su 63212 (setup), 0 su 63209 (hold) \\ Endpoint falliti & 0 (setup e hold) \\
LUT utilizzate & 12535 / 63400 (19.77\%) \\ LUT utilizzate & 12536 / 63400 (19.77\%) \\
Registri utilizzati & 19902 / 126800 (15.70\%) \\ Registri utilizzati & 19910 / 126800 (15.70\%) \\
DSP48E1 utilizzati & 64 / 240 (26.7\%) \\ DSP48E1 utilizzati & 64 / 240 (26.7\%) \\
Parallelismo reale & 8 elementi di elaborazione paralleli \\ Parallelismo reale & 8 elementi di elaborazione paralleli \\
\bottomrule \bottomrule
\end{tabular} \end{tabular}
\caption{Segno di temporizzazione reale, definitivo, per \texttt{n8\_system\_ddr3\_top.v} (EXP-0095/0096).} \caption{Segno di temporizzazione reale per \texttt{n8\_system\_ddr3\_top.v} con la correzione pipeline MAC (EXP-0097) --- la configurazione fisicamente in fabbricazione. Il segno originale (pre-correzione, EXP-0095/0096) chiudeva a WNS=0.000\,ns esatto, senza margine di riserva; questa correzione lo migliora senza alcuna regressione funzionale.}
\end{table} \end{table}
\textbf{Nota importante, dichiarata onestamente}: il margine WNS=0.000\,ns \subsection{Il nuovo segno reale: N=16 ORA CHIUDE}
\`e \emph{esattamente} zero --- reale e chiuso, ma senza alcuno slack di \label{sec:n16-closed}
riserva. Qualunque futura modifica RTL a questo top-level o ai suoi
moduli dipendenti richiede un nuovo, vero place-and-route (con la
stessa sequenza di direttive: \texttt{opt\_design -directive Explore},
\texttt{place\_design -directive ExtraNetDelay\_high},
\texttt{phys\_opt\_design -directive AggressiveExplore},
\texttt{route\_design -directive AggressiveExplore}) prima di
fidarsi nuovamente della temporizzazione.
\subsection{Il collo di bottiglia reale trovato a N=16} Il percorso critico che inizialmente impediva la chiusura del timing a
N=16 \`e stato tracciato realmente (non ipotizzato) fino all'interno
Il percorso critico che impedisce la chiusura del timing a N=16 del datapath MAC di \texttt{neural\_processor\_packed.v} --- lo stesso
\`e stato tracciato realmente (non ipotizzato) fino all'interno del
datapath MAC di \texttt{neural\_processor\_packed.v} --- lo stesso
nucleo di calcolo descritto nel Capitolo~\ref{chap:architecture}, nucleo di calcolo descritto nel Capitolo~\ref{chap:architecture},
invariato dalla prima versione del progetto, che a N=2 chiudeva gi\`a invariato dalla prima versione del progetto, che a N=2 chiudeva gi\`a
con un margine estremamente sottile (+0.0999\,ns). A N=16 la maggiore con un margine estremamente sottile (+0.0999\,ns). A N=16 la maggiore
occupazione complessiva del die (31\% LUT) aumenta la congestione di occupazione complessiva del die (31\% LUT) aumentava la congestione di
piazzamento a sufficienza da erodere quel margine gi\`a minimo --- piazzamento a sufficienza da erodere quel margine gi\`a minimo.
un problema diffuso di congestione, non un singolo collo di bottiglia
strutturale come quello, gi\`a risolto, dell'arbitro. \textbf{Questo collo di bottiglia \`e stato risolto} (\S\ref{sec:mac-pipeline}
del Capitolo~\ref{chap:architecture}) aggiungendo un vero stadio di
pipeline nel datapath MAC. Risultato reale, su un branch di sviluppo
dedicato (\texttt{n16-timing-closure}), isolato dalla scheda fisica in
produzione:
\begin{table}[h]
\centering
\begin{tabular}{ll}
\toprule
\textbf{Metrica} & \textbf{Valore reale} \\
\midrule
WNS (setup slack) & \textbf{+0.269\,ns} --- \textbf{CHIUSO} \\
WHS (hold slack) & +0.026\,ns \\
Endpoint falliti & 0 su 106087 (setup), 0 su 106084 (hold) \\
LUT utilizzate & 19903 / 63400 (31.39\%) \\
Registri utilizzati & 35409 / 126800 (27.93\%) \\
DSP48E1 utilizzati & 128 / 240 (53.33\%) \\
Parallelismo reale & 16 elementi di elaborazione paralleli \\
Verifica funzionale & 32/32 PASS su DDR3 reale (\texttt{tb\_n16\_system\_ddr3.v}) \\
\bottomrule
\end{tabular}
\caption{Segno di temporizzazione reale per \texttt{n16\_system\_ddr3\_top.v} con la correzione pipeline MAC (EXP-0097). Margine reale migliore di quello storico originale di N=2 (+0.0999\,ns).}
\end{table}
Dopo questa correzione, il percorso critico si \`e spostato di nuovo
(prova reale che la correzione ha effettivamente risolto il proprio
problema specifico): ora si trova nella logica di aggiornamento della
coda del Director (\texttt{neural\_director\_grouped.v}), ancora
dominato dal routing fisico (73\%) pi\`u che dalla profondit\`a
logica. Un secondo tentativo reale con direttive di place-and-route
alternative ha dato un risultato \emph{peggiore} (+0.168\,ns), confermando
che il margine attuale (+0.269\,ns) \`e gi\`a il migliore risultato
reale ottenuto con il solo tuning delle direttive, senza ulteriori
modifiche RTL.
\textbf{Stato reale onesto}: N=16 non \`e pi\`u limitato da un vincolo
tecnico di temporizzazione. La sua adozione come scheda fisica \`e ora
una decisione dell'utente, non un problema RTL ancora aperto. La
scheda fisica attualmente in fabbricazione resta N=8 (fabbricazione
gi\`a avviata al momento di questa correzione); N=16 \`e un candidato
reale, verificato, per una futura revisione.
\section{Confronto con un riferimento reale: ESP32-S3} \section{Confronto con un riferimento reale: ESP32-S3}
\label{sec:esp32-comparison} \label{sec:esp32-comparison}
@@ -153,22 +187,69 @@ come attuale un dato ormai superato.
\textbf{Stato onesto attuale}: non esiste ancora una misura reale, \textbf{Stato onesto attuale}: non esiste ancora una misura reale,
diretta, di throughput aggregato (inferenze al secondo, o MAC/s diretta, di throughput aggregato (inferenze al secondo, o MAC/s
sostenuti) della configurazione N=8 definitiva confrontata con un sostenuti) di nessuna configurazione (N=8 o N=16) confrontata con un
benchmark ESP32-S3 aggiornato. Il picco teorico calcolabile per N=8 benchmark ESP32-S3 aggiornato. Il picco teorico calcolabile,
\`e: puramente computazionale, \`e:
\[ \[
8~\text{PE} \times 16~\text{MAC/ciclo} \times 155.039\times10^6~\text{cicli/s} = 19.84~\text{GMAC/s di picco teorico aggregato} 8~\text{PE} \times 16~\text{MAC/ciclo} \times 155.039\times10^6~\text{cicli/s} = 19.84~\text{GMAC/s di picco teorico aggregato (N=8)}
\]
\[
16~\text{PE} \times 16~\text{MAC/ciclo} \times 155.039\times10^6~\text{cicli/s} = 39.69~\text{GMAC/s di picco teorico aggregato (N=16)}
\] \]
ma questo \`e un limite superiore puramente computazionale: dato il ma questo \`e un limite superiore puramente computazionale: dato il
vincolo reale di banda DDR3 (2.48\,GB/s fisici, condivisi tra tutti vincolo reale di banda DDR3 (2.48\,GB/s fisici, condivisi tra tutti i
gli 8 PE), il throughput realmente sostenibile in un carico di lavoro PE attivi), il throughput realmente sostenibile in un carico di lavoro
reale sar\`a inferiore, nella stessa misura gi\`a documentata a N=1/2 reale sar\`a inferiore, nella stessa misura gi\`a documentata a N=1/2
(\S~\ref{sec:bottleneck}) --- una misura reale e diretta di questo (\S~\ref{sec:bottleneck}) --- e la condivisione dello stesso canale
throughput sostenuto a N=8, e un nuovo confronto onesto con un DDR3 tra il doppio dei PE a N=16 rende questo vincolo ANCORA PI\`U
benchmark ESP32-S3 aggiornato, restano lavoro futuro non ancora stringente rispetto a N=8, non meno.
eseguito.
\subsection{Una nuova stima onesta per N=16 (proiezione, non misura)}
Su richiesta esplicita, ecco una stima reale --- costruita da dati
reali gi\`a misurati, ma essa stessa \textbf{una proiezione, non una
nuova misura diretta} --- pi\`u solida della vecchia proiezione
$\sim$55--85$\times$ ormai superata (sopra), perch\'e confronta lo
\textbf{stesso numero di core} (16 vs 16) invece di ipotizzare 30 core
limitati solo dai DSP:
\begin{table}[h]
\centering
\begin{tabular}{lll}
\toprule
\textbf{Fattore di scala reale} & \textbf{Base} & \textbf{Valore} \\
\midrule
Numero di core & 16 (ECP5) vs 16 (Artix-7 N=16) & $1\times$ (confronto diretto) \\
Rapporto di clock & 155.039\,MHz vs 64--97\,MHz (reale, ECP5) & $1.60\times$--$2.42\times$ \\
Packing INT8 per DSP & 2 MAC/DSP (Artix-7) vs 1 MAC/DSP (ECP5, non impacchettato) & $2\times$ \\
\midrule
\textbf{Fattore di scala combinato} & & \textbf{$3.20\times$--$4.85\times$} \\
\bottomrule
\end{tabular}
\caption{Fattori di scala reali (misurati) tra l'architettura ECP5 (v2) e Artix-7 N=16 (v3), a parit\`a di numero di core.}
\end{table}
Applicando questo fattore di scala al dato reale, misurato,
$\sim$9.5--14$\times$ su ESP32-S3 (stesso workload MobileFaceNet-class):
\[
9.5\times3.20 \approx \mathbf{30\times} \qquad\qquad 14\times4.85 \approx \mathbf{68\times}
\]
\textbf{Stima onesta}: N=16, su un carico di lavoro simile a quello
del confronto originale, \`e plausibilmente nell'ordine di
\textbf{30--68$\times$ ESP32-S3} (punto centrale $\sim$45$\times$) ---
una proiezione di puro throughput di calcolo, costruita da fattori di
scala reali, \textbf{non una misura diretta}. Il dato esatto reale
richiede un vero benchmark end-to-end (stesso workload, hardware
reale o simulazione a livello di sistema completo) --- non ancora
eseguito. Come per N=8, il vincolo reale di banda DDR3
(\S~\ref{sec:bottleneck}) former\`a probabilmente un tetto reale
inferiore a questa proiezione puramente computazionale su un carico
di lavoro reale, specialmente a N=16 dove pi\`u PE condividono lo
stesso canale fisico.
\section{Lezioni reali dal processo di verifica} \section{Lezioni reali dal processo di verifica}
+227
View File
@@ -6784,3 +6784,230 @@ docs/PINOUT.md, BOM confirmation from docs/BOM.md -- both already real
and unaffected by the N_GROUPS choice, since DDR3/SPI/flash/config and unaffected by the N_GROUPS choice, since DDR3/SPI/flash/config
pins are package-level, not internal-core-count-dependent) is the real pins are package-level, not internal-core-count-dependent) is the real
next milestone now that a real, closed, deployable RTL target exists. next milestone now that a real, closed, deployable RTL target exists.
EXP-0097 -- N=16 REAL TIMING CLOSED: extra MAC pipeline stage (branch
`n16-timing-closure`), the previously-deferred real fix now built and
verified (2026-09-21/22, user's own explicit direction: "creare una
branch del progetto e lavora per scoprire come fare funzionare il
timing", physical board fabrication continues in parallel on the
already-fixed N=8 design, unaffected by this branch)
CONTEXT: EXP-0094's own real, traced remaining N=16 bottleneck (after
the hierarchical arbiter + P&R directive tuning already closed most of
the gap, WNS -0.913ns -> -0.338ns) was inside `neural_processor_
packed.v`'s own DSP48E1 MAC datapath -- a pre-existing, N=2-era design
(unchanged since EXP-0059) with an already razor-thin real margin
(+0.099962ns) that N=16's own higher real die congestion eroded past
zero. EXP-0094's own `next_action` flagged real MAC-datapath pipelining
as the most direct remaining fix, deliberately not attempted then
(shared, load-bearing module, needed explicit direction + isolation
from the definitive N=8 signoff -- hence the real, separate branch).
METHOD: real, traced worst-violated-path analysis (EXP-0094's own real
post-route report) pinpointed the exact real gap: a DSP48E1's own
(Vivado-auto-retimed) product register feeding STRAIGHT THROUGH the
real carry-heavy INT8-unpack logic (`pb_comb`'s own shift + conditional
+1 carry-propagate add, CARRY4-dominated) into `proda1`/`prodb1` in a
SINGLE real cycle. Real fix: split the original single "Stage 1" into
two real stages -- **Stage 1a** registers the RAW DSP48E1 product with
zero logic in between (`product_reg`, a real, explicit register
boundary immediately after the multiply); **Stage 1b** does the
carry-heavy unpack FROM the already-registered `product_reg` and
registers the result into `proda1`/`prodb1` (unchanged real math,
now one real cycle later). Real, deliberate consequence: end-to-end
per-tile latency grows by exactly ONE real clock cycle; throughput is
unaffected (still accepts one new operand per cycle, real valid/ready
handshaking throughout, no fixed-latency assumption anywhere
downstream). `pipeline_busy`/`valid_tree`'s own level-0 input and the
module's own header comment updated to match.
REAL BUG FOUND AND FIXED IN THE TESTBENCH BEFORE A TRUSTWORTHY RESULT
WAS POSSIBLE (not an RTL bug): `tb_neural_processor_packed.v`'s own
comparison logic required all three cores (2 real reference `neural_
processor.v` instances + the DUT) to assert `result_valid`
SIMULTANEOUSLY -- correct only when all three share the exact same
real pipeline depth. Since `result_valid` is a genuine ONE-SHOT pulse
in every one of these FSMs (self-clears the cycle after `result_ready`
is seen, identical pattern in both v2 and v3 cores), and the DUT is
now deliberately one real cycle deeper than the reference cores, the
reference cores' own `result_valid` had already dropped by the time
the DUT's own pulse arrived -- the original three-way AND never
triggered again, a real 18/18 watchdog-timeout false-failure, not an
actual DUT bug (confirmed via a real, controlled A/B: the SAME failure
does NOT reproduce against the unmodified reference-only comparison
path). Fixed by latching each core's own result independently the
cycle its own `result_valid` first pulses, then comparing the three
LATCHED values once all three have arrived -- correct regardless of
real relative pipeline depth.
Also hit and root-caused (real, not guessed): `tb_np_packed_layer_
reuse.v` fails (3/16 PASS) identically against BOTH the modified AND
the original, unmodified `neural_processor_packed.v` (confirmed via a
real, direct A/B comparison) -- a real, PRE-EXISTING, already-broken/
stale testbench (real port-width mismatch warning on `layer_prefetch_
ctrl.v`'s own `ctrl_wdata`/`ctrl_rdata`, 128 bits wired against a
256-bit real port -- dates from before EXP-0084's own 32-bit DDR3
widening, apparently never updated), unrelated to this real fix, out
of scope for this branch's own task.
Also hit and root-caused (real Vivado project-state quirk, not an RTL
bug): a first real P&R attempt on this branch elaborated with
`N_GROUPS` bound to 2, not the RTL's own real default of 4, despite no
`-generic` override on the actual `synth_design` command line, an
empty real `GENERIC` property on the `synth_1` run, the correct real
`top` property, and no stale imported copy of `n16_system_ddr3_top.v`
anywhere in the project (all confirmed via direct real queries, not
assumed) -- most likely Vivado's own "Incremental synthesis strategy
default" silently carrying forward a parameter binding from this
session's own earlier `-generic N_GROUPS=2` sweep run (EXP-0095),
despite an intervening `reset_run`. Real fix: pass `-generic
N_GROUPS=4` explicitly on the `synth_design` command line rather than
relying on the RTL's own default resolving correctly -- confirmed via
a real, explicit post-synth DSP48E1 cell-count check (128, matching
real N=16) before trusting anything downstream this time.
REAL RESULT: (1) isolated bit-exact verification,
`tb_neural_processor_packed.v` (real Icarus xsim, against 2x real
`hardware/v2/rtl/neural_processor.v`): **18/18 PASS, 0 errors**. (2)
real, full-system functional xsim, `tb_n16_system_ddr3.v` (real DDR3
model, real Vivado xsim): **32/32 PASS, 0 errors**, `$finish` at
197735.6335ns (same real completion time as the pre-fix EXP-0094
result -- the extra real pipeline cycle is fully absorbed by DDR3's
own already-dominant real latency, no observable end-to-end slowdown
at this scale). (3) real, full P&R (`n16_system_ddr3_top.v`, real
XC7A100T-CSG324-2, `Explore`/`ExtraNetDelay_high`/`AggressiveExplore`
directive stack, EXP-0094's own real N_GROUPS=4 explicitly confirmed
via a real post-synth DSP48E1 count of 128): **WNS=+0.269ns,
WHS=+0.026ns, TNS=0.000ns, 0 FAILING SETUP OR HOLD ENDPOINTS --
TIMING CONSTRAINTS ARE MET.** Real utilization: 19903 LUTs (31.39%),
35409 registers (27.93%, up from 19936/27.92%... i.e. genuinely more
registers than the pre-fix EXP-0094 result, matching the real,
expected cost of the added pipeline stage across 16 real PE
instances), 128 DSP48E1 (53.33%).
DECISION: N=16 hybrid systolic (`n16_system_ddr3_top.v`) is now REAL,
functionally verified, AND timing-CLOSED, on this real, isolated
branch (`n16-timing-closure`) -- does not touch or affect the physical
board fabrication already underway on N=8 (`v3-artix7`, unmodified).
This is a real, significant milestone: it confirms the N=16 hybrid
systolic architecture is fundamentally viable at full real scale, not
just "close" -- the earlier N=8-as-definitive decision was a real,
reasonable engineering choice under the "ship something real now"
constraint (WNS=0.000ns exact-zero margin vs. this fix's own real,
more comfortable +0.269ns), not a permanent architectural ceiling.
next_action: (1) real, dedicated functional xsim + P&R re-confirmation
specifically for the definitive N=8 configuration WITH this same MAC
pipeline fix applied (verify it does not regress N=8's own real,
already-closed signoff, and ideally IMPROVES its own already-thin
future margin) -- not yet done on this branch. (2) a real, explicit
decision with the user on whether/when to promote this fix back to
`v3-artix7` (the physical board's own branch) -- given the board is
already in fabrication as the UNMODIFIED N=8 design, this is a real
question about a FUTURE board revision, not the current one. (3) this
branch's own real bug findings (the testbench latching fix, the
Vivado incremental-synthesis generic-binding quirk) are worth folding
into CLAUDE.md's own hard-won-lessons section regardless of the
promotion decision.
EXP-0097 (continued) -- real N=8 re-verification with the same MAC
pipeline fix (improves, does not regress) + a real margin-hunt attempt
for N=16 (2026-09-21/22, user's own explicit direction: "verifica
anche su N=8 e verifica se possiamo guadagnare qualcosina ancora su
N16 perche' io implemento N16 se funziona")
REAL RESULT (1), N=8 with the same pipelined `neural_processor_
packed.v`: real functional xsim (`tb_n8_system_ddr3.v`, real DDR3
model) **16/16 PASS, 0 errors**, identical real completion time to the
pre-fix result (131145.8335ns) -- no functional regression. Real, full
P&R (`n8_system_ddr3_top.v`, `-generic N_GROUPS=2` explicit, confirmed
via a real post-synth 64 DSP48E1 count): **WNS=+0.108ns (up from the
pre-fix exact-zero 0.000ns), WHS=+0.036ns, 0 failing endpoints, 64
DSP48E1 (26.7%), 12536 LUTs (19.77%)**. The same fix that closes N=16
also gives N=8 real, comfortable margin instead of the exact-zero
margin its original (unmodified-core) signoff had -- a real, additive
improvement, not a tradeoff.
REAL RESULT (2), N=16 margin-hunt: the real worst path has moved AGAIN
(confirming the MAC-datapath fix genuinely resolved ITS OWN real
bottleneck) -- now inside `neural_director_grouped.v`'s own queue
update logic (`u_dir/q_head_reg[3]` -> `q_count_reg[0]/CE`), still
real route-dominated (73%), not logic-depth-dominated. A second real
P&R attempt with alternate directives (`opt_design -directive
ExploreWithRemap`, `place_design -directive Explore`, same
`phys_opt_design`/`route_design -directive AggressiveExplore`) gave
**WNS=+0.168ns -- WORSE than the first attempt's own real +0.269ns**,
confirming real run-to-run/directive-to-directive P&R variance, not a
systematic further improvement available from directive-tuning alone.
DECISION: the original directive stack (`Explore`/`ExtraNetDelay_
high`/`AggressiveExplore`) remains the best real N=16 result found
(WNS=+0.269ns). Further real margin would require touching `neural_
director_grouped.v`'s own queue RTL (a new, separate, not-yet-scoped
piece of real engineering) -- NOT attempted, given the current real
margin is already comfortably closed (better than N=2's own original
historical +0.0999962ns real signoff margin) and blind further P&R-
directive search already showed diminishing/negative real returns.
Real, cumulative state on this branch (`n16-timing-closure`), both
configurations using the SAME real pipelined `neural_processor_
packed.v`:
N=8 (`n8_system_ddr3_top.v`): WNS=+0.108ns, 64 DSP48E1, 16/16 functional PASS
N=16 (`n16_system_ddr3_top.v`): WNS=+0.269ns, 128 DSP48E1, 32/32 functional PASS
next_action: real, explicit decision with the user on whether to (1)
promote this fix to `v3-artix7` for a FUTURE board revision (the
current physical board is already in fabrication as the unmodified
N=8 design, unaffected), and/or (2) actually build the NEXT physical
board as N=16 instead of N=8, given N=16 is now real, functionally
verified, AND timing-closed with a real, comfortable margin -- a real,
consequential hardware decision, not an RTL one.
EXP-0098 -- consolidating N=16 to the same real rigor as N=8: N=2
re-verified too (the pipelined MAC core is SHARED across the whole
real family), full real family state now closed at every N
(2026-09-22, user's own explicit direction: "consolidare N=16 come
faccio con N=8")
CONTEXT: EXP-0097 verified the pipelined `neural_processor_packed.v`
against N=8 and N=16, but never against N=2 -- a real, disclosed gap,
since N=2 (`n2_system_ddr3_top.v`) is the SAME shared core and remains
this project's own documented real fallback signoff (EXP-0088).
Consolidating N=16 to N=8's own level of rigor means confirming the
WHOLE real family, not just the two configurations directly asked
about.
REAL RESULT: functional xsim (`tb_n2_system_ddr3.v`, real DDR3 model):
**8/8 PASS, 0 errors**, `$finish` at the EXACT SAME real completion
time as the pre-fix baseline (101204.9335ns) -- zero observable
end-to-end effect at this scale, same as N=16's own real finding.
Real, full P&R (`n2_system_ddr3_top.v`, real XC7A100T-CSG324-2, same
directive stack): **WNS=+0.389ns (up from the original real
+0.099962ns), WHS=+0.017ns, 0 failing endpoints, 16 DSP48E1 (6.67%),
6645 LUTs (10.48%)** -- a real, substantial margin improvement, no
regression.
REAL, CONSOLIDATED FAMILY STATE (branch `n16-timing-closure`, all
three real top-levels sharing the SAME pipelined `neural_processor_
packed.v`, all real, functionally verified AND timing-closed):
N=2 (`n2_system_ddr3_top.v`): WNS=+0.389ns, 16 DSP48E1, 8/8 functional PASS
N=8 (`n8_system_ddr3_top.v`): WNS=+0.108ns, 64 DSP48E1, 16/16 functional PASS
N=16 (`n16_system_ddr3_top.v`): WNS=+0.269ns, 128 DSP48E1, 32/32 functional PASS
DECISION: the real MAC-pipeline fix (EXP-0097) is a pure, unconditional
improvement across the entire real product family -- every real
configuration this project has ever built a dedicated top-level for
now closes with real, comfortable, positive margin, not just N=16. No
real regression found anywhere. This is now a real, trustworthy,
fully-consolidated state for this branch, at the same level of rigor
EXP-0096 established for N=8 alone.
next_action: update the project's own primary real docs
(`docs/PHYSICAL_REALIZATION.md` §3, `docs/ARCHITECTURE_ANALYSIS.md`
§5.6 and its own top-of-document pointer) on this branch to reflect
this consolidated real family state -- not yet done, the LaTeX
deliverables (`docs/latex/*.tex`) were updated first per the user's
own more immediate request, but the markdown docs are this project's
own real, authoritative source of truth per CLAUDE.md's own "Read
first" section and deserve the same update.
+85 -24
View File
@@ -12,12 +12,20 @@
// to keep this module's own pipeline depth/stage count identical to // to keep this module's own pipeline depth/stage count identical to
// the V2 original for a direct structural comparison). // the V2 original for a direct structural comparison).
// //
// Pipeline stages match V2's neural_processor.v exactly, just doubled // Pipeline stages, originally matched V2's neural_processor.v exactly
// on the accumulator side (one accumulate/bias/activation/saturation // (one accumulate/bias/activation/saturation path per job, A and B,
// path per job, A and B, sharing the SAME multiply/adder-tree stages // sharing the SAME multiply/adder-tree stages since they consume the
// since they consume the SAME weight stream): // SAME weight stream). EXTENDED BY ONE REAL STAGE on the
// Stage 0 input alignment (x0_a, x0_b, w0 -- ONE shared weight) // n16-timing-closure branch (real fix for EXP-0094's own real, traced
// Stage 1 P_IN packed-MAC lanes: p0[i]=x0_a[i]*w0[i], p1[i]=x0_b[i]*w0[i] // N=16 P&R timing failure -- see Stage 1a/1b's own header comments for
// the full real root-cause story):
// Stage 0 input alignment (x0_a, x0_b, w0 -- ONE shared weight)
// Stage 1a P_IN real DSP48E1 packed multiplies, registered raw
// (product_reg) -- NEW real stage
// Stage 1b unpack the two packed INT8 products from product_reg:
// p0[i]=x0_a[i]*w0[i], p1[i]=x0_b[i]*w0[i] -- same real
// math as the original single "Stage 1", now one real
// cycle later
// Stage 2..(1+TREE_LEVELS) TWO balanced adder trees (A and B) // Stage 2..(1+TREE_LEVELS) TWO balanced adder trees (A and B)
// Stage (2+TREE_LEVELS) TWO accumulators // Stage (2+TREE_LEVELS) TWO accumulators
// Stage (3+TREE_LEVELS) bias add (shared bias/activation -- same // Stage (3+TREE_LEVELS) bias add (shared bias/activation -- same
@@ -25,6 +33,14 @@
// + activation, per job // + activation, per job
// Stage (4+TREE_LEVELS) INT8 saturation / output register, per job // Stage (4+TREE_LEVELS) INT8 saturation / output register, per job
// //
// Real, deliberate consequence: end-to-end per-tile latency grows by
// exactly ONE real clock cycle versus the original design (throughput
// is unaffected -- the pipeline still accepts one new operand per
// cycle in steady state). Functional behavior (the actual packed-MAC
// arithmetic) is byte-for-byte unchanged -- verified bit-exact against
// the same real reference used since EXP-0059 (2x real
// hardware/v2/rtl/neural_processor.v), `tb_neural_processor_packed.v`.
//
// job_bias/job_activation are SHARED between A and B (same resident // job_bias/job_activation are SHARED between A and B (same resident
// neuron), matching this project's own weight-reuse semantics (a // neuron), matching this project's own weight-reuse semantics (a
// neuron/filter's bias and activation type don't vary by spatial // neuron/filter's bias and activation type don't vary by spatial
@@ -116,17 +132,28 @@ module neural_processor_packed #(
end end
// ============================================================ // ============================================================
// STAGE 1 -- P_IN packed-MAC lanes (mac2_dsp_packed.v's own // STAGE 1a -- P_IN real DSP48E1 packed multiplies, registered RAW
// verified combinational formula, inlined per lane) // (n16-timing-closure branch, real fix for EXP-0094's own real,
// traced N=16 critical path). EXP-0093/0094's own real post-route
// reports found the worst violated path running from a DSP48E1's
// own (Vivado-auto-retimed) product register straight through the
// pb_comb unpack logic below (a real, CARRY4-heavy shift + carry-
// propagate add) into proda1/prodb1 in a SINGLE cycle -- already
// razor-thin at N=2 (WNS=+0.0999962ns, EXP-0088) and pushed
// negative by N=16's own extra real placement congestion (EXP-
// 0093/0094). This stage makes the DSP's own real output register
// explicit in RTL (captures the WHOLE raw packed product, zero
// logic in between) instead of relying on the tool to retime one
// in automatically -- the actual, additional real pipeline stage
// this fix needs is STAGE 1b below, which now has its own full
// real clock period to do the unpack work in.
// ============================================================ // ============================================================
reg valid1, last1; reg valid1, last1;
reg signed [ACC_WIDTH-1:0] proda1 [0:P_IN-1];
reg signed [ACC_WIDTH-1:0] prodb1 [0:P_IN-1];
localparam A_WIDTH = 3*DATA_WIDTH + 1; localparam A_WIDTH = 3*DATA_WIDTH + 1;
localparam PRODUCT_WIDTH = A_WIDTH + DATA_WIDTH;
wire signed [PROD_WIDTH-1:0] pa_comb [0:P_IN-1]; wire signed [PRODUCT_WIDTH-1:0] product_comb [0:P_IN-1];
wire signed [PROD_WIDTH-1:0] pb_comb [0:P_IN-1]; reg signed [PRODUCT_WIDTH-1:0] product_reg [0:P_IN-1];
genvar gm; genvar gm;
generate generate
@@ -134,12 +161,7 @@ module neural_processor_packed #(
wire signed [A_WIDTH-1:0] x0_sext25 = {{(A_WIDTH-DATA_WIDTH){xa0[gm][DATA_WIDTH-1]}}, xa0[gm]}; wire signed [A_WIDTH-1:0] x0_sext25 = {{(A_WIDTH-DATA_WIDTH){xa0[gm][DATA_WIDTH-1]}}, xa0[gm]};
wire signed [A_WIDTH-1:0] x1_shifted = $signed(xb0[gm]) <<< (2*DATA_WIDTH); wire signed [A_WIDTH-1:0] x1_shifted = $signed(xb0[gm]) <<< (2*DATA_WIDTH);
wire signed [A_WIDTH-1:0] packed_a = x1_shifted + x0_sext25; wire signed [A_WIDTH-1:0] packed_a = x1_shifted + x0_sext25;
wire signed [A_WIDTH+DATA_WIDTH-1:0] product = packed_a * w0[gm]; assign product_comb[gm] = packed_a * w0[gm];
assign pa_comb[gm] = product[PROD_WIDTH-1:0];
wire signed [A_WIDTH+DATA_WIDTH-2*DATA_WIDTH-1:0] pb_raw =
$signed(product) >>> (2*DATA_WIDTH);
assign pb_comb[gm] = pb_raw[PROD_WIDTH-1:0] + (pa_comb[gm][PROD_WIDTH-1] ? 1'b1 : 1'b0);
end end
endgenerate endgenerate
@@ -150,6 +172,45 @@ module neural_processor_packed #(
end else begin end else begin
valid1 <= valid0; valid1 <= valid0;
last1 <= last0; last1 <= last0;
for (gi = 0; gi < P_IN; gi = gi + 1)
product_reg[gi] <= product_comb[gi];
end
end
// ============================================================
// STAGE 1b -- unpack the two packed INT8 products from the
// ALREADY-REGISTERED product_reg (real, added pipeline stage --
// the actual timing fix). pa_comb/pb_comb's own math is byte-for-
// byte IDENTICAL to the original single-stage version, only the
// source (product_reg, a real register) and the register that
// captures the result (proda1/prodb1, now one real cycle later)
// changed -- functional behavior is unchanged, only latency grows
// by exactly one real clock cycle.
// ============================================================
reg valid1b, last1b;
reg signed [ACC_WIDTH-1:0] proda1 [0:P_IN-1];
reg signed [ACC_WIDTH-1:0] prodb1 [0:P_IN-1];
wire signed [PROD_WIDTH-1:0] pa_comb [0:P_IN-1];
wire signed [PROD_WIDTH-1:0] pb_comb [0:P_IN-1];
genvar gp;
generate
for (gp = 0; gp < P_IN; gp = gp + 1) begin : GEN_UNPACK
assign pa_comb[gp] = product_reg[gp][PROD_WIDTH-1:0];
wire signed [A_WIDTH+DATA_WIDTH-2*DATA_WIDTH-1:0] pb_raw =
$signed(product_reg[gp]) >>> (2*DATA_WIDTH);
assign pb_comb[gp] = pb_raw[PROD_WIDTH-1:0] + (pa_comb[gp][PROD_WIDTH-1] ? 1'b1 : 1'b0);
end
endgenerate
always @(posedge clk) begin
if (rst) begin
valid1b <= 1'b0;
last1b <= 1'b0;
end else begin
valid1b <= valid1;
last1b <= last1;
for (gi = 0; gi < P_IN; gi = gi + 1) begin for (gi = 0; gi < P_IN; gi = gi + 1) begin
proda1[gi] <= {{(ACC_WIDTH-PROD_WIDTH){pa_comb[gi][PROD_WIDTH-1]}}, pa_comb[gi]}; proda1[gi] <= {{(ACC_WIDTH-PROD_WIDTH){pa_comb[gi][PROD_WIDTH-1]}}, pa_comb[gi]};
prodb1[gi] <= {{(ACC_WIDTH-PROD_WIDTH){pb_comb[gi][PROD_WIDTH-1]}}, pb_comb[gi]}; prodb1[gi] <= {{(ACC_WIDTH-PROD_WIDTH){pb_comb[gi][PROD_WIDTH-1]}}, pb_comb[gi]};
@@ -183,8 +244,8 @@ module neural_processor_packed #(
valid_tree[gl] <= 1'b0; valid_tree[gl] <= 1'b0;
last_tree[gl] <= 1'b0; last_tree[gl] <= 1'b0;
end else begin end else begin
valid_tree[gl] <= (gl == 0) ? valid1 : valid_tree[gl-1]; valid_tree[gl] <= (gl == 0) ? valid1b : valid_tree[gl-1];
last_tree[gl] <= (gl == 0) ? last1 : last_tree[gl-1]; last_tree[gl] <= (gl == 0) ? last1b : last_tree[gl-1];
end end
end end
for (gn = 0; gn < (P_IN >> (gl+1)); gn = gn + 1) begin : GEN_TREE_NODE for (gn = 0; gn < (P_IN >> (gl+1)); gn = gn + 1) begin : GEN_TREE_NODE
@@ -203,8 +264,8 @@ module neural_processor_packed #(
end end
endgenerate endgenerate
wire valid_tree_out = (TREE_LEVELS == 0) ? valid1 : valid_tree[TREE_LEVELS-1]; wire valid_tree_out = (TREE_LEVELS == 0) ? valid1b : valid_tree[TREE_LEVELS-1];
wire last_tree_out = (TREE_LEVELS == 0) ? last1 : last_tree[TREE_LEVELS-1]; wire last_tree_out = (TREE_LEVELS == 0) ? last1b : last_tree[TREE_LEVELS-1];
wire signed [ACC_WIDTH-1:0] tile_sum_a = (TREE_LEVELS == 0) ? proda1[0] : treea[TREE_LEVELS][0]; wire signed [ACC_WIDTH-1:0] tile_sum_a = (TREE_LEVELS == 0) ? proda1[0] : treea[TREE_LEVELS][0];
wire signed [ACC_WIDTH-1:0] tile_sum_b = (TREE_LEVELS == 0) ? prodb1[0] : treeb[TREE_LEVELS][0]; wire signed [ACC_WIDTH-1:0] tile_sum_b = (TREE_LEVELS == 0) ? prodb1[0] : treeb[TREE_LEVELS][0];
@@ -294,7 +355,7 @@ module neural_processor_packed #(
end end
end end
wire pipeline_busy = valid0 || valid1 || (|valid_tree) || valid5 || valid6 || valid7; wire pipeline_busy = valid0 || valid1 || valid1b || (|valid_tree) || valid5 || valid6 || valid7;
assign job_ready = (np_state == NP_IDLE) && !pipeline_busy; assign job_ready = (np_state == NP_IDLE) && !pipeline_busy;
// ============================================================ // ============================================================
+44 -16
View File
@@ -148,29 +148,57 @@ module tb;
v2_tile_last = 0; v2_tile_last = 0;
tile_last = 0; tile_last = 0;
// real fix (n16-timing-closure branch): result_valid is a
// real ONE-SHOT pulse in every one of these FSMs (`NP_
// WRITE_RESULT: if (result_valid && result_ready)
// result_valid<=0`, identical in neural_processor.v and
// neural_processor_packed.v) -- with result_ready already
// held high before this wait begins, each core's own
// result_valid self-clears the very next cycle after it
// first asserts, independent of whether the OTHER cores
// have caught up yet. The original three-way simultaneous
// AND assumed all three cores share the exact same real
// pipeline depth -- true before this branch's own real
// extra pipeline stage in neural_processor_packed.v (added
// to fix EXP-0094's own real N=16 timing failure), no
// longer true now that the DUT is deliberately one real
// cycle deeper than the reference cores. Real fix: latch
// each core's own result independently the cycle its own
// result_valid pulses, then compare the three LATCHED
// values once all three have arrived -- correct regardless
// of real relative pipeline depth between DUT and
// reference.
v2_result_ready = 1; v2_result_ready = 1;
result_ready = 1; result_ready = 1;
watchdog = 0; begin : capture
while (!(v2a_result_valid && v2b_result_valid && result_valid) && watchdog < 300) begin reg v2a_got, v2b_got, dut_got;
@(posedge clk); reg signed [DATA_WIDTH-1:0] v2a_val, v2b_val, dut_val_a, dut_val_b;
watchdog = watchdog + 1; v2a_got = 0; v2b_got = 0; dut_got = 0;
end watchdog = 0;
while (!(v2a_got && v2b_got && dut_got) && watchdog < 300) begin
@(posedge clk);
if (!v2a_got && v2a_result_valid) begin v2a_got = 1; v2a_val = v2a_result_data; end
if (!v2b_got && v2b_result_valid) begin v2b_got = 1; v2b_val = v2b_result_data; end
if (!dut_got && result_valid) begin dut_got = 1; dut_val_a = result_data_a; dut_val_b = result_data_b; end
watchdog = watchdog + 1;
end
if (!v2a_result_valid || !v2b_result_valid || !result_valid) begin if (!v2a_got || !v2b_got || !dut_got) begin
$display("FAIL n=%0d: watchdog timeout waiting for results (v2a=%b v2b=%b dut=%b)", $display("FAIL n=%0d: watchdog timeout waiting for results (v2a_got=%b v2b_got=%b dut_got=%b)",
n, v2a_result_valid, v2b_result_valid, result_valid); n, v2a_got, v2b_got, dut_got);
errors = errors + 1;
end else begin
if (result_data_a !== v2a_result_data || result_data_b !== v2b_result_data) begin
$display("FAIL n=%0d bias=%0d act=%0d: v2a=%0d v2b=%0d dut_a=%0d dut_b=%0d MISMATCH",
n, bias, activation, v2a_result_data, v2b_result_data, result_data_a, result_data_b);
errors = errors + 1; errors = errors + 1;
end else begin end else begin
$display("PASS n=%0d bias=%0d act=%0d: a=%0d b=%0d (bit-exact vs 2x real neural_processor.v)", if (dut_val_a !== v2a_val || dut_val_b !== v2b_val) begin
n, bias, activation, result_data_a, result_data_b); $display("FAIL n=%0d bias=%0d act=%0d: v2a=%0d v2b=%0d dut_a=%0d dut_b=%0d MISMATCH",
n, bias, activation, v2a_val, v2b_val, dut_val_a, dut_val_b);
errors = errors + 1;
end else begin
$display("PASS n=%0d bias=%0d act=%0d: a=%0d b=%0d (bit-exact vs 2x real neural_processor.v)",
n, bias, activation, dut_val_a, dut_val_b);
end
end end
@(posedge clk);
end end
@(posedge clk);
while (!job_ready || np_state !== 4'd0 || !v2a_job_ready || !v2b_job_ready) @(posedge clk); while (!job_ready || np_state !== 4'd0 || !v2a_job_ready || !v2b_job_ready) @(posedge clk);
end end