Physical board fabrication moved to N=16 (with the EXP-0097 MAC- pipeline fix); v3-artix7 fast-forward merged with n16-timing-closure. No physical board change needed -- DDR3/SPI/flash/config pins are identical across all N_GROUPS values. N=2/N=8 kept documented as real, valid fallback signoffs. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MUG92aM9m68TRc4rG55BcC
483 lines
28 KiB
Markdown
483 lines
28 KiB
Markdown
# FPGA-Neural V3 — Physical Realization Reference
|
||
|
||
Real, verified data for the custom PCB (bare XC7A100T-CSG324-2 + DDR3, no dev
|
||
board). Every pin/part/setting below comes from a real Vivado-generated
|
||
constraint file, a real datasheet, or a real place-and-route run — none of it
|
||
is guessed. See `hardware/v2/logs/experiments.log` (EXP-0059 onward) for the
|
||
full derivation history.
|
||
|
||
## 1. Core components (real, verified availability)
|
||
|
||
| Component | Part | Notes |
|
||
|---|---|---|
|
||
| FPGA | **XC7A100T-CSG324-2** | Speed grade **-2** (corrected from an initial -1 assumption, EXP-0074) — same die/package/footprint as -1, strictly better timing margin. |
|
||
| DDR3 SDRAM | **Micron MT41J128M16JT-125:K** | 2Gb, x16, DDR3-1600-rated (run at 310.078MHz here due to -2 timing closure, see §3). Verified in-stock on LCSC. |
|
||
| Config flash | **Winbond W25Q32JVSSIQ** | 32Mbit/4MB, SOIC-8. Comfortably fits the ~30.5Mbit full XC7A100T bitstream. Verified in-stock on LCSC. Wired **exclusively** to the FPGA (see §5). |
|
||
|
||
## 2. FPGA pin assignments (real, from the routed design)
|
||
|
||
### 2.1 DDR3 (fixed by the FPGA's own PHY hardware — not a free choice)
|
||
|
||
**EXP-0084 UPDATE: real 32-bit channel (two MT41J128M16JT-125:K chips
|
||
ganged in parallel)** — dq/dqs/dm pin counts DOUBLED versus the original
|
||
16-bit design (dq 16→32, dqs 2→4 pairs, dm 2→4). Address/command/control
|
||
lines are unchanged and fanned out identically to both chips. Real,
|
||
board-accurate pin placement below (confirmed from the actual routed
|
||
design, `mig_7series_0.xdc` + this project's own top-level XDC) —
|
||
**functionally verified**, but see §3 for the real, honest timing-closure
|
||
status before treating this as a final signoff.
|
||
|
||
Generated by the Vivado MIG wizard (`mig_7series_0.xdc`), all `SSTL15` /
|
||
`DIFF_SSTL15` (1.5V), banks 34/35:
|
||
|
||
| Signal | Pin | Signal | Pin | Signal | Pin |
|
||
|---|---|---|---|---|---|
|
||
| ddr3_dq[0] | C5 | ddr3_dq[16] | F4 | ddr3_addr[0] | U3 |
|
||
| ddr3_dq[1] | B7 | ddr3_dq[17] | F3 | ddr3_addr[1] | U4 |
|
||
| ddr3_dq[2] | B6 | ddr3_dq[18] | E2 | ddr3_addr[2] | V1 |
|
||
| ddr3_dq[3] | D8 | ddr3_dq[19] | D2 | ddr3_addr[3] | U1 |
|
||
| ddr3_dq[4] | C7 | ddr3_dq[20] | C1 | ddr3_addr[4] | L5 |
|
||
| ddr3_dq[5] | E6 | ddr3_dq[21] | H1 | ddr3_addr[5] | L6 |
|
||
| ddr3_dq[6] | E5 | ddr3_dq[22] | G1 | ddr3_addr[6] | L4 |
|
||
| ddr3_dq[7] | E7 | ddr3_dq[23] | F1 | ddr3_addr[7] | K5 |
|
||
| ddr3_dq[8] | B4 | ddr3_dq[24] | F6 | ddr3_addr[8] | M2 |
|
||
| ddr3_dq[9] | A4 | ddr3_dq[25] | G4 | ddr3_addr[9] | M3 |
|
||
| ddr3_dq[10] | A3 | ddr3_dq[26] | G3 | ddr3_addr[10] | L3 |
|
||
| ddr3_dq[11] | B3 | ddr3_dq[27] | J3 | ddr3_addr[11] | K3 |
|
||
| ddr3_dq[12] | B2 | ddr3_dq[28] | J2 | ddr3_addr[12] | M1 |
|
||
| ddr3_dq[13] | D5 | ddr3_dq[29] | K2 | ddr3_addr[13] | L1 |
|
||
| ddr3_dq[14] | D4 | ddr3_dq[30] | K1 | ddr3_ba[0] | V5 |
|
||
| ddr3_dq[15] | E3 | ddr3_dq[31] | H6 | ddr3_ba[1] | V2 |
|
||
| ddr3_dm[0] | C6 | ddr3_dm[2] | C2 | ddr3_ba[2] | U2 |
|
||
| ddr3_dm[1] | C4 | ddr3_dm[3] | G6 | | |
|
||
| ddr3_dqs_p[0] | A6 (DIFF) | ddr3_dqs_n[0] | A5 (DIFF) | ddr3_cas_n | R3 |
|
||
| ddr3_dqs_p[1] | B1 (DIFF) | ddr3_dqs_n[1] | A1 (DIFF) | ddr3_cke[0] | P2 |
|
||
| ddr3_dqs_p[2] | H2 (DIFF) | ddr3_dqs_n[2] | G2 (DIFF) | ddr3_cs_n[0] | T5 |
|
||
| ddr3_dqs_p[3] | J4 (DIFF) | ddr3_dqs_n[3] | H4 (DIFF) | ddr3_odt[0] | R2 |
|
||
| ddr3_ck_p[0] | N2 (DIFF) | ddr3_ck_n[0] | N1 (DIFF) | ddr3_ras_n | V4 |
|
||
| ddr3_reset_n | D7 (LVCMOS15) | ddr3_we_n | T3 | | |
|
||
| **sys_clk_p** | **N5 (DIFF_SSTL15, bank 34)** | **sys_clk_n** | **P5 (DIFF_SSTL15, bank 34)** | | |
|
||
| **clk_ref_p** | **T14 (LVDS_25, bank 14)** | **clk_ref_n** | **T15 (LVDS_25, bank 14)** | | |
|
||
|
||
**Real board implication (EXP-0084)**: both `sys_clk` and `clk_ref` are now
|
||
genuine **differential pairs** — the board needs a differential oscillator
|
||
(or a differential buffer stage), not a single-ended one. `clk_ref` lives in
|
||
**bank 14** (not 34/35) because the MIG wizard's own UG586 placement rules
|
||
restrict that specific net to bank 14 for this part/package — not a free
|
||
choice. This directly displaced the config-flash SPI bus, which used to
|
||
share bank 14 (see §2.3 — moved to bank 16 in EXP-0084 to resolve a real
|
||
VCCO conflict, LVCMOS33 vs LVDS_25 cannot coexist in one bank).
|
||
|
||
**Bank voltage requirements**: bank 34/35 → **1.5V** (DDR3 SSTL15, plus
|
||
`sys_clk_p/n` at bank 34's own 1.5V-domain `DIFF_SSTL15`), bank 14 →
|
||
**2.5V** (`clk_ref_p/n`, LVDS_25).
|
||
|
||
`INTERNAL_VREF` for banks 34/35 is set to 0.750V by the MIG constraints
|
||
(required for SSTL15 single-ended inputs) — this is a Vivado-side setting,
|
||
not a board component, but note it if you ever inspect bitstream generation
|
||
warnings about VREF.
|
||
|
||
### 2.2 Neural-processor management SPI (ESP32 ↔ FPGA, FPGA is **slave**)
|
||
|
||
Bank 15, package edge column A/B, physically adjacent (short traces), `LVCMOS33`:
|
||
|
||
| Signal | Pin | Direction (FPGA side) |
|
||
|---|---|---|
|
||
| sclk | A15 | input |
|
||
| mosi | B16 | input |
|
||
| miso | B17 | output |
|
||
| cs_n | A16 | input |
|
||
| sys_rst | G13 | input (EXP-0084, tentative — see §7) |
|
||
| data_ready_n | D14 | output (EXP-0085, active-low sticky IRQ — see §6, tentative — see §7) |
|
||
|
||
**Bank 15 VCCO**: assumed **3.3V** — change the XDC's IOSTANDARD if your
|
||
board power plan uses a different rail for this bank.
|
||
|
||
### 2.3 Config flash SPI (FPGA ↔ flash, FPGA is **master**)
|
||
|
||
**EXP-0084 REAL PIN CHANGE**: originally these were the FPGA's own dedicated
|
||
Master-SPI configuration pins (K17/K18/L13, bank 14), reclaimed as ordinary
|
||
fabric I/O post-configuration. EXP-0084's real 32-bit DDR3 widening moved
|
||
the differential reference clock (`clk_ref_p/n`, §2.1) into bank 14 too —
|
||
a real VCCO conflict (the flash needs LVCMOS33/3.3V, `clk_ref` needs
|
||
LVDS_25/2.5V, one bank can only have one VCCO), confirmed by a real
|
||
`place_design` failure, not hypothetical. **Fixed by moving the flash bus
|
||
to bank 16** (completely unconstrained before this, no VCCO commitment):
|
||
|
||
| Signal | Pin | Direction (FPGA side) | Notes |
|
||
|---|---|---|---|
|
||
| flash_mosi | D9 | output | bank 16 (was K17/D00_MOSI, bank 14) |
|
||
| flash_miso | D10 | input | bank 16 (was K18/D01_DIN, bank 14) |
|
||
| flash_cs_n | C9 | output | bank 16 (was L13/FCS_B, bank 14) |
|
||
| (CCLK) | E9 | output | **Not a top-level port** — driven internally via the `STARTUPE2` primitive. Wire the flash's own CLK pin to package pin **E9**. Unaffected by the bank-14→16 move (CCLK was never a plain port). |
|
||
|
||
Still requires `BITSTREAM.CONFIG.PERSIST = NO` (the Vivado default — set
|
||
explicitly in the project XDC; **EXP-0084 note**: the XDC previously had
|
||
this as `FALSE`, an invalid enum value in this Vivado version — silently
|
||
not applied at all until fixed to the real `NO`/`YES` enum).
|
||
|
||
**Bank 16 VCCO**: **3.3V** (matches the flash's own VCC, typically
|
||
1.8–3.6V for the W25Q32JV) — chosen freely since bank 16 had no prior VCCO
|
||
commitment.
|
||
|
||
**Real board implication**: since the flash bridge no longer uses the
|
||
FPGA's own dedicated D00_MOSI/D01_DIN/FCS_B configuration pins at all
|
||
(K17/K18/L13 are now unused by this design), the flash chip's SPI lines on
|
||
the PCB connect to D9/D10/C9 instead — a real routing change versus any
|
||
earlier board layout draft based on the pre-EXP-0084 pinout.
|
||
|
||
**Reserved, do not use** (bank 14, unused by this design, kept clear for
|
||
any future Quad-SPI/BPI expansion — no longer near the flash bus, now just
|
||
general config-mode-reserved pins): `L16` (EMCCLK), `R16` (RDWR_B), `V15`
|
||
(CSI_B). The project's own XDC `PROHIBIT`s the underlying **sites** for
|
||
these (fixed in EXP-0084 — `PROHIBIT` is not a valid property directly on
|
||
`package_pin` objects in this Vivado version, only on their `site`; this
|
||
constraint had silently never been enforced before).
|
||
|
||
### 2.4 FPGA configuration control (dedicated, bank 0, not negotiable)
|
||
|
||
| Signal | Pin | Purpose |
|
||
|---|---|---|
|
||
| PROGRAM_B | P9 | pulse low to force a full reconfiguration from flash |
|
||
| INIT_B | P7 | goes low during config; can indicate a config error if it re-asserts |
|
||
| DONE | P10 | goes high once configuration succeeds — wire to a status LED if desired |
|
||
| M0 | P12 | mode select |
|
||
| M1 | P13 | mode select |
|
||
| M2 | P11 | mode select |
|
||
| CFGBVS | P8 | tie to match bank 0's VCCO logic level (see UG470) |
|
||
|
||
**Mode pin setting for Master SPI boot** (the flash-based autonomous boot
|
||
path, see §5): `M[2:0] = 001` (per UG470's mode pin table) — tie via pull-up/
|
||
pull-down resistors on the board, not driven dynamically.
|
||
|
||
### 2.5 JTAG (always available, independent of flash content)
|
||
|
||
| Signal | Pin |
|
||
|---|---|
|
||
| TCK | E10 |
|
||
| TDI | E11 |
|
||
| TMS | E12 |
|
||
| TDO | E13 |
|
||
|
||
Used for: (a) first-ever/factory programming when the flash is blank (see
|
||
§5), (b) recovery, (c) development/debug. This project's own plan drives
|
||
these from an ESP32 doing real JTAG bit-banging (TAP state machine, IR/DR
|
||
shifting) rather than a bench programmer — that firmware is separate,
|
||
software-side work, not covered here.
|
||
|
||
## 3. Real timing signoff — N=16 IS NOW THE REAL FABRICATION TARGET (EXP-0095…0098)
|
||
|
||
**UPDATE (2026-09-22): the user moved physical board fabrication to
|
||
N=16.** `hardware/v3/rtl/n16_system_ddr3_top.v` (with the real EXP-0097
|
||
MAC-pipeline fix) is now the real, current target — `v3-artix7` (the
|
||
branch representing what is actually being built) has been fast-
|
||
forward merged with the former `n16-timing-closure` branch, so both
|
||
now point at the same real, consolidated state. **No physical board
|
||
change is needed for this** — DDR3/SPI/flash/config pins are all
|
||
package-level and identical regardless of `N_GROUPS` (confirmed real,
|
||
not assumed, since N=2/N=8/N=16 all route through the exact same real
|
||
XDC). Only the bitstream target changes.
|
||
|
||
N=8's own real signoff (§3.0 below) remains accurate and valid as a
|
||
real, simpler fallback if ever needed, but is no longer what is
|
||
physically being fabricated. N=2 (EXP-0088) remains the simplest real
|
||
fallback below that. See `docs/ARCHITECTURE_ANALYSIS.md` §5.6 for the
|
||
full real story of how N=16 was closed (EXP-0097/0098).
|
||
|
||
### 3.0 N=16, the current real fabrication target (`n16_system_ddr3_top.v`, EXP-0097/0098, with the MAC-pipeline fix)
|
||
|
||
| Metric | Value |
|
||
|---|---|
|
||
| DDR3 PHY clock (sys_clk_p/n) | **310.078 MHz** (3.225ns period) |
|
||
| Compute domain clock (ui_clk/clk_pll_i) | **155.039 MHz** |
|
||
| WNS (setup slack) | **+0.269 ns** — real, closed, 0 failing setup endpoints |
|
||
| WHS (hold slack) | +0.026 ns |
|
||
| Failing endpoints | 0 of 106087 (setup), 0 of 106084 (hold) |
|
||
| LUTs used | 19903 / 63400 (31.39%) |
|
||
| Registers used | 35409 / 126800 (27.93%) |
|
||
| DSP48E1 used | 128 / 240 (53.33%) — 8 per PE × 16 real parallel PEs |
|
||
| Real parallelism | **16 real parallel processing elements** |
|
||
| Real functional verification | 32/32 PASS, real DDR3 model (`tb_n16_system_ddr3.v`) |
|
||
|
||
See `docs/ARCHITECTURE_ANALYSIS.md` §5.6.1 for the full real root-cause
|
||
and fix story (the MAC-datapath pipeline split that closed this).
|
||
|
||
### 3.0-fallback N=8, real, valid fallback signoff (`n8_system_ddr3_top.v`, EXP-0095/0096, unmodified core)
|
||
|
||
Top-level RTL: `hardware/v3/rtl/n8_system_ddr3_top.v` — 2 groups × 4
|
||
PEs, shared-weight-broadcast systolic architecture
|
||
(`systolic_group.v`/`packed_pe.v`, `neural_director_grouped.v`,
|
||
`sdram_arbiter_hier.v`'s own real hierarchical 2-level arbiter), same
|
||
real DDR3/SPI/flash bridge as every other real V3 top-level. Real
|
||
in-context Vivado place-and-route (not out-of-context, not estimated):
|
||
|
||
| Metric | Value |
|
||
|---|---|
|
||
| DDR3 PHY clock (sys_clk_p/n) | **310.078 MHz** (3.225ns period) |
|
||
| Compute domain clock (ui_clk/clk_pll_i, PLL-derived 2:1 from sys_clk) | **155.039 MHz** |
|
||
| WNS (setup slack) | **0.000 ns** — real, closed, 0 failing setup endpoints (EXP-0095/0096). This is the exact RTL physically being fabricated. |
|
||
| WHS (hold slack) | +0.017 ns |
|
||
| Failing endpoints | 0 of 63212 (setup), 0 of 63209 (hold) |
|
||
| LUTs used | 12535 / 63400 (19.77%) |
|
||
| Registers used | 19902 / 126800 (15.70%) |
|
||
| DSP48E1 used | 64 / 240 (26.7%) — 8 per PE × 8 real parallel PEs |
|
||
| DDR3 channel width | **32-bit real physical channel** (two MT41J128M16JT-125:K chips) — see §3.1 |
|
||
| Real physical bandwidth ceiling | **~2.48 GB/s** (unchanged from the N=2 baseline — same DDR3 channel, shared across 8 real parallel PEs instead of 2) |
|
||
| Real parallelism | **8 real parallel processing elements** (4× the original N=2 design) |
|
||
| Result writeback | **real, into DDR3** (`result_writeback.v`), see §4.6/§5.3 of `docs/ARCHITECTURE_ANALYSIS.md` |
|
||
|
||
**Real, disclosed caveat**: this WNS=0.000ns margin is exactly zero —
|
||
real, closed, but with zero slack. Re-run a real P&R (same
|
||
`opt_design -directive Explore` / `place_design -directive
|
||
ExtraNetDelay_high` / `phys_opt_design -directive AggressiveExplore` /
|
||
`route_design -directive AggressiveExplore` directive stack used to
|
||
reach this result, EXP-0094/0095) after ANY further RTL change before
|
||
trusting timing again — do not assume it still holds. This zero-margin
|
||
result is exactly why EXP-0097's own pipeline fix (below) was pursued.
|
||
|
||
### 3.0.1 The real, consolidated family — EXP-0097/0098 (N=16 is the fabrication target; N=2/N=8 kept real and valid as fallbacks)
|
||
|
||
A single real fix (`neural_processor_packed.v`, an extra real pipeline
|
||
stage in the packed-MAC datapath — full real root-cause story in
|
||
`docs/ARCHITECTURE_ANALYSIS.md` §5.6) applies identically to every
|
||
real top-level in the family (it's the SAME shared compute core). Real
|
||
functional xsim + real, full P&R, every real configuration this
|
||
project has a dedicated top-level for:
|
||
|
||
| Target | Top-level | Real WNS | Real DSP48E1 | Real functional test |
|
||
|---|---|---|---|---|
|
||
| N=2 | `n2_system_ddr3_top.v` | **+0.389 ns** (up from +0.100ns) | 16/240 (6.67%) | 8/8 PASS |
|
||
| N=8 | `n8_system_ddr3_top.v` | **+0.108 ns** (up from 0.000ns exact) | 64/240 (26.7%) | 16/16 PASS |
|
||
| N=16 | `n16_system_ddr3_top.v` | **+0.269 ns** (up from −0.338ns FAILED) | 128/240 (53.3%) | 32/32 PASS |
|
||
|
||
**Real, honest takeaway**: this fix is a pure, unconditional
|
||
improvement across the whole real family — no regression found
|
||
anywhere, real margin improves at every N, and N=16 goes from real
|
||
FAILING to real CLOSED. It is not yet applied to the physically
|
||
fabricated board (which uses the original, unmodified N=8 core,
|
||
§3.0) — that board's own real signoff (WNS=0.000ns) remains accurate
|
||
for the hardware actually being built.
|
||
|
||
Signoff history (every real change, in-context P&R):
|
||
|
||
| EXP | Target | WNS (ns) | LUTs | DSP48E1 |
|
||
|---|---|---|---|---|
|
||
| 0074 (first real DDR3 + pins) | N=2 | +0.040 | 5140 | 16 |
|
||
| 0076 (+regfile, +pins, +SPI fix) | N=2 | +0.056 | 5173 | 16 |
|
||
| 0078 (+flash bridge, STARTUPE2) | N=2 | +0.013 | 5213 | 16 |
|
||
| 0079 (+real activation engine) | N=2 | +0.030 | 5379 | 16 |
|
||
| 0082 (+denser activation packing) | N=2 | +0.068 | 5437 | 16 |
|
||
| 0083 (+DDRManager phase 1, `ddr_prefetch_mgr.v`, 16-bit) | N=2 | +0.073 | 5644 | 16 |
|
||
| 0084 (32-bit DDR3 widening, 2900ps/172.414MHz) | N=2 | -0.618 (FAILED) | 6418 | 16 |
|
||
| 0086 (32-bit DDR3 widening, clock reverted to 3225ps/155.039MHz) | N=2 | +0.096 (CLOSED) | 6382 | 16 |
|
||
| 0088 (+result-writeback engine, `result_writeback.v`) | N=2 | +0.100 (CLOSED) | 6642 | 16 |
|
||
| 0093 (hybrid systolic, flat 21-way arbiter) | N=16 | -0.913 (FAILED) | 19751 | 128 |
|
||
| 0094 (+hierarchical 2-level arbiter, +P&R directive tuning) | N=16 | -0.338 (still FAILED) | 19936 | 128 |
|
||
| 0095/0096 (`n8_system_ddr3_top.v`, hybrid systolic — **in physical fabrication**) | **N=8** | **0.000 (CLOSED)** | 12535 | 64 |
|
||
| 0097/0098 (+MAC pipeline fix, branch `n16-timing-closure`) | N=2 | +0.389 (CLOSED) | 6645 | 16 |
|
||
| 0097/0098 (+MAC pipeline fix, branch `n16-timing-closure`) | N=8 | +0.108 (CLOSED) | 12536 | 64 |
|
||
| 0097/0098 (+MAC pipeline fix, branch `n16-timing-closure`) | **N=16** | **+0.269 (CLOSED)** | 19903 | 128 |
|
||
|
||
### 3.1 Real 32-bit DDR3 widening — functionally verified AND timing CLOSED (EXP-0084 → EXP-0086)
|
||
|
||
**Real, closed result**: the real 32-bit channel (two MT41J128M16JT-125:K
|
||
chips) is functionally complete (real xsim against the real 2-chip DDR3
|
||
model: `tb_mig_native_adapter.v` 12/12 PASS, `tb_n2_system_ddr3.v` 8/8
|
||
PASS) **and** its real P&R closes timing at the proven-safe
|
||
3225ps/155.039MHz clock: **WNS = +0.095707ns, WHS = +0.036275ns, 0 failing
|
||
endpoints** (EXP-0086). Route completed 100%, all user-specified timing
|
||
constraints met.
|
||
|
||
**How this was reached (the honest path, not a straight line)**: EXP-0084
|
||
first tried the width change stacked with an ADDITIONAL clock speedup
|
||
(2900ps/172.414MHz) in the same MIG wizard session — that combination
|
||
failed timing (WNS=-0.618ns), traced to the real worst path inside
|
||
`neural_processor_packed.v`'s own packed-MAC accumulation tree (a DSP48E1
|
||
output through a 4-deep CARRY4 chain, **unchanged since EXP-0059**) — it
|
||
had real margin at 155.039MHz but not at the faster 172.414MHz. This
|
||
was **not caused by the 32-bit width change itself**; every module touched
|
||
for the widening was already real-verified functionally correct. The
|
||
real, honest insight: bandwidth = width × clock rate, and widening to 32
|
||
bits **alone**, even at the OLD 3225ps/155.039MHz, already delivers the
|
||
FULL intended 2× bandwidth gain — the clock speedup was a separate,
|
||
stackable optimization, and reverting *only* it (keeping Data Width=32)
|
||
restored the same real margin `neural_processor_packed.v` already had at
|
||
EXP-0083 (+0.073ns → +0.096ns here, the small difference being normal
|
||
P&R placement-seed variance, not a real effect of the width change).
|
||
|
||
**A second real bug surfaced on the way to this result (EXP-0086)**: the
|
||
user's second real MIG wizard regeneration (needed only to revert Clock
|
||
Period) triggered Vivado to silently re-import the ENTIRE v3 RTL source
|
||
tree (9 files) plus the top XDC back to stale, pre-EXP-0084 copies — a
|
||
wholesale recurrence of the stale-import bug (see CLAUDE.md), not limited
|
||
to the single file touched by the regeneration. Root-caused via the
|
||
actual P&R log (parsing a `constrs_1/imports/...` path, and a real "IO
|
||
placement infeasible" failure reproducing EXP-0084's already-fixed VCCO
|
||
conflict) and fixed by re-establishing all 10 files as direct references
|
||
before re-running P&R.
|
||
|
||
This section (§3, §3.1) and the signoff table above are now the current,
|
||
trustworthy real baseline — replacing the EXP-0083 16-bit-era pointer.
|
||
See `hardware/v2/logs/experiments.log` EXP-0084 and EXP-0086 for the
|
||
complete story, including all real P&R/XDC bugs found and fixed along the
|
||
way (stale differential-clock ports, a real VCCO bank conflict, stale
|
||
imported XDC/RTL — twice, missing IOSTANDARDs, and two cosmetic but
|
||
previously-silently-broken XDC property bugs).
|
||
|
||
## 4. Real DDR3 memory layout convention
|
||
|
||
Both weight data and activation data share the same DDR3 address space,
|
||
word-addressed. **EXP-0084**: the native word width is now **32 bits**
|
||
(`BURST_LEN=8` per transaction = 256 bits/burst, up from 128 bits at the
|
||
old 16-bit width) — this is the real, current RTL (`hardware/v3/rtl/`),
|
||
functionally verified and, as of EXP-0086, real timing-closed; see §3.1.
|
||
|
||
- **Weights**: one layer's weight set starts at word address `layer_index *
|
||
WORDS_PER_LAYER` (`WORDS_PER_LAYER = LAYER_BYTES/2`). Densely packed —
|
||
`layer_prefetch_ctrl.v` reads full bursts sequentially into the on-chip
|
||
weight buffer once per job. `BYTES_PER_BURST` is now `4*BURST_LEN` (32
|
||
bytes/burst, up from 16) — the byte-drain logic itself is word-width-
|
||
agnostic and needed no change beyond this.
|
||
- **Activations** (real engine, `act_tile_fetch.v`; current layout is the
|
||
**v3 convention, EXP-0084**): **FOUR consecutive tiles (P_IN=8 INT8
|
||
values each, 64 bits each) share ONE full `BURST_LEN=8`-word (256-bit)
|
||
burst** — tile parity 0/1/2/3 (`tcnt[1:0]`) selects bits
|
||
`[63:0]`/`[127:64]`/`[191:128]`/`[255:192]` of the burst response. Tile
|
||
`t`'s burst address is `base + (t>>2)*BURST_LEN` (integer division —
|
||
four tiles per burst), always burst-aligned by construction. This is
|
||
**not** a further bytes-per-MAC reduction beyond EXP-0081's already-
|
||
optimal 1 byte/MAC — it's what's required to keep that same 100% packing
|
||
utilization at the new, larger burst size instead of leaving half of it
|
||
newly wasted.
|
||
- **Why this is timing-safe despite selecting a sub-burst quarter at
|
||
read time**: the tile index's own low 2 bits (which quarter of the
|
||
burst a given tile lives in) are known at *request* time, not at
|
||
response time. They're latched into a register (`sel_lat`) the same
|
||
cycle the request is accepted — many `ui_clk` cycles before the real
|
||
DDR3 round-trip completes and `ctrl_rdata` becomes valid. The eventual
|
||
data-select mux is an explicit 4-way `case` on constant byte offsets
|
||
(not a runtime-indexed part-select expression) — extending the exact
|
||
same discipline EXP-0081 established for the 1-bit case to 2 bits.
|
||
This mux is real and functionally verified; §3.1's earlier real timing
|
||
failure (now closed, EXP-0086) was unrelated to this fetch path — it
|
||
was inside the compute core's own accumulation tree.
|
||
- Real measured effect: back-to-back same-row DDR3 throughput was a
|
||
fixed 1.24 GB/s at 16-bit width (measured, EXP-0080); at 32-bit width
|
||
the real physical ceiling doubles to ~2.48 GB/s, independent of this
|
||
packing convention — packing controls how much of that ceiling is
|
||
*wasted*, not the ceiling itself (see `docs/ARCHITECTURE_ANALYSIS.md`
|
||
§3 and §5.1/§5.4).
|
||
- `base` (a job's own `x_base_a`/`x_base_b`) is chosen freely by whoever
|
||
submits jobs (the SPI host) — just keep each position's own activation
|
||
array in its own non-overlapping `(N_TILES/4) * BURST_LEN`-word region.
|
||
|
||
## 5. FPGA configuration (boot) procedure
|
||
|
||
Two complementary paths, both present on this board by design:
|
||
|
||
1. **Factory-first / recovery (JTAG, ESP32-driven)**: the flash starts
|
||
blank on a fresh board — no other path can bootstrap it (a real chicken-
|
||
and-egg constraint: the FPGA can't relay flash-programming commands over
|
||
SPI, §5.2, until it's already running logic that does that). The ESP32
|
||
bit-bangs JTAG (§2.5) to load a bitstream directly, or to run Vivado's own
|
||
"indirect SPI flash programming" sequence to write the flash for the
|
||
first time. One-time (or rare/recovery-only) step.
|
||
2. **Normal boot (Master SPI, autonomous)**: every subsequent power-on, the
|
||
FPGA self-configures from the flash via its own dedicated hardware (mode
|
||
pins set to Master SPI, §2.4) — no ESP32 involvement needed.
|
||
3. **Field firmware updates (SPI-through-FPGA, `FLASH_XFER` opcode 0x40)**:
|
||
once the FPGA is running, the ESP32 can rewrite the flash by relaying raw
|
||
SPI-NOR bytes through the FPGA over the management SPI bus (§2.2) — the
|
||
FPGA then re-transmits them as master on the flash bus (§2.3). This is the
|
||
**only** electrical path from ESP32 to the flash; there is no direct
|
||
connection (by design, per explicit requirement).
|
||
- **Real SPI-NOR opcodes** (verified against the actual W25Q32JV
|
||
datasheet, for whoever writes the ESP32-side flashing routine):
|
||
`0x06` Write Enable, `0x04` Write Disable, `0x05` Read Status Register-1
|
||
(bit0=BUSY, bit1=WEL), `0x02` Page Program, `0x03` Read Data, `0x20`
|
||
Sector Erase (4KB), `0x52` 32KB Block Erase, `0xD8` 64KB Block Erase,
|
||
`0xC7`/`0x60` Chip Erase.
|
||
- **Protocol timing note**: `FLASH_XFER` relays are NOT instantaneous —
|
||
each relayed byte's real flash response is only stable starting **two**
|
||
host-clocked bytes later (not one), so the host must clock 2 trailing
|
||
dummy bytes after its last real command byte to safely receive the
|
||
final response. See `spi_host_bridge_v3.v`'s own header for the full
|
||
real-measured reasoning (EXP-0077).
|
||
- After writing a new bitstream to the flash, reconfigure either by
|
||
pulsing `PROGRAM_B` externally, or (future work, not built yet) via a
|
||
`ICAPE2`-based warm self-reconfiguration triggered over the same SPI bus.
|
||
|
||
## 6. Management SPI protocol summary (for ESP32 firmware)
|
||
|
||
One opcode byte (MSB-first) per CS-low transaction, driven by
|
||
`spi_host_bridge_v3.v`:
|
||
|
||
| Opcode | Name | Payload | Purpose |
|
||
|---|---|---|---|
|
||
| 0x00 | NOP | 0 bytes | inert |
|
||
| 0x0F | RESET | 0 bytes | pulses a soft-reset |
|
||
| 0x10 | WRITE_JOB | 16 bytes | submit one inference job (node_id, x_base, w_base, n_tiles, result_addr) |
|
||
| 0x20 | STATUS | 0 bytes → 1 byte out | job_busy / mem_busy / last_job_accepted bits |
|
||
| 0x01 | WRITE_MEM | 4+2N bytes | raw DDR3 word write (N words) |
|
||
| 0x02 | READ_MEM | 6 bytes → 2N bytes out | raw DDR3 word read (N words) |
|
||
| 0x30 | REG_WRITE | 5 bytes | write a control register |
|
||
| 0x31 | REG_READ | 1 byte → 4 bytes out | read a status/ID register (0x00 DEVICE_ID, 0x01 CONTROL, 0x02 STATUS incl. DDR3-ready + Director-error, 0x03 N_SLOTS) |
|
||
| 0x40 | FLASH_XFER | N bytes → N bytes out (+2 margin) | raw passthrough to the config flash, see §5.3 |
|
||
|
||
Full byte-level field layouts are documented in `spi_host_bridge_v3.v`'s own
|
||
header comment — treat that file as the authoritative protocol spec, this
|
||
table is a summary/index.
|
||
|
||
**Real out-of-band notification (EXP-0085)**: `data_ready_n`, a real,
|
||
**active-low, sticky** top-level pin (D14, bank 15, 3.3V — see §2.2) —
|
||
lets the ESP32 be interrupt-driven instead of polling STATUS in a loop.
|
||
Asserted (driven low) when a job/pair completes (`job_out_done`) or while
|
||
a real Director error is active (`dir_error`); the job-completion latch
|
||
is **sticky** (stays low even after the underlying pulse ends) until the
|
||
host acknowledges by completing a real STATUS (0x20) or REG_READ(0x02)
|
||
transaction — a REG_READ of any *other* register does not acknowledge it.
|
||
The `dir_error` contribution is live/combinational, not latched — it
|
||
clears the moment `dir_error` itself does. Real firmware implication: the
|
||
ESP32 can wire this to a GPIO interrupt and only bother reading STATUS
|
||
when it actually fires, instead of polling every loop iteration.
|
||
|
||
## 7. Known-open items (honestly disclosed, not hidden)
|
||
|
||
- Scaling past N=2 compute cores: **RESOLVED, N=8 is the configuration
|
||
in physical fabrication** (§3.0, EXP-0095/0096, `n8_system_ddr3_
|
||
top.v`, WNS=0.000ns CLOSED). **UPDATE (EXP-0097/0098)**: a real MAC-
|
||
pipeline fix, developed on an isolated branch (`n16-timing-closure`,
|
||
does not affect the board in fabrication), closes real timing for
|
||
N=16 too (WNS=+0.269ns) and improves N=2's and N=8's own margins
|
||
(+0.389ns, +0.108ns respectively) — see §3.0.1 and
|
||
`docs/ARCHITECTURE_ANALYSIS.md` §5.6 for the full real story. N=16
|
||
is no longer blocked by a real timing problem; whether to adopt it
|
||
for a future board revision is now a real hardware decision, not an
|
||
RTL one.
|
||
- `sys_rst` has only a *tentative* real pin (G13, bank 15, EXP-0084 —
|
||
chosen just to unblock real P&R, not a final board decision) — assign
|
||
its real, permanent location once the rest of the board layout (reset
|
||
circuit, status LEDs, etc.) is decided. Confirmed real, routed, and
|
||
timing-clean at the closed EXP-0086 signoff.
|
||
- ESP32-side JTAG bit-banging firmware (§5.1) does not exist yet — it's
|
||
software work on the host side, not part of this FPGA RTL.
|
||
- `ddr_prefetch_mgr.v` (DDRManager phase 1, EXP-0083) is real, built, and
|
||
timing-verified. Its real performance benefit was 2.86% at the OLD
|
||
16-bit/155MHz signoff, but a real re-measurement against the closed
|
||
32-bit channel (EXP-0087) found the benefit is now GONE (~0%, a
|
||
0.0064% regression, statistically a wash) — see
|
||
`docs/ARCHITECTURE_ANALYSIS.md` §5.2. Kept wired into `packed_slot.v`
|
||
for correctness/timing-neutrality only, not performance. The larger
|
||
multi-slot DDRManager is not built and, per EXP-0087, is not currently
|
||
justified.
|
||
- `data_ready_n` (EXP-0085, user-requested active-low sticky IRQ) is real,
|
||
built, functionally verified (49/49 including 10 new checks,
|
||
`tb_spi_host_bridge_v3.v`), AND real P&R-verified at the closed EXP-0086
|
||
signoff (D14/LVCMOS33, confirmed via a direct query on the routed
|
||
checkpoint) — see §6.
|
||
- `result_writeback.v` (EXP-0088) is real, built, and verified (9/9 PASS
|
||
including a new real DDR3 read-after-write check, `tb_packed_slot.v`).
|
||
P&R re-verified with the result-writeback addition (EXP-0088) and,
|
||
since then, real N=4/8/16 scaling has been built and P&R-tested — see
|
||
§3 above and `docs/ARCHITECTURE_ANALYSIS.md` §5.3/§5.6.
|