Files
FPGA-Neural/docs/PHYSICAL_REALIZATION.md
micheleandClaude Sonnet 5 0303aacb9f docs: consolidate N=16 to N=8's own level of rigor (EXP-0098) -- whole family real and closed
N=2 re-verified with the same shared pipelined neural_processor_packed.v
(EXP-0097's fix): real functional xsim 8/8 PASS (identical completion
time to the pre-fix baseline), real full P&R WNS=+0.389ns (up from
+0.099962ns), 0 failing endpoints, 16 DSP48E1 -- no regression.

Real, consolidated family state, all sharing the same pipelined core:
  N=2:  WNS=+0.389ns, 16 DSP48E1,  8/8 functional PASS
  N=8:  WNS=+0.108ns, 64 DSP48E1,  16/16 functional PASS
  N=16: WNS=+0.269ns, 128 DSP48E1, 32/32 functional PASS

Updated the project's own primary real docs (docs/PHYSICAL_REALIZATION.md
S3, docs/ARCHITECTURE_ANALYSIS.md S5.6) to the same depth EXP-0096 gave
N=8 alone -- not just the docs/latex/ deliverables. Clear two-track
framing throughout: the physically fabricated board still uses N=8's
original, unmodified core (WNS=0.000ns exact); this consolidated,
improved family lives on the isolated n16-timing-closure branch,
pending a real hardware decision on whether/when to adopt it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MUG92aM9m68TRc4rG55BcC
2026-09-22 01:15:15 +02:00

27 KiB
Raw Permalink Blame History

FPGA-Neural V3 — Physical Realization Reference

Real, verified data for the custom PCB (bare XC7A100T-CSG324-2 + DDR3, no dev board). Every pin/part/setting below comes from a real Vivado-generated constraint file, a real datasheet, or a real place-and-route run — none of it is guessed. See hardware/v2/logs/experiments.log (EXP-0059 onward) for the full derivation history.

1. Core components (real, verified availability)

Component Part Notes
FPGA XC7A100T-CSG324-2 Speed grade -2 (corrected from an initial -1 assumption, EXP-0074) — same die/package/footprint as -1, strictly better timing margin.
DDR3 SDRAM Micron MT41J128M16JT-125:K 2Gb, x16, DDR3-1600-rated (run at 310.078MHz here due to -2 timing closure, see §3). Verified in-stock on LCSC.
Config flash Winbond W25Q32JVSSIQ 32Mbit/4MB, SOIC-8. Comfortably fits the ~30.5Mbit full XC7A100T bitstream. Verified in-stock on LCSC. Wired exclusively to the FPGA (see §5).

2. FPGA pin assignments (real, from the routed design)

2.1 DDR3 (fixed by the FPGA's own PHY hardware — not a free choice)

EXP-0084 UPDATE: real 32-bit channel (two MT41J128M16JT-125:K chips ganged in parallel) — dq/dqs/dm pin counts DOUBLED versus the original 16-bit design (dq 16→32, dqs 2→4 pairs, dm 2→4). Address/command/control lines are unchanged and fanned out identically to both chips. Real, board-accurate pin placement below (confirmed from the actual routed design, mig_7series_0.xdc + this project's own top-level XDC) — functionally verified, but see §3 for the real, honest timing-closure status before treating this as a final signoff.

Generated by the Vivado MIG wizard (mig_7series_0.xdc), all SSTL15 / DIFF_SSTL15 (1.5V), banks 34/35:

Signal Pin Signal Pin Signal Pin
ddr3_dq[0] C5 ddr3_dq[16] F4 ddr3_addr[0] U3
ddr3_dq[1] B7 ddr3_dq[17] F3 ddr3_addr[1] U4
ddr3_dq[2] B6 ddr3_dq[18] E2 ddr3_addr[2] V1
ddr3_dq[3] D8 ddr3_dq[19] D2 ddr3_addr[3] U1
ddr3_dq[4] C7 ddr3_dq[20] C1 ddr3_addr[4] L5
ddr3_dq[5] E6 ddr3_dq[21] H1 ddr3_addr[5] L6
ddr3_dq[6] E5 ddr3_dq[22] G1 ddr3_addr[6] L4
ddr3_dq[7] E7 ddr3_dq[23] F1 ddr3_addr[7] K5
ddr3_dq[8] B4 ddr3_dq[24] F6 ddr3_addr[8] M2
ddr3_dq[9] A4 ddr3_dq[25] G4 ddr3_addr[9] M3
ddr3_dq[10] A3 ddr3_dq[26] G3 ddr3_addr[10] L3
ddr3_dq[11] B3 ddr3_dq[27] J3 ddr3_addr[11] K3
ddr3_dq[12] B2 ddr3_dq[28] J2 ddr3_addr[12] M1
ddr3_dq[13] D5 ddr3_dq[29] K2 ddr3_addr[13] L1
ddr3_dq[14] D4 ddr3_dq[30] K1 ddr3_ba[0] V5
ddr3_dq[15] E3 ddr3_dq[31] H6 ddr3_ba[1] V2
ddr3_dm[0] C6 ddr3_dm[2] C2 ddr3_ba[2] U2
ddr3_dm[1] C4 ddr3_dm[3] G6
ddr3_dqs_p[0] A6 (DIFF) ddr3_dqs_n[0] A5 (DIFF) ddr3_cas_n R3
ddr3_dqs_p[1] B1 (DIFF) ddr3_dqs_n[1] A1 (DIFF) ddr3_cke[0] P2
ddr3_dqs_p[2] H2 (DIFF) ddr3_dqs_n[2] G2 (DIFF) ddr3_cs_n[0] T5
ddr3_dqs_p[3] J4 (DIFF) ddr3_dqs_n[3] H4 (DIFF) ddr3_odt[0] R2
ddr3_ck_p[0] N2 (DIFF) ddr3_ck_n[0] N1 (DIFF) ddr3_ras_n V4
ddr3_reset_n D7 (LVCMOS15) ddr3_we_n T3
sys_clk_p N5 (DIFF_SSTL15, bank 34) sys_clk_n P5 (DIFF_SSTL15, bank 34)
clk_ref_p T14 (LVDS_25, bank 14) clk_ref_n T15 (LVDS_25, bank 14)

Real board implication (EXP-0084): both sys_clk and clk_ref are now genuine differential pairs — the board needs a differential oscillator (or a differential buffer stage), not a single-ended one. clk_ref lives in bank 14 (not 34/35) because the MIG wizard's own UG586 placement rules restrict that specific net to bank 14 for this part/package — not a free choice. This directly displaced the config-flash SPI bus, which used to share bank 14 (see §2.3 — moved to bank 16 in EXP-0084 to resolve a real VCCO conflict, LVCMOS33 vs LVDS_25 cannot coexist in one bank).

Bank voltage requirements: bank 34/35 → 1.5V (DDR3 SSTL15, plus sys_clk_p/n at bank 34's own 1.5V-domain DIFF_SSTL15), bank 14 → 2.5V (clk_ref_p/n, LVDS_25).

INTERNAL_VREF for banks 34/35 is set to 0.750V by the MIG constraints (required for SSTL15 single-ended inputs) — this is a Vivado-side setting, not a board component, but note it if you ever inspect bitstream generation warnings about VREF.

2.2 Neural-processor management SPI (ESP32 ↔ FPGA, FPGA is slave)

Bank 15, package edge column A/B, physically adjacent (short traces), LVCMOS33:

Signal Pin Direction (FPGA side)
sclk A15 input
mosi B16 input
miso B17 output
cs_n A16 input
sys_rst G13 input (EXP-0084, tentative — see §7)
data_ready_n D14 output (EXP-0085, active-low sticky IRQ — see §6, tentative — see §7)

Bank 15 VCCO: assumed 3.3V — change the XDC's IOSTANDARD if your board power plan uses a different rail for this bank.

2.3 Config flash SPI (FPGA ↔ flash, FPGA is master)

EXP-0084 REAL PIN CHANGE: originally these were the FPGA's own dedicated Master-SPI configuration pins (K17/K18/L13, bank 14), reclaimed as ordinary fabric I/O post-configuration. EXP-0084's real 32-bit DDR3 widening moved the differential reference clock (clk_ref_p/n, §2.1) into bank 14 too — a real VCCO conflict (the flash needs LVCMOS33/3.3V, clk_ref needs LVDS_25/2.5V, one bank can only have one VCCO), confirmed by a real place_design failure, not hypothetical. Fixed by moving the flash bus to bank 16 (completely unconstrained before this, no VCCO commitment):

Signal Pin Direction (FPGA side) Notes
flash_mosi D9 output bank 16 (was K17/D00_MOSI, bank 14)
flash_miso D10 input bank 16 (was K18/D01_DIN, bank 14)
flash_cs_n C9 output bank 16 (was L13/FCS_B, bank 14)
(CCLK) E9 output Not a top-level port — driven internally via the STARTUPE2 primitive. Wire the flash's own CLK pin to package pin E9. Unaffected by the bank-14→16 move (CCLK was never a plain port).

Still requires BITSTREAM.CONFIG.PERSIST = NO (the Vivado default — set explicitly in the project XDC; EXP-0084 note: the XDC previously had this as FALSE, an invalid enum value in this Vivado version — silently not applied at all until fixed to the real NO/YES enum).

Bank 16 VCCO: 3.3V (matches the flash's own VCC, typically 1.83.6V for the W25Q32JV) — chosen freely since bank 16 had no prior VCCO commitment.

Real board implication: since the flash bridge no longer uses the FPGA's own dedicated D00_MOSI/D01_DIN/FCS_B configuration pins at all (K17/K18/L13 are now unused by this design), the flash chip's SPI lines on the PCB connect to D9/D10/C9 instead — a real routing change versus any earlier board layout draft based on the pre-EXP-0084 pinout.

Reserved, do not use (bank 14, unused by this design, kept clear for any future Quad-SPI/BPI expansion — no longer near the flash bus, now just general config-mode-reserved pins): L16 (EMCCLK), R16 (RDWR_B), V15 (CSI_B). The project's own XDC PROHIBITs the underlying sites for these (fixed in EXP-0084 — PROHIBIT is not a valid property directly on package_pin objects in this Vivado version, only on their site; this constraint had silently never been enforced before).

2.4 FPGA configuration control (dedicated, bank 0, not negotiable)

Signal Pin Purpose
PROGRAM_B P9 pulse low to force a full reconfiguration from flash
INIT_B P7 goes low during config; can indicate a config error if it re-asserts
DONE P10 goes high once configuration succeeds — wire to a status LED if desired
M0 P12 mode select
M1 P13 mode select
M2 P11 mode select
CFGBVS P8 tie to match bank 0's VCCO logic level (see UG470)

Mode pin setting for Master SPI boot (the flash-based autonomous boot path, see §5): M[2:0] = 001 (per UG470's mode pin table) — tie via pull-up/ pull-down resistors on the board, not driven dynamically.

2.5 JTAG (always available, independent of flash content)

Signal Pin
TCK E10
TDI E11
TMS E12
TDO E13

Used for: (a) first-ever/factory programming when the flash is blank (see §5), (b) recovery, (c) development/debug. This project's own plan drives these from an ESP32 doing real JTAG bit-banging (TAP state machine, IR/DR shifting) rather than a bench programmer — that firmware is separate, software-side work, not covered here.

3. Real timing signoff — N=8 in physical fabrication; N=2/N=8/N=16 all real, closed on branch n16-timing-closure (EXP-0095…0098)

Real, current, two-track status (2026-09-22): the physical board currently being fabricated uses N=8 (hardware/v3/rtl/ n8_system_ddr3_top.v, unmodified, EXP-0095/0096 signoff below) — this was the real, definitive decision made on 2026-09-21 after a real, measured N=4/8/16 timing curve found N=8 the largest configuration that closed cleanly at the time. Separately, on an isolated real development branch (n16-timing-closure, does NOT affect the board already in fabrication), a real fix (EXP-0097 — an extra real pipeline stage in neural_processor_packed.v's own MAC datapath, see docs/ARCHITECTURE_ANALYSIS.md §5.6) has since closed real timing for N=16 too, and improved N=2's and N=8's own margins — the whole real product family (N=2/N=8/N=16) now closes with positive margin under that fix. Whether/when to adopt this fix for a future board revision (potentially N=16) is a real, open hardware decision, not an RTL blocker — see docs/ARCHITECTURE_ANALYSIS.md §5.6 for the full story.

3.0 N=8, in physical fabrication (n8_system_ddr3_top.v, EXP-0095/0096, unmodified core)

Top-level RTL: hardware/v3/rtl/n8_system_ddr3_top.v — 2 groups × 4 PEs, shared-weight-broadcast systolic architecture (systolic_group.v/packed_pe.v, neural_director_grouped.v, sdram_arbiter_hier.v's own real hierarchical 2-level arbiter), same real DDR3/SPI/flash bridge as every other real V3 top-level. Real in-context Vivado place-and-route (not out-of-context, not estimated):

Metric Value
DDR3 PHY clock (sys_clk_p/n) 310.078 MHz (3.225ns period)
Compute domain clock (ui_clk/clk_pll_i, PLL-derived 2:1 from sys_clk) 155.039 MHz
WNS (setup slack) 0.000 ns — real, closed, 0 failing setup endpoints (EXP-0095/0096). This is the exact RTL physically being fabricated.
WHS (hold slack) +0.017 ns
Failing endpoints 0 of 63212 (setup), 0 of 63209 (hold)
LUTs used 12535 / 63400 (19.77%)
Registers used 19902 / 126800 (15.70%)
DSP48E1 used 64 / 240 (26.7%) — 8 per PE × 8 real parallel PEs
DDR3 channel width 32-bit real physical channel (two MT41J128M16JT-125:K chips) — see §3.1
Real physical bandwidth ceiling ~2.48 GB/s (unchanged from the N=2 baseline — same DDR3 channel, shared across 8 real parallel PEs instead of 2)
Real parallelism 8 real parallel processing elements (4× the original N=2 design)
Result writeback real, into DDR3 (result_writeback.v), see §4.6/§5.3 of docs/ARCHITECTURE_ANALYSIS.md

Real, disclosed caveat: this WNS=0.000ns margin is exactly zero — real, closed, but with zero slack. Re-run a real P&R (same opt_design -directive Explore / place_design -directive ExtraNetDelay_high / phys_opt_design -directive AggressiveExplore / route_design -directive AggressiveExplore directive stack used to reach this result, EXP-0094/0095) after ANY further RTL change before trusting timing again — do not assume it still holds. This zero-margin result is exactly why EXP-0097's own pipeline fix (below) was pursued.

3.0.1 The real, consolidated family — EXP-0097/0098, branch n16-timing-closure (not yet on the fabricated board)

A single real fix (neural_processor_packed.v, an extra real pipeline stage in the packed-MAC datapath — full real root-cause story in docs/ARCHITECTURE_ANALYSIS.md §5.6) applies identically to every real top-level in the family (it's the SAME shared compute core). Real functional xsim + real, full P&R, every real configuration this project has a dedicated top-level for:

Target Top-level Real WNS Real DSP48E1 Real functional test
N=2 n2_system_ddr3_top.v +0.389 ns (up from +0.100ns) 16/240 (6.67%) 8/8 PASS
N=8 n8_system_ddr3_top.v +0.108 ns (up from 0.000ns exact) 64/240 (26.7%) 16/16 PASS
N=16 n16_system_ddr3_top.v +0.269 ns (up from 0.338ns FAILED) 128/240 (53.3%) 32/32 PASS

Real, honest takeaway: this fix is a pure, unconditional improvement across the whole real family — no regression found anywhere, real margin improves at every N, and N=16 goes from real FAILING to real CLOSED. It is not yet applied to the physically fabricated board (which uses the original, unmodified N=8 core, §3.0) — that board's own real signoff (WNS=0.000ns) remains accurate for the hardware actually being built.

Signoff history (every real change, in-context P&R):

EXP Target WNS (ns) LUTs DSP48E1
0074 (first real DDR3 + pins) N=2 +0.040 5140 16
0076 (+regfile, +pins, +SPI fix) N=2 +0.056 5173 16
0078 (+flash bridge, STARTUPE2) N=2 +0.013 5213 16
0079 (+real activation engine) N=2 +0.030 5379 16
0082 (+denser activation packing) N=2 +0.068 5437 16
0083 (+DDRManager phase 1, ddr_prefetch_mgr.v, 16-bit) N=2 +0.073 5644 16
0084 (32-bit DDR3 widening, 2900ps/172.414MHz) N=2 -0.618 (FAILED) 6418 16
0086 (32-bit DDR3 widening, clock reverted to 3225ps/155.039MHz) N=2 +0.096 (CLOSED) 6382 16
0088 (+result-writeback engine, result_writeback.v) N=2 +0.100 (CLOSED) 6642 16
0093 (hybrid systolic, flat 21-way arbiter) N=16 -0.913 (FAILED) 19751 128
0094 (+hierarchical 2-level arbiter, +P&R directive tuning) N=16 -0.338 (still FAILED) 19936 128
0095/0096 (n8_system_ddr3_top.v, hybrid systolic — in physical fabrication) N=8 0.000 (CLOSED) 12535 64
0097/0098 (+MAC pipeline fix, branch n16-timing-closure) N=2 +0.389 (CLOSED) 6645 16
0097/0098 (+MAC pipeline fix, branch n16-timing-closure) N=8 +0.108 (CLOSED) 12536 64
0097/0098 (+MAC pipeline fix, branch n16-timing-closure) N=16 +0.269 (CLOSED) 19903 128

3.1 Real 32-bit DDR3 widening — functionally verified AND timing CLOSED (EXP-0084 → EXP-0086)

Real, closed result: the real 32-bit channel (two MT41J128M16JT-125:K chips) is functionally complete (real xsim against the real 2-chip DDR3 model: tb_mig_native_adapter.v 12/12 PASS, tb_n2_system_ddr3.v 8/8 PASS) and its real P&R closes timing at the proven-safe 3225ps/155.039MHz clock: WNS = +0.095707ns, WHS = +0.036275ns, 0 failing endpoints (EXP-0086). Route completed 100%, all user-specified timing constraints met.

How this was reached (the honest path, not a straight line): EXP-0084 first tried the width change stacked with an ADDITIONAL clock speedup (2900ps/172.414MHz) in the same MIG wizard session — that combination failed timing (WNS=-0.618ns), traced to the real worst path inside neural_processor_packed.v's own packed-MAC accumulation tree (a DSP48E1 output through a 4-deep CARRY4 chain, unchanged since EXP-0059) — it had real margin at 155.039MHz but not at the faster 172.414MHz. This was not caused by the 32-bit width change itself; every module touched for the widening was already real-verified functionally correct. The real, honest insight: bandwidth = width × clock rate, and widening to 32 bits alone, even at the OLD 3225ps/155.039MHz, already delivers the FULL intended 2× bandwidth gain — the clock speedup was a separate, stackable optimization, and reverting only it (keeping Data Width=32) restored the same real margin neural_processor_packed.v already had at EXP-0083 (+0.073ns → +0.096ns here, the small difference being normal P&R placement-seed variance, not a real effect of the width change).

A second real bug surfaced on the way to this result (EXP-0086): the user's second real MIG wizard regeneration (needed only to revert Clock Period) triggered Vivado to silently re-import the ENTIRE v3 RTL source tree (9 files) plus the top XDC back to stale, pre-EXP-0084 copies — a wholesale recurrence of the stale-import bug (see CLAUDE.md), not limited to the single file touched by the regeneration. Root-caused via the actual P&R log (parsing a constrs_1/imports/... path, and a real "IO placement infeasible" failure reproducing EXP-0084's already-fixed VCCO conflict) and fixed by re-establishing all 10 files as direct references before re-running P&R.

This section (§3, §3.1) and the signoff table above are now the current, trustworthy real baseline — replacing the EXP-0083 16-bit-era pointer. See hardware/v2/logs/experiments.log EXP-0084 and EXP-0086 for the complete story, including all real P&R/XDC bugs found and fixed along the way (stale differential-clock ports, a real VCCO bank conflict, stale imported XDC/RTL — twice, missing IOSTANDARDs, and two cosmetic but previously-silently-broken XDC property bugs).

4. Real DDR3 memory layout convention

Both weight data and activation data share the same DDR3 address space, word-addressed. EXP-0084: the native word width is now 32 bits (BURST_LEN=8 per transaction = 256 bits/burst, up from 128 bits at the old 16-bit width) — this is the real, current RTL (hardware/v3/rtl/), functionally verified and, as of EXP-0086, real timing-closed; see §3.1.

  • Weights: one layer's weight set starts at word address layer_index * WORDS_PER_LAYER (WORDS_PER_LAYER = LAYER_BYTES/2). Densely packed — layer_prefetch_ctrl.v reads full bursts sequentially into the on-chip weight buffer once per job. BYTES_PER_BURST is now 4*BURST_LEN (32 bytes/burst, up from 16) — the byte-drain logic itself is word-width- agnostic and needed no change beyond this.
  • Activations (real engine, act_tile_fetch.v; current layout is the v3 convention, EXP-0084): FOUR consecutive tiles (P_IN=8 INT8 values each, 64 bits each) share ONE full BURST_LEN=8-word (256-bit) burst — tile parity 0/1/2/3 (tcnt[1:0]) selects bits [63:0]/[127:64]/[191:128]/[255:192] of the burst response. Tile t's burst address is base + (t>>2)*BURST_LEN (integer division — four tiles per burst), always burst-aligned by construction. This is not a further bytes-per-MAC reduction beyond EXP-0081's already- optimal 1 byte/MAC — it's what's required to keep that same 100% packing utilization at the new, larger burst size instead of leaving half of it newly wasted.
    • Why this is timing-safe despite selecting a sub-burst quarter at read time: the tile index's own low 2 bits (which quarter of the burst a given tile lives in) are known at request time, not at response time. They're latched into a register (sel_lat) the same cycle the request is accepted — many ui_clk cycles before the real DDR3 round-trip completes and ctrl_rdata becomes valid. The eventual data-select mux is an explicit 4-way case on constant byte offsets (not a runtime-indexed part-select expression) — extending the exact same discipline EXP-0081 established for the 1-bit case to 2 bits. This mux is real and functionally verified; §3.1's earlier real timing failure (now closed, EXP-0086) was unrelated to this fetch path — it was inside the compute core's own accumulation tree.
    • Real measured effect: back-to-back same-row DDR3 throughput was a fixed 1.24 GB/s at 16-bit width (measured, EXP-0080); at 32-bit width the real physical ceiling doubles to ~2.48 GB/s, independent of this packing convention — packing controls how much of that ceiling is wasted, not the ceiling itself (see docs/ARCHITECTURE_ANALYSIS.md §3 and §5.1/§5.4).
    • base (a job's own x_base_a/x_base_b) is chosen freely by whoever submits jobs (the SPI host) — just keep each position's own activation array in its own non-overlapping (N_TILES/4) * BURST_LEN-word region.

5. FPGA configuration (boot) procedure

Two complementary paths, both present on this board by design:

  1. Factory-first / recovery (JTAG, ESP32-driven): the flash starts blank on a fresh board — no other path can bootstrap it (a real chicken- and-egg constraint: the FPGA can't relay flash-programming commands over SPI, §5.2, until it's already running logic that does that). The ESP32 bit-bangs JTAG (§2.5) to load a bitstream directly, or to run Vivado's own "indirect SPI flash programming" sequence to write the flash for the first time. One-time (or rare/recovery-only) step.
  2. Normal boot (Master SPI, autonomous): every subsequent power-on, the FPGA self-configures from the flash via its own dedicated hardware (mode pins set to Master SPI, §2.4) — no ESP32 involvement needed.
  3. Field firmware updates (SPI-through-FPGA, FLASH_XFER opcode 0x40): once the FPGA is running, the ESP32 can rewrite the flash by relaying raw SPI-NOR bytes through the FPGA over the management SPI bus (§2.2) — the FPGA then re-transmits them as master on the flash bus (§2.3). This is the only electrical path from ESP32 to the flash; there is no direct connection (by design, per explicit requirement).
    • Real SPI-NOR opcodes (verified against the actual W25Q32JV datasheet, for whoever writes the ESP32-side flashing routine): 0x06 Write Enable, 0x04 Write Disable, 0x05 Read Status Register-1 (bit0=BUSY, bit1=WEL), 0x02 Page Program, 0x03 Read Data, 0x20 Sector Erase (4KB), 0x52 32KB Block Erase, 0xD8 64KB Block Erase, 0xC7/0x60 Chip Erase.
    • Protocol timing note: FLASH_XFER relays are NOT instantaneous — each relayed byte's real flash response is only stable starting two host-clocked bytes later (not one), so the host must clock 2 trailing dummy bytes after its last real command byte to safely receive the final response. See spi_host_bridge_v3.v's own header for the full real-measured reasoning (EXP-0077).
    • After writing a new bitstream to the flash, reconfigure either by pulsing PROGRAM_B externally, or (future work, not built yet) via a ICAPE2-based warm self-reconfiguration triggered over the same SPI bus.

6. Management SPI protocol summary (for ESP32 firmware)

One opcode byte (MSB-first) per CS-low transaction, driven by spi_host_bridge_v3.v:

Opcode Name Payload Purpose
0x00 NOP 0 bytes inert
0x0F RESET 0 bytes pulses a soft-reset
0x10 WRITE_JOB 16 bytes submit one inference job (node_id, x_base, w_base, n_tiles, result_addr)
0x20 STATUS 0 bytes → 1 byte out job_busy / mem_busy / last_job_accepted bits
0x01 WRITE_MEM 4+2N bytes raw DDR3 word write (N words)
0x02 READ_MEM 6 bytes → 2N bytes out raw DDR3 word read (N words)
0x30 REG_WRITE 5 bytes write a control register
0x31 REG_READ 1 byte → 4 bytes out read a status/ID register (0x00 DEVICE_ID, 0x01 CONTROL, 0x02 STATUS incl. DDR3-ready + Director-error, 0x03 N_SLOTS)
0x40 FLASH_XFER N bytes → N bytes out (+2 margin) raw passthrough to the config flash, see §5.3

Full byte-level field layouts are documented in spi_host_bridge_v3.v's own header comment — treat that file as the authoritative protocol spec, this table is a summary/index.

Real out-of-band notification (EXP-0085): data_ready_n, a real, active-low, sticky top-level pin (D14, bank 15, 3.3V — see §2.2) — lets the ESP32 be interrupt-driven instead of polling STATUS in a loop. Asserted (driven low) when a job/pair completes (job_out_done) or while a real Director error is active (dir_error); the job-completion latch is sticky (stays low even after the underlying pulse ends) until the host acknowledges by completing a real STATUS (0x20) or REG_READ(0x02) transaction — a REG_READ of any other register does not acknowledge it. The dir_error contribution is live/combinational, not latched — it clears the moment dir_error itself does. Real firmware implication: the ESP32 can wire this to a GPIO interrupt and only bother reading STATUS when it actually fires, instead of polling every loop iteration.

7. Known-open items (honestly disclosed, not hidden)

  • Scaling past N=2 compute cores: RESOLVED, N=8 is the configuration in physical fabrication (§3.0, EXP-0095/0096, n8_system_ddr3_ top.v, WNS=0.000ns CLOSED). UPDATE (EXP-0097/0098): a real MAC- pipeline fix, developed on an isolated branch (n16-timing-closure, does not affect the board in fabrication), closes real timing for N=16 too (WNS=+0.269ns) and improves N=2's and N=8's own margins (+0.389ns, +0.108ns respectively) — see §3.0.1 and docs/ARCHITECTURE_ANALYSIS.md §5.6 for the full real story. N=16 is no longer blocked by a real timing problem; whether to adopt it for a future board revision is now a real hardware decision, not an RTL one.
  • sys_rst has only a tentative real pin (G13, bank 15, EXP-0084 — chosen just to unblock real P&R, not a final board decision) — assign its real, permanent location once the rest of the board layout (reset circuit, status LEDs, etc.) is decided. Confirmed real, routed, and timing-clean at the closed EXP-0086 signoff.
  • ESP32-side JTAG bit-banging firmware (§5.1) does not exist yet — it's software work on the host side, not part of this FPGA RTL.
  • ddr_prefetch_mgr.v (DDRManager phase 1, EXP-0083) is real, built, and timing-verified. Its real performance benefit was 2.86% at the OLD 16-bit/155MHz signoff, but a real re-measurement against the closed 32-bit channel (EXP-0087) found the benefit is now GONE (~0%, a 0.0064% regression, statistically a wash) — see docs/ARCHITECTURE_ANALYSIS.md §5.2. Kept wired into packed_slot.v for correctness/timing-neutrality only, not performance. The larger multi-slot DDRManager is not built and, per EXP-0087, is not currently justified.
  • data_ready_n (EXP-0085, user-requested active-low sticky IRQ) is real, built, functionally verified (49/49 including 10 new checks, tb_spi_host_bridge_v3.v), AND real P&R-verified at the closed EXP-0086 signoff (D14/LVCMOS33, confirmed via a direct query on the routed checkpoint) — see §6.
  • result_writeback.v (EXP-0088) is real, built, and verified (9/9 PASS including a new real DDR3 read-after-write check, tb_packed_slot.v). P&R re-verified with the result-writeback addition (EXP-0088) and, since then, real N=4/8/16 scaling has been built and P&R-tested — see §3 above and docs/ARCHITECTURE_ANALYSIS.md §5.3/§5.6.