New ddr_prefetch_mgr.v wraps act_tile_fetch.v with a depth-2 ping-pong buffer, issuing the next tile's DDR3 fetch as soon as the fetch engine is free instead of waiting for packed_slot.v to finish consuming the current tile. Wired into packed_slot.v's tile loop (job-level start instead of per-tile req), simplifying the S_TILEWAIT join in the process (ddrpf_tile_valid is level-held, no separate act_seen latch needed). Verification: new tb_ddr_prefetch_mgr.v (25/25 PASS after fixing a real testbench polling race found via iteration-tagged tracing, not an RTL bug), tb_packed_slot.v re-run unmodified (9/9 PASS, bit-identical results), tb_n2_system_ddr3.v re-run via real xsim against real ddr3_model.sv (8/8 PASS). Real P&R: WNS +0.073ns (up from EXP-0082's +0.068ns), LUTs 5644, DSP48E1 16 unchanged, 0 failing endpoints. Honest result: real A/B on the actual DDR3 backend (same testbench, before/after) shows a real but modest 2.86% reduction in total simulated time - smaller than the original hypothesis suggested, because neural_processor_packed.v already accepts one operand per cycle, so the per-tile dead time being removed was already small relative to real DDR3 fetch latency. Docs updated to report this honestly rather than oversell it; the larger multi-slot DDRManager is deferred pending re-measurement against the (still pending, user-gated) 32-bit channel widening. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MUG92aM9m68TRc4rG55BcC
268 lines
14 KiB
Markdown
268 lines
14 KiB
Markdown
# FPGA-Neural V3 — Physical Realization Reference
|
||
|
||
Real, verified data for the custom PCB (bare XC7A100T-CSG324-2 + DDR3, no dev
|
||
board). Every pin/part/setting below comes from a real Vivado-generated
|
||
constraint file, a real datasheet, or a real place-and-route run — none of it
|
||
is guessed. See `hardware/v2/logs/experiments.log` (EXP-0059 onward) for the
|
||
full derivation history.
|
||
|
||
## 1. Core components (real, verified availability)
|
||
|
||
| Component | Part | Notes |
|
||
|---|---|---|
|
||
| FPGA | **XC7A100T-CSG324-2** | Speed grade **-2** (corrected from an initial -1 assumption, EXP-0074) — same die/package/footprint as -1, strictly better timing margin. |
|
||
| DDR3 SDRAM | **Micron MT41J128M16JT-125:K** | 2Gb, x16, DDR3-1600-rated (run at 310.078MHz here due to -2 timing closure, see §3). Verified in-stock on LCSC. |
|
||
| Config flash | **Winbond W25Q32JVSSIQ** | 32Mbit/4MB, SOIC-8. Comfortably fits the ~30.5Mbit full XC7A100T bitstream. Verified in-stock on LCSC. Wired **exclusively** to the FPGA (see §5). |
|
||
|
||
## 2. FPGA pin assignments (real, from the routed design)
|
||
|
||
### 2.1 DDR3 (fixed by the FPGA's own PHY hardware — not a free choice)
|
||
|
||
Generated by the Vivado MIG wizard (`mig_7series_0.xdc`), all `SSTL15` /
|
||
`DIFF_SSTL15` (1.5V), banks 34/35:
|
||
|
||
| Signal | Pin | Signal | Pin | Signal | Pin |
|
||
|---|---|---|---|---|---|
|
||
| ddr3_dq[0] | G4 | ddr3_dq[8] | M1 | ddr3_addr[0] | B1 |
|
||
| ddr3_dq[1] | G3 | ddr3_dq[9] | K3 | ddr3_addr[1] | A3 |
|
||
| ddr3_dq[2] | J3 | ddr3_dq[10] | L3 | ddr3_addr[2] | A4 |
|
||
| ddr3_dq[3] | J2 | ddr3_dq[11] | M3 | ddr3_addr[3] | B4 |
|
||
| ddr3_dq[4] | K2 | ddr3_dq[12] | M2 | ddr3_addr[4] | C4 |
|
||
| ddr3_dq[5] | K1 | ddr3_dq[13] | K5 | ddr3_addr[5] | E7 |
|
||
| ddr3_dq[6] | H6 | ddr3_dq[14] | L4 | ddr3_addr[6] | E5 |
|
||
| ddr3_dq[7] | H5 | ddr3_dq[15] | L6 | ddr3_addr[7] | E6 |
|
||
| ddr3_addr[8] | C7 | ddr3_addr[9] | D8 | ddr3_addr[10] | B6 |
|
||
| ddr3_addr[11] | B7 | ddr3_addr[12] | C5 | ddr3_addr[13] | C6 |
|
||
| ddr3_ba[0] | B2 | ddr3_ba[1] | B3 | ddr3_ba[2] | A1 |
|
||
| ddr3_ras_n | D5 | ddr3_cas_n | D4 | ddr3_we_n | E3 |
|
||
| ddr3_reset_n | F6 (LVCMOS15) | ddr3_cke[0] | D7 | ddr3_odt[0] | H2 |
|
||
| ddr3_cs_n[0] | D3 | ddr3_dm[0] | G6 | ddr3_dm[1] | L1 |
|
||
| ddr3_dqs_p[0] | J4 (DIFF) | ddr3_dqs_n[0] | H4 (DIFF) | | |
|
||
| ddr3_dqs_p[1] | N2 (DIFF) | ddr3_dqs_n[1] | N1 (DIFF) | | |
|
||
| ddr3_ck_p[0] | A6 (DIFF) | ddr3_ck_n[0] | A5 (DIFF) | | |
|
||
| sys_clk_i | E2 (SSTL15, bank 35) | clk_ref_i | C9 (LVCMOS25, bank 16) | | |
|
||
|
||
**Bank voltage requirements**: bank 34/35 → **1.5V** (DDR3 SSTL15), bank 16 →
|
||
**2.5V** (clk_ref_i, LVCMOS25).
|
||
|
||
`INTERNAL_VREF` for banks 34/35 is set to 0.750V by the MIG constraints
|
||
(required for SSTL15 single-ended inputs) — this is a Vivado-side setting,
|
||
not a board component, but note it if you ever inspect bitstream generation
|
||
warnings about VREF.
|
||
|
||
### 2.2 Neural-processor management SPI (ESP32 ↔ FPGA, FPGA is **slave**)
|
||
|
||
Bank 15, package edge column A/B, physically adjacent (short traces), `LVCMOS33`:
|
||
|
||
| Signal | Pin | Direction (FPGA side) |
|
||
|---|---|---|
|
||
| sclk | A15 | input |
|
||
| mosi | B16 | input |
|
||
| miso | B17 | output |
|
||
| cs_n | A16 | input |
|
||
|
||
**Bank 15 VCCO**: assumed **3.3V** — change the XDC's IOSTANDARD if your
|
||
board power plan uses a different rail for this bank.
|
||
|
||
### 2.3 Config flash SPI (FPGA ↔ flash, FPGA is **master**)
|
||
|
||
These are the FPGA's own dedicated Master-SPI configuration pins, **reclaimed
|
||
as ordinary fabric I/O after configuration completes** (requires
|
||
`BITSTREAM.CONFIG.PERSIST = FALSE`, the Vivado default — already set
|
||
explicitly in the project XDC). Bank 14, `LVCMOS33`:
|
||
|
||
| Signal | Pin | Direction (FPGA side) | Notes |
|
||
|---|---|---|---|
|
||
| flash_mosi | K17 | output | = D00_MOSI (config pin, reclaimed) |
|
||
| flash_miso | K18 | input | = D01_DIN (config pin, reclaimed) |
|
||
| flash_cs_n | L13 | output | = FCS_B (config pin, reclaimed) |
|
||
| (CCLK) | E9 | output | **Not a top-level port** — driven internally via the `STARTUPE2` primitive. Wire the flash's own CLK pin to package pin **E9**. |
|
||
|
||
**Bank 14 VCCO**: assumed **3.3V** (matches the flash's own VCC, typically
|
||
1.8–3.6V for the W25Q32JV — check its datasheet's exact operating range
|
||
against whatever VCCO you choose for bank 14).
|
||
|
||
**Reserved, do not use** (bank 14, same reasons as above but unused by this
|
||
design — kept clear for any future Quad-SPI/BPI expansion): `L16` (EMCCLK),
|
||
`R16` (RDWR_B), `V15` (CSI_B). The project's own XDC `PROHIBIT`s these so
|
||
Vivado's auto-placement never claims them by accident.
|
||
|
||
### 2.4 FPGA configuration control (dedicated, bank 0, not negotiable)
|
||
|
||
| Signal | Pin | Purpose |
|
||
|---|---|---|
|
||
| PROGRAM_B | P9 | pulse low to force a full reconfiguration from flash |
|
||
| INIT_B | P7 | goes low during config; can indicate a config error if it re-asserts |
|
||
| DONE | P10 | goes high once configuration succeeds — wire to a status LED if desired |
|
||
| M0 | P12 | mode select |
|
||
| M1 | P13 | mode select |
|
||
| M2 | P11 | mode select |
|
||
| CFGBVS | P8 | tie to match bank 0's VCCO logic level (see UG470) |
|
||
|
||
**Mode pin setting for Master SPI boot** (the flash-based autonomous boot
|
||
path, see §5): `M[2:0] = 001` (per UG470's mode pin table) — tie via pull-up/
|
||
pull-down resistors on the board, not driven dynamically.
|
||
|
||
### 2.5 JTAG (always available, independent of flash content)
|
||
|
||
| Signal | Pin |
|
||
|---|---|
|
||
| TCK | E10 |
|
||
| TDI | E11 |
|
||
| TMS | E12 |
|
||
| TDO | E13 |
|
||
|
||
Used for: (a) first-ever/factory programming when the flash is blank (see
|
||
§5), (b) recovery, (c) development/debug. This project's own plan drives
|
||
these from an ESP32 doing real JTAG bit-banging (TAP state machine, IR/DR
|
||
shifting) rather than a bench programmer — that firmware is separate,
|
||
software-side work, not covered here.
|
||
|
||
## 3. Real timing signoff (EXP-0083, the current, trustworthy number)
|
||
|
||
Real in-context Vivado place-and-route (not out-of-context, not estimated):
|
||
|
||
| Metric | Value |
|
||
|---|---|
|
||
| DDR3 PHY clock (sys_clk_i) | **310.078 MHz** (3.225ns period) |
|
||
| Compute domain clock (ui_clk, PLL-derived 2:1 from sys_clk_i) | **155.039 MHz** |
|
||
| WNS (setup slack) | **+0.073 ns** — real, still thin but improved vs every prior real run. Re-verify with a fresh P&R after ANY further logic addition. |
|
||
| WHS (hold slack) | +0.036 ns |
|
||
| Failing endpoints | 0 (setup), 0 (hold) |
|
||
| LUTs used | 5644 / 63400 (8.90%) |
|
||
| DSP48E1 used | 16 / 240 (6.67%) — 8 per compute core × 2 cores, unchanged since EXP-0059 |
|
||
| Block RAM used | 0 |
|
||
| STARTUPE2 used | 1 / 1 (100%) — the config-flash bridge |
|
||
|
||
Signoff history (every real change, same target, in-context P&R):
|
||
|
||
| EXP | WNS (ns) | LUTs | DSP48E1 |
|
||
|---|---|---|---|
|
||
| 0074 (first real DDR3 + pins) | +0.040 | 5140 | 16 |
|
||
| 0076 (+regfile, +pins, +SPI fix) | +0.056 | 5173 | 16 |
|
||
| 0078 (+flash bridge, STARTUPE2) | +0.013 | 5213 | 16 |
|
||
| 0079 (+real activation engine) | +0.030 | 5379 | 16 |
|
||
| 0082 (+denser activation packing) | +0.068 | 5437 | 16 |
|
||
| 0083 (+DDRManager phase 1, `ddr_prefetch_mgr.v`) | **+0.073** | 5644 | 16 |
|
||
|
||
## 4. Real DDR3 memory layout convention
|
||
|
||
Both weight data and activation data share the same DDR3 address space
|
||
(word-addressed, 16-bit words, `BURST_LEN=8` per transaction = 128 bits/burst).
|
||
|
||
- **Weights**: one layer's weight set starts at word address `layer_index *
|
||
WORDS_PER_LAYER` (`WORDS_PER_LAYER = LAYER_BYTES/2`). Densely packed —
|
||
`layer_prefetch_ctrl.v` reads full bursts sequentially into the on-chip
|
||
weight buffer once per job.
|
||
- **Activations** (real engine, `act_tile_fetch.v`; current layout is the
|
||
**v2 convention, EXP-0081/0082**): **two consecutive tiles (P_IN=8 INT8
|
||
values each) share ONE full `BURST_LEN=8`-word (128-bit) burst** — the
|
||
even-indexed tile occupies the low 64 bits, the odd-indexed tile occupies
|
||
the high 64 bits. Tile `t`'s burst address is `base + (t>>1)*BURST_LEN`
|
||
(integer division — two tiles per burst), always burst-aligned by
|
||
construction. This **halves** real DDR3 bytes-moved-per-useful-byte versus
|
||
the original EXP-0079 "1 tile = 1 burst" layout, which wasted the upper 64
|
||
bits of every burst as padding.
|
||
- **Why this is timing-safe despite selecting a sub-burst half at read
|
||
time**: the tile index's LSB (which half of the burst a given tile
|
||
lives in) is known at *request* time, not at response time. It's
|
||
latched into a register (`sel_lat`) the same cycle the request is
|
||
accepted — many `ui_clk` cycles before the real DDR3 round-trip
|
||
completes and `ctrl_rdata` becomes valid. The eventual data-select mux
|
||
therefore always selects on an already-long-stable registered bit,
|
||
never a bit racing live read data — this is NOT the runtime-indexed-
|
||
part-select-on-the-critical-path pattern flagged as a real Fmax risk in
|
||
EXP-0061 (that pattern is about a select signal arriving *late*,
|
||
simultaneously with the data it gates). Confirmed timing-safe by real
|
||
P&R (EXP-0082): margin *improved* from +0.030ns to +0.068ns, not
|
||
degraded.
|
||
- Real measured effect: back-to-back same-row DDR3 throughput is a fixed
|
||
1.24 GB/s (measured, EXP-0080) regardless of packing — this convention
|
||
doesn't change that ceiling, it changes how much of it is *wasted* on
|
||
padding, doubling the real achievable useful fraction (see
|
||
`docs/ARCHITECTURE_ANALYSIS.md` §3 and §5.1).
|
||
- `base` (a job's own `x_base_a`/`x_base_b`) is chosen freely by whoever
|
||
submits jobs (the SPI host) — just keep each position's own activation
|
||
array in its own non-overlapping `(N_TILES/2) * BURST_LEN`-word region.
|
||
|
||
## 5. FPGA configuration (boot) procedure
|
||
|
||
Two complementary paths, both present on this board by design:
|
||
|
||
1. **Factory-first / recovery (JTAG, ESP32-driven)**: the flash starts
|
||
blank on a fresh board — no other path can bootstrap it (a real chicken-
|
||
and-egg constraint: the FPGA can't relay flash-programming commands over
|
||
SPI, §5.2, until it's already running logic that does that). The ESP32
|
||
bit-bangs JTAG (§2.5) to load a bitstream directly, or to run Vivado's own
|
||
"indirect SPI flash programming" sequence to write the flash for the
|
||
first time. One-time (or rare/recovery-only) step.
|
||
2. **Normal boot (Master SPI, autonomous)**: every subsequent power-on, the
|
||
FPGA self-configures from the flash via its own dedicated hardware (mode
|
||
pins set to Master SPI, §2.4) — no ESP32 involvement needed.
|
||
3. **Field firmware updates (SPI-through-FPGA, `FLASH_XFER` opcode 0x40)**:
|
||
once the FPGA is running, the ESP32 can rewrite the flash by relaying raw
|
||
SPI-NOR bytes through the FPGA over the management SPI bus (§2.2) — the
|
||
FPGA then re-transmits them as master on the flash bus (§2.3). This is the
|
||
**only** electrical path from ESP32 to the flash; there is no direct
|
||
connection (by design, per explicit requirement).
|
||
- **Real SPI-NOR opcodes** (verified against the actual W25Q32JV
|
||
datasheet, for whoever writes the ESP32-side flashing routine):
|
||
`0x06` Write Enable, `0x04` Write Disable, `0x05` Read Status Register-1
|
||
(bit0=BUSY, bit1=WEL), `0x02` Page Program, `0x03` Read Data, `0x20`
|
||
Sector Erase (4KB), `0x52` 32KB Block Erase, `0xD8` 64KB Block Erase,
|
||
`0xC7`/`0x60` Chip Erase.
|
||
- **Protocol timing note**: `FLASH_XFER` relays are NOT instantaneous —
|
||
each relayed byte's real flash response is only stable starting **two**
|
||
host-clocked bytes later (not one), so the host must clock 2 trailing
|
||
dummy bytes after its last real command byte to safely receive the
|
||
final response. See `spi_host_bridge_v3.v`'s own header for the full
|
||
real-measured reasoning (EXP-0077).
|
||
- After writing a new bitstream to the flash, reconfigure either by
|
||
pulsing `PROGRAM_B` externally, or (future work, not built yet) via a
|
||
`ICAPE2`-based warm self-reconfiguration triggered over the same SPI bus.
|
||
|
||
## 6. Management SPI protocol summary (for ESP32 firmware)
|
||
|
||
One opcode byte (MSB-first) per CS-low transaction, driven by
|
||
`spi_host_bridge_v3.v`:
|
||
|
||
| Opcode | Name | Payload | Purpose |
|
||
|---|---|---|---|
|
||
| 0x00 | NOP | 0 bytes | inert |
|
||
| 0x0F | RESET | 0 bytes | pulses a soft-reset |
|
||
| 0x10 | WRITE_JOB | 16 bytes | submit one inference job (node_id, x_base, w_base, n_tiles, result_addr) |
|
||
| 0x20 | STATUS | 0 bytes → 1 byte out | job_busy / mem_busy / last_job_accepted bits |
|
||
| 0x01 | WRITE_MEM | 4+2N bytes | raw DDR3 word write (N words) |
|
||
| 0x02 | READ_MEM | 6 bytes → 2N bytes out | raw DDR3 word read (N words) |
|
||
| 0x30 | REG_WRITE | 5 bytes | write a control register |
|
||
| 0x31 | REG_READ | 1 byte → 4 bytes out | read a status/ID register (0x00 DEVICE_ID, 0x01 CONTROL, 0x02 STATUS incl. DDR3-ready + Director-error, 0x03 N_SLOTS) |
|
||
| 0x40 | FLASH_XFER | N bytes → N bytes out (+2 margin) | raw passthrough to the config flash, see §5.3 |
|
||
|
||
Full byte-level field layouts are documented in `spi_host_bridge_v3.v`'s own
|
||
header comment — treat that file as the authoritative protocol spec, this
|
||
table is a summary/index.
|
||
|
||
## 7. Known-open items (honestly disclosed, not hidden)
|
||
|
||
- Scaling past N=2 compute cores (silicon budget allows up to ~30 per the
|
||
DSP48E1 count) is not yet built or timing-verified. A result-writeback
|
||
engine (currently `result_data_a/b` are literal top-level pins, fine at
|
||
N=2 but a real pin-budget blocker beyond that) must land first — see
|
||
`docs/ARCHITECTURE_ANALYSIS.md` §5.3.
|
||
- The reset pin and other very-low-pin-count signals have no fixed PCB
|
||
location yet — assign once the rest of the board layout (reset circuit,
|
||
status LEDs, etc.) is decided.
|
||
- The §3 timing margin (+0.073ns) is real but thin — do not add logic
|
||
without a fresh real P&R to confirm it still closes.
|
||
- ESP32-side JTAG bit-banging firmware (§5.1) does not exist yet — it's
|
||
software work on the host side, not part of this FPGA RTL.
|
||
- A 32-bit-wide DDR3 channel (decided, see `docs/ARCHITECTURE_ANALYSIS.md`
|
||
§5.4) requires the user's own real interactive Vivado MIG "Customize IP"
|
||
wizard session (Data Width + Input Clock Period changed together) —
|
||
not yet done; every timing/bandwidth number in this document still
|
||
reflects the current 16-bit channel.
|
||
- `ddr_prefetch_mgr.v` (DDRManager phase 1, EXP-0083) is real, built, and
|
||
timing-verified, but its measured benefit (2.86% real reduction in total
|
||
simulated time) is modest — see `docs/ARCHITECTURE_ANALYSIS.md` §5.2 for
|
||
the honest number and why. The larger multi-slot DDRManager is not built.
|