docs: Phase 7 placement-seed sweep + hardware design document
Phase 7 (docs/FPGA-NeuralNetwork-Engine.md): re-ran nextpnr-ecp5 on the already-synthesized Phase 5 spi_neuron_top netlists (top.json reused, only placement re-seeded) at --seed 1/2/3 for both P8 and P2. Both land in a tight band regardless of seed (P8: 39.5-40.6 MHz, 2.6% spread; P2: 42.5-45.0 MHz, 5.8% spread) -- confirms the Phase 5 timing shortfall is a real structural bottleneck, not placement noise, unlike the much smaller same-tier benchmark design (<2% utilization, huge placer freedom, genuinely noisy). Corrected the earlier "pipeline the saturate stage" candidate fix, which targeted Phase 4's critical path and not the one Phase 5's logic actually shifted to; block RAM for x_mem/w_mem remains the leading candidate, not yet implemented. New docs/FPGA-Neural-Hardware-Design.md: draft hardware design doc for a board carrying the project's actual target device (LFE5U-45F-8BG381C) plus the parallel PSRAM rtl/psram_controller.v is written for. Covers: why not the basic-ecp5-pcb reference board (wrong package/speed grade, no RAM), a real I/O pin budget from Lattice's own CABGA381 pinout table, a researched PSRAM part (ISSI IS66WVE4M16EBLL-70BLI -- 70ns access matches the controller's timing assumption exactly, with a note on the byte/word address shift in int8_memory_access.v so the chip's top address line is correctly left as spare headroom, not a wiring error), clock (16 MHz, no PLL exists yet so CLK_FREQ_MHZ must match whatever oscillator is fitted), power/config reusing the reference board's proven circuitry and errata (config-SPI pin can't double as the application SPI interface), and a BOM/open-items list. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQV3vS9TXaGDJ5cRfnfidt
This commit is contained in:
@@ -0,0 +1,256 @@
|
|||||||
|
# FPGA-Neural — Hardware Design Document
|
||||||
|
|
||||||
|
Status: draft, pre-schematic. Component choices below are researched
|
||||||
|
against current distributor listings (2026-09-02) but not yet
|
||||||
|
ordered/prototyped. No PCB layout exists yet.
|
||||||
|
|
||||||
|
Goal: a board carrying the project's actual target device
|
||||||
|
(`LFE5U-45F-8BG381C`) plus the parallel PSRAM the current RTL
|
||||||
|
(`rtl/psram_controller.v`) is written for, so real hardware exists
|
||||||
|
to run everything already synthesized/benchmarked in this repo.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Why a new board (not `basic-ecp5-pcb`)
|
||||||
|
|
||||||
|
A reference ECP5 dev board (Matt Venn's `basic-ecp5-pcb`,
|
||||||
|
OSHWA-approved, in this workspace at `../basic-ecp5-pcb`) exists and
|
||||||
|
is a useful source of **proven power/config circuitry** — but it
|
||||||
|
carries the wrong chip for this project and has no RAM at all:
|
||||||
|
|
||||||
|
| | `basic-ecp5-pcb` | This project's target |
|
||||||
|
|---|---|---|
|
||||||
|
| Device | `LFE5U-45F-6BG256C` | `LFE5U-45F-8BG381C` |
|
||||||
|
| Package | 256-ball CABGA | 381-ball CABGA |
|
||||||
|
| Speed grade | -6 (slowest ECP5 grade) | -8 (fastest ECP5 grade) |
|
||||||
|
| RAM | none (6 PMODs, no memory chip) | parallel PSRAM required |
|
||||||
|
|
||||||
|
Same die (`LFE5U-45F`, same 44K LUT / 72 DSP), different package and
|
||||||
|
a materially slower speed grade. All Fmax numbers measured so far in
|
||||||
|
this repo (`docs/FPGA-NeuralNetwork-Engine.md` §15 "Phase 7 — Optimization") target
|
||||||
|
the -8 grade; they do not directly transfer to a -6 part.
|
||||||
|
|
||||||
|
**What we reuse from it anyway:** the power tree and bitstream-config
|
||||||
|
approach (§4, §5) are package-independent and already validated on
|
||||||
|
real, shipped hardware — no reason to redesign those from scratch.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. I/O pin budget (real ball data, CABGA381)
|
||||||
|
|
||||||
|
Extracted from Lattice's own ECP5-45 pinout table (`../basic-ecp5-pcb/docs/ECP5Upinouts.ods`,
|
||||||
|
sheet `ECP5U45Pinout`, `CABGA381` column) — not estimated:
|
||||||
|
|
||||||
|
| Bank | Usable I/O balls |
|
||||||
|
|---|---|
|
||||||
|
| 0 | 29 |
|
||||||
|
| 1 | 35 |
|
||||||
|
| 2 | 35 |
|
||||||
|
| 3 | 36 |
|
||||||
|
| 6 | 36 |
|
||||||
|
| 7 | 35 |
|
||||||
|
| 8 | 22 |
|
||||||
|
| 40 (config-related) | 4 |
|
||||||
|
| **Total usable** | **~232** |
|
||||||
|
| Power/ground/NC (remaining of 381 balls) | 149 |
|
||||||
|
|
||||||
|
Signal budget this design actually needs:
|
||||||
|
|
||||||
|
| Function | Pins |
|
||||||
|
|---|---|
|
||||||
|
| PSRAM (`psram_a` 22b worst case, `psram_dq` 16b, `ce_n/oe_n/we_n/lb_n/ub_n/zz_n` 6b) | up to 44 (real usage likely less — see §3, address lines can be trimmed to match actual chip density) |
|
||||||
|
| Application SPI (`sclk/mosi/miso/cs_n`) | 4 |
|
||||||
|
| `clk`, `rst` | 2 |
|
||||||
|
| Config SPI (to onboard FLASH) | 4 |
|
||||||
|
| JTAG (recommended, for bring-up/debug) | 4 |
|
||||||
|
| **Total** | **~58** |
|
||||||
|
|
||||||
|
~58 of ~232 usable I/O used — **plenty of headroom** (~170+ spare
|
||||||
|
pins) for LEDs, buttons, a debug PMOD-style header, or a second SPI
|
||||||
|
host, without any pin-count pressure. This board does not need to be
|
||||||
|
pin-constrained the way a 256-ball/PMOD-only design would.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. PSRAM subsystem (the piece `basic-ecp5-pcb` doesn't have)
|
||||||
|
|
||||||
|
`rtl/psram_controller.v` implements a plain **asynchronous parallel**
|
||||||
|
interface — address bus, 16-bit data bus, `ce_n`/`oe_n`/`we_n` and
|
||||||
|
byte-lane `lb_n`/`ub_n`, plus `zz_n` — and its timing already
|
||||||
|
hardcodes a **70 ns access latency** assumption
|
||||||
|
(`ACCESS_CYCLES = ceil(70ns × CLK_FREQ_MHZ / 1000)`). This is a
|
||||||
|
classic async-SRAM-style bus, not QSPI — most "PSRAM" sold today
|
||||||
|
(including what's on typical ESP32 boards) is serial/QSPI and **will
|
||||||
|
not** plug into this controller without a rewrite.
|
||||||
|
|
||||||
|
**Recommended part: ISSI IS66WVE4M16EBLL-70BLI**
|
||||||
|
- 64 Mbit (4M × 16), parallel pseudo-SRAM, async, **70 ns
|
||||||
|
access** — matches the controller's timing assumption exactly, no
|
||||||
|
RTL change needed.
|
||||||
|
- TSOP-44/48 package — hand-solderable-adjacent, real distributor
|
||||||
|
listings (DigiKey, Mouser) at time of writing.
|
||||||
|
- **Address bus note:** the chip is 4M×16 words (8 MB total,
|
||||||
|
needs a real 22-bit word address, A0–A21). The current RTL's
|
||||||
|
`ADDR_WIDTH=22` is a **byte** address (4 MiB space) that
|
||||||
|
`int8_memory_access.v` right-shifts by 1 (`addr >> 1`) into a word
|
||||||
|
address before it reaches `psram_controller` — so only **21**
|
||||||
|
word-address bits are actually driven today. Wire all 22 chip
|
||||||
|
address balls, but the chip's topmost line (A21) stays unused/tied
|
||||||
|
low until `ADDR_WIDTH` is widened to 23 to use the chip's full
|
||||||
|
8 MB instead of today's 4 MiB. Free headroom, not a defect.
|
||||||
|
|
||||||
|
**Fallback: ISSI IS61WV6416DBLL / IS61WV102416BLL** (true async
|
||||||
|
SRAM, not pseudo-SRAM) — electrically drop-in on the same
|
||||||
|
`ce_n/oe_n/we_n/lb_n/ub_n` signals, no internal refresh (so `zz_n`
|
||||||
|
can just be tied inactive), faster than needed (~10 ns), useful
|
||||||
|
if the ISSI PSRAM specifically is out of stock. Smaller density
|
||||||
|
(1–16 Mbit depending on exact part) — fine for this
|
||||||
|
project's current memory footprint (weights/biases/activations for
|
||||||
|
the networks exercised so far are well under 1 MB).
|
||||||
|
|
||||||
|
Real part numbers, not yet ordered — verify current stock/pricing
|
||||||
|
before BOM lock.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Clock
|
||||||
|
|
||||||
|
`basic-ecp5-pcb` uses a fixed **16 MHz** MEMS oscillator
|
||||||
|
(SiTime SiT2001B family) — no crystal driver on the ECP5, the clock
|
||||||
|
input must come from an oscillator IC into a `PCLK` pad.
|
||||||
|
|
||||||
|
**Recommendation: keep 16 MHz**, same SiT2001B family (or
|
||||||
|
SiT1602/SiT8008, same vendor, also in stock). Rationale, not just
|
||||||
|
"reuse what worked":
|
||||||
|
|
||||||
|
- No PLL exists anywhere in this project's RTL yet — `CLK_FREQ_MHZ`
|
||||||
|
is a **timing parameter**, not a clock generator. Whatever
|
||||||
|
oscillator is fitted drives `clk` directly.
|
||||||
|
- Every Fmax measured so far for the *full* integrated system
|
||||||
|
(`spi_neuron_top`, Phase 5) sits at 39.5–45 MHz across a
|
||||||
|
seed sweep (`docs/FPGA-NeuralNetwork-Engine.md` §15 "Phase 7 — Optimization") —
|
||||||
|
confirmed structural, not placement luck. 16 MHz sits well
|
||||||
|
under that with real margin.
|
||||||
|
- **`CLK_FREQ_MHZ` must be set to match whatever oscillator is
|
||||||
|
actually fitted** (16, if this recommendation is taken) — it feeds
|
||||||
|
the PSRAM access-timing formulas directly (§3); using the RTL's
|
||||||
|
default of 80 with a 16 MHz real clock would under-time the
|
||||||
|
PSRAM by 5×.
|
||||||
|
|
||||||
|
A higher oscillator (e.g. 25 or 32 MHz) is possible with margin
|
||||||
|
to spare, but revisit once the Phase 7 timing-closure work
|
||||||
|
(`docs/FPGA-NeuralNetwork-Engine.md`) lands rather than guessing a
|
||||||
|
number now.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Power
|
||||||
|
|
||||||
|
Reuse `basic-ecp5-pcb`'s proven three-rail tree as-is (same device
|
||||||
|
family, same rail requirements regardless of package):
|
||||||
|
|
||||||
|
| Rail | Value | Part | Load | Status |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Core | 1.1 V | TLV62568 (buck) | ≥600 mA | Confirmed in production, DigiKey/Mouser listed |
|
||||||
|
| I/O | 3.3 V | TLV62568 (buck) | 1 A (all banks + PSRAM + PMODs share this) | Confirmed in production |
|
||||||
|
| Auxiliary | 2.5 V | TLV73325 (LDO) | 10 mA | Confirmed in production |
|
||||||
|
|
||||||
|
Decoupling: one cap per I/O bank minimum, per Lattice's ECP5
|
||||||
|
Hardware Checklist (referenced by `basic-ecp5-pcb`, not re-derived
|
||||||
|
here).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Configuration (bitstream load)
|
||||||
|
|
||||||
|
Reuse `basic-ecp5-pcb`'s SPI-FLASH-boot approach:
|
||||||
|
|
||||||
|
- **W25Q128JV** SPI NOR flash (16 MB) — confirmed in production,
|
||||||
|
multiple package options (WSON, SOIC) currently listed.
|
||||||
|
- ECP5 reads its bitstream from this flash at power-on (`sysCONFIG`
|
||||||
|
SPI master mode); no external programmer needed for normal
|
||||||
|
power-up, only for the initial flash write.
|
||||||
|
|
||||||
|
**Lessons reused from `basic-ecp5-pcb`'s errata (do not re-discover
|
||||||
|
these the hard way):**
|
||||||
|
|
||||||
|
- Config-mode select pins should tie directly to GND, not through a
|
||||||
|
10 k resistor — the ECP5 test point is ~1 V, too close to
|
||||||
|
the 3.3 V bank's input threshold through a resistor divider.
|
||||||
|
- Not every SPI flash that claims QSPI actually has a usable QE
|
||||||
|
(quad-enable) bit in practice — `basic-ecp5-pcb` hit this with an
|
||||||
|
IS25LP016D and switched to the W25Q12x family instead. Stick with
|
||||||
|
W25Q128JV rather than substituting on price alone.
|
||||||
|
- **The dedicated config-SPI clock pin cannot be reused as a general
|
||||||
|
input** post-configuration without extra board-level workaround
|
||||||
|
(`basic-ecp5-pcb` needed a bodge wire to let a Raspberry Pi talk
|
||||||
|
SPI to the FPGA over the *same* physical pin used for flash boot).
|
||||||
|
**This project's application SPI** (`spi_neuron_top`'s
|
||||||
|
`sclk`/`mosi`/`miso`/`cs_n`, the host-facing protocol in
|
||||||
|
`docs/FPGA-NeuralNetwork-Engine.md` §8.1) **must land on separate,
|
||||||
|
ordinary I/O pins — never the config-SPI pins** — precisely to
|
||||||
|
avoid needing that same workaround.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Signal map (draft — not yet a real LPF)
|
||||||
|
|
||||||
|
No `.lpf` pin constraints exist for this device/package combination
|
||||||
|
yet (all `.lpf` files in `synth/` are currently empty — nextpnr has
|
||||||
|
been auto-placing I/O for every synthesis run so far, fine for
|
||||||
|
Fmax/resource benchmarking, **not** sufficient for a real board).
|
||||||
|
Before schematic capture, someone needs to:
|
||||||
|
|
||||||
|
1. Pick actual CABGA381 ball numbers for each signal below from the
|
||||||
|
pinout table referenced in §2 (bank-aware: keep the PSRAM data/
|
||||||
|
address bus in one or two adjacent banks to ease layout and
|
||||||
|
timing).
|
||||||
|
2. Write a real `.lpf` with those assignments and re-run
|
||||||
|
`nextpnr-ecp5` with it (current benchmark runs deliberately
|
||||||
|
skipped this — see `tools/fpga_benchmark.py`).
|
||||||
|
3. Confirm bank voltage compatibility (all banks are 3.3 V I/O
|
||||||
|
in this design, per §5 — fine for both the PSRAM candidates in §3
|
||||||
|
and standard SPI-level signaling).
|
||||||
|
|
||||||
|
| Signal group | Port(s) | Count | Target bank (TBD) |
|
||||||
|
|---|---|---|---|
|
||||||
|
| PSRAM address | `psram_a[21:0]` | 22 | one bank |
|
||||||
|
| PSRAM data | `psram_dq[15:0]` | 16 | same or adjacent bank |
|
||||||
|
| PSRAM control | `psram_ce_n/oe_n/we_n/lb_n/ub_n/zz_n` | 6 | same bank as above |
|
||||||
|
| Application SPI | `sclk/mosi/miso/cs_n` | 4 | any bank, NOT the config-SPI bank (§6) |
|
||||||
|
| Clock/reset | `clk`, `rst` | 2 | `clk` must land on a `PCLK`-capable pad |
|
||||||
|
| Config SPI | to onboard flash | 4 | dedicated config bank (bank "40" balls, §2) |
|
||||||
|
| JTAG (debug) | TCK/TMS/TDI/TDO | 4 | dedicated JTAG balls |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Bill of materials (draft)
|
||||||
|
|
||||||
|
| Ref | Part | Function | Availability |
|
||||||
|
|---|---|---|---|
|
||||||
|
| U1 | LFE5U-45F-8BG381C | FPGA | Already the project's confirmed target (see main docs, price/stock table) |
|
||||||
|
| U2 | ISSI IS66WVE4M16EBLL-70BLI | Parallel PSRAM, 64Mb, 70ns | Verified listed, DigiKey/Mouser |
|
||||||
|
| U3, U4 | TLV62568 | Buck converter, core + IO rails | Confirmed in production |
|
||||||
|
| U5 | TLV73325 | LDO, 2.5V aux rail | Confirmed in production |
|
||||||
|
| U6 | W25Q128JV | SPI NOR flash, config | Confirmed in production, multiple packages |
|
||||||
|
| Y1 | SiT2001B, 16 MHz | System clock oscillator | Confirmed in production |
|
||||||
|
|
||||||
|
Not yet specified: exact package/footprint per part, decoupling cap
|
||||||
|
values, JTAG header, PSRAM address-bus trim if a smaller/cheaper
|
||||||
|
density than 4M×16 turns out to be sufficient once real network
|
||||||
|
sizes are decided.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Open items before schematic capture
|
||||||
|
|
||||||
|
- [ ] Decide real PSRAM density needed (drives whether `ADDR_WIDTH`
|
||||||
|
stays 22 or can shrink, and whether the fallback true-SRAM
|
||||||
|
part in §3 is sufficient instead of the pseudo-SRAM)
|
||||||
|
- [ ] Real `.lpf` pin assignment (§7) and a synthesis run against it
|
||||||
|
(current benchmark results all use auto-placed I/O)
|
||||||
|
- [ ] Confirm PSRAM/SPI signal integrity at whatever clock is
|
||||||
|
actually fitted (§4) — no signal integrity analysis done yet
|
||||||
|
- [ ] JTAG header footprint choice
|
||||||
|
- [ ] KiCad (or other) schematic capture — none exists yet for this
|
||||||
|
device/package combination
|
||||||
@@ -919,24 +919,51 @@ neither Phase 4's nor Phase 5's single P2/P8 runs establish that.
|
|||||||
**Candidate directions (not implemented, need a decision before
|
**Candidate directions (not implemented, need a decision before
|
||||||
touching the validated core further):**
|
touching the validated core further):**
|
||||||
|
|
||||||
|
**Placement-seed sweep (2026-09-02) — resolved: it's structural, not
|
||||||
|
noise.** Re-ran `nextpnr-ecp5` on the already-synthesized Phase 5
|
||||||
|
netlists (`top.json` reused, only placement re-seeded — no re-synth)
|
||||||
|
at `--seed 1/2/3` for both P8 and P2:
|
||||||
|
|
||||||
|
| Build | seed (default) | seed 1 | seed 2 | seed 3 | spread |
|
||||||
|
|---|---|---|---|---|---|
|
||||||
|
| P8 | 40.57 MHz | 39.70 MHz | 39.54 MHz | 40.00 MHz | 1.03 MHz (2.6%) |
|
||||||
|
| P2 | 42.54 MHz | 42.82 MHz | 42.72 MHz | 45.01 MHz | 2.47 MHz (5.8%) |
|
||||||
|
|
||||||
|
Both builds land in a tight band regardless of seed — nothing close
|
||||||
|
to the P2/P4/P8 same-tier benchmark's own non-monotonic swings (that
|
||||||
|
benchmark's design uses <2% of the device, so its placer has enormous
|
||||||
|
freedom; `spi_neuron_top`'s fuller design does not). **This confirms
|
||||||
|
a real structural bottleneck, not a lucky/unlucky placement** — a
|
||||||
|
seed sweep or floorplan constraint will not close a ~2x gap to
|
||||||
|
80 MHz by itself; an RTL change is needed.
|
||||||
|
|
||||||
|
**Candidate directions (not implemented, need a decision before
|
||||||
|
touching the validated core further):**
|
||||||
|
|
||||||
- Move `x_mem`/`w_mem` onto `DP16KD` block RAM (0% used) instead of
|
- Move `x_mem`/`w_mem` onto `DP16KD` block RAM (0% used) instead of
|
||||||
LUT fabric — likely the highest-leverage single change, since it
|
LUT fabric — the leading candidate. The critical path runs through
|
||||||
is directly implicated in the Phase 5 critical path and currently
|
what looks like `w_mem`'s *write*-address decode (`neuron_memory`
|
||||||
wastes a free, purpose-built resource.
|
loads one byte per cycle into a variable index during
|
||||||
- Add a pipeline register between the final MAC accumulation and the
|
`STATE_READ_X`/`STATE_READ_W`, which needs a full-width mux/demux
|
||||||
activation/saturate stage in `neuron_parallel.v`, trading +1 cycle
|
in LUT fabric to steer that byte to the right register) feeding
|
||||||
of per-neuron latency (negligible next to the existing
|
forward into the accumulator's input path — not the arithmetic
|
||||||
N_INPUTS/PARALLEL group-count cycles) for a shorter combinational
|
itself. A block-RAM read/write port replaces that scattered
|
||||||
path. No existing testbench hardcodes an exact cycle count for
|
LUT-based decode with a single hard macro. This is a genuine
|
||||||
neuron completion (`wait(done)`/polling patterns throughout), so
|
redesign, not a drop-in swap: `x_bus`/`w_bus` are currently
|
||||||
this should be low regression-risk — but it is still a change to
|
exposed to `neuron_parallel` as one fully-parallel N_INPUTS-wide
|
||||||
the explicitly "validated, don't touch" datapath (`mac8`/`mac_unit`
|
combinational bus (built by unrolling every `x_mem`/`w_mem`
|
||||||
/`neuron_parallel`), so deserves an explicit go-ahead rather than
|
element every cycle); block RAM has a registered, address-in/
|
||||||
being done opportunistically.
|
data-out-next-cycle read port, so `neuron_parallel`'s group
|
||||||
- A `nextpnr` seed/effort sweep to check how much of the shortfall is
|
processing would need to become RAM-latency-aware instead of
|
||||||
placement variance (see P4's own 75.01/76.41 MHz two-number report
|
assuming the whole bus is already valid. Real potential upside,
|
||||||
in the same run, or P2/P4/P8's non-monotonic Fmax in the
|
real design effort — needs a go-ahead, not something to do
|
||||||
same-tier benchmark) vs. a structural bottleneck.
|
opportunistically.
|
||||||
|
- A pipeline register between MAC accumulation and the
|
||||||
|
activation/saturate stage (this session's original guess) does
|
||||||
|
**not** target the actual Phase 5 bottleneck — that finding was
|
||||||
|
Phase 4's critical path (the saturation comparator), and the path
|
||||||
|
moved once Phase 5's logic was wired in (see above). Keeping this
|
||||||
|
noted for Phase 4-only builds; not a fix for the current problem.
|
||||||
|
|
||||||
## Phase 8 — Optional Hardware Training
|
## Phase 8 — Optional Hardware Training
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user