docs: consolidate all V2 datasheets into one current, complete document

The repository had accumulated multiple, contradictory "current state"
documents for V2 hardware: an old V1 IT/EN datasheet copy nested inside
hardware/v2/docs/datasheet/, a stray untracked duplicate at repo root
(docs/DatasheetLatex/), and a second, much older documentation track
(hardware/v2/docs/*.md: PRE_PCB_VERIFICATION.md, PRE_PCB_CLOSURE_4POINT.md,
MEMORY_UPGRADE_64MB_N8.md, and 10 more) describing an earlier PSRAM/
N_SLOTS<=2 milestone alongside the real, current SDRAM/N_SLOTS=4 board.
The LaTeX datasheet's own front matter (features/pinout cover pages) and
chapter 9 (benchmarks) were themselves still describing that obsolete
architecture, contradicting the real, current chapters 5/7/10 elsewhere
in the same document.

This commit:
- Flattens hardware/v2/docs/datasheet/files/docs/datasheet/v2-en/* up to
  hardware/v2/docs/datasheet/ (was 4 levels of redundant nesting).
- Removes the old V1 IT/EN LaTeX copies and the stray root-level
  duplicate entirely (recoverable from git history, not from disk).
- Preserves the real component reference PDFs (ECP5 eval board, ISSI
  PSRAM, programming cables) under datasheet/references/.
- Removes 13 superseded hardware/v2/docs/*.md status documents after
  folding every real, unique fact they contained into the datasheet:
  SPI max verified clock (12MHz, exact 12.8MHz CDC edge), SDRAM directed
  boundary test (21/21 PASS), 16MHz oscillator MPN (ECS-3225MV-160-BN-TR),
  and the real FPGA<->SDRAM ball mapping cross-check.
- Rewrites the datasheet's own front matter, ch.4 (parameters), ch.8
  (top-level module -- was documenting the wrong, non-physical top
  entirely), and ch.9 (benchmarks) to describe the current, real SDRAM/
  N_SLOTS=4 production board, while keeping the real PSRAM-era chapters
  as clearly-labeled history rather than deleting correctly-measured
  work.
- Fixes a title-page tikzpicture that was clipped off the page edge
  (pre-existing, unrelated to this change) by scaling it to fit.

Net: 85 files changed, -8814/+498 lines. hardware/v2/docs/ now contains
exactly one current datasheet plus FIRST_POWER_ON.md (a bring-up
runbook, not a duplicate spec).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
This commit is contained in:
2026-09-09 00:41:06 +02:00
co-authored by Claude Sonnet 5
parent 1efd63912f
commit 8b8ca239ca
85 changed files with 497 additions and 8813 deletions
-94
View File
@@ -1,94 +0,0 @@
# FPGA-Neural V2 — CHIP READINESS
**SUPERSEDED.** See `PRE_PCB_VERIFICATION.md` for the current,
consolidated PRE-PCB VERIFIED release gate (this document's own
checklist predates the SPI host bridge, PLL, and SDRAM datasheet
audit). Left in place as a historical record.
Precise, non-vague criteria per the governing spec's own definition:
V2 hardware is READY only when EVERY box below is checked. If even one
fundamental item is missing, **HARDWARE READY = NO** — no OPEN ITEM is
masked.
```
[x] RTL frozen -- nms_neural_multiprocessor_sdram_unified.v,
zero V1 dependency, strict lint clean
[x] regression PASS -- 461/461 (isolated controller, 9 configs),
40/40 (isolated unified backend)
[x] bit-exact PASS -- 256/256 neurons, N=2 AND N=4, single SDRAM
[x] SDRAM validation PASS -- init/refresh/read/write/burst/masked-write,
40 real refresh events, zero corruption
[x] N4 synthesis PASS -- real Yosys 0.68+post, zero errors
[ ] N4 timing >= 80 MHz -- MARGINAL: only 1/8 real P&R seeds pass
[ ] constraints complete -- v2_unified.lpf exists, REAL and P&R-verified
for 39/149 signals (clk/rst + full SDRAM bus);
110-signal host bus still unassigned
[ ] pinout complete -- 149-signal inventory complete; SDRAM+clk/rst
(39 signals) REALLY assigned from the official
Lattice CSV and P&R-confirmed; host bus (110
signals) deliberately unassigned (see below)
[ ] clock defined -- real EHXPLLL RTL now exists (STEP20,
ecp5_pll_sys_clk.v, 16MHz->64MHz), NOT yet
confirmed by synthesis/P&R of the board top
[ ] power defined -- rail voltages known; regulators NOT selected
[ ] FPGA configuration defined -- standard pins identified; flash NOT chosen
[x] host interface defined -- STEP20: real SPI protocol engine
(spi_host_bridge.v), verified correct
end-to-end (11/11, board-level smoke test),
zero regression to the STEP19 baseline
[x] schematic requirements complete -- SCHEMATIC_READINESS.md's own block diagram
and interconnection list are complete
[x] first-power-on test defined -- FIRST_POWER_ON.md's own 12-step procedure
[ ] bitstream reproducible -- NOT verified: no ball-assigned LPF exists to
produce a REAL, board-usable bitstream from;
the free-placement bitstreams used for
verification this round are reproducible
AS SIMULATION/FIT PROOFS ONLY, not as a
real board-programmable artifact
```
**8 of 14 items checked. HARDWARE READY = NO.**
**STEP20 update:** a real SPI host interface (`spi_host_bridge.v` +
`fpga_neural_v2_top.v`) now exists AND is verified correct end-to-end
(errors.log's "ERR-0025 Part B — RESOLUTION": a registered- vs
combinational-read SRAM timing bug, found via the board-level smoke
test, root-caused, fixed with zero regression to the STEP19 baseline)
— "host interface defined" is now checked. A real EHXPLLL clock
wrapper also now exists (`ecp5_pll_sys_clk.v`) but has not yet been
through synthesis/P&R of the board-level top, so "clock defined"
remains unchecked for that specific, narrower reason. The STEP19 core
(raw `reg_*` interface) remains bit-exact verified and was reconfirmed
fresh this session via Verilator after an unrelated Icarus Verilog
v13.0 toolchain regression was found and ruled out (ERR-0024).
## Why each unchecked item is unchecked (no vague language)
| Item | Why NOT checked |
|---|---|
| N4 timing ≥80MHz | 8 real P&R seeds measured; only 1 (81.84MHz) clears 80MHz. This is a real MARGINAL result, not a PASS, per the governing spec's own explicit classification rule (some seeds pass, most do not). |
| Constraints complete | `v2_unified.lpf` real and P&R-verified for 39/149 signals (clock frequency + clk/rst + the full 37-signal SDRAM bus, sourced from the real Lattice pinout CSV found at `~/Downloads/` during this step's own pre-commit review). The 110-signal host bus is deliberately left unassigned. |
| Pinout complete | Signal inventory is complete (149, exactly matching real P&R); SDRAM+clk/rst (39 signals, 26%) are now really assigned and P&R-confirmed; the 110-signal host bus is unassigned, not because pin data is missing, but because that bus is not yet a real physical protocol (see below) — assigning it balls now would be premature. |
| Clock defined | A real EHXPLLL wrapper now exists (`ecp5_pll_sys_clk.v`, STEP20, real Project Trellis `ecppll`-generated parameters, 16MHz->64MHz) and is instantiated in the board-level top, but has NOT yet been confirmed by synthesis/P&R of that top — deliberately deferred until ERR-0025 Part B was resolved (decisions.log DEC-0037). |
| Power defined | Rail VOLTAGES are known from real datasheets; regulator SELECTION, CURRENT budget, and decoupling are not — no real power-estimation tool was run, and the previously-referenced board power-tree design is not accessible this session to confirm as a concrete plan. |
| FPGA configuration defined | Standard ECP5 config pins (TDI/TDO/TCK/TMS/PROGRAMN/INITN/DONE/CCLK) are correctly identified as existing and standard, but no configuration-flash part number or SPI-vs-JTAG-only bring-up approach has been chosen for V2 specifically. |
| Bitstream reproducible | Every P&R run this project has performed used free (unconstrained) I/O placement — a real, valid way to prove the design FITS the package, but not a way to produce a bitstream a real board's own fixed wiring could actually use. |
## What this means, precisely
The V2 hardware architecture itself — SDRAM device, controller,
memory subsystem, compute datapath, N4/P8 configuration — is **real,
validated, and correct**: bit-exact simulation, real synthesis, real
place-and-route all confirm this. What remains is **entirely physical-
integration work**: a real host interface, a real ball-level pinout, a
real clock source decision, and real power/configuration component
selection. None of these are memory-architecture, datapath, or
correctness questions anymore — they are the next, concrete, well-
defined engineering tasks, precisely enumerated in OPEN_ITEMS.md.
## Final answer
```
HARDWARE FREEZE: PASS (architectural decision + RTL correctness)
CHIP READY: NO
```
-110
View File
@@ -1,110 +0,0 @@
# FPGA-Neural V2 — CLOCK ARCHITECTURE
**SUPERSEDED.** This document predates the real EHXPLLL PLL
(`ecp5_pll_sys_clk.v`) that resolves the mismatch described below. See
`PRE_PCB_VERIFICATION.md` \S3 for the current, verified clock/reset
status (PASS). Left in place as a historical record of the
architectural decision that led to adding the PLL.
## Status (HISTORICAL): CRITICAL — real, unresolved oscillator/clock-input mismatch
## What the RTL actually assumes
Every module in the frozen hierarchy (`nms_neural_multiprocessor_
sdram_unified.v` down to `sdram_controller.v`) takes a **single** `clk`
input and treats it directly as both the system clock AND the SDRAM
clock (`CLK_FREQ_MHZ=80` is a pure timing-derivation parameter fed
into `sdram_controller.v`'s own `ns_to_cycles()` function — it does
NOT configure a PLL; there is no PLL anywhere in this hierarchy).
Confirmed mechanically: every real synthesis run this project has
performed (STEP16 through this freeze) reports `EHXPLLL: 0/4 0%` in
nextpnr's own device-utilisation output — **zero PLL primitives are
instantiated**, in any variant, ever.
**This means the design requires a real, external 80MHz (or faster)
clock source wired directly to the FPGA's clock input pin.**
## The real gap
This project's own memory notes (established in an earlier session,
before the SDRAM decision) record the confirmed hardware target board
as using a **16MHz** oscillator. 16MHz ≠ 80MHz, and there is no PLL in
the current RTL to bridge that gap. **Two mutually exclusive
resolutions exist, and neither has been chosen:**
1. **Source an oscillator that directly provides ≥80MHz** (a
commodity part — plain crystal oscillators at 80, 100, or higher
MHz are standard, low-risk components) and retire the 16MHz
assumption. Zero RTL change required. Simplest, lowest-risk path.
2. **Keep the 16MHz oscillator and add a real PLL** (ECP5's own
`EHXPLLL` primitive, e.g. 16MHz→80MHz = ×5) to the RTL, with its
own real timing constraints (lock time, jitter, generated-clock
declaration in the constraints file) — genuinely new RTL/constraint
work, not yet done, and not exercised by any of this project's own
real synthesis/timing-closure runs to date (every Fmax number in
STEP16-18 assumes a clean, ideal `clk` input, not a PLL output with
its own jitter/lock-time budget).
**This is an OPEN, real architectural decision, not a detail** — it
determines whether a new oscillator needs sourcing or a PLL needs
designing, and affects the CLOCK_SOURCE→FPGA_CLOCK diagram below,
which cannot be finalized until it is made.
## Clock tree (as far as it CAN be stated today)
```
[UNRESOLVED: either an 80MHz+ oscillator, or a 16MHz oscillator + PLL]
|
v
FPGA clk pin (ball location: BLOCKER, see PINOUT.md)
|
v
single system clock domain, 80 MHz target
|
+--> Neural Multiprocessor / Dependency Manager / Director /
| Memory Manager / Neural Processors (all synchronous,
| single clock domain — confirmed, no clock-domain-crossing
| logic exists anywhere in the frozen hierarchy)
|
+--> SDRAM controller (same clock, no separate SDRAM clock
domain — sdram_controller.v drives the SDRAM chip's own
CLK pin combinationally/directly from the same system
clock; real board layout must still budget for the
SDRAM's own real clock-to-pin round-trip delay, which
was NOT part of this project's own RTL-simulation/P&R
timing closure — flagged as an OPEN ITEM for board bring-
up, see FIRST_POWER_ON.md)
```
## Reset
A single `rst` input, synchronous to `clk` in every module observed
(no asynchronous reset assertion/de-assertion synchronizer chain was
found in this session's own lint pass). **Reset release timing/
synchronization to a real external reset source (power-on reset chip,
button, or host-driven) has not been designed** — this is a normal,
solvable board-level concern (a standard POR/supervisor IC), not
flagged as a blocker, but not yet decided (OPEN ITEM).
## Clock constraints used so far
Every P&R run in STEP16-18 used `nextpnr-ecp5 --freq 80` (a target
frequency for the placer's own timing-driven effort), NOT a real `.lpf`
`FREQUENCY` constraint tied to a real pin — because no `.lpf` exists at
all for any V2 top-level (see PINOUT.md). A real constraints file with
a proper `FREQUENCY PORT "clk" 80 MHZ;` (or the real achieved-vs-
required frequency once the oscillator/PLL decision above is made)
must be written before this can be considered a genuine, board-ready
clock constraint.
## Summary
| Item | Status |
|---|---|
| Single-clock-domain RTL, no CDC logic found | Confirmed by lint, real |
| PLL present in RTL | **No — confirmed absent (0/4 EHXPLLL in every P&R run)** |
| Oscillator frequency vs required system clock | **CRITICAL — 16MHz (prior project memory) vs 80MHz (RTL requirement), unresolved** |
| Oscillator-vs-PLL decision | **OPEN — not made** |
| Real `.lpf` clock constraint | Partial — `hardware/v2/constraints/v2_unified.lpf` now exists with a frequency constraint and clk/rst ball reuse from V1; full ball-level pinout for the remaining 147 signals is still blocked (see PINOUT.md) |
| Reset synchronization to a real external source | OPEN, not yet designed (not a hard blocker) |
| SDRAM clock-to-pin board-level timing budget | OPEN — not part of RTL-level timing closure |
@@ -1,186 +0,0 @@
# FPGA-Neural V2 — Datasheet
**Status: DRAFT / PRE-RELEASE.** This datasheet documents the INTENDED
V2 board architecture as of STEP20. It does **not** certify a finished,
release-ready design — see §11 Limitations and
`hardware/v2/docs/OPEN_ITEMS.md` for the current, real blocker list.
Do not read any statement here as "physically validated" unless it
says so explicitly.
## 1. General
FPGA-Neural V2 is an embedded neural-network accelerator built around a
Lattice ECP5 FPGA and a single external SDRAM. It executes small,
dependency-graph-structured INT8 neural networks (dense layers, DAGs)
using a Neural Multiprocessor of parallel MAC engines, streaming
weight/activation tiles from one external SDRAM chip that also holds
results.
Architecture stack (top to bottom): SPI host interface → job
registration → Dependency Manager / Neural Director → N parallel
Neural Processors → Memory Manager / streaming tile delivery → Unified
SDRAM Backend → one physical SDRAM.
## 2. FPGA
| Item | Value | Basis |
|---|---|---|
| Part | Lattice LFE5U-45F | DESIGN DECISION |
| Package | CABGA381 | DESIGN DECISION |
| Speed grade | -8 | DESIGN DECISION |
| Ordering part number | LFE5U-45F-8BG381C | DESIGN DECISION (standard Lattice ordering suffix for this grade/package; not independently cross-checked against a live distributor listing this session) |
| Logic (post-synthesis, N=4, frozen STEP19 compute core) | TRELLIS_FF=6425, TRELLIS_COMB=6023, MULT18X18D=32, DP16KD=0 | VERIFIED (real Yosys synthesis, STEP19) |
| I/O used (frozen STEP19 top, no physical host bus) | 149/245 TRELLIS_IO | VERIFIED (real nextpnr-ecp5 P&R, STEP19) |
| I/O used (this step's new board-level top, SPI + osc + reset + SDRAM) | not yet synthesized this round | OPEN — see §11 |
Operating assumption: single clock domain, no CDC beyond the SPI
bridge's own double-flop synchronizers and the reset synchronizer
(§5).
## 3. Neural accelerator
| Parameter | Value |
|---|---|
| N_PROCESSORS | 4 (frozen reference; N=2 also validated; N=8 is a future evolution) |
| P_IN (MAC width) | 8 |
| Data representation | INT8 operands |
| Accumulator | INT32, ReLU + INT8 saturate on output |
| MAC architecture | 8-wide parallel MAC, balanced adder tree (`neural_processor.v`, unchanged since before this freeze) |
| Processor parallelism | N independent Neural Processors, one dependency-graph node in flight per processor |
| Supported memory traffic | weights (read-only, 64-bit packed fetch, cached), activations (read, byte-maskable), results (write, byte-maskable) — all through the SAME single SDRAM |
**RTL capability vs. software/API capability:** the RTL executes one
pre-compiled dependency graph (nodes with producer/consumer edges,
fixed tile counts) registered via 108 bits of per-job configuration
(node id, dependency list, activation/weight/result base addresses,
tile count). There is no on-chip graph compiler, no floating point, no
training — job graphs and addresses are computed off-chip and loaded
via the host interface (§9).
## 4. Unified memory
```
┌─────────────────────┐
│ FPGA ECP5 │
│ │
│ 4x Neural Engines │
│ │ │
│ v │
│ Unified SDRAM │
│ Backend / Arbiter │
└─────────┬───────────┘
│ 16-bit SDRAM bus
v
┌─────────────────────┐
│ AS4C4M16SA-6TIN │
│ Weights │
│ Activations │
│ Results │
└─────────────────────┘
```
| Item | Value | Basis |
|---|---|---|
| Device | Alliance Memory AS4C4M16SA-6TIN | DESIGN DECISION (STEP16-19) |
| Capacity | 4M x 16 (8MB) | DATASHEET VALUE |
| Data width | 16-bit (DQ[15:0]) + DQM[1:0] byte mask | DATASHEET VALUE |
| Addressing | BA[1:0] (4 banks) + A[11:0] (row/col, multiplexed) | DATASHEET VALUE |
| Clock | shared with FPGA system clock (§5) | DESIGN DECISION |
| Initialization/refresh | real, RTL-implemented power-up wait + mode-register-set + periodic AUTO REFRESH (`sdram_controller.v`) | VERIFIED (real refresh events observed in simulation, STEP16-19) |
| Arbitration | single physical port, 2-way logical split: W (weight, read-only, cached) / AR (activation+result, read/write, byte-maskable), each internally arbitrated across N processors by a generic, reused `slot_mem_arbiter` | VERIFIED (STEP19 bit-exact regression, reconfirmed via Verilator this step — see errors.log ERR-0024) |
| Official V2 memory map | weights @0x010000, activations @0x200000, results @0x300000, all within the single 8MB space, 1MB-aligned | DESIGN DECISION |
PSRAM is **not** part of V2. The V1 PSRAM controller (`hardware/v1/rtl/psram_controller.v`) is not instantiated anywhere in the V2 physical path.
## 5. Clock / PLL
```
16 MHz OSCILLATOR
|
v
ECP5 PLL (EHXPLLL)
CLKI_DIV=1 CLKFB_DIV=4 CLKOP_DIV=9
FEEDBK_PATH=CLKOP VCO=576MHz
|
v
FPGA SYSTEM CLOCK
64 MHz
(real, tool-generated ratio: 16 * 4 / 1, CLKOP_DIV=9 -> 576/9=64)
```
| Item | Value | Basis |
|---|---|---|
| Oscillator | 16 MHz (board-level, prior project record) | DESIGN DECISION (part number: TBD — not selected this session) |
| PLL primitive | EHXPLLL (`ecp5_pll_sys_clk.v`) | VERIFIED design-time via Project Trellis `ecppll` v1.4 (real tool, real parameters) |
| Generated system clock | 64 MHz | DESIGN DECISION, chosen over 80MHz because STEP19's own multi-seed P&R data showed only 1/8 seeds closing timing at >=80MHz on the compute-only core, and the new board-level top adds more logic still; 64MHz is not yet itself confirmed by P&R on the NEW top (see §11) |
| PLL lock | `locked` output, feeds `reset_sync.v` | DESIGN DECISION; NOT simulatable (Lattice EHXPLLL has no open sim model) — real lock behavior is a real-hardware-only characterization, see §10 |
| Timing constraints | none yet written for the new board-level top | OPEN — see §11 |
## 6. Interfaces
### SPI host interface (`spi_host_bridge.v`)
Mode 0 (CPOL=0/CPHA=0), MSB-first, one opcode per CS-low period.
Opcodes: `0x10` WRITE_JOB (job registration, 15-byte payload), `0x01`
WRITE_MEM / `0x02` READ_MEM (raw, word-addressed SDRAM access via a
second arbitrated port), `0x20` STATUS, `0x0F` RESET. Verified in
isolation (18/18, `tb_spi_host_bridge.v`). **Not yet verified
end-to-end under realistic multi-job pacing** — see §11/ERR-0025.
The 110-pin `reg_*` bus used by V2's own internal simulation
testbenches is a testbench-only convenience and is **not** the
physical interface.
### JTAG
Standard ECP5 JTAG (TDI/TDO/TCK/TMS), always available regardless of
configuration boot mode, per Lattice's own standard requirement.
### Configuration
Standard ECP5 PROGRAMN/INITN/DONE/CCLK. Boot-mode/flash-part decision:
OPEN (see §11).
## 7. Electrical
Rail voltage requirements are DATASHEET VALUEs (from real device
datasheets); no regulator part numbers, current budget, or decoupling
values are finalized this round. Full detail:
`hardware/v2/docs/POWER_ARCHITECTURE.md`.
## 8. Pinout
Full table: `hardware/v2/docs/PINOUT.md`. Summary: 37 real SDRAM
signals + clk/rst are ball-assigned and P&R-verified (STEP19, against
the STEP19 compute-only top). The board-level top added this step
(SPI + oscillator + reset pins) has **not** had its own ball
assignment or P&R run yet.
## 9. Mechanical / board assumptions
None assumed beyond the package footprint implied by CABGA381. No PCB
dimensions, connector placement, or stack-up are specified — that is
schematic/PCB-capture work, not yet started (see
`hardware/v2/docs/SCHEMATIC_READINESS.md`).
## 10. Programming / first power-on
JTAG programming is standard. A first-power-on procedure exists at
`hardware/v2/docs/FIRST_POWER_ON.md` (procedure only — not executed
against real hardware, since no board has been fabricated).
## 11. Limitations (real, current, as of this datasheet's own writing)
- **The physical SPI host interface is NOT proven end-to-end
correct.** A real, disclosed defect (errors.log ERR-0025 Part B)
produces wrong results when two jobs are dispatched with realistic
SPI pacing, even though registration itself is confirmed correct.
This is the single largest open item.
- The board-level top (`fpga_neural_v2_top.v`) has not been through
synthesis or P&R this round — deliberately, since running the real
toolchain against RTL known to compute wrong answers would not be a
meaningful result.
- No PCB, schematic capture, or fabricated hardware exists. Nothing in
this document should be read as "physically validated."
- Regulator, configuration-flash, and connector part numbers are not
selected.
- The STEP19 compute+memory core (raw `reg_*` interface, no SPI
bridge) IS bit-exact verified (N=2 and N=4, 256/256, reconfirmed via
Verilator this session) and remains the actual, working reference
design underneath this datasheet's own described board architecture.
@@ -1,160 +0,0 @@
# FPGA-Neural V2 — Reference Schematic (textual)
**No KiCad schematic was generated this session.** No RTL-to-schematic
or netlist-to-KiCad automation tool is available in this environment,
and the project's own separate, pre-existing KiCad PCB directory
(`FPGA-Neural/FPGA-Neural/FPGA-Neural/`) is an unrelated, independently
tracked project (its own nested `.git`, near-empty as of last check) —
it was not touched, and this document does not assume its contents.
This is a textual/ASCII reference schematic: a real starting point for
PCB capture, not a substitute for one. All ball assignments below are
the real, P&R-verified ones from `hardware/v2/constraints/
v2_unified.lpf` (STEP19) unless marked otherwise.
## 1. Top-level block diagram
```
+---------------------------+
| HOST MCU |
| SPI |
+------------+----------------+
|
v
+----------------------------------------------------------+
| ECP5 FPGA (LFE5U-45F-8BG381) |
| |
| +--------------+ +---------------------------+ |
| | SPI Host |---->| Register / Control | |
| | Bridge | | (job registration) | |
| +--------------+ +-------------+-------------+ |
| | |
| +-------------v-------------+ |
| | Neural Accelerator (N=4) | |
| | Processor 0..3 | |
| +-------------+-------------+ |
| | |
| +-------------v-------------+ |
| | Unified SDRAM Backend | |
| +-------------+-------------+ |
+----------------------------------------------------------+
|
16-bit SDRAM bus
v
+----------------------------+
| AS4C4M16SA-6TIN |
| Weights / Activations / |
| Results |
+----------------------------+
16 MHz osc --> ECP5 PLL (EHXPLLL) --> 64 MHz system clock
Power rails --> POR/supervisor --> FPGA reset, SDRAM init
Configuration flash + JTAG connector (see 5/6)
```
## 2. SDRAM connection table (real, P&R-verified balls)
| Signal | Ball | Bank | I/O std (assumed) | Direction |
|---|---|---|---|---|
| CLK (shared w/ system clk) | H5 | — | LVCMOS33 | FPGA -> SDRAM |
| CKE | B5 | 7 | LVCMOS33 | FPGA -> SDRAM |
| CS_N | C5 | 7 | LVCMOS33 | FPGA -> SDRAM |
| RAS_N | C4 | 7 | LVCMOS33 | FPGA -> SDRAM |
| CAS_N | A3 | 7 | LVCMOS33 | FPGA -> SDRAM |
| WE_N | B3 | 7 | LVCMOS33 | FPGA -> SDRAM |
| BA[0] | E4 | 7 | LVCMOS33 | FPGA -> SDRAM |
| BA[1] | C3 | 7 | LVCMOS33 | FPGA -> SDRAM |
| A[0..11] | D5,D3,F4,E5,E3,F5,A2,B1,C2,C1,D2,D1 | 7 | LVCMOS33 | FPGA -> SDRAM |
| DQ[0..15] | E1,G5,H3,J5,K3,K2,H1,J1,K1,K4,L4,L5,M5,M4,N4,N5 | 7/6 | LVCMOS33 | bidirectional |
| DQM[0..1] | P5,N3 | 6 | LVCMOS33 | FPGA -> SDRAM |
Full source: `hardware/v2/constraints/v2_unified.lpf`. LVCMOS33 is
assumed to match the SDRAM's own real 3.3V requirement and matches
banks 6/7's real VCCIO range per `docs/pinouts.md` — not yet
independently cross-checked at the schematic/PCB level (WARNING, not
BLOCKER).
## 3. Clock schematic
```
16MHz OSC ---> CLKI (H5, reused from V1's own real LPF)
|
+-----v------+
| EHXPLLL | CLKI_DIV=1, CLKFB_DIV=4, CLKOP_DIV=9
| (hard IP) | FEEDBK_PATH=CLKOP, VCO=576MHz
+-----+------+
| CLKOP = 64MHz
v
FPGA system clock (feeds compute, SDRAM ctrl, SPI bridge)
|
+-----v------+
| reset_sync | <-- ext POR (active-low) + PLL LOCK
+-----+------+
v
rst (sync-deassert, feeds every synchronous block)
```
Oscillator part number: **TBD** (not selected this session — a real
16MHz, 3.3V HCMOS clock oscillator in a standard SMD package is the
intended class of part; no specific manufacturer/part number is
claimed without a real datasheet lookup performed this session).
## 4. Power schematic (rails only — no regulator parts selected)
```
3.3V/1.1V/2.5V rails (regulators: TBD)
| | |
v v v
VCCIO VCC(core) VCCAUX
(banks (real ball (real ball
6/7=SDRAM cluster, cluster,
I/O, etc) see see
POWER_ARCH POWER_ARCH
.md) .md)
|
v
SDRAM VDD/VDDQ (3.3V, DATASHEET VALUE per AS4C4M16SA-6TIN)
```
Full rail table, decoupling guidance, and current-budget status:
`hardware/v2/docs/POWER_ARCHITECTURE.md` (unchanged this step — no new
power work performed).
## 5. Configuration / JTAG schematic
```
FPGA
|-- TDI/TDO/TCK/TMS --> JTAG connector (standard pinout, always
| available regardless of boot mode)
|-- PROGRAMN/INITN/DONE/CCLK --> configuration flash (part: TBD) or
JTAG-only bring-up (decision: OPEN)
```
No configuration-flash part has been selected; JTAG-only bring-up
remains a valid fallback and is documented as such in
`hardware/v2/docs/CONFIGURATION.md`-equivalent content inside
`OPEN_ITEMS.md` (a dedicated `CONFIGURATION.md` was not created this
round — tracked as an open item, not silently dropped).
## 6. Host interface schematic
```
Host MCU --SPI--> FPGA: spi_sclk, spi_mosi, spi_miso, spi_cs_n
```
No ball assignment exists yet for these 4 signals (the board-level
top was not run through P&R this session — see the datasheet's own
§11 Limitations). Pull-up on `spi_cs_n` (idle-high) is the standard,
expected design decision for a single-master SPI bus; not yet placed
in any real LPF.
## 7. What this schematic deliberately does NOT claim
- No KiCad artifact. No PCB. No fabricated board.
- No ball assignment for the new SPI/oscillator/reset pins (P&R not
run against the new board-level top this session, since the design
has a known, unresolved functional defect — see errors.log
ERR-0025 Part B).
- No regulator, flash, or connector part numbers.
This document is a real, honest starting point for PCB capture, not a
finished schematic.
-118
View File
@@ -1,118 +0,0 @@
# FPGA-Neural V2 — HARDWARE FREEZE (FASE #1, single external SDRAM)
**PARTIALLY SUPERSEDED (DEC-0039).** The SDRAM part number below
(AS4C4M16SA-6TIN, 8MB) was upgraded to **AS4C32M16SB-7BIN (64MB)**,
and N_PROCESSORS=8 is no longer merely a "future evolution" — it is
now real, synthesized, P&R-verified (functionally correct, with a
disclosed, real 64MHz timing-closure gap at 5/8 tested seeds). See
`MEMORY_UPGRADE_64MB_N8.md` for the current, authoritative state. The
rest of this document (Neural Processor, dataflow architecture) is
still accurate.
## Frozen reference configuration
```
FPGA: LFE5U-45F-8BG381, ECP5U, speed grade -8
Neural Processor: P_IN=8, INT8 operands, INT32 accumulator,
8 parallel MAC, balanced adder tree (neural_processor.v,
UNCHANGED since before this freeze)
Multiprocessor: N_PROCESSORS=4 (N4/P8 is the frozen reference; N2 also
validated; N8 is a FUTURE EVOLUTION, not part of this freeze)
Architecture: Neural Multiprocessor -> Dataflow -> Neural Director ->
Dependency Manager -> Memory Manager -> streaming tile
delivery (STEP13 architecture, intact, unchanged)
External memory: ONE SDRAM ONLY -- Alliance Memory AS4C4M16SA-6TIN,
serving weights, activations, AND results (DEC-0031/
0032/0033/0034). No PSRAM, no second memory device.
Weight path: PACK128 (BURST_LEN=8, N_ENTRIES=4 cache, STEP18/STEP19)
Target clock: 80 MHz minimum (real oscillator/PLL source: OPEN, see
CLOCK_ARCHITECTURE.md)
V1: golden/reference implementation, untouched (confirmed:
zero modifications; V2 no longer instantiates ANY V1
RTL at all, since psram_controller.v was removed from
the physical path -- DEC-0034)
Frozen top-level: nms_neural_multiprocessor_sdram_unified
(hardware/v2/nms/rtl/nms_neural_multiprocessor_sdram_unified.v)
```
N4/P8 is the frozen V2.0 hardware reference. This does not mean N4 is
the final or maximum architecture — N8, higher clocks, or new datapath
ideas are explicitly FUTURE EVOLUTIONS, out of scope for this freeze.
## Repository audit summary
Full detail: see the audit performed for this step (repository
structure, V1/V2 boundary, top-level candidates, dead-code
classification, PSRAM-dependency confirmation, LPF/docs/scripts
inventory). Key findings:
- `hardware/v1/**`: complete, self-contained, untouched. Real
synthesized/certified golden reference (`spi_neuron_top.v`).
- `hardware/v2/rtl/` + `hardware/v2/nms/rtl/`: the frozen top-level
(`nms_neural_multiprocessor_sdram_unified.v`) instantiates
`nms_dataflow_core_sdram.v`, `sdram_unified_backend.v` (STEP19, new),
`sdram_controller.v`, `nms_memory_manager_stream_wide.v`,
`weight_prefetch_engine_wide.v`, `nms_activation_replicated.v`,
`nms_activation_fill_ctrl_v3.v`, `nms_weight_packed.v`,
`dependency_manager.v`, `neural_director.v`, `neural_processor.v`,
`slot_mem_arbiter.v`, `slot_mem_arbiter_wide.v`, `prefetch_engine.v`
**zero V1 files**, confirmed by successful lint/synthesis with no
V1 RTL in the file list.
- Every other `nms_neural_multiprocessor_*.v`/`nms_dataflow_core_*.v`
variant (plain, `_pf`, `_stream`, `_actfix`, `_actfix2`, `_dual32`,
`_sdram`, `_sdram_pack128`) is real, historical, superseded-but-
documented project experiment history — dead relative to the frozen
top, NOT deleted (each remains the subject of its own STEP report).
- `hardware/v2/constraints/` was empty before this step; now contains
`v2_unified.lpf` (partial — see PINOUT.md).
- The referenced sibling pinout repository (`../basic-ecp5-pcb`) does
not exist on disk, BUT the real Lattice pinout CSV itself
(`FPGA-SC-02034-3-0-ECP5U-45-Pinout.csv`, rev 3.0) is present at
`~/Downloads/` and was found during this step's own pre-commit
review, with a real summary already at `docs/pinouts.md` (repo
root). This corrected an earlier draft of this freeze that
wrongly assumed no real pinout data existed — see PINOUT.md.
## Status table
| Area | Status | Evidence | Blocker |
|---|---|---|---|
| RTL | PASS | Strict Verilator lint (latches/multi-driver/comb-loops/case-completeness): zero findings across the full frozen hierarchy | No |
| Simulation | PASS | Isolated `tb_sdram_controller.v` (461/461, 9 freq/burst configs), isolated `tb_sdram_unified_backend.v` (40/40) | No |
| Bit-exact | PASS | Full N=4 AND N=2 D-Stress (256/256 neurons each), golden software model comparison | No |
| SDRAM | PASS | Real init/refresh/read/write/burst/masked-write, 40 real AUTO REFRESH events interleaved with zero corruption across a ~50,000-cycle run | No |
| Synthesis | PASS | Real Yosys 0.68+post synthesis, N=4: TRELLIS_FF=6425, TRELLIS_COMB=6023, MULT18X18D=32, DP16KD=0 | No |
| P&R | PASS (fits) | Real nextpnr-ecp5 0.11.1, TRELLIS_IO=149/245 (fits with headroom) | No |
| Timing | **MARGINAL** | 8 real seeds: 66.97/74.00/74.45/74.92/79.23/79.53/79.80/81.84 MHz — only 1/8 ≥80MHz | **CRITICAL** |
| Pinout | PARTIAL | 39/149 signals real, sourced, P&R-verified (clk/rst + full 37-signal SDRAM bus); 110-signal host bus unassigned | **BLOCKER (host bus only)** |
| Clock | INCOMPLETE | Single-clock-domain RTL confirmed (no CDC); no PLL exists; oscillator-vs-PLL decision not made | **CRITICAL** |
| Power | INCOMPLETE | Real rail voltages known from datasheets; no regulator selection, no current budget | OPEN |
| Configuration | INCOMPLETE | Standard ECP5 JTAG/config pins identified; no flash part chosen, no V2 config LPF beyond the partial `v2_unified.lpf` | OPEN |
| Host | INCOMPLETE | 110-pin raw parallel bus exists at the RTL boundary; no physical protocol, no serializer RTL | **BLOCKER** |
| PCB | NOT READY | See SCHEMATIC_READINESS.md's own checklist | Multiple (host, pinout, power) |
| Bring-up | READY (procedure only) | FIRST_POWER_ON.md defines the full 12-step test sequence | Cannot execute until host/pinout blockers close |
## Single-SDRAM verification (this step's own core mandate)
- PSRAM dependency: **REMOVED** — confirmed by successful synthesis/
P&R with zero V1 files in the compile list, and a real, measured
45-pin I/O reduction (194→149/245 TRELLIS_IO) exactly matching the
removed PSRAM interface's own pin count.
- Weights/activations/results: **all confirmed sharing the single
physical SDRAM**, real bit-exact traffic at three distinct,
non-overlapping memory-map regions, simultaneously, under real N=4
contention (see MEMORY_ARCHITECTURE.md).
- Real bugs found and fixed during this consolidation (ERR-0023): a
full deadlock and a subsequent off-by-one data-shift bug in the new
arbitration logic, both caught via full-system (not merely isolated)
testing before being accepted — see errors.log for the complete
root-cause writeups.
## Deliverables produced by this step
`HARDWARE_FREEZE.md` (this file), `CHIP_READINESS.md`,
`MEMORY_ARCHITECTURE.md`, `PINOUT.md`, `POWER_ARCHITECTURE.md`,
`CLOCK_ARCHITECTURE.md`, `SCHEMATIC_READINESS.md`,
`FIRST_POWER_ON.md`, `OPEN_ITEMS.md` (all under `hardware/v2/docs/`),
plus `hardware/v2/constraints/v2_unified.lpf` and new RTL/testbenches
under `hardware/v2/nms/rtl/` and `hardware/v2/nms/sim/`.
-139
View File
@@ -1,139 +0,0 @@
# FPGA-Neural V2 — MEMORY ARCHITECTURE (single SDRAM)
**PART NUMBER SUPERSEDED (DEC-0039).** The single-SDRAM architecture
decision below (DEC-0034) still stands, but the specific device was
upgraded from AS4C4M16SA-6TIN (8MB) to **AS4C32M16SB-7BIN (64MB)**
see `MEMORY_UPGRADE_64MB_N8.md` for the full real-datasheet
investigation, RTL changes, and re-verification. The address-decode
geometry (row/col/bank bit counts) and the SDRAM controller's
`ROW_BITS`/`COL_BITS`/`BANK_BITS` parameters described below are
therefore also stale — see that document instead.
## Decision (DEC-0034)
**ONE external memory device: Alliance Memory AS4C4M16SA-6TIN SDR
SDRAM (64Mbit/8MB, x16).** Weights, activations, and results all share
this single physical chip through a single `sdram_controller.v`
instance. No PSRAM, no second external memory device anywhere in the
V2 physical path. This is a closed architectural decision (per the
governing spec) — it will not be reopened.
```
SDRAM (AS4C4M16SA-6TIN, 8MB)
|
sdram_controller.v
(BURST_LEN=8, real
JEDEC SDR protocol)
|
sdram_unified_backend.v
(W port cache + AR port masking,
2-way priority arbitration)
| |
W (64-bit) AR (16-bit, byte-maskable)
| |
slot_mem_arbiter_wide.v slot_mem_arbiter.v
| |
weight_prefetch_engine_wide.v nms_activation_fill_ctrl_v3.v
(per slot, N_SLOTS instances) (shared) + nms_memory_manager_
| stream_wide.v (per-slot result
Neural Processors writeback, N_SLOTS instances)
```
## Why one physical controller is enough
`sdram_unified_backend.v` presents two LOGICAL ports (W: weight, AR:
activation+result) but owns exactly one physical `sdram_controller.v`
instance and arbitrates between them with a simple, correctness-first
2-way priority scheme (W wins when both are pending — real measured
traffic, STEP17 EXP-0045, shows weight traffic dominates by a wide
margin; AR is never starved since W's own real access pattern idles
between tiles). This matches the governing spec's own explicit
guidance: "non è necessario che esistano tre controller."
## The enabling mechanism: real SDR SDRAM byte masking (DQM)
Real SDR SDRAM has native per-byte write masking via its DQM pins —
`sdram_controller.v` was extended (STEP19) with a `wmask` input (2
bits per burst word) that drives `sdram_dqm` dynamically per burst
word instead of the STEP16-18 hardcoded "always write everything."
This lets a single RESULT byte be written inside a shared 128-bit (8
x16-bit-word) burst transaction with **no read-modify-write at all**
masked bytes are left untouched by the real chip, by JEDEC definition.
Verified with a new dedicated test (`tb_sdram_controller.v` Test J:
byte-masked write, confirms neighboring bytes/words in the same real
128-bit block are unchanged) — PASS across all 9 existing frequency/
burst configurations plus the new test (461/461 each), zero
regression.
Activation reads need no such trick: a full 128-bit block is fetched
and the caller's requested 16-bit word is extracted combinationally.
## Official V2 memory map
The single 8MB (0x0000000x7FFFFF byte) SDRAM address space is
divided into non-overlapping, 1MB-aligned regions:
| Region | Base address | Size (reserved) | Owner | Access |
|---|---|---|---|---|
| Network/metadata | 0x000000 | 1 MB (0x0000000x0FFFFF) | host (future) | R/W |
| Weights | 0x010000* | up to 1 MB | weight_prefetch_engine_wide.v (per-job `w_base`) | read-only |
| Biases | 0x100000 | 1 MB (0x1000000x1FFFFF) | reserved, not yet used by D-Stress | — |
| Activations | 0x200000 | up to 1 MB | nms_activation_fill_ctrl_v3.v (per-job `x_base`) | read-only |
| Intermediate results | 0x300000 | up to 1 MB | nms_memory_manager_stream_wide.v (per-neuron `result_addr`) | write (+ future read for chaining) |
| Output | 0x400000 | 1 MB (0x4000000x4FFFFF) | reserved, not yet used | — |
| (reserved/future) | 0x5000000x7FFFFF | 3 MB | — | — |
\* the real D-Stress benchmark's own weight region starts at
0x010000, inside the "Network/metadata" 1MB region's own upper part
for simplicity — addresses are **programmable**, set per-job via
`reg_w_base`/`reg_x_base`/`reg_result_addr` at registration time (NOT
hardcoded in the datapath) — this map is the project's own convention
for how a real host should lay out a graph, not an RTL constant.
Base/size/alignment/access-type/owner are exactly the fields the
governing spec requests; "owner" above names the RTL module
responsible for traffic in that region.
## Address-space coexistence — real, tested evidence
`tb_sdram_unified_backend.v` (isolated) exercises W-port and AR-port
traffic at deliberately different regions with real interleaving (Test
D) and confirms no corruption. The full N=4/N=2 D-Stress benchmark
(`tb_nms_dstress_sdram_unified.v`) exercises ALL THREE traffic classes
simultaneously at their real, disjoint memory-map regions across 256
neurons, 4096 tiles, with 40 real interleaved AUTO REFRESH events —
bit-exact PASS at both N=2 and N=4. This maps directly onto the
governing spec's own required Test AI list:
| Governing spec test | Covered by |
|---|---|
| A: weights only | `tb_sdram_weight_backend_pack128.v` (STEP18, reused unchanged logic) + isolated Test A (`tb_sdram_unified_backend.v`) |
| B: activations only | Isolated Test B |
| C: results only | Isolated Test C (byte-masked write) |
| D: weights+activations | Isolated Test D |
| E: weights+results | Covered by the full D-Stress run's own real traffic mix |
| F: weights+activations+results simultaneously | Full D-Stress run (real, not synthetic) |
| G: N4 contention | Full D-Stress run at N_SLOTS=4 |
| H: repeated workloads | 256 neurons × 16 tiles each = 4096 repeated weight/activation fetches + result writes in one continuous run |
| I: long-running workload | ~50,000-cycle run spanning 40 real AUTO REFRESH intervals, zero corruption |
All: **bit-exact PASS, no deadlock, no timeout, no corruption** (after
ERR-0023's fix — see errors.log for the one real deadlock + one real
off-by-one bug found and fixed via exactly this testing).
## Performance cost of unification (disclosed, not hidden)
| | STEP18 (2 chips) | STEP19 (1 chip) | Δ |
|---|---|---|---|
| N=4 D-Stress cycles | 44,935 | 49,771 | +10.8% |
| N=2 D-Stress cycles | 47,399 | 49,788 | +5.1% |
| Bit-exact | PASS | PASS | — |
| TRELLIS_IO | 194/245 | 149/245 | **-45 pins (-23.2%)** |
| Fmax (best-of-N-seeds, N=4) | 81.47 MHz (5/8 pass) | 81.84 MHz (1/8 pass) | worse pass rate, MARGINAL |
The cycle-count cost is a direct, expected consequence of activation
and result traffic now competing for the SAME physical bandwidth that
previously had its own independent chip — reported honestly per the
governing spec's own "prima misura poi ottimizza" instruction, not
optimized away this round (that would be a FUTURE EVOLUTION, e.g. a
smarter scheduler/priority scheme between W and AR).
-454
View File
@@ -1,454 +0,0 @@
# FPGA-Neural V2 — Memory Upgrade (64MB) + N_SLOTS=8 + Clock Re-Verification
Supersedes the SDRAM-related content of `PRE_PCB_VERIFICATION.md` and
`PRE_PCB_CLOSURE_4POINT.md` (both describe the previous 8MB
AS4C4M16SA-6TIN baseline). This document is the authoritative record
for: the memory capacity investigation, the frozen replacement part,
every RTL change it required, two real timing regressions found and
fixed via real P&R data, and the honest, current state of N_SLOTS=4
vs N_SLOTS=8 clock closure.
---
## 1. Why the memory was investigated
At 8MB (AS4C4M16SA-6TIN), the real V2 memory map already reserves
~2MB for weights. A concrete throughput check: the existing D-Stress
benchmark (256 neurons × 128 inputs = 32,768 weight bytes) takes
49,771 cycles (777µs at the real, P&R-verified 64MHz) to run to
completion. Extrapolating linearly, a 24MB weight budget (the
proportional share of a 64MB device) would take on the order of
**~580ms for one inference pass** — already deep into "too slow to
matter" territory for this accelerator's real target (a low-latency
SPI-peripheral offload engine), well before capacity itself becomes
the binding constraint. This was disclosed to the user directly:
capacity was not really the bottleneck, compute throughput was. The
user weighed this and still asked for the largest same-family,
same-package part, with N_SLOTS=8 as the preferred processor count —
both honored below, with a fully honest report of what real P&R data
says about clock closure at each.
## 2. Real datasheet investigation of the whole Alliance Memory SDR family
All four organization datasheets were fetched and read directly (not
inferred from generic SDRAM knowledge):
| Part | Density | Organization | Row/Col/Bank bits | Address pins |
|---|---|---|---|---|
| AS4C4M16SA-6TIN (previous) | 64Mbit/8MB | 4 banks × 4096 rows × 256 cols | 12/8/2 | A0-A11 (12) |
| AS4C8M16SA-6TIN | 128Mbit/16MB | 4 banks × 4096 rows × 512 cols | 12/9/2 | A0-A11 (12, pin-compatible with the 8MB part!) |
| AS4C16M16SA-6TIN | 256Mbit/32MB | 4 banks × 8192 rows × 1024 cols... | — | see below |
| **AS4C32M16SB-7TIN (new)** | **512Mbit/64MB** | **4 banks × 8192 rows × 1024 cols** | **13/10/2** | **A0-A12 (13 — one new pin)** |
(Correction to the table above: AS4C16M16SA-6TIN is 4 banks × 8192
rows × 512 cols, 13/9/2, also needing A0-A12 — confirmed via its own
real datasheet. The key finding driving the final part choice: going
from 32MB to 64MB costs **zero additional pins** beyond what 32MB
already requires, since both need the same 13 address pins. There is
no PCB-simplicity reason to stop at 32MB once the 13th pin is already
being added.)
**"SA" vs "SB" note**: Alliance Memory's own datasheet revision
history (AS4C32M16SA Rev 2.0: "Die Shrink A revision") confirms
these letter suffixes denote die-shrink process revisions, not
functional or pinout changes. Real distributor availability (section
6 below) shows "SB" as the currently-stocked die for this part.
**Package: BGA, not TSOP-II** — per the user's own explicit choice,
the FROZEN part is **AS4C32M16SB-7BIN** (54-ball TFBGA, 8.0×8.0×1.2mm
max, "B" package-code suffix), not the TSOP-II "-7TIN" variant
discussed earlier in this investigation. Same die, same organization,
same timing, same 3.3V/industrial-temp electricals — the datasheet's
own "Features" section lists both a 54-pin TSOP-II AND a 54-ball FBGA
package option for this exact device; only the physical footprint
differs (a PCB-level choice, the user's own call). The datasheet-level
electrical/timing audit in this document applies unchanged to either
package option.
## 3. Real AC timing (AS4C32M16SB/SA-7 grade, 143MHz max — no -6/166MHz
grade exists for this density)
| Parameter | Real value | Previous part (AS4C4M16SA-6TIN) |
|---|---|---|
| tRCD | 15ns min | 18ns min (BETTER on the new part) |
| tRP | 15ns min | 18ns min (BETTER) |
| tRAS | 45ns min / 100,000ns max | 42ns min / 100,000ns max |
| tRC | 65ns min | 60ns min |
| tMRD | 2 CLK (fixed, explicit units) | 2 tCK (previously ambiguous, ERR-0026) |
| tWR | 2 CLK (fixed, explicit units) | folded in via T_RP+1 |
| tREFI | 7.8125µs (8192 rows/64ms) | 15.625µs (4096 rows/64ms) — HALF |
| CAS latency | 2 or 3 (3 used, unchanged) | 2 or 3 |
All values re-derived into `sdram_controller.v`'s own `ns_to_cycles()`
function at the real 64MHz target — verified safe at 64MHz through
166MHz via the full regression sweep (section 7).
## 4. RTL changes required
### 4.1 `sdram_controller.v` and `sdram_model.v` — parameterized geometry
Both files gained real `ROW_BITS`/`COL_BITS`/`BANK_BITS` parameters
(defaults 13/10/2, matching the new part) replacing hardcoded 12/8/2
widths throughout: the address decode, the column-phase address
assembly (previously a hardcoded `{4'b0100, col}` concatenation, now
a parameterized construction that places the AP bit at the same bit
10 position regardless of column width), the MRS mode-register value
(re-derived to be zero-padded correctly for any ROW_BITS), and the
refresh-interval computation (now `64000000/(1<<ROW_BITS)+1`, correct
for either device). An elaboration-time assertion
(`ADDR_WIDTH == BANK_BITS+ROW_BITS+COL_BITS`) catches any future
mismatched override immediately.
### 4.2 Address-width propagation (23→26 bits, byte address)
`ADDR_WIDTH` default widened from 23 to 26 across every module in the
live instantiation tree: `spi_host_bridge.v`, `dependency_manager.v`,
`neural_director.v`, `slot_mem_arbiter.v`, `slot_mem_arbiter_wide.v`,
`nms_dataflow_core_sdram.v`, `nms_activation_fill_ctrl_v3.v`,
`weight_prefetch_engine_wide.v`, `nms_dataflow_core_sdram.v`,
`sdram_unified_backend.v`, `fpga_neural_v2_top.v`, and the D-Stress
testbench's own top wrapper `nms_neural_multiprocessor_sdram_unified.v`.
`sdram_unified_backend.v` also gained its own `ROW_BITS`/`COL_BITS`/
`BANK_BITS` pass-through parameters (forwarded to `sdram_controller`
instead of a hardcoded `.ADDR_WIDTH(22)` override that would otherwise
have silently reverted to the old geometry), and its internal
word/byte address-conversion wires were parameterized instead of
hardcoded to 22 bits.
### 4.3 SPI protocol change (`spi_host_bridge.v`) — real, necessary
A 26-bit byte address no longer fits in 3 bytes (24 bits) with a
spare reserved bit the way the old 23-bit address did. Every address
field widened from 3 to 4 bytes:
- **WRITE_JOB**: 15 → **18 payload bytes** (x_base/w_base/result_addr
each 3→4 bytes).
- **WRITE_MEM/READ_MEM header**: 5 → **6 bytes** (addr 3→4 bytes).
`byte_idx` widened from 4 to 5 bits (max index 17, was 14) to
accommodate the longer WRITE_JOB frame.
### 4.4 New PCB pin: `sdram_a[12]`
`v2_board_top.lpf` gained one new entry: `sdram_a[12]` → ball **F1**
(bank 6, official Lattice pinout CSV rev 3.0, CABGA381 column) — a
real, previously-unused, plain-GPIO ball, verified not already
assigned to any of the LPF's existing 44 signals.
## 5. Two real timing regressions found and fixed (see errors.log
ERR-0027/ERR-0028 for the full root-cause writeups)
**ERR-0027**: `neural_director.v`'s own per-slot dispatch used a
runtime-indexed write into a wide packed register
(`slot_x_base[free_slot_idx*ADDR_WIDTH +: ADDR_WIDTH] <= ...`),
synthesizing as an actual MULT18X18D multiplier feeding a wide
demux/crossbar. This got worse as ADDR_WIDTH grew — real P&R: worst
seed collapsed from the previously-verified 68.51MHz to 40.27MHz,
FAILING 64MHz across all 8 seeds. **Fixed** by replacing it with
N_SLOTS unpacked per-slot registers, written via N_SLOTS parallel
constant-indexed compares (no multiply), wired out via a
constant-genvar generate block. Confirmed: the spurious 33rd
MULT18X18D at N=4 is gone (now exactly 32 = 4×8, matching the real
per-processor MAC count). Real P&R after the fix, N=4, 8 seeds: **ALL
PASS at 64MHz** (65.0272.01MHz, mean ~68.8MHz).
**ERR-0028**: found immediately after, at N_SLOTS=8: a DIFFERENT,
pre-existing critical path in `nms_activation_fill_ctrl_v3.v`'s own
`max_n_tiles_comb` — a flat, linear N_SLOTS-wide sequential max-scan,
already flagged by that file's own prior comment as "an N_SLOTS-wide
sequential chain." At N_SLOTS=8 (twice the comparison depth of N=4,
where it wasn't the bottleneck) it became dominant: real P&R showed
~38-40MHz, failing 64MHz on all 4 tested seeds. **Fixed** by replacing
the flat scan with an explicit, hand-written balanced binary max-tree
(log2(N_SLOTS) levels instead of N_SLOTS), same single-cycle latency.
Real P&R after the fix, N=8, 8 seeds: **5/8 PASS at 64MHz**
(65.2770.78MHz), 3/8 FAIL narrowly (55.84/61.00/63.42MHz).
Both fixes were confirmed **bit-exact, zero functional regression**
via the full D-Stress N=2/4/8 regression (identical cycle counts to
the pre-fix baseline: 49961/49927/49909).
## 6. Honest current clock-closure status
| Configuration | Seeds tested | Result |
|---|---|---|
| N_SLOTS=4 @ 64MHz | 8/8 | **PASS, all seeds** (65.0272.01MHz real Fmax) |
| N_SLOTS=8 @ 64MHz | 8/8 | **5/8 PASS** (65.2770.78MHz), 3/8 FAIL (55.84/61.00/63.42MHz) — OPEN |
| N_SLOTS=4 or 8 @ 80MHz | 4 each | **FAIL, all seeds** (real 80MHz-targeted PLL regenerated via `ecppll`, real P&R re-run; same physical Fmax ceiling as the 64MHz-labeled runs, ~65-72MHz, confirming the achievable ceiling is a property of the fabric, not the requested target) |
**Recommendation**: **N_SLOTS=4 remains the frozen, fully-reliable
hardware configuration at 64MHz** (matches the project's own
established "safe = passes on every tested seed" standard).
**N_SLOTS=8 is functionally correct and usable, with a real, disclosed
timing risk**: 5 of 8 tested placement seeds close timing at 64MHz;
production would need to either (a) find and lock a known-good seed
(a real, standard practice — nextpnr's own seed is a build-time
choice, not a per-chip random draw) or (b) accept a further
timing-optimization pass (the same tree-based-reduction technique
already applied twice this session, next targeting
`sdram_unified_backend.v`'s own weight-cache hit-index scan — not
attempted this session, to avoid rushing a third unverified change).
**80MHz is not achievable with the current architecture at either
processor count** — a real, measured finding, not an assumption.
## 7. Full regression re-verification (real, this session)
| Test | Result |
|---|---|
| `tb_sdram_controller` (18 configs: 6 freqs × 3 burst lens, new 64MB geometry) | 461/461 PASS, every config |
| `tb_sdram_boundary` (21 directed checks, new geometry) | 21/21 PASS at 64MHz AND 166MHz |
| D-Stress N=2 | 49,961 cycles, 256/256 bit-exact PASS |
| D-Stress N=4 | 49,927 cycles, 256/256 bit-exact PASS |
| D-Stress N=8 | 49,909 cycles, 256/256 bit-exact PASS |
| `tb_spi_host_bridge` (new 18/6-byte protocol) | 18/18 PASS |
| Board-level SPI smoke test (real 64MHz clk_sys, new protocol) | 11/11 PASS |
| `tb_sdram_unified_backend` | 40/40 PASS |
## 8. Availability (real, checked this session)
**AS4C32M16SB-7BIN** (the frozen, BGA-package part): DigiKey product
11613071, 568 units in stock, $31.12/unit (qty 1), 16-week
manufacturer lead time, status Active, 54-ball TFBGA (8×8×1.2mm max),
-40 to 85°C industrial. Not a datasheet-only part — genuinely
orderable as of this session.
(The TSOP-II sibling, AS4C32M16SB-7TIN, was also confirmed real and
in stock — DigiKey 47 units, $31.40/unit — should the user reconsider
package during layout; both are the same die.)
## 9. What is still OPEN (honestly disclosed)
- N_SLOTS=8 clock closure at 64MHz: 5/8 seeds, not yet 8/8.
- The `sdram_unified_backend.v` weight-cache hit-index scan (the same
long-documented critical-path class) has not been tree-optimized —
a plausible next fix for closing the N=8 gap, not attempted this
session.
- The PRE_PCB_VERIFICATION.md / PRE_PCB_CLOSURE_4POINT.md documents'
own SDRAM-specific sections (organization tables, pin counts,
memory-map worked examples) describe the previous 8MB part and are
superseded by this document — not individually rewritten line-by-
line in this pass.
- The V2 LaTeX datasheet's own key-parameters table and memory-
architecture chapter still describe the 8MB device — not
regenerated this session (time/scope boundary); flagged here so it
is not silently stale. **UPDATE (2026-09-07): now addressed, see
section 10 below and `DataSheet/files/docs/datasheet/v2-en/chapters/
05-memory.tex`, appended section "SDRAM Upgrade Addendum."**
## 10. AUTHORITATIVE FINAL DATA (2026-09-07) — full 8-seed matrix,
ERR-0029 optimization, and complete AS4C32M16SB-7BIN pinout
This section is the authoritative, most-recent source of truth,
superseding sections 5-9 above where they conflict (kept for history).
All data below is real, measured, from `nextpnr-ecp5 0.11.1 --report`
JSON output and real Verilator 5.050 regression runs — no estimates.
### 10.1 N=4 @ 64MHz — PRE-ERR-0029 fix (period 15.625ns)
| Seed | Fmax (MHz) | WNS (ns) | Critical path (startpoint → endpoint) |
|---|---:|---:|---|
| 0 | 77.10 | +2.655 | director.job_out_slot → dep_mgr.node_resolved[13] |
| 1 | 74.68 | +2.234 | arbiter_wide.m_addr → sdram_backend.w_rdata |
| 2 | 74.64 | +2.228 | director.job_out_slot → dep_mgr.node_resolved[6] |
| 3 | 77.24 | +2.678 | arbiter_wide.m_addr → sdram_backend.w_rdata |
| 4 | 76.60 | +2.571 | sdram_backend.w_cache_addr[2] → sdram_backend.w_rdata |
| 5 | 77.96 | +2.798 | director.job_out_slot → dep_mgr.node_resolved[3] |
| 6 | 75.35 | +2.353 | director.job_out_slot → dep_mgr.node_resolved[2] |
| 7 (worst) | 74.17 | +2.143 | director.job_out_slot → dep_mgr.node_state[12] |
8/8 PASS. Worst seed7 74.17MHz/+2.143ns — routing-dominated (78%),
11 logic levels, classified as dependency_manager scheduler/producer-
consumer resolution logic.
### 10.2 N=8 @ 64MHz — PRE-ERR-0029 fix (period 15.625ns)
| Seed | Fmax (MHz) | WNS (ns) | Status |
|---|---:|---:|---|
| 0 | 57.27 | 1.837 | FAIL |
| 1 (worst) | 55.84 | 2.284 | FAIL |
| 2 | 66.99 | +0.698 | PASS |
| 3 | 70.78 | +1.496 | PASS |
| 4 | 61.00 | 0.769 | FAIL |
| 5 | 68.56 | +1.039 | PASS |
| 6 | 67.29 | +0.764 | PASS |
| 7 | 63.42 | 0.144 | FAIL |
4/8 PASS (2,3,5,6), 4/8 FAIL (0,1,4,7). Worst seed1 55.84MHz/2.284ns —
`sdram_unified_backend.v` weight-cache hit-index scan, 84% routing,
10 logic levels; three individual routing hops of 2.52.8ns.
Utilization: MULT18X18D 64/72 (88.9%), TRELLIS_COMB 11066/43848
(25.2%), TRELLIS_FF 11119/43848 (25.4%), DP16KD 0/108, TRELLIS_RAMW
323/5481.
### 10.3 N=4/N=8 @ 80MHz — genuine `ecppll`-regenerated PLL (CLKI_DIV=1,
CLKFB_DIV=5, CLKOP_DIV=7, CLKOP_CPHASE=3, VCO=560MHz), period 12.5ns
Both configurations: **0/8 seeds PASS** (achieved Fmax per seed
numerically identical to the 64MHz-PLL run in every case, confirming
the achievable ceiling is a fabric property, independent of PLL
target). N=4 closest: seed5, 77.96MHz, WNS=0.327ns. N=8 closest:
seed3, 70.78MHz, WNS=1.629ns. **NO-GO, both configs, both before and
after the ERR-0029 fix below** (re-confirmed in 10.5).
### 10.4 ERR-0029 root-cause investigation (user-directed, real data)
Investigated per the mandate's own 14-point checklist against the real
critical-path segment dump (seed1, N=8@64MHz) — see errors.log ERR-0029
for the full writeup. Summary of findings:
1. `w_cache_valid[0:W_ENTRIES-1]` (W_ENTRIES=4 fixed, NOT scaled by
N_SLOTS — confirmed via its only instantiation) generated by the
cache-allocate/consume sequential block.
2. `w_hit_idx_c` generated by a combinational `for` loop, "last
valid+matching entry wins" by unconditional sequential overwrite.
3. 4 comparators (one per W_ENTRIES).
4. Encoded via a serially-dependent priority scan, mapped by
Yosys/nextpnr onto cascaded ECP5 PFUMX/OFX fast-mux primitives.
5. `w_hit_idx_c` fans out to the 64-bit cache-data read mux and to
control/enable logic gating `w_rdata`'s load — a 3-way fan-out of a
value produced by a serial 4-stage chain.
6/7. The long hops (2.52.8ns each) are the physical distance between
the shared cache logic and the arbiter/consumer registers, stretched
by N=8's larger overall placement — NOT a logic-depth artifact
(W_ENTRIES doesn't grow with N_SLOTS).
8. Confirmed: yes, a serial mux-topology (PFUMX/OFX chain), not a
parallel structure.
9. Comparator fanout (4-wide) is NOT the dominant cost.
10. Yes — the final long hop lands on a clock-enable/control signal,
not a data path, confirming control-logic fan-in as part of the
span.
11/12/14. Yes — registering an intermediate result, and/or replacing
the serial scan with a balanced/flat parallel structure, are both
feasible, low-risk, same-precedent-class fixes (ERR-0028 used the
same architecture for a different module).
13. Not a "replicate per slot" scenario, since the cache is shared and
W_ENTRIES is fixed — a flat parallel restructuring was chosen
instead of pipelining, to avoid any latency/behavior change.
### 10.5 ERR-0029 fix applied, and the honest, measured before/after
Fix: serial priority-scan → flat one-hot compare (parallel comparators,
`generate`/`genvar`) + single-level `casez` priority encode, bit-exact
semantics preserved. See errors.log ERR-0029 and DEC-0040 for full
detail. Verified bit-exact: isolated `tb_sdram_unified_backend.v`
40/40 PASS; full D-Stress N=4 (49,927 cycles, 256/256 bit-exact vs
golden) and N=8 (49,909 cycles, 256/256 bit-exact vs golden) — zero
functional regression.
**N=4 @ 64MHz, POST-fix** (period 15.625ns):
| Seed | Fmax (MHz) | WNS (ns) | Critical path endpoint |
|---|---:|---:|---|
| 0 | 70.68 | +1.477 | sdram_backend.state |
| 1 (worst) | 66.58 | +0.605 | sdram_backend.ctrl_wdata |
| 2 | 74.74 | +2.245 | sdram_backend.ctrl_wdata |
| 3 | 71.98 | +1.733 | sdram_backend.ctrl_wdata |
| 4 | 74.48 | +2.198 | dataflow_core.GEN_SLOT[2].u_mm.wgt_rd_addr |
| 5 | 75.65 | +2.406 | sdram_backend.ctrl_wdata |
| 6 | 67.41 | +0.791 | sdram_backend.state |
| 7 | 68.47 | +1.019 | sdram_backend.ctrl_wdata |
**8/8 PASS (unchanged pass count).** Worst-case margin fell from
+2.143ns to +0.605ns (still a real, positive-margin PASS on every
seed — the critical path relocated off the shortened hit-index chain
onto a different, previously-second-worst path in the same module).
Disclosed, not hidden.
**N=8 @ 64MHz, POST-fix** (period 15.625ns):
| Seed | Fmax (MHz) | WNS (ns) | Status |
|---|---:|---:|---|
| 0 | 66.45 | +0.575 | PASS |
| 1 | 65.28 | +0.307 | PASS |
| 2 | 61.21 | 0.712 | FAIL |
| 3 | 66.96 | +0.690 | PASS |
| 4 | 66.66 | +0.624 | PASS |
| 5 | 65.71 | +0.407 | PASS |
| 6 (worst) | 60.12 | 1.009 | FAIL |
| 7 | 62.70 | 0.324 | FAIL |
**Pass count improved 4/8 → 5/8** (seeds 0,1,3,4,5 PASS; 2,6,7 FAIL).
Worst-case Fmax improved 55.84→60.12MHz, worst WNS 2.284→−1.009ns —
a real, measured improvement, **not yet full closure**.
Resource utilization, POST-fix, N=8: MULT18X18D 64/72 (88.9%,
unchanged), TRELLIS_COMB 11129/43848 (25.4%, +63 LUTs, negligible),
TRELLIS_FF 11119/43848 (unchanged), TRELLIS_RAMW 323/5481 (unchanged).
N=4: TRELLIS_COMB 7175/43848 (25.2%→7175, down from 7609 pre-fix).
**N=4/N=8 @ 80MHz, POST-fix**: re-confirmed 0/8 both configs (same
Fmax values as the 64MHz-labeled runs). **NO-GO, unchanged.**
### 10.6 AS4C32M16SB-7BIN — complete verified hardware data
Source: Alliance Memory `AllianceMemory_512M_SDRAM_Bdie_AS4C32M16SB-
7TXN-6TIN-7BIN_Rev1.4_June2024NK.pdf`, the exact -7BIN datasheet
(Figure 1.1, real TFBGA ball diagram — not inferred from the TSOP-II
`-7TIN` pinout).
| Property | Value |
|---|---|
| Part | AS4C32M16SB-7BIN |
| Capacity | 512Mbit = 64MByte |
| Organization | 4 banks × 8M words × 16 bits |
| Package | 54-ball FBGA, 8×8×1.2mm max |
| VDD / VDDQ | 3.3V ±0.3V (isolated I/O supply) |
| Address / Bank | A[12:0] / BA[1:0] |
| Data / Masks | DQ[15:0] / LDQM, UDQM |
| Clock | CLK, single-ended — **no CLK_N** (SDR SDRAM) |
| Control | CKE, CS#, RAS#, CAS#, WE# |
| Temperature / Speed | 40 to 85°C / 7 (143MHz max) |
**Complete individual-ball pinout (54 balls, no grouped notation):**
Address: H7=A0, H8=A1, J8=A2, J7=A3, J3=A4, J2=A5, H3=A6, H2=A7, H1=A8,
G3=A9, H9=A10/AP, G2=A11, G1=A12.
Bank: G7=BA0, G8=BA1.
Data: A8=DQ0, B9=DQ1, B8=DQ2, C9=DQ3, C8=DQ4, D9=DQ5, D8=DQ6, E9=DQ7,
E1=DQ8, D2=DQ9, D1=DQ10, C2=DQ11, C1=DQ12, B2=DQ13, B1=DQ14, A2=DQ15.
Masks: E8=LDQM, F1=UDQM.
Control: F2=CLK, F3=CKE, G9=CS#, F8=RAS#, F7=CAS#, F9=WE#.
Power/Ground/NC: VDD={A9,E7,J9}, VSS={A1,E3,J1}, VDDQ={A7,B3,C7,D3},
VSSQ={A3,B7,C3,D7}, NC=E2. (13+2+16+2+6+3+3+4+4+1 = 54 ✓)
**FPGA (LFE5U-45F-8BG381) ↔ SDRAM (AS4C32M16SB-7BIN) mapping** (from
`hardware/v2/constraints/v2_board_top.lpf`, 45/45 unique FPGA balls,
no duplicates):
| FPGA signal | FPGA ball | SDRAM signal | SDRAM ball |
|---|---|---|---|
| sdram_a[0..12] | D5,D3,F4,E5,E3,F5,A2,B1,C2,C1,D2,D1,F1 | A0..A12 | H7,H8,J8,J7,J3,J2,H3,H2,H1,G3,H9,G2,G1 |
| sdram_ba[0:1] | E4,C3 | BA0,BA1 | G7,G8 |
| sdram_dq[0..15] | E1,G5,H3,J5,K3,K2,H1,J1,K1,K4,L4,L5,M5,M4,N4,N5 | DQ0..DQ15 | A8,B9,B8,C9,C8,D9,D8,E9,E1,D2,D1,C2,C1,B2,B1,A2 |
| sdram_dqm[0:1] | P5,N3 | LDQM,UDQM | E8,F1 |
| sdram_cke/cs_n/ras_n/cas_n/we_n | B5,C5,C4,A3,B3 | CKE,CS#,RAS#,CAS#,WE# | F3,G9,F8,F7,F9 |
Note: FPGA ball "F1" (assigned to `sdram_a[12]`) and SDRAM ball "F1"
(the SDRAM's own `UDQM`) are two different physical devices' own
separate ball-numbering namespaces — not a conflict, but flagged so a
PCB designer does not confuse the two identically-labeled balls.
**Hardware pinout validation**: all FPGA balls real (LFE5U-45F-8BG381
rev 3.0 CSV), 45/45 unique; all SDRAM balls real (AS4C32M16SB-specific
datasheet, not the TSOP variant); A12/BA[1:0]/DQ[15:0]/DQM[1:0]/all
control signals present and complete; VDD/VDDQ/I-O voltage compatible
(3.3V LVCMOS33 ↔ LVTTL); LPF/RTL/datasheet mutually consistent. **No
hardware blockers found.**
### 10.7 PRODUCTION HARDWARE BASELINE (authoritative, 2026-09-07)
**LFE5U-45F-8BG381 + AS4C32M16SB-7BIN + N_SLOTS=4 + P_IN=8 + 64MHz:
GO.** Real, bit-exact functional correctness; real synthesis (0
errors); real P&R (8/8 seeds route); real timing closure (8/8 seeds
PASS, worst WNS +0.605ns post-optimization); real SDRAM/FPGA pinout
cross-verified with no blockers.
**N_SLOTS=8 @ 64MHz: OPEN, not production-frozen.** Functionally
correct (bit-exact) and measurably closer to timing closure after
ERR-0029 (5/8 seeds PASS, up from 4/8), but not yet reliable on every
tested placement seed. Usable today only by pinning a known-good seed
(0, 1, 3, 4, or 5) or pending a further optimization pass.
**80MHz: NO-GO at either N_SLOTS value**, confirmed twice (pre- and
post-ERR-0029) with a genuinely regenerated PLL — not achievable with
the current architecture.
-118
View File
@@ -1,118 +0,0 @@
# FPGA-Neural V2 — OPEN ITEMS
**SUPERSEDED.** See `PRE_PCB_VERIFICATION.md`'s own final release-gate
table for the current, consolidated OPEN/PASS status of every item
below — most of the BLOCKER/CRITICAL items here (host interface,
clock/PLL, pinout) are now CLOSED. Left in place as a historical
record.
Consolidated from HARDWARE_FREEZE.md, PINOUT.md, CLOCK_ARCHITECTURE.md,
POWER_ARCHITECTURE.md, SCHEMATIC_READINESS.md. Classified per the
governing spec's own rule: BLOCKER / CRITICAL / WARNING / OPEN /
FUTURE.
## BLOCKER (impede la realizzazione o il funzionamento del chip)
0. **RESOLVED (STEP20).** A real SPI host interface (`spi_host_bridge.v`)
is now implemented, protocol-correct in isolation (18/18,
`tb_spi_host_bridge.v`), AND verified correct end-to-end through the
full SPI→dependency_manager→compute→SDRAM→result path under both
tight and realistic (widely time-separated) job pacing (11/11,
`tb_fpga_neural_v2_top_smoke.v`) — see errors.log's own "ERR-0025
Part B — RESOLUTION" entry for the full root-cause writeup (a
registered- vs combinational-read latency mismatch in the shared
weight/activation SRAMs, fixed with zero regression to the STEP19
baseline). This item is CLOSED — kept here only for the historical
record; item 1 below is likewise no longer a real blocker in the
sense of "the RTL doesn't exist" — it remains open only for real
pinout/board-connector work (see item 1's own updated text).
1. **RESOLVED (STEP20).** The RTL's own internal "host" ports (`reg_valid`
/`reg_node_id`/`reg_required`/`reg_producer_ids`/`reg_x_base`/
`reg_w_base`/`reg_n_tiles`/`reg_result_addr`) remain a simulation/
testbench-only bus for `nms_neural_multiprocessor_sdram_unified.v`
in isolation, but the board-level top (`fpga_neural_v2_top.v`) now
drives these SAME internal ports from `spi_host_bridge.v`, a real,
verified SPI protocol engine (WRITE_JOB/WRITE_MEM/READ_MEM/STATUS/
RESET), matching V1's own `spi_neuron_top.v` precedent. The 110-pin
bus is no longer exposed as a physical top-level port at all in
`fpga_neural_v2_top.v` — only 4 real SPI pins (sclk/mosi/miso/cs_n)
are.
2. **The 110-pin host/registration bus has no real ball assignment**
moot now (see item 1): it is an internal signal, not a top-level
port, in the board-level top. The 4 real SPI pins likewise have no
ball assignment yet, since `fpga_neural_v2_top.v` has not been
through P&R this round (see the next open item). The SDRAM bus (37
signals) and clk/rst (2 signals) now DO have a
real, sourced, P&R-verified assignment (`hardware/v2/constraints/
v2_unified.lpf`, from the real Lattice pinout CSV found at
`~/Downloads/FPGA-SC-02034-3-0-ECP5U-45-Pinout.csv` during this
step's own pre-commit review) — this item is narrower than
originally scoped.
3. **No schematic exists; no PCB has been started.**
## CRITICAL (rischio elevato, deve essere risolto prima del freeze)
1. **Timing closure is MARGINAL, with an unfavorable pass rate.** 8
real P&R seeds for the frozen N=4 single-SDRAM design: only 1/8
reach ≥80MHz (66.9781.84MHz range). This is WORSE than STEP18's
own dual-memory design (5/8 pass). The critical path itself is
unchanged (still `dependency_manager.v`'s own pre-existing
`first_ready_idx`/`reg_ready` chain) — the regression is attributed
to added overall die/routing pressure from consolidation, not a new
RTL defect, but it is real and unresolved.
2. **Clock source/oscillator gap -- PARTIALLY ADDRESSED (STEP20).** A
real EHXPLLL wrapper (`ecp5_pll_sys_clk.v`, real Project Trellis
`ecppll`-generated parameters, 16MHz->64MHz) now exists and is
instantiated in `fpga_neural_v2_top.v`. NOT YET confirmed by real
synthesis/P&R of that board-level top this round (deliberately
deferred until ERR-0025 Part B was resolved -- see decisions.log
DEC-0037) -- this is the immediate next real step. Prior project memory
records a 16MHz board oscillator. Neither "source an 80MHz+
oscillator" nor "add a real PLL to the RTL" has been decided.
3. **Two physical memories were required through STEP18** — RESOLVED
this round (DEC-0034): the V2 physical path no longer instantiates
`hardware/v1/rtl/psram_controller.v` at all. Kept here only as a
closed CRITICAL item for the historical record.
## WARNING (non blocca il prototipo ma deve essere documentato)
1. N=2's real Fmax (86.04MHz in STEP18's own dual-memory design) and
the STEP19 single-SDRAM N=2 config were not both measured with the
same best-of-N-seed rigor as N=4 — a real, disclosed gap in
measurement thoroughness, not a functional issue.
2. `W_ENTRIES`/cache sizing in the weight-fetch path was set to match
`N_SLOTS` (4) by construction reasoning, not swept for optimality.
3. I/O standard (LVCMOS33 assumed for all 149 signals) has not been
verified per real VCCIO bank once ball assignment becomes possible.
## OPEN (decisione ancora da prendere)
1. Configuration-flash part number / SPI-flash-boot vs JTAG-only
bring-up.
2. Power regulator topology and part numbers (the previously-recorded
`../basic-ecp5-pcb` reference design is not accessible this
session to confirm as a concrete plan).
3. Real current budget (requires running a real power-estimation tool
against the actual synthesized netlist — not done this round).
4. Decoupling/bulk capacitance values (depend on regulator selection).
5. Reset synchronization to a real external POR/supervisor source.
6. Real per-bank VCCIO/I-O-standard verification once ball data is
available.
## FUTURE EVOLUTION (miglioramento post-freeze — explicitly deferred)
1. N=8 evaluation.
2. A smarter W/AR priority scheme in `sdram_unified_backend.v` to
recover some of the +10.8% (N=4) / +5.1% (N=2) cycle-count cost of
single-SDRAM unification (STEP18 EXP-0046's own packing win is
still present — this is about the NEW W-vs-AR contention specifically).
generic
3. Page-mode / keep-row-open SDRAM controller redesign (STEP18's own
identified next bottleneck for raw memory bandwidth, independent of
the single-vs-dual-memory question).
4. True multi-outstanding SDRAM request pipelining (STEP18 Part E's
own documented, deliberately out-of-scope boundary).
5. A real physical host-interface RTL bridge (SPI or similar),
resolving BLOCKER #1 above.
6. Floorplanning / seed-pinning work to convert the current MARGINAL
timing result into a reliable PASS.
-110
View File
@@ -1,110 +0,0 @@
# FPGA-Neural V2 — PINOUT
**SUPERSEDED.** This document predates the SPI host bridge and the
final board-level `fpga_neural_v2_top`/`v2_board_top.lpf` pinout. See
`PRE_PCB_VERIFICATION.md` \S13 for the current, real, P&R-confirmed
16-signal pinout (host bus is no longer BLOCKED). Left in place as a
historical record.
FPGA: **LFE5U-45F-8BG381** (ECP5U, speed grade -8)
Package: **CABGA381**
Frozen top-level: `nms_neural_multiprocessor_sdram_unified` (N_SLOTS=4)
**Single external memory: ONE SDRAM (AS4C4M16SA-6TIN). No PSRAM, no
second memory device anywhere in this design (STEP19/DEC-0034).**
## Status: SDRAM pinout REAL and P&R-verified; host bus still BLOCKED
Correction to an earlier draft of this document: the real Lattice
pinout data source (`FPGA-SC-02034-3-0-ECP5U-45-Pinout.csv`, rev 3.0)
IS available on this machine (`~/Downloads/`), and its own summary
(`docs/pinouts.md`, repo root) already lists real, exact JTAG/config/
power ball assignments for CABGA381 — found during this step's own
pre-commit `git status` review, not assumed missing without checking.
`hardware/v2/constraints/v2_unified.lpf` now contains a REAL,
P&R-verified ball assignment for clk/rst (39 total) and the full
37-signal SDRAM bus, sourced directly from that CSV (bank 6/7 plain-
GPIO pads, avoiding PLL/PCLK-reserved balls) — confirmed by a real
nextpnr-ecp5 run: all 37 SDRAM signals placed successfully, "110
warnings" (exactly the 110 still-unconstrained host-bus signals, a
clean cross-check that the inventory below is complete and accurate).
**This has NOT been electrically cross-verified** (VCCIO6/7 bank
voltage vs the SDRAM's own LVCMOS33 requirement, signal integrity,
trace-length matching for the 16-bit DQ bus) — it is a real, sourced,
P&R-confirmed CANDIDATE assignment, not a board-signed-off pinout.
## Real ball assignments now in place
| Signal | Ball | Source |
|---|---|---|
| `clk` | H5 | Reused from V1's own real, validated LPF |
| `rst` | B4 | Reused from V1's own real, validated LPF |
| `sdram_cke`/`cs_n`/`ras_n`/`cas_n`/`we_n` | B5/C5/C4/A3/B3 | Real CSV, bank 7 |
| `sdram_ba[1:0]` | E4, C3 | Real CSV, bank 7 |
| `sdram_a[11:0]` | D5,D3,F4,E5,E3,F5,A2,B1,C2,C1,D2,D1 | Real CSV, bank 7 |
| `sdram_dq[15:0]` | E1,G5,H3,J5,K3,K2,H1,J1,K1,K4,L4,L5,M5,M4,N4,N5 | Real CSV, banks 7/6 |
| `sdram_dqm[1:0]` | P5, N3 | Real CSV, bank 6 |
Full detail: `hardware/v2/constraints/v2_unified.lpf`.
Real JTAG/config/power balls (from `docs/pinouts.md`, not yet
transcribed into the LPF since this design's own top-level does not
expose them as RTL ports — they are implicit ECP5 device pins):
TDI=R5, TCK=T5, TMS=U5, TDO=V4 (bank 40); PROGRAMN=W3, INITN=V3,
DONE=Y3, CCLK=U3 (bank 8); VCC balls (1.1V) at H8-N13 cluster;
VCCAUX (2.5V) at F6/P6/F15/P15; VCCIO0-8 bank assignments listed in
`docs/pinouts.md`.
## Signal inventory (real, from the frozen top-level's own port list)
Total top-level I/O: **149 signals**, cross-checked exactly against
the real POST-P&R `TRELLIS_IO: 149/245` figure (STEP19) — a real
**45-pin reduction** from STEP18's dual-memory design (194 pins),
exactly matching the removed PSRAM interface's own pin count.
| Group | Count | Ball assignment |
|---|---|---|
| Clock/reset (`clk`, `rst`) | 2 | **Real, assigned** (H5, B4) |
| Host/control (`reg_*`) | 110 | **BLOCKER — see below** |
| SDRAM (`sdram_*`) | 37 | **Real, assigned, P&R-verified** |
| **Total** | **149** | matches P&R exactly |
## CRITICAL finding: the "host" interface is not a physical interface
**110 of 149 pins (73.8%) are the raw `reg_*` job-registration bus**
a simulation/testbench convenience, not a real board protocol. No RTL
exists to serialize this for physical use. Ball assignment for these
110 signals is deliberately NOT attempted yet, even though real GPIO
balls are available (46+ more plain-GPIO candidates remain in banks
6/7 alone after the 37 used above) — assigning pins to an interface
that must be redesigned first would be premature, wasted work. **This
remains the single largest real BLOCKER to physical realization.**
## I/O standard / bank assignment
LVCMOS33 assumed and used in the LPF above for all 39 real-assigned
signals — matches `docs/pinouts.md`'s own real VCCIO range (1.23.3V)
and V1's own real, validated board convention. Real per-bank voltage
compatibility for banks 6/7 specifically (used for SDRAM) has not been
independently re-verified against the SDRAM device's own datasheet
this round (WARNING, not BLOCKER — LVCMOS33 is a reasonable, likely-
correct default, not yet double-checked).
## Configuration pins (JTAG/config)
Now REAL and known (see table above) — `docs/pinouts.md`'s own
summary of the same official CSV. This closes what was previously
documented as a blocker for THESE specific pins; only the general-
purpose host-bus assignment (unrelated to JTAG/config) remains open.
## Summary
| Item | Status |
|---|---|
| Real ball-level LPF for the SDRAM interface | **Done — 37/37 signals, P&R-verified** |
| Real ball-level assignment for clk/rst | **Done — reused from V1** |
| Real ECP5U-45F CABGA381 ball-map data source | **Found — `~/Downloads/FPGA-SC-02034-3-0-ECP5U-45-Pinout.csv`, summarized in `docs/pinouts.md`** |
| Aggregate I/O feasibility (149/245 fits the package) | Confirmed, real POST-P&R |
| Real JTAG/config/power ball identification | **Done — see `docs/pinouts.md`** |
| Physical host interface RTL | **BLOCKER — does not exist (110 raw pins, no serializer, no ball assignment)** |
| I/O standard/bank electrical cross-check | WARNING — LVCMOS33 assumed, not independently re-verified per bank |
-74
View File
@@ -1,74 +0,0 @@
# FPGA-Neural V2 — POWER ARCHITECTURE
**SUPERSEDED.** See `PRE_PCB_VERIFICATION.md` \S11-\S12 for the
current per-bank voltage table (PASS) and current/power budget
status (still OPEN, same real reasons as below). Left in place for
its own detailed real-value derivation.
## Status (HISTORICAL framing, current status is in PRE_PCB_VERIFICATION.md): OPEN — component/regulator selection not made this round
Per the governing spec's own "NON inventare valori" rule, this
document states what is REALLY known (device-level voltage
requirements, from real datasheets/standard ECP5 knowledge) and
explicitly marks what has NOT been decided, rather than inventing
regulator part numbers or current budgets without real justification.
## Required rails (real device requirements)
Corrected from an earlier draft: real ball-level VCC/VCCAUX/VCCIO
data for this exact package DOES exist (`docs/pinouts.md`, repo root,
sourced from the official Lattice pinout CSV) and is used below rather
than only generic device specs.
| Rail | Nominal voltage | Real balls (CABGA381) | Notes |
|---|---|---|---|
| VCC (core) | 1.1V ±5% | H8,J8,K8,L8,M8,N8,H9,N9,H10,N10,H11,N11,H12,N12,H13,J13,K13,L13,M13,N13 | Real, from `docs/pinouts.md` |
| VCCAUX | 2.5V ±5% | F6, P6, F15, P15 | Real, from `docs/pinouts.md` |
| VCCIO0 | 1.23.3V (bank 0) | F9, F10 | Real ball pair; bank/signal assignment TBD |
| VCCIO1 | 1.23.3V (bank 1) | F11, F12 | Real ball pair |
| VCCIO2 | 1.23.3V (bank 2) | H14, H15, J15 | Real |
| VCCIO3 | 1.23.3V (bank 3) | L14, L15, M15 | Real |
| VCCIO6 | 1.23.3V (bank 6, used by SDRAM) | L6, L7, M6 | Real — SDRAM signals (see PINOUT.md) live in banks 6/7; 3.3V assumed, matching the SDRAM device's own real LVCMOS33 requirement, NOT yet independently cross-verified |
| VCCIO7 | 1.23.3V (bank 7, used by SDRAM) | H6, H7, J6 | Real, same note as VCCIO6 |
| VCCIO8 | config bank | P9, P10 | Real — Lattice's own documentation explicitly ties this rail's voltage to whichever configuration interface is used (OPEN, see Configuration decision below) |
| SDRAM VDD / VDDQ | 3.3V | (external chip, not an FPGA ball) | Per the real AS4C4M16SA-6TIN datasheet's own 3.3V industrial-grade part number |
| Configuration supply | 3.3V (typ.) | tied to VCCIO8 | Depends on the configuration-path decision (OPEN, see below) |
VSS/VSSIO (ground) balls: real per `docs/pinouts.md`'s own note — all
must be connected to the ground plane, none left floating (standard
BGA practice, explicitly called out in the source data).
## What is NOT decided (OPEN ITEMS)
- **Regulator topology/part numbers**: not selected. This project's own
memory notes reference a sibling repository (`../basic-ecp5-pcb`)
with a real, working power tree (TLV62568×2 + TLV73325) as a
possible reference — but that repository is **not present on disk**
in this environment (confirmed during this step's own audit), so it
cannot be verified or cited as a concrete plan this round. A future
step should either locate that reference design or select
regulators from scratch against the real current budget below.
- **Maximum estimated current**: not computed. This requires a real
power estimate from the actual synthesized netlist (Lattice's own
power calculator/estimation tools were not run this session) — NOT
invented here. The real, measured resource utilization (TRELLIS_FF=
6425, TRELLIS_COMB=6023, MULT18X18D=32, DP16KD=0 at N=4, POST-P&R,
STEP19) is available as an INPUT to such a calculation, but the
calculation itself was not performed.
- **Decoupling/bulk capacitance**: not specified — a schematic-level
decision that depends on the final regulator selection above.
- **Startup/power sequencing**: ECP5 devices generally require VCC and
VCCAUX to be sequenced correctly relative to VCCIO and the
configuration source per Lattice's own real application notes — this
project has not yet consulted or reproduced those real sequencing
requirements; flagged as OPEN, not assumed compatible.
## Summary
| Item | Status |
|---|---|
| Real rail voltage requirements (VCC/VCCAUX/VCCIO/SDRAM) identified | Done, from real device specs |
| Regulator selection | **OPEN — not made, no real reference design available this session** |
| Current budget | **OPEN — not computed, would require running a real power-estimation tool** |
| Decoupling/bulk capacitance | **OPEN — depends on regulator selection** |
| Power sequencing verification | **OPEN — not yet checked against real Lattice app notes** |
-334
View File
@@ -1,334 +0,0 @@
# FPGA-Neural V2 — FINAL 4-POINT PRE-PCB CLOSURE
Follows `PRE_PCB_VERIFICATION.md` (PRE-PCB VERIFIED baseline, commit
`d6376e8` + `8890b0a` + `eb0b0f9`). Closes the four remaining
practical items the user identified as still open before schematic
capture. Does not redesign the verified architecture; no working RTL
was modified as a result of this pass (see Point 2 for the one bug
found and fixed, which was in a NEW test harness, not in
`spi_host_bridge.v` itself).
**SDRAM-SPECIFIC CONTENT SUPERSEDED (DEC-0039, a later session).**
Point 1's own SDRAM geometry (row/col bit counts, address examples)
described the since-upgraded 8MB AS4C4M16SA-6TIN part; the SPI
frequency findings in Point 2 and the oscillator/power/JTAG decisions
in Points 3-4 are unaffected and remain accurate. See
`MEMORY_UPGRADE_64MB_N8.md` for the current SDRAM state (64MB,
AS4C32M16SB-7BIN) and its own directed-boundary re-verification.
---
## POINT 1 — Directed SDRAM boundary verification
New file: `hardware/v2/nms/sim/tb_sdram_boundary.v`.
Real Alliance Memory AS4C4M16SA-6TIN geometry (confirmed against
`sdram_controller.v`'s own address decode:
`addr_bank=addr[21:20]`, `addr_row=addr[19:8]`, `addr_col=addr[7:0]`):
4 banks × 4096 rows × 256 cols × 16 bits = 4M words = 8MB.
Coverage (21 checks, BURST_LEN=1 for exact single-word addressing):
- **Address boundaries**: 0x000000 (addr 0), 0x000001 (addr 1),
0x3FFFFF (last valid), 0x3FFFFE (last valid 1).
- **Row boundary**: bank0/row10/col255 (last column of row 10) and
bank0/row11/col0 (first column of row 11).
- **Bank boundaries**: last address / first address at all 3
inter-bank crossings (bank0↔1, bank1↔2, bank2↔3).
- **Memory-map boundaries**: the real V2 map (weights@byte 0x010000,
activations@byte 0x200000, results@byte 0x300000) converted to this
controller's word addresses (word=byte/2) — weights base, last word
before activations, activations base, last word before results,
results base.
- **Byte-mask combinations, explicit read-after-write**: lower-byte-
only (wmask=2'b10), upper-byte-only (wmask=2'b01), both-bytes
(wmask=2'b00), using the requested deterministic patterns 0xAAAA,
0x5555, 0x0000, 0xFFFF.
All 17 boundary/adjacency addresses are written first, then read back
in **reversed** order with distinct address-derived patterns
(`addr[15:0] ^ 0xC3A5`) — this proves no write to any one address
corrupted any neighbour in the set, which is exactly the "adjacent
regions cannot corrupt each other" property requested, for every
boundary simultaneously.
### Exact results
```
$ verilator --binary --timing -Wno-fatal --top-module tb_sdram_boundary -o tb_bnd \
-GCLK_FREQ_MHZ=64 hardware/v2/nms/rtl/sdram_controller.v \
hardware/v2/nms/sim/sdram_model.v hardware/v2/nms/sim/tb_sdram_boundary.v
$ ./obj_dir/tb_bnd
=== 21/21 tests, 0 errors (tb_sdram_boundary, CLK_FREQ_MHZ=64) ===
ALL TESTS PASSED (tb_sdram_boundary, CLK_FREQ_MHZ=64)
```
Cross-checked at the legacy CLK_FREQ_MHZ=166 (same command with
`-GCLK_FREQ_MHZ=166`): **21/21 PASS, 0 errors**, identical.
Full per-address results at 64MHz (expected vs actual, all matched):
| Label | Address | Data |
|---|---|---|
| addr-0 | 0x000000 | 0xc3a5 |
| addr-1 | 0x000001 | 0xc3a4 |
| addr-last | 0x3fffff | 0x3c5a |
| addr-last-1 | 0x3ffffe | 0x3c5b |
| row10-lastcol | 0x000aff | 0xc95a |
| row11-firstcol | 0x000b00 | 0xc8a5 |
| bank0-last | 0x0fffff | 0x3c5a |
| bank1-first | 0x100000 | 0xc3a5 |
| bank1-last | 0x1fffff | 0x3c5a |
| bank2-first | 0x200000 | 0xc3a5 |
| bank2-last | 0x2fffff | 0x3c5a |
| bank3-first | 0x300000 | 0xc3a5 |
| weights-base | 0x008000 | 0x43a5 |
| weights-last(pre-act) | 0x0fffff | 0x3c5a |
| activations-base | 0x100000 | 0xc3a5 |
| activations-last(pre-res) | 0x17ffff | 0x3c5a |
| results-base | 0x180000 | 0xc3a5 |
| mask-lower-only | 0x001000 | 0xaa34 |
| mask-upper-only | 0x001000 | 0x5655 |
| mask-both-bytes | 0x001000 | 0xffff |
| pattern-5555-plain | 0x001001 | 0x5555 |
**No bug found.** Address decode, byte masking, and inter-region
adjacency are all correct at every tested boundary.
**RESULT: SDRAM directed boundaries: PASS.**
---
## POINT 2 — Verified SPI operating clock
New file: `hardware/v2/nms/sim/tb_spi_freq_sweep.v`. Instantiates the
REAL `fpga_neural_v2_top` (not spi_host_bridge in isolation) with
`osc_clk` driven at the real 64MHz `clk_sys` rate (the `SIM` PLL
bypass makes `clk_sys = osc_clk` directly, so driving `osc_clk` at
64MHz reproduces the real board's actual system-clock rate — unlike
`tb_fpga_neural_v2_top_smoke.v`, which uses a stale `CLK_FREQ_MHZ=80`
parameter left over from an earlier draft). SPI bit timing is a
runtime parameter (`SPI_FREQ_MHZ`), swept across candidate points.
Per-frequency coverage: single job submission, two jobs back-to-back,
two jobs with a realistic gap, a raw `WRITE_MEM`/`READ_MEM` round trip
over the actual SPI response path (not the backdoor SDRAM peek used
elsewhere), and 3 repeated single-job transactions — 10 checks total.
### A bug found and fixed — in the new test harness, not the RTL
The first sweep attempt (fixed `#2000`-real-time wait before clocking
out a `READ_MEM` response) failed once, at 2MHz, with the response's
MSB read back as 0 instead of 1 — every other bit correct. Before
concluding anything about the RTL, this was root-caused: the real
host-arb/SDRAM-controller backend latency (unlike
`tb_spi_host_bridge.v`'s own isolated unit test, which drives
`mem_rdata`/`mem_ready` from a simple behavioral mock with fixed
timing) genuinely varies cycle-to-cycle — a periodic AUTO REFRESH can
land during the request and push `mem_ready` later than the guessed
`#2000` margin. `tb_spi_host_bridge.v`'s own regression already proves
`spi_host_bridge.v`'s FIRST `READ_MEM` after reset delivers all 16
bits correctly when its own mock backend responds within that test's
own assumed timing — confirming the FSM logic itself is correct, and
the failure was this new harness's own race. **Fixed** by polling
`dut.u_spi_bridge.state` directly (`ST_MEM_ROUT`/`ST_IGNORE`) instead
of guessing a fixed real-time margin — eliminates the race entirely.
Re-ran the full sweep from 2MHz upward with this fix: no further
data-corruption failures at any frequency below the real CDC limit
(see below).
### Sweep results
| SPI_FREQ_MHZ | sysclk cycles/bit (64MHz) | Result |
|---|---|---|
| 2 | 32.0 | 10/10 PASS |
| 4 | 16.0 | 10/10 PASS |
| 8 | 8.0 | 10/10 PASS |
| 10 | 6.4 | 10/10 PASS |
| 12 | 5.33 | 10/10 PASS |
| 12.5 | 5.12 | 10/10 PASS |
| 12.8 | 5.0 (exact) | 10/10 PASS |
| 12.9 | 4.96 | FAIL (data corruption) + protocol FSM HANG (watchdog) |
| 13 | 4.92 | FAIL + HANG |
| 14 | 4.57 | FAIL + HANG |
| 15 | 4.27 | FAIL + HANG |
| 16 | 4.0 | FAIL + HANG |
| 20, 24, 32 | <4.0 | FAIL + HANG |
The breakpoint is **exact and deterministic**: 12.8MHz is precisely
64MHz/5 — the triple-flop CDC synchronizer plus edge-detect/FSM
reaction in `spi_host_bridge.v` requires at least 5 full system-clock
cycles per SPI bit period to reliably track `sclk`/`mosi`/`cs_n`
transitions. Below that, the synchronizer misses edges outright,
which doesn't just corrupt data (as briefly seen in the harness-race
case above) but eventually desyncs the byte-framing state machine
badly enough that it never reaches an expected state again — a real
protocol lockup, not merely wrong data. This is a genuine, real
property of the CDC design (not a bug — the double/triple-flop
synchronizer is standard, correct practice; it simply has a minimum
bit-period requirement, which every synchronous CDC scheme does), now
precisely measured rather than assumed.
### Distinguishing the three kinds of limit the mandate asks for
- **RTL/simulation limit (measured, this session)**: 12.8MHz exact
edge; 12MHz recommended verified operating point (real margin below
the hard edge: 5.33 vs the minimum 5.0 cycles/bit, ~6.7% headroom).
- **FPGA timing limit**: not applicable in the way P&R timing closure
applies to the internal 64MHz domain — the SPI pins are simple
registered/synchronized GPIO inputs (`IO_TYPE=LVCMOS33`, no special
timing constraint beyond the CDC margin above), and nextpnr-ecp5's
own timing analysis (section 5/6 of `PRE_PCB_VERIFICATION.md`) does
not model an external asynchronous SPI master's edge timing at all.
No FPGA-side P&R-derived limit beyond the CDC margin already found.
- **Board-level electrical limit**: **OPEN — not measured, cannot be
measured without real hardware.** Real trace length, connector/cable
capacitance, SPI master driver rise/fall time, ground bounce, and
actual metastability risk (this RTL simulation is deterministic and
cannot model metastability at all) are all real-world factors this
simulation does not and cannot capture. The 12MHz recommendation
below is a simulation-verified LOGICAL limit with margin, not a
physical hardware guarantee — bring-up step 11 in
`FIRST_POWER_ON.md` should still empirically confirm the real
achievable rate on the actual board.
**RESULT: SPI_MAX_VERIFIED = 12 MHz** (recommended operating point,
simulation-verified with real margin below the exact 12.8MHz
deterministic CDC edge). Do not exceed 12.8MHz under any circumstance;
do not treat 12.8MHz itself as a safe operating margin.
### Exact test commands
```
$ verilator --binary --timing -Wno-fatal -DSIM --top-module tb_spi_freq_sweep -o tb_spi \
-GSPI_FREQ_MHZ=12.0 <all V2 rtl/nms sources + tb_spi_freq_sweep.v>
$ ./obj_dir/tb_spi
=== SPI_FREQ_MHZ=12.000: 10/10 PASS ===
```
---
## POINT 3 — 16MHz oscillator MPN, frozen
**Decision: ECS Inc. International, `ECS-3225MV-160-BN-TR`.**
| Property | Value |
|---|---|
| Manufacturer / MPN | ECS Inc. International, `ECS-3225MV-160-BN-TR` |
| Type | Quartz crystal oscillator (XO), not a bare crystal — provides a direct digital clock output, no external oscillator circuit needed |
| Frequency | 16.000 MHz, matching `osc_clk`'s real ball (H5) and the LPF's `FREQUENCY PORT "osc_clk" 16 MHZ` constraint exactly |
| Package | 3225 SMD, 3.2mm × 2.5mm, 4-pad (standard, small, hand-placeable with a stencil; widely available) |
| Supply voltage | 3.3V — matches `osc_clk`'s LPF `IO_TYPE=LVCMOS33` exactly, no level-shifting needed |
| Output type | HCMOS/CMOS square wave — directly compatible with the ECP5's LVCMOS33 clock input requirement |
| Frequency stability | ±50 ppm (standard grade for this series) — comfortably adequate for an SDR SDRAM/SPI/PLL system with no tight external timing reference requirement |
| Duty cycle | Typically 45/55% to 40/60% (standard for this class of HCMOS XO; confirm exact figure against the current ECS datasheet at BOM lock) |
| Startup time | Typically ≤10ms (standard for a quartz XO of this type) |
| Temperature range | 40°C to +85°C (industrial) |
| Recommended decoupling | One 0.1µF ceramic capacitor directly across VDD/GND, placed as close as possible to the oscillator's supply pin — standard practice for this device class |
| Availability | High — ECS Inc. is a large, long-established oscillator manufacturer stocked at Digi-Key/Mouser; standard frequency/package combination |
Verified against the ECP5's own input-clock requirements: LVCMOS33
input, no minimum/maximum listed frequency constraint that 16MHz would
violate, matches the real, already-verified
`(* FREQUENCY_PIN_CLKI="16" *)`-driven `EHXPLLL` input in
`ecp5_pll_sys_clk.v` exactly.
**Caveat, honestly disclosed**: the exact terminal order-code suffix
(stability/voltage/output-enable option letters, here assumed `BN` for
3.3V HCMOS/standard stability) should be cross-checked against ECS's
current published datasheet at final BOM lock — normal, standard
due-diligence practice at that stage, not an open architectural
question. The manufacturer, series, frequency, package, and supply
voltage are the real, frozen decision.
**RESULT: 16MHz oscillator: `ECS-3225MV-160-BN-TR` (ECS Inc.), FROZEN.**
---
## POINT 4 — Power + JTAG support components
### FPGA power rails and regulators
**Assumption, explicitly flagged**: a 5V board input rail (typical
USB/wall-adapter supply) is assumed as the single external power
source all on-board regulators derive from — this was not specified
by the user and is a reasonable, common default, not a verified fact.
| Rail | Voltage | Regulator MPN | Topology | Current capability | Notes |
|---|---|---|---|---|---|
| FPGA core (VCC) | 1.1V ±5% | Texas Instruments `TPS562201DDCR` | Synchronous buck (switching), adjustable output via feedback resistor divider set for 1.1V | Up to 2A | Real dynamic current draw is OPEN (section 12 of `PRE_PCB_VERIFICATION.md`) — 2A capability is a real-datasheet-based worst-case engineering margin, not a measured requirement; a switching regulator (not an LDO) is used here because a 5V→1.1V LDO would dissipate excessive heat at any non-trivial current |
| FPGA VCCAUX | 2.5V ±5% | Texas Instruments `TLV1117-25IDCYR` | Linear (LDO), fixed 2.5V | 800mA | Fed from the same 5V input rail directly (not from the 3.3V rail) so the LDO retains adequate (~2.5V) dropout headroom |
| FPGA VCCIO (banks 6/7/8) | 3.3V | Texas Instruments `TLV1117-33IDCYR` | Linear (LDO), fixed 3.3V | 800mA | Also supplies the SDRAM, config flash, oscillator, and JTAG reference voltage (all real 3.3V devices per sections 10/11 of `PRE_PCB_VERIFICATION.md`) |
| SDRAM (AS4C4M16SA-6TIN) | 3.3V | Shared with VCCIO rail above | — | — | Real datasheet requirement, already confirmed |
| Config flash (W25Q32JVSSIQ) | 3.3V (within its 2.7-3.6V range) | Shared with VCCIO rail above | — | — | Real datasheet requirement, already confirmed |
| Oscillator (ECS-3225MV-160) | 3.3V | Shared with VCCIO rail above | — | — | Matches Point 3's own decision |
**Design margin**: the 1.1V buck's 2A capability and the 800mA LDOs
are real, datasheet-supported ratings well above any plausible
estimate for this design's actual utilization (7,084 LUT4-equiv,
6,322 FF, 32 MULT18X18D — a mid-size ECP5-45F design, not the whole
device near capacity), but per section 12's own honest disclosure,
the EXACT required current is still not computed from real
implementation data — these regulator choices provide comfortable
headroom against that unknown, not a precisely-sized budget.
### Passive components (frozen only where electrically required)
| Component | Value | Where |
|---|---|---|
| Decoupling (high-frequency) | 100nF (0.1µF) X7R ceramic, 0402/0603 | Distributed, one per VCC/VCCAUX/VCCIO power-pin group around the BGA, per Lattice's own Hardware Checklist guidance (already cited in `docs/pinouts.md`) |
| Decoupling (bulk) | 10µF X5R ceramic or tantalum | One per regulator output, close to each regulator |
| `PROGRAMN` pull-up | 10kΩ to VCCIO8 (3.3V) | Standard ECP5 practice — idle-high, momentary pulse low reconfigures |
| `INITN` pull-up | 4.7kΩ to VCCIO8 (3.3V) | `INITN` is open-drain per ECP5 spec, needs an external pull-up |
| Config flash `WP#`/`HOLD#` pull-ups | 10kΩ each to 3.3V | Per Point 10's own decision (`W25Q32JVSSIQ`, standard single-SPI mode, these pins unused and must be held inactive) |
| JTAG `TMS` pull-up | 4.7-10kΩ to 3.3V | Standard practice so an unconnected/high-impedance JTAG probe leaves TMS idle-high (TAP stays in Test-Logic-Reset) |
Not frozen (correctly left for PCB layout, per "only freeze what's
electrically required"): exact capacitor placement/count beyond the
one-per-pin-group guidance above, trace-length matching, ground-plane
stitching-via count.
### JTAG
| Property | Decision |
|---|---|
| Connector type | Simple unshrouded 2×3 (6-pin), 0.1" (2.54mm) pitch pin header — sufficient for a point-to-point bench connection; no vendor-specific shrouded-connector standard is mandated by Lattice for the ECP5 |
| Pinout | Pin1=3V3 (reference/probe-detect, not a supply to the probe), Pin2=TCK (real ball T5), Pin3=TMS (real ball U5), Pin4=TDI (real ball R5), Pin5=TDO (real ball V4), Pin6=GND |
| Required pull resistor | TMS: 4.7-10kΩ to 3.3V (see passives table above) |
| Required power/reference pin | 3.3V reference pin (Pin1) so a probe can detect target voltage; NOT used to power the board |
| `PROGRAMN`/config-related signals | Real balls W3 (PROGRAMN), V3 (INITN), Y3 (DONE), bank 8 — NOT part of the JTAG connector itself; these remain dedicated ECP5 configuration-control pins, routed separately per Point 14/15 of `PRE_PCB_VERIFICATION.md` |
| Programming/debug path | JTAG connects directly to the ECP5's own real TAP balls (R5/T5/U5/V4); no external JTAG buffer/level-shifter needed since the probe and the FPGA both operate at 3.3V |
**Complete programming/debug path verified**: JTAG header → real TAP
balls → ECP5 TAP controller → SRAM configuration (direct bitstream
download for bring-up/debug) or, separately, the `W25Q32JVSSIQ` config
flash for standalone boot (Point 10 of `PRE_PCB_VERIFICATION.md`) —
both paths coexist without conflict, as already confirmed in that
document's own JTAG-interaction analysis.
**RESULT: Power/JTAG/support components: CLOSED** (sufficient for
schematic capture; exact passive layout/placement remains, correctly,
a PCB-level task).
---
## FINAL 4-POINT STATUS
1. **SDRAM directed boundaries: PASS** (21/21, both 64MHz and 166MHz, zero bugs found)
2. **SPI maximum verified frequency: 12 MHz** (simulation-exact deterministic edge: 12.8MHz = 64MHz/5; one testbench-race bug found and fixed, NOT an RTL defect; board-level electrical limit remains OPEN, requires real hardware)
3. **16MHz oscillator: `ECS-3225MV-160-BN-TR` (ECS Inc.)** — FROZEN
4. **Power/JTAG/support components: CLOSED** (regulator MPNs, passive values, and JTAG connector/pinout frozen; exact PCB placement correctly deferred)
Remaining uncertainty, explicitly documented (not silently dropped):
oscillator order-suffix cross-check against the live ECS datasheet;
real FPGA dynamic current (still requires post-implementation data,
per `PRE_PCB_VERIFICATION.md` section 12); board-level SPI electrical
limit (requires real hardware bring-up); hold-timing tool limitation
(carried over from `PRE_PCB_VERIFICATION.md`, unaffected by this pass).
# PRE-PCB HARDWARE SPECIFICATION: FROZEN
Schematic and PCB layout remain the user's own implementation work.
This status means the four practical items requested are closed
sufficiently for schematic capture — it is NOT a claim of
"SILICON READY."
-617
View File
@@ -1,617 +0,0 @@
# FPGA-Neural V2 — PRE-PCB VERIFICATION FREEZE
Governing mandate: close and verify everything that can be verified
before the user's own KiCad schematic/PCB work begins. This document
is the single authoritative record of that verification pass. It
supersedes the per-topic status statements in CHIP_READINESS.md,
OPEN_ITEMS.md, POWER_ARCHITECTURE.md, PINOUT.md, CLOCK_ARCHITECTURE.md
and SCHEMATIC_READINESS.md, which predate the SPI host bridge, the
real PLL, the ball-assigned LPF, and this session's SDRAM datasheet
audit, and are marked SUPERSEDED with a pointer back here rather than
individually rewritten.
Baseline commit: `d6376e8` (user-designated engineering reference).
This session's own fix on top of it: `8890b0a` (ERR-0026, SDRAM tMRD).
**SDRAM-SPECIFIC CONTENT SUPERSEDED (DEC-0039, a later session).** The
SDRAM was upgraded from AS4C4M16SA-6TIN (8MB) to AS4C32M16SB-7BIN
(64MB), and N_SLOTS=8 was added as a real, verified configuration
alongside N_SLOTS=4. Every SDRAM organization table, pin count, and
memory-map worked example below describing the 8MB part is stale —
see `MEMORY_UPGRADE_64MB_N8.md` for the current, authoritative state.
Sections unrelated to SDRAM specifics (RTL freeze, synthesis warning
classification methodology, SPI protocol *structure* though not its
exact byte counts, config flash, power/pinout for non-SDRAM signals)
remain accurate.
---
## 1. RTL functional freeze — audit result
Re-inspected `fpga_neural_v2_top.v`'s full port list and instantiation
tree this session (not assumed from prior reports):
- 16 top-level ports: `osc_clk`, `ext_rst_n`, `spi_sclk`, `spi_mosi`,
`spi_miso`, `spi_cs_n`, `sdram_cke`, `sdram_cs_n`, `sdram_ras_n`,
`sdram_cas_n`, `sdram_we_n`, `sdram_ba[1:0]`, `sdram_a[11:0]`,
`sdram_dq[15:0]` (inout), `sdram_dqm[1:0]`, `pll_locked`. Zero
`reg_*`/testbench-only ports on the physical top.
- Instantiation tree: `nms_dataflow_core_sdram``sdram_unified_backend`
`slot_mem_arbiter` / `slot_mem_arbiter_wide``spi_host_bridge`.
No V1 module anywhere in this tree.
- No PSRAM reference anywhere in the V2 compile list (`grep -ri psram
hardware/v2/` returns nothing outside historical log/doc commentary
explaining why it was removed).
- No stale host-bus (`reg_*`) driver active on the physical top; the
only place `reg_*` signals exist is internal, between
`spi_host_bridge` and `nms_dataflow_core_sdram`, which is the
intended internal protocol-translation boundary, not a leftover
interface.
- No simulation-only initialization required for correctness: SDRAM
power-up/init is a real FSM in `sdram_controller.v`
(`S_INIT_*` states), not a `$readmemh`/testbench force.
**STATUS: PASS.**
## 2. ERR-0025 — final closure (re-verified this session)
Re-confirmed via direct source inspection (not assumed) that the
combinational-read fix is present, unregressed, in
`nms_weight_packed.v` and `nms_activation_replicated.v`, and that
`nms_memory_manager_stream_wide.v`'s `rd_pending` read-ahead pipeline
is unchanged from the fixed baseline. Full regression re-run fresh
from current source (Verilator, DEC-0004):
| Test | Result |
|---|---|
| N=2 D-Stress (`tb_nms_dstress_sdram_unified.v`) | 49,788 cycles, 256/256 bit-exact PASS |
| N=4 D-Stress | 49,771 cycles, 256/256 bit-exact PASS |
| Board-level smoke (`tb_fpga_neural_v2_top_smoke.v`) | 11/11 PASS (single-neuron, wide-gap, back-to-back, gap100ns/5000ns/50000ns) |
| SPI host bridge (`tb_spi_host_bridge.v`) | 18/18 PASS |
| Unified SDRAM backend (`tb_sdram_unified_backend.v`) | 40/40 PASS |
| SDRAM controller (`tb_sdram_controller.v`), 9-config legacy sweep | 461/461 PASS, all 9 configs (100/133/166MHz × BURST_LEN 1/4/8) |
| SDRAM controller, NEW 64MHz/BURST_LEN=4 config | 461/461 PASS |
**STATUS: CLOSED. All numbers identical to the pre-ERR-0026-fix
baseline (T_MRD only affects the one-time init sequence).**
## 3. Clock and reset verification
- `ecp5_pll_sys_clk.v` instantiates a real `EHXPLLL` primitive, real
Project Trellis `ecppll`-derived parameters: CLKI_DIV=1,
CLKFB_DIV=4, CLKOP_DIV=9, VCO=576MHz, exact 64MHz output from a
16MHz input. `(* FREQUENCY_PIN_CLKOP="64" *)` is present on the
output net.
- Re-verified this session (prior phase, re-confirmed not re-run this
round since no RTL affecting the PLL changed): P&R run WITHOUT a
`--freq 64` CLI flag still reports "PASS at 64.00 MHz" — the
RTL-embedded attribute alone drives nextpnr's generated-clock timing
analysis, not a fragile external flag.
- `reset_sync.v`: asynchronous assert, synchronous deassert, gated by
`ext_rst_n` AND `pll_locked` (confirmed by source inspection: reset
is held asserted until both the external POR and the PLL lock
signal are satisfied).
- Confirmed the generated 64MHz clock is the ONLY clock driving the
compute/memory datapath (`sdram_controller`, `nms_dataflow_core_sdram`,
`dependency_manager`, `neural_processor` all take the PLL's `CLKOP`
output, not `osc_clk` directly).
**STATUS: PASS.**
## 4. Synthesis (re-confirmed from prior real Yosys run, unchanged
this session since no synthesis-affecting RTL changed beyond
ERR-0026's single localparam, which does not change resource
counts)
| Resource | Count |
|---|---|
| TRELLIS_FF | 6,322 |
| TRELLIS_COMB (LUT4-equiv) | 7,084 |
| MULT18X18D | 32 (4 processors × 8-wide MAC) |
| EHXPLLL | 1 |
| DP16KD (block RAM) | 0 (all small SRAMs synthesize to distributed RAM) |
38 unique warnings (43 total). Each category re-classified this
session by reading the actual flagged RTL, not by matching a
historical baseline:
- `neural_processor.v \gi` multi-driver warning — **benign, confirmed**:
`gi` is a plain `integer` loop variable (not a genvar) reused across
two separate `always` blocks; a cosmetic Yosys elaboration artifact,
not a real multi-driver hazard.
- "Replacing memory with list of registers" (small weight/activation/
result buffers) — **benign, confirmed**: these are small,
fully-parallel-access pipeline arrays, correctly synthesized as
discrete FFs, not a genuine memory-inference miss.
- SDRAM `dq[15:0]` tristate inference — **expected, correct**: this is
the real bidirectional SDRAM data bus; Yosys/nextpnr correctly infer
a real `TRELLIS_IO` tristate buffer per bit.
- No inferred latches, no width-truncation warnings, no signed/
unsigned mismatch warnings found in this run.
**STATUS: PASS. Zero CHECK-pass problems. No warning classified as
"must fix" or "potentially dangerous."**
## 5. Place and route — 8-seed timing table (unchanged this session;
T_MRD is a single localparam value, not a structural RTL change,
so a full 8-seed re-run was not repeated — re-running P&R was not
warranted since the change cannot affect placement/routing/timing
of the compute or SDRAM-transaction datapath)
| Seed | Fmax (MHz) | Result | Slack @ 64MHz |
|---|---|---|---|
| 1 | 73.17 | PASS | +1.958 ns |
| 2 | 68.90 | PASS | +1.111 ns |
| 3 | 72.10 | PASS | +1.755 ns |
| 4 | 68.51 | PASS | +1.029 ns (worst) |
| 5 | 69.29 | PASS | +1.193 ns |
| 6 | 73.03 | PASS | +1.931 ns |
| 7 | 74.17 | PASS | +2.143 ns (best) |
| 8 | 70.10 | PASS | +1.360 ns |
8/8 seeds PASS at 64MHz. Worst 68.51MHz, best 74.17MHz, mean 71.16MHz.
`TRELLIS_IO`=44/245 (17%), zero unrouted nets, zero placement/routing
errors, all 8 seeds. Critical path routing-dominated (~80-85%
routing/15-20% logic), alternating between `dependency_manager.v`'s
priority-encoder scan and `sdram_unified_backend.v`'s weight-cache
hit-index logic — a long-documented, pre-existing pattern.
**STATUS: PASS.**
## 6. Setup and hold timing
- **Setup: PASS** — see section 5 (8/8 seeds, worst case +1.029ns
slack @ 64MHz, real nextpnr-ecp5 timing analysis, not a bare
Fmax-vs-target comparison).
- **Hold: HOLD VERIFICATION OPEN — TOOL LIMITATION.** Directly
investigated this session's prior phase: nextpnr-ecp5's
`--report <json> --detailed-timing-report` output was generated and
inspected in full; it contains `critical_paths` (setup-side,
posedge→posedge max-delay only), `detailed_net_timings`, `fmax`, and
`utilization` — no hold/min-delay data anywhere in either the JSON
or the text log. No standalone Project Trellis hold-timing tool
(`ecptime`) exists in this environment; no `pytrellis` Python module
is installed. This is a genuine, disclosed tool-chain limitation,
not an omission. Hold-time closure requires either a `pytrellis`-based
min-delay analysis pass or vendor-tool (Lattice Diamond/Radiant)
static timing analysis against the final routed netlist — neither
is available in this environment.
**STATUS: SETUP VERIFIED / HOLD VERIFICATION OPEN — TOOL LIMITATION.**
## 7. SDRAM datasheet-level audit
Source: real Alliance Memory AS4C4M16SA-6TIN datasheet, Rev 5.0,
October 2018, Table 17 (Electrical Characteristics / AC Operating
Conditions, -6 speed grade) and Note 11 (power-up sequence).
| Datasheet parameter | Required value | RTL value (`sdram_controller.v`) | Status |
|---|---|---|---|
| Organization | 4M×16, x16, 8MB | `sdram_dq[15:0]`, single 8MB (0x0000000x7FFFFF) address space | PASS |
| Command truth table | Standard SDR SDRAM (NOP/ACT/READ/WRITE/PRE/REF/MRS) | FSM issues exactly these commands via `{ras_n,cas_n,we_n}` encoding | PASS (re-traced this session) |
| CAS latency | Fixed, device-configured via MRS (this design uses CL=2 or CL=3 per MRS programming) | `localparam CAS_LATENCY` — fixed value, matches MRS-programmed CL | PASS |
| tCK (clock period) | ≥ 1/166MHz at -6 grade (min cycle time varies by CL) | 64MHz (15.625ns) — well within the -6 grade's supported range at either CL | PASS |
| tRCD (ACT→READ/WRITE) | 18 ns min | `T_RCD = ns_to_cycles(18)` → 2 cycles @ 64MHz (31.25ns ≥ 18ns) | PASS |
| tRP (PRE→ACT) | 18 ns min | `T_RP = ns_to_cycles(18)` → 2 cycles @ 64MHz (31.25ns ≥ 18ns) | PASS |
| tRAS (ACT→PRE) | 42 ns min, 100,000 ns max | Not an explicit counter — satisfied by construction: the fixed tRCD+CAS_LATENCY+BURST_LEN dispatch sequence is always ≥6 cycles (93.75ns ≥ 42ns @ 64MHz); max is not a real constraint at these transaction rates | PASS (verified by direct calculation, not merely cited) |
| tRC (ACT→ACT, same bank) | 60 ns min | Governed by tRAS+tRP sequencing in the FSM; ≥ 125ns @ 64MHz (8 cycles) ≥ 60ns | PASS |
| tWR (write recovery) | 2 tCK min | Folded in conservatively via `T_RP + 1` after burst writes → 3 cycles ≥ 2-cycle requirement @ 64MHz | PASS |
| tMRD (MRS→any command) | 2 tCK, fixed | **Was `ns_to_cycles(12)` → rounds to 1 cycle @ 64MHz (ERR-0026, FIXED to `localparam T_MRD = 2` this session)** | **PASS (post-fix)** |
| tREFI (refresh interval) | 15.6 µs max | `T_REFI` = 15625ns = 15.625µs | PASS |
| Initialization sequence | 100µs+ power-stable wait, NOP/PRE-ALL, ≥2 AUTO-REFRESH, MRS | `S_INIT_*` FSM chain implements this exact sequence (re-traced this session) | PASS |
| Byte mask (DQM) behavior | `dqm` high = mask that byte lane on read/write | `sdram_dqm[1:0]` driven from `mem_lb_n`/`mem_ub_n`, verified via the SDRAM controller's own `J-mask` regression test (byte-masked write, bit-exact, all 10 configs incl. 64MHz) | PASS |
| Power-up requirement | Stable clock + 100µs wait before any command except NOP/DESELECT | `S_INIT_WAIT` FSM state enforces the wait before issuing PRE-ALL | PASS |
**Only discrepancy found: ERR-0026 (tMRD), now fixed and re-verified
with zero regression (section 2).**
**STATUS: CLOSED.** (Revises the prior "OPEN, sim-level only" status
in CHIP_READINESS.md/OPEN_ITEMS.md — see DEC-0038.)
## 8. SDRAM address/memory-map boundary verification
Official V2 memory map (unchanged): weights @0x010000, activations
@0x200000, results @0x300000, all within the single 8MB
(0x0000000x7FFFFF) SDRAM space, host-programmable per job (not
hard-coded in the datapath).
Boundary coverage actually exercised by the existing regression suite
(re-examined this session, not merely asserted):
- `tb_sdram_controller.v`'s randomized-address sweep (9 legacy configs
+ the new 64MHz config) exercises addresses spanning the full
22-bit word-address range, including addresses within a few words of
0x000000 and within a few words of the 8MB top (e.g. addr=4194300 ≈
0x3FFFFC observed in the 64MHz run), and crosses multiple
bank/row boundaries as a side effect of pseudo-random addressing —
not a directed first/last-address or exact-bank-boundary test.
- Byte-masked writes (`J-mask` test) confirmed bit-exact in every
config.
- Simultaneous read/write traffic under realistic load is exercised by
the D-Stress N=2/N=4 regressions (concurrent weight reads + result
writes across multiple slots via the arbiter), not by an isolated
directed test.
**No directed test exists for the EXACT first address (0x000000),
EXACT last address (0x7FFFFF), or an EXACT bank/row boundary
crossing.** Given the controller's address decode is a uniform,
parameterized bit-slice (no special-cased boundary logic to fail), and
the randomized sweep already exercises addresses adjacent to both
extremes without failure, the residual risk is assessed as low — but
per the mandate's own "do not invent margins" rule, this is disclosed
as a genuine, narrow **OPEN** item rather than claimed closed by
inference.
**STATUS: PASS (randomized coverage, high confidence) / OPEN (no
directed first/last-address or exact-boundary-crossing test exists).**
## 9. SPI host bridge — protocol documentation
Source: `hardware/v2/rtl/spi_host_bridge.v` (re-read in full this
session).
- **Mode/polarity/phase**: SPI mode 0 (CPOL=0, CPHA=0), MSB-first,
one opcode byte per CS-low period. Triple-flop CDC synchronizer on
`sclk`/`mosi`/`cs_n` (metastability-safe crossing into the 64MHz
system-clock domain).
- **Max tested clock**: the board-level smoke test
(`tb_fpga_neural_v2_top_smoke.v`) drives SPI at a 500ns bit period
(~2MHz effective SCLK rate). **This is the only rate actually
exercised in simulation.** The CDC synchronizer's own latency
(3 system-clock cycles ≈ 46.9ns @ 64MHz) bounds a theoretical
maximum SPI rate well above 2MHz, but no empirical test exists above
2MHz — **max real operating SPI clock is OPEN, to be characterized
at bring-up** (this is exactly what `FIRST_POWER_ON.md` step 11
already exists to determine).
- **Command set** (opcode, MSB-first byte, one CS-low transaction
each): `0x00 NOP` (0 payload), `0x0F RESET` (0 payload, pulses
`soft_rst_pulse` one cycle after CS rises), `0x10 WRITE_JOB` (15
payload bytes: node_id, required, producer_ids[15:0], x_base[22:0],
w_base[22:0], n_tiles[15:0], result_addr[22:0] — all MSB-first,
23-bit address fields packed as byte,byte,byte with the top byte's
MSB reserved/zero), `0x20 STATUS` (0 payload, 1 response byte:
bit0=job_busy, bit1=mem_busy, bit2=last_job_accepted [sticky,
cleared by next WRITE_JOB], bits[7:3]=0), `0x01 WRITE_MEM` (5 header
bytes [addr[22:0], len_words[15:0]] + 2×len_words payload bytes,
WORD address not byte address), `0x02 READ_MEM` (5 header bytes,
same shape, 0 further MOSI payload; 2×len_words response bytes
clocked out on MISO). Any other opcode is treated as NOP (0 payload,
MISO drives 0x00) — confirmed inert, never wedges the bus.
- **Response latency**: `WRITE_JOB` holds `reg_valid` until
`reg_ready` (same-cycle valid&&ready acceptance, never a blind
pulse) — latency is whatever `dependency_manager`'s own
`reg_ready` takes to assert (job-queue-dependent, not fixed).
`WRITE_MEM`/`READ_MEM` each issue one `mem_req`/`mem_ready` handshake
per word — latency is the backend arbiter's per-word grant latency
(see MEMORY_ARCHITECTURE.md), not a fixed cycle count either.
- **Reset behavior**: `0x0F RESET` pulses `soft_rst_pulse` for one
system-clock cycle after CS deasserts; this is a soft, protocol-level
reset pulse distinct from the board's own `ext_rst_n`/PLL-lock-gated
hardware reset (section 3).
- **Framing / back-to-back transactions**: a new CS assertion normally
restarts the opcode state machine — EXCEPT when the previous
transaction is still pending a backend handshake (`ST_JOB_WAIT`,
`ST_MEM_WISS`, `ST_MEM_RISS`), in which case state is deliberately
NOT reset, preventing a new WRITE_JOB's incoming bytes from
corrupting the still-pending previous job's fields through the same
registers (a real bug found and fixed during this project's own
STEP20 development, documented in the module's own header comment
and re-confirmed present in the current source this session).
Back-to-back WRITE_JOB transactions are exercised and PASS in the
board-level smoke test (`C-back-to-back-A/B`, 11/11 PASS overall).
- **No reliance on testbench-only timing**: the synchronizer and FSM
operate purely on `posedge clk` and edge-detected `sclk`/`cs_n`
transitions; nothing in the design depends on a specific testbench
delay value, only on real edges crossing the CDC boundary.
**STATUS: PASS (documented, protocol-correct, end-to-end verified at
the one tested rate) / max operating clock rate OPEN pending bring-up
characterization.**
## 10. FPGA configuration flash — FROZEN (not left OPEN)
**Decision: Winbond `W25Q32JVSSIQ`.**
| Property | Value |
|---|---|
| Manufacturer / MPN | Winbond Electronics, `W25Q32JVSSIQ` |
| Capacity | 32 Mbit (4 MB) — the LFE5U-45F's own uncompressed bitstream is well under 1MB, giving >4x margin even uncompressed, more with `ecppack` compression |
| Package | SOIC-8, 208-mil body (standard, hand-solder/hobby-friendly, widely stocked) |
| Supply voltage | 2.73.6V (VCC), matches the bank-8 (config bank) VCCIO which this design sets to 3.3V, matching the SDRAM's own 3.3V LVCMOS33 I/O already used throughout banks 6/7 |
| Protocol | Standard/Dual/Quad SPI, JEDEC-standard command set; ECP5's own "Master SPI" configuration boot mode uses only standard single-line SPI reads, which this part supports natively |
| Pull resistors | `WP#` and `HOLD#` (pins 3 and 7 of the standard 8-SOIC pinout) must be pulled to VCC (or tied directly) since this design uses standard single-SPI mode only, not the quad I/O functions those pins double as — unused-active-low-pin convention, standard practice |
| Reset/hold/WP behavior | No dedicated `RESET#` pin on this part (some competing devices have one; this part does not) — `HOLD#` pauses the bus mid-transaction when asserted low, tied inactive (high) here since this design never needs to pause a config read |
| Config clock requirement | ECP5 Master SPI mode drives its own `CCLK` output during configuration at a rate set by the `ecppack --freq` option at bitstream-generation time; this part supports standard SPI reads up to 104MHz, far above any practical `ecppack` config-clock setting |
| Boot-mode requirement | Must be wired for ECP5's "Master SPI" (also called "SPI Flash") boot mode — mode selection is via the ECP5's own dedicated CFG mode-strap balls (distinct from JTAG/PROGRAMN/INITN/DONE); **exact CFG-strap ball numbers for this specific package are not yet extracted from the pinout CSV and remain a schematic-level lookup, OPEN** (the component decision itself does not depend on this) |
| JTAG interaction | JTAG (TDI/TCK/TMS/TDO, real balls R5/T5/U5/V4, bank 40) remains available in parallel with SPI-flash boot for direct bitstream download/debug without touching the flash — standard ECP5 dual-boot-path behavior, no conflict |
| DONE/INITN/PROGRAMN | Real balls Y3 (DONE), V3 (INITN), W3 (PROGRAMN), all bank 8 — these are configuration-control signals common to every ECP5 boot mode, not specific to the flash choice |
| ECP5-flow support | `ecppack` (Project Trellis) natively supports generating SPI-flash-compatible bitstream images (`.bit`/raw binary) with a selectable config-clock frequency; Winbond W25Qxx-series parts are a standard, widely-used choice in the ECP5/Project-Trellis open-source ecosystem (used on multiple real, shipped ECP5 boards) |
| Availability confidence | High — standard, long-lived, multi-source JEDEC part, stocked at major distributors (Digi-Key, Mouser); not a claim of real-time stock levels, which were not checked |
**STATUS: CLOSED. Concrete, purchasable, technically appropriate part
frozen.** (One narrow sub-item — the exact CFG mode-strap ball
numbers — remains a schematic-level CSV lookup, not a blocker to this
component decision.)
## 11. FPGA power requirements — real per-bank table
Source: official Lattice pinout CSV (`FPGA-SC-02034-3-0-ECP5U-45-
Pinout.csv`, rev 3.0) and Lattice's own published LFE5U voltage
requirements (VCC=1.1V±5%, VCCAUX=2.5V±5%, VCCIO=1.23.3V
per-bank-selectable, VCCIO8=configuration-bank, voltage must match the
chosen config interface).
| Bank | VCCIO | Used signals | Function | Status |
|---|---|---|---|---|
| Core (VCC) | 1.1V | internal fabric/PLL core | FPGA core logic supply | Real, required, all `VCC` balls (H8N13 region) must connect |
| VCCAUX | 2.5V | PLL analog/aux circuitry | Required for `EHXPLLL` operation | Real, required, all 4 `VCCAUX` balls (F6/P6/F15/P15) must connect |
| Bank 6 | 3.3V (LVCMOS33, per LPF) | `spi_sclk`, `spi_mosi`, `spi_miso`, `spi_cs_n` (some), SDRAM bus (some) | SPI host + SDRAM I/O | Real, matches SDRAM's own 3.3V requirement |
| Bank 7 | 3.3V (LVCMOS33, per LPF) | `pll_locked`, SDRAM bus (remainder), `osc_clk`, `ext_rst_n` | Clock/reset/debug + SDRAM I/O | Real, matches SDRAM's own 3.3V requirement |
| Bank 8 | 3.3V (must match config interface) | `CCLK` (U3), `PROGRAMN` (W3), `INITN` (V3), `DONE` (Y3) + config-flash SPI lines (mode-strap balls not yet extracted, see section 10) | FPGA configuration | Real for CCLK/PROGRAMN/INITN/DONE; flash SPI-line ball numbers OPEN |
| Bank 40 | (JTAG, standard 3.3V/1.8V-tolerant per ECP5 JTAG spec) | `TDI` (R5), `TCK` (T5), `TMS` (U5), `TDO` (V4) | JTAG programming/debug | Real balls, standard JTAG voltage compliance (not independently re-verified against the exact chosen VCCIO this session) |
| Banks 0/1/2/3 | 1.23.3V (unused this design) | none | Unused general-purpose I/O | Not used by this design; no signals assigned |
**Note (unchanged from the prior draft, re-confirmed real, not yet
independently cross-verified at the schematic/PCB level): all
banks 6/7/8 signals are assumed LVCMOS33 — a disclosed WARNING to
double-check at schematic capture, not a blocker.**
**STATUS: PASS (voltage requirements and bank/signal mapping are
real and sourced) — current/decoupling BUDGET remains a separate,
explicitly OPEN item (section 12).**
## 12. Power budget
Per the mandate's own explicit rule ("do not pretend to know FPGA
dynamic power exactly without implementation data"), this section
states only what is genuinely known and marks the rest OPEN rather
than inventing numbers:
- **Known real values**: rail voltages (section 11) and each part's
own datasheet-stated supply-voltage range (SDRAM 3.3V±0.3V per
AS4C4M16SA-6TIN Table 17; config flash 2.73.6V per section 10).
- **NOT known / OPEN**: exact static and dynamic current draw for the
ECP5-45F at this design's actual utilization (7,084 LUT4-equiv,
6,322 FF, 32 MULT18X18D, 1 PLL) and actual 64MHz toggle rate. This
requires either the Lattice Power Calculator tool (not available in
this Yosys/nextpnr-only environment) or the vendor's own published
ECP5-45F datasheet current tables cross-referenced against the real
post-P&R netlist — neither was performed this session, and no
number is invented in their place.
- **SDRAM/flash/oscillator current**: each part's own datasheet
states typical operating currents (SDRAM: on the order of tens of
mA active, per AS4C4M16SA-6TIN Table 17 — not re-quoted here to
avoid restating a number from memory rather than re-reading the
table; re-read the datasheet directly if an exact figure is needed
for schematic-stage regulator sizing).
- Regulator selection itself is explicitly out of scope for this
document (that is PCB/schematic-level component selection, the
user's own stated responsibility).
**STATUS: OPEN (voltage requirements known and real; current/power
budget genuinely not computable without post-implementation data or
tools not present in this environment — explicitly disclosed, not
fabricated).**
## 13. I/O and pinout freeze
All 16 top-level signals of `fpga_neural_v2_top.v` carry a real ball
assignment in `v2_board_top.lpf`, sourced from the official Lattice
pinout CSV (rev 3.0):
| Signal | Ball | Bank | Direction | Function | Status |
|---|---|---|---|---|---|
| `osc_clk` | H5 | — | in | 16MHz board oscillator input | Real, reused from V1's validated LPF |
| `ext_rst_n` | B4 | — | in | active-low external reset | Real, reused from V1's validated LPF |
| `spi_sclk` | L3 | 6/7 | in | SPI host clock | Real, plain GPIO |
| `spi_mosi` | M3 | 6/7 | in | SPI host data in | Real, plain GPIO |
| `spi_miso` | L2 | 6/7 | out | SPI host data out | Real, plain GPIO |
| `spi_cs_n` | N2 | 6/7 | in | SPI host chip-select | Real, plain GPIO |
| `pll_locked` | L1 | 6/7 | out | PLL lock status (bring-up/debug) | Real, plain GPIO |
| `sdram_cke` | B5 | 6/7 | out | SDRAM clock enable | Real |
| `sdram_cs_n` | C5 | 6/7 | out | SDRAM chip select | Real |
| `sdram_ras_n` | C4 | 6/7 | out | SDRAM RAS | Real |
| `sdram_cas_n` | A3 | 6/7 | out | SDRAM CAS | Real |
| `sdram_we_n` | B3 | 6/7 | out | SDRAM WE | Real |
| `sdram_ba[1:0]` | E4, C3 | 6/7 | out | SDRAM bank address | Real |
| `sdram_a[11:0]` | D5,D3,F4,E5,E3,F5,A2,B1,C2,C1,D2,D1 | 6/7 | out | SDRAM row/column address | Real |
| `sdram_dq[15:0]` | E1,G5,H3,J5,K3,K2,H1,J1,K1,K4,L4,L5,M5,M4,N4,N5 | 6/7 | inout | SDRAM data bus | Real |
| `sdram_dqm[1:0]` | P5, N3 | 6/7 | out | SDRAM byte mask | Real |
Duplicate/illegal/incompatible-assignment check (re-verified this
session by direct LPF inspection): 44/44 ball assignments are
distinct sites, all IOBUF entries specify `IO_TYPE=LVCMOS33`
consistently, no ball appears twice, no config-reserved ball (CCLK/
PROGRAMN/INITN/DONE/JTAG, section 10/11) is accidentally reused by any
design signal.
**STATUS: PASS. Real, P&R-confirmed, no placeholders, no conflicts.**
## 14. Configuration/JTAG/boot strategy
- **JTAG connector**: standard 4-wire JTAG (TDI=R5, TCK=T5, TMS=U5,
TDO=V4, bank 40) plus the board's own GND/VCC reference — a
standard 2×5 or 2×7 JTAG header is a schematic-level choice, not
frozen here (connector part number is a BOM item, section 15).
- **Config flash**: Winbond `W25Q32JVSSIQ` (section 10), wired for
ECP5 "Master SPI" boot mode.
- **PROGRAMN/INITN/DONE**: real balls W3/V3/Y3, bank 8. Standard ECP5
behavior: pulsing `PROGRAMN` low re-triggers configuration;
`INITN` low indicates a configuration error (or is held during the
init-wait window); `DONE` goes high once configuration completes
successfully and the fabric is released from configuration reset.
- **Boot mode**: SPI-flash boot (Master SPI) is the primary path;
JTAG remains available in parallel for direct bitstream download
during bring-up/debug without touching the flash (section 10).
- **Pull resistors**: `PROGRAMN` typically needs a pull-up (idle-high,
momentary-pulse-low to reconfigure) per standard ECP5 practice;
`INITN` is open-drain, needs a pull-up; exact resistor values are a
schematic-level detail, not fixed here.
- **Reset interaction**: `ext_rst_n`/`pll_locked`-gated internal reset
(section 3) is entirely independent of the FPGA's own configuration
reset (PROGRAMN/INITN/DONE cycle) — the design's internal reset
logic only takes effect after configuration completes and the
fabric is live.
- **First-programming and recovery**: initial bring-up should use
JTAG direct-to-SRAM configuration first (fastest iteration, no flash
programming risk); once verified, program the SPI flash via JTAG
(using nextpnr/Project-Trellis-generated `.bit` converted to a flash
image) for standalone power-on boot. Recovery from a bad flash image
is via JTAG direct configuration, which does not depend on flash
content.
**STATUS: PASS (strategy defined with real ball/part data) — exact
CFG mode-strap ball numbers and connector/pull-resistor values remain
schematic-level detail, consistent with this mandate's own scope
boundary (user does schematic/PCB).**
## 15. Preliminary BOM (not PCB — component decisions only)
| Component | Manufacturer / MPN | Package | Voltage | Role | Mandatory/Optional | Availability confidence |
|---|---|---|---|---|---|---|
| FPGA | Lattice `LFE5U-45F-8BG381C` | CABGA381 | 1.1V core / 2.5V aux / 1.2-3.3V I/O per bank | Compute | Mandatory | Not independently checked this session (real, standard part number, previously confirmed target) |
| SDRAM | Alliance Memory `AS4C4M16SA-6TIN` | TSOP-II-54 (standard for this part family) | 3.3V | Unified weight/activation/result memory | Mandatory | Not independently checked this session (real datasheet on file, previously confirmed target) |
| Config flash | Winbond `W25Q32JVSSIQ` | SOIC-8 | 2.7-3.6V | FPGA configuration boot | Mandatory | High (standard, multi-source JEDEC part) — see section 10 |
| Oscillator | 16MHz, real device MPN not re-selected this session | — | 3.3V (typical) | System clock source | Mandatory | **OPEN — no specific MPN frozen this session; only the frequency (16MHz) and its ball (H5) are fixed by the RTL/LPF** |
| JTAG connector | not selected this session | — | — | Programming/debug | Mandatory for bring-up | **OPEN — schematic-level choice** |
| Pull resistors (PROGRAMN, INITN, WP#, HOLD#) | generic, values not specified | 0402/0603 | — | Config-signal biasing | Mandatory | OPEN — standard values (e.g. 4.7kΩ-10kΩ), exact value is schematic-level |
| Decoupling capacitors | generic, per Lattice Hardware Checklist guidance (distributed network, not one-cap-per-ball) | 0402/0603 | — | Power integrity | Mandatory | OPEN — exact count/placement is PCB-level |
| Voltage regulators (1.1V core, 2.5V aux, 3.3V I/O) | not selected this session | — | — | Power supply | Mandatory | **OPEN — depends on the still-open current budget (section 12)** |
**STATUS: PARTIAL.** FPGA, SDRAM, and config flash are frozen, real,
purchasable parts. Oscillator MPN, JTAG connector, regulators, and
passive values are explicitly left OPEN — genuinely not decidable
without either a prior explicit decision (oscillator) or the current
budget this session could not fabricate (regulators), consistent with
"do not invent stock availability" and "do not invent margins."
## 16. First-board bring-up spec
Already exists at `hardware/v2/docs/FIRST_POWER_ON.md` (14+ step
procedure with measurable PASS/FAIL criteria: power rails → FPGA
configuration → DONE → JTAG detection → clock → SDRAM init → SPI host
comm → memory test → neural test). Re-read this session and confirmed
its sequencing and pass/fail criteria remain consistent with the
current design (SPI host interface, real PLL, real pinout) — no
update needed beyond noting here that this document's own prior
"cannot be executed until BLOCKER items close" caveat is now
significantly narrowed: the SPI host interface and ball-level pinout
BLOCKERs it references are CLOSED as of this session; the only
genuine hardware-domain blockers remaining are schematic/PCB/BOM
completion (sections 12, 15) and the max-SPI-clock characterization
noted in section 9.
**STATUS: PASS (procedure exists, real criteria, consistent with
current design).**
## 17. Benchmark finalization
Real, current-source benchmark results (Verilator, this session):
| Config | Cycles | Result |
|---|---|---|
| N=2 D-Stress | 49,788 | 256/256 bit-exact PASS |
| N=4 D-Stress | 49,771 | 256/256 bit-exact PASS |
| Board-level SPI (single job) | 99-100 cycles/job | PASS |
| Board-level SPI (back-to-back) | 88-100 cycles/job | PASS |
| Board-level SPI (gap100ns/5000ns/50000ns) | 88-100 cycles/job (steady-state unaffected by gap) | PASS |
At the real, P&R-verified 64MHz system clock: N=4 D-Stress (49,771
cycles) corresponds to 49,771 / 64,000,000 = **777.7 µs** wall-clock
for the full 256-neuron D-Stress workload. Throughput scaling from
N=2→N=4 is essentially flat in total cycle count (49,788→49,771,
<0.1% difference) because D-Stress's own workload shape keeps the
SDRAM/arbiter bandwidth as the binding constraint at this tile size,
not per-processor compute — consistent with this project's own prior
scaling analysis (STEP17/STEP18 reports), not a new finding.
**No embedded-target (ESP32-class) physical baseline is available —
this remains explicitly OPEN, not fabricated.** No comparison against
an unrelated desktop CPU is made here.
**STATUS: PASS (real cycle counts, real 64MHz-derived wall-clock
time) — embedded-baseline comparison OPEN (no hardware available).**
## 18. Datasheet (LaTeX) — status
`hardware/v2/docs/DatasheetLatex/` chapters were re-read this session
(00-features, 02-architecture, 05-pinout-timing, 08-status-roadmap).
Content is current and accurate against this session's own findings
EXCEPT the readiness checklist in `08-status-roadmap.tex`, which
predates this session's SDRAM-datasheet-audit closure (section 7) and
config-flash freeze (section 10). That chapter is updated as part of
this same change (see the diff to `08-status-roadmap.tex`) to move
"SDRAM datasheet-parameter cross-check" and "Configuration flash
selection" from OPEN to closed/decided, and the PDF is rebuilt and
confirmed to compile cleanly.
**STATUS: PASS (updated and rebuilt this session).**
## 19. Cross-domain consistency audit
Checked this session:
- RTL (`fpga_neural_v2_top.v` port list) ↔ LPF (`v2_board_top.lpf`):
all 16 ports have exactly one LPF entry each, no orphaned port, no
orphaned LPF entry. **Consistent.**
- LPF ↔ FPGA device: all sites are real CABGA381 balls per the
official Lattice CSV; IO_TYPE=LVCMOS33 throughout banks 6/7,
consistent with the SDRAM's 3.3V requirement. **Consistent.**
- RTL SDRAM timing constants ↔ real SDRAM datasheet: closed this
session (section 7), one discrepancy found and fixed (ERR-0026).
**Consistent (post-fix).**
- SPI host protocol (section 9) ↔ config-flash SPI (section 10): two
functionally and physically SEPARATE interfaces — the host SPI uses
banks 6/7 GPIO (L3/M3/L2/N2), the config flash uses bank-8
dedicated config-mode balls — confirmed no ball overlap. **Consistent.**
- Power requirements (section 11) ↔ BOM (section 15): SDRAM and
config-flash voltage requirements (3.3V, 2.7-3.6V) are both
satisfiable by a single 3.3V I/O rail choice; no contradiction found.
**Consistent.**
- Datasheet (section 18) ↔ this document: reconciled by this same
session's edit to `08-status-roadmap.tex`. **Consistent.**
- No stale V1 component name, no stale PSRAM reference, no
inconsistent memory-size/timing/performance number found across any
of the documents re-read this session.
**STATUS: PASS.**
---
## FINAL RELEASE GATE
| Item | Status |
|---|---|
| RTL functional freeze | PASS |
| ERR-0025 closure | PASS |
| Regression (full suite, this session) | PASS |
| Synthesis | PASS |
| Place & route (8 seeds) | PASS |
| Setup timing | PASS |
| Hold timing | **OPEN — TOOL LIMITATION** |
| PLL / clock generation | PASS |
| Reset architecture | PASS |
| SDRAM functional (sim) | PASS |
| SDRAM datasheet audit | PASS (closed this session, ERR-0026 fixed) |
| SDRAM address/memory-map boundary | PASS (randomized) / OPEN (no directed first/last/exact-boundary test) |
| SPI host protocol | PASS (documented, verified at tested rate) |
| Configuration flash | **PASS — FROZEN (Winbond W25Q32JVSSIQ)** |
| FPGA power requirements (voltage/bank mapping) | PASS |
| Power budget (current/decoupling) | OPEN (no implementation-level current data available) |
| I/O / pinout | PASS |
| Configuration / JTAG / boot strategy | PASS (strategy defined; mode-strap ball numbers schematic-level) |
| Preliminary BOM | PARTIAL (FPGA/SDRAM/flash frozen; oscillator MPN/regulators/connector/passives OPEN) |
| First-board bring-up spec | PASS |
| Benchmark | PASS (embedded baseline OPEN, no hardware) |
| Datasheet | PASS (updated, rebuilt) |
| Cross-domain audit | PASS |
**CLASSIFICATION: PRE-PCB VERIFIED**, with the following items
explicitly and honestly OPEN (not silently dropped): hold-time
verification (tool limitation), exact first/last-address and
bank-boundary directed SDRAM tests, max operating SPI clock rate,
FPGA power/current budget, oscillator MPN, JTAG connector, pull
resistor/decoupling values, voltage regulator selection, and an
embedded-target (ESP32-class) benchmark baseline.
**SCHEMATIC: USER IMPLEMENTATION PENDING.**
**PCB: USER IMPLEMENTATION PENDING.**
**SILICON READY: NO — SCHEMATIC AND PCB NOT YET IMPLEMENTED.**
-154
View File
@@ -1,154 +0,0 @@
# FPGA-Neural V2 — stato roadmap
Fonte del mandato: `docs/v2-description.md` (root del repository). Baseline
funzionale/numerica/bit-exact: `hardware/v1/` (frozen, sola lettura — vedi
`hardware/v1/README.md`).
Legenda: `[ ]` non iniziato · `[~]` in corso · `[x]` completo (sim+synth+timing
reali, non solo scritto).
- [x] **M1 — Neural Processor** (`hardware/v2/rtl/neural_processor.v`, P8).
Bit-exact vs V1 (7/7 test, Verilator), pipeline a 8 stadi
funzionante, throughput reale (1 tile/ciclo). Sintesi reale: 0
problemi CHECK, Fmax 183.12 MHz (ACC_WIDTH=32) — vedi
`logs/experiments.log` EXP-0001/EXP-0002, `logs/errors.log` per 3
bug reali trovati e risolti (2 del toolchain Icarus, 1 RTL).
- [x] **M2 — Processor Array** (`neural_processor_array.v`). 1/2/4/8
processor testati (sim concorrenza reale + sintesi/P&R reali).
Fmax sempre PASS a 80MHz (159.11→134.70 MHz). Scoperta: il DSP
(MULT18X18D), non LUT/FF, satura per primo (88% a N=8) — vedi
`logs/decisions.log` DEC-0005.
- [x] **M3 — Buffers** (`activation_buffer.v`, `weight_buffer.v`,
`result_buffer.v`). Tutti inferiscono DP16KD reale (10/10 test,
6/6 config sintetizzate 0 problemi). Scoperta: il costo BRAM di
weight_buffer e' guidato da P_IN (larghezza), non da DEPTH.
- [x] **M4 — Memory Manager** (`memory_manager.v`, `prefetch_engine.v`),
backend PSRAM V1 riusato SENZA MODIFICHE. End-to-end reale (3/3
job PASS) con vero neural_processor + vera catena PSRAM V1.
3 bug RTL trovati/risolti (`logs/errors.log` ERR-0006). Fmax
165.86 MHz.
- [x] **M5 — Neural Director** (`neural_director.v`), scheduling
first-free. 4/4 test PASS (dispatch + coda + backpressure reale
su N_SLOTS=2). FSM ridotta a 4 stati, dependency rimandata a M6
(`logs/decisions.log` DEC-0007). Fmax 250.50 MHz.
- [x] **M6 — Dependency Manager** (`dependency_manager.v`), ready/waiting
queue, dependency counters, wake-up, producer tracking. 4/4 test
PASS (dipendenze multiple + produttore condiviso/piu' consumer).
Fmax 155.30 MHz. Forwarding di valori e riuso slot rimandati
(`logs/decisions.log` DEC-0008).
- [x] **M7 — Dataflow Core** (`dataflow_core.v`), prima integrazione
completa: Dependency Manager (M6) -> Neural Director (M5) ->
N_SLOTS x (Memory Manager (M4) + Neural Processor (M1)), loop di
wake-up chiuso end-to-end. 4/4 test PASS su un DAG a 3 nodi (node2
dipende da entrambi node0+node1, dispatch confermato solo dopo che
ENTRAMBI completano davvero). Sintesi reale 0 problemi a
N_SLOTS=2 e N_SLOTS=4. Fmax reale (harness): 165.15 MHz
(N_SLOTS=2), 133.19 MHz (N_SLOTS=4). Buffer M3 e arbitraggio PSRAM
condiviso rimandati esplicitamente a M8 (`logs/decisions.log`
DEC-0009).
- [x] **M8 — PSRAM integration** (`neural_multiprocessor.v`,
`slot_mem_arbiter.v`), controller V1 riusato SENZA MODIFICHE,
condiviso tra N_SLOTS memory_manager concorrenti reali. Trovato e
risolto un bug RTL reale: il primo arbitro perdeva silenziosamente
una richiesta arrivata durante la contesa (protocollo byte-level
"fire-and-forget", mai esposto da M4 che collega un solo master
direttamente) — vedi `logs/errors.log` ERR-0008. Dopo il fix: 4/4
test PASS (2 slot in vera contesa concorrente sulla stessa PSRAM
reale). Sintesi reale 0 problemi (nessun harness necessario — pin
reali PSRAM tengono il top-level a 157 pin). Fmax reale 142.45
MHz. Politica di arbitraggio a priorità fissa, non ancora fair
(`logs/decisions.log` DEC-0010).
- [x] **M9 — Full benchmark**, tabella V1 vs V2 (§32 del mandato) —
confronto full-system, stesso PARALLEL/P_IN=8, stesso backend
PSRAM reale V1 in entrambi. Fmax POST-P&R: V2 142.45 MHz (PASS
@80MHz) vs V1 68.65 MHz (FAIL @80MHz). Cicli/neurone SIMULATED
(1 neurone, 8 input, PSRAM reale): V2 166 vs V1 209 (2.6x
speedup wall-clock reale). MAC/cycle di picco: V2 16 (N_SLOTS=2 x
P_IN=8, concorrenza reale) vs V1 8 (core sequenziale singolo).
LUT/FF: V2 4191/3659 vs V1 8907/4900. 9/12 righe con dati reali
misurati; stall %/memory utilization/processor utilization
esplicitamente NON misurati questo milestone (`logs/decisions.log`
DEC-0011), rimandati a M10. Tabella completa in
`logs/benchmark.log`.
- [x] **M10 — Optimization**, solo sulla base dei dati raccolti in
M1-M9. N_SLOTS=8 sintetizzato e P&R reale (92.63 MHz, PASS
@80MHz, DSP 64/72=88.9%) — tetto pratico raccomandato per P_IN=8
su LFE5U-45F (`logs/decisions.log` DEC-0012). Sweep reale a 6
seed ACC_WIDTH 24 vs 32 (riusando i netlist gia' sintetizzati):
ACC_WIDTH=24 vince sia in Fmax medio (+6.2%, 180.71 vs 170.12
MHz) che in varianza (~3.4x piu' stretta) — risolve
l'inconcludenza a singolo seed di EXP-0002, nuovo default
raccomandato (DEC-0013). Strumentazione di conteggio cicli
(solo testbench, nessun RTL toccato) chiude la lacuna
stall%/utilization di DEC-0011 con dati reali: porta PSRAM
condivisa all'81.7% di utilizzo, slot0 95.2%, slot1 65.2%.
## Roadmap completa (§33)
Tutte e 10 le milestone del mandato (`docs/v2-description.md` §33) sono
complete: simulazione reale (Verilator), sintesi reale (Yosys), place &
route reale (nextpnr-ecp5) per ognuna, con log completi in
`hardware/v2/logs/` (EXP-0001..EXP-0013, DEC-0001..DEC-0013,
ERR-0001..ERR-0008). Elementi esplicitamente rimandati (non
dimenticanze, ognuno con la propria motivazione in `decisions.log`):
riuso slot in dependency_manager (DEC-0008), fairness dell'arbitro PSRAM
sotto contesa piu' estesa (DEC-0010), sweep P_IN<8 per N_SLOTS ancora
piu' alto (DEC-0012), riuso dei buffer M3 come cache condivisa
(DEC-0009), strumentazione stall%/utilization completa anche lato V1
(DEC-0011).
## Final Benchmark Campaign (post-roadmap, richiesta utente)
Dopo il completamento di M1-M10, una campagna di benchmark finale
completa e reale (`hardware/v2/sim/tb_benchmark_suite.v`, EXP-0014) ha
caratterizzato V2 end-to-end su 6 workload realistici (16-256 neuroni
indipendenti, un layer multilivello con forwarding reale via PSRAM, un
DAG a 6 nodi con risveglio a 2 hop) su N_SLOTS=1/2/4/8. 24/24 PASS
bit-exact dopo aver trovato e risolto 3 problemi reali (1 bug RTL in
`neural_director.v` mai testato a N_SLOTS=1, 2 bug nel testbench --
`logs/errors.log` ERR-0009). Scoperta principale: lo scaling parallelo
reale e' sostanzialmente PIATTO oltre N_SLOTS=2 (la vera porta PSRAM
condivisa satura al 91%, non il numero di processori) -- N_SLOTS=4 e'
misurabilmente PIU' LENTO in wall-clock reale di N_SLOTS=1 per il
workload Stress una volta considerato il vero Fmax POST-P&R.
**N_SLOTS=2 raccomandato come default** (`logs/decisions.log`
DEC-0014). Report completo (21 sezioni, THEORETICAL/SIMULATED/
POST-P&R/DERIVED classificati): `hardware/v2/docs/benchmarks/
final-benchmark.md`.
### Ottimizzazioni post-campagna (su richiesta utente)
Implementate entrambe le raccomandazioni #1/#2 del report:
1. **Burst a livello di parola** (`prefetch_engine.v`/`memory_manager.v`
parlano direttamente il protocollo a 16 bit di `memory_interface.v`,
bypassando `int8_memory_access.v` -- ancora congelato, semplicemente
non piu' istanziato in questo percorso). Reale: -49/-56% cicli sui
job singoli, 2.24-2.37x speedup wall-clock reale sull'intera
campagna. `logs/decisions.log` DEC-0015.
2. **Cache condivisa on-chip per il vettore di attivazione** (nuovo
`activation_cache.v`, evita che neuroni con lo stesso `x_base`
rileggano X da PSRAM). Reale: ulteriore -1.66/-2.00x cicli. MA costo
Fmax reale molto piu' ripido del previsto: N_SLOTS=4 ora FALLISCE il
target 80MHz (65.01 MHz, prima passava). N_SLOTS=2 (default
raccomandato) resta valido con margine piu' sottile (87.72 MHz).
Speedup wall-clock reale combinato (#1+#2) vs baseline originale:
N=1 3.86x, N=2 2.45x. `logs/decisions.log` DEC-0016.
Trovati e risolti 3 bug RTL reali durante l'implementazione
(`logs/errors.log` ERR-0009, ERR-0010). 24/24 combinazioni
workload/config ancora bit-exact dopo entrambe le ottimizzazioni.
## Log
Vedi `hardware/v2/logs/` (`development.log` per la cronologia di sessione,
`decisions.log` per le decisioni architetturali con motivazione,
`experiments.log` per ogni EXP-XXXX end-to-end).
## Regole non negoziabili attive (§34 del mandato, per riferimento rapido)
1. V1 (`hardware/v1/`) rimane intatta — mai modificata.
2. V2 vive esclusivamente sotto `hardware/v2/`.
3. Nessun risultato inventato: THEORETICAL vs SIMULATED vs SYNTHESIZED vs
POST-P&R sempre etichettati esplicitamente.
4. Ogni modifica/esperimento/decisione registrata nei log, mai persa.
5. Ogni esperimento ha un ID univoco, mai riutilizzato — anche i FAIL restano.
-114
View File
@@ -1,114 +0,0 @@
# FPGA-Neural V2 — SCHEMATIC READINESS
**SUPERSEDED.** This document predates the SPI host bridge, the real
PLL, and the final pinout. See `PRE_PCB_VERIFICATION.md` for the
current pre-schematic verification state. Left in place as a
historical record of the earlier block-diagram planning.
## Status (HISTORICAL): NOT READY
## Block diagram (what a hardware designer needs to know)
```
┌─────────────┐ ┌──────────────────────────────┐
│ Clock │ clk │ │
│ BLOCK ├───────►│ │
│ (OPEN item: │ │ │
│ 16MHz osc │ rst │ FPGA BLOCK │
│ vs 80MHz ├───────►│ LFE5U-45F-8BG381/CABGA381 │
│ needed -- │ │ │
│ see CLOCK_ │ │ nms_neural_multiprocessor_ │
│ ARCHITECTURE│ │ sdram_unified (N_SLOTS=4) │
│ .md) │ │ │
└─────────────┘ │ ┌────────────────────────┐ │ ┌───────────────┐
│ │ SDRAM interface (37 pins)├──────►│ SDRAM BLOCK │
│ │ sdram_cke/cs_n/ras_n/ │ │ │ AS4C4M16SA-6TIN│
│ │ cas_n/we_n/ba/a/dq/dqm │ │ │ (ONE chip -- │
│ └────────────────────────┘ │ │ weights+ │
│ │ │ activations+ │
│ ┌────────────────────────┐ │ │ results ALL │
│ │ Host bus (110 pins, │ │ │ here) │
│ │ BLOCKER -- raw parallel, │ │ └───────────────┘
│ │ not a real protocol yet) │ │
│ └────────────────────────┘ │
│ │
┌─────────────┐ │ ┌────────────────────────┐ │
│ CONFIG BLOCK │ JTAG │ │ TDI/TDO/TCK/TMS/ │ │
│ (OPEN: no ├────────►│ │ PROGRAMN/INITN/DONE/ │ │
│ flash part │ SPI │ │ CCLK (standard ECP5, │ │
│ chosen) ├────────►│ │ ball location BLOCKED) │ │
└─────────────┘ │ └────────────────────────┘ │
└──────────────────────────────┘
┌───────────┴───────────┐
│ POWER BLOCK │
│ VCC 1.1V / VCCAUX 2.5V / │
│ VCCIO 3.3V / SDRAM 3.3V │
│ (OPEN: regulators not │
│ selected) │
└──────────────────────────┘
┌─────────────┐
│ HOST BLOCK │ <-- BLOCKER: does not exist yet as real RTL.
│ (a real MCU/ │ Must serialize the 110-pin reg_* bus into
│ SPI/UART │ a real physical protocol (SPI, matching V1's
│ interface) │ own spi_neuron_top.v precedent, or similar)
└─────────────┘
┌─────────────┐
│ DEBUG/JTAG │ <-- standard ECP5 JTAG chain; no V2-specific
│ BLOCK │ debug infrastructure beyond that identified
└─────────────┘ this round.
```
## Interconnections a schematic designer needs (real, from the RTL)
- **FPGA ↔ SDRAM**: 37 real signals (`sdram_cke`, `sdram_cs_n`,
`sdram_ras_n`, `sdram_cas_n`, `sdram_we_n`, `sdram_ba[1:0]`,
`sdram_a[11:0]`, `sdram_dq[15:0]` bidirectional, `sdram_dqm[1:0]`) —
a single-chip, direct point-to-point connection (no bus sharing, no
second memory device). Real bank/ball assignment is BLOCKED (see
PINOUT.md) but the SIGNAL LIST itself is complete and final.
- **FPGA ↔ Clock**: one clock input pin (ball H5, reused from V1's own
real, validated assignment) — the SOURCE feeding that pin (direct
80MHz+ oscillator, or 16MHz oscillator + internal PLL) is an OPEN
decision (see CLOCK_ARCHITECTURE.md); the schematic cannot be
finalized for this block until that choice is made.
- **FPGA ↔ Reset**: one reset input pin (ball B4, reused from V1) —
real synchronization to a power-on-reset supervisor or button is
OPEN (not designed).
- **FPGA ↔ Configuration**: standard ECP5 JTAG/config pins exist by
device definition; whether the board ALSO includes an SPI
configuration flash (for standalone, non-JTAG boot) is an OPEN
decision (see CHIP_READINESS.md and OPEN_ITEMS.md).
- **FPGA ↔ Host**: **BLOCKER**. The real RTL currently exposes a
110-pin raw parallel bus with no serializing interface. A schematic
cannot meaningfully route "the host connection" until a real
physical protocol (and its own RTL bridge) exists.
- **FPGA ↔ Power**: standard ECP5 rail requirements (VCC/VCCAUX/VCCIO)
plus the SDRAM's own 3.3V rail — real regulator selection is OPEN
(see POWER_ARCHITECTURE.md).
## What IS ready
- The FPGA/package/speed-grade target is fixed and unambiguous
(LFE5U-45F-8BG381, CABGA381, -8).
- The external memory device is fixed and unambiguous (ONE
AS4C4M16SA-6TIN, no second chip).
- The complete, real signal list for the SDRAM interface is final (37
signals, confirmed by real POST-P&R synthesis).
- Real rail VOLTAGES (not currents) are known from device datasheets.
## What blocks starting the schematic today
1. Host interface: no real physical protocol exists (BLOCKER).
2. Ball-level pinout: no real assignment exists for SDRAM or host
signals (BLOCKER, same root cause as PINOUT.md's own finding).
3. Clock source decision: oscillator-only vs oscillator+PLL (CRITICAL,
OPEN).
4. Power regulator selection and current budget (OPEN).
5. Configuration-flash decision (OPEN).
**Conclusion: NOT READY.** A hardware designer could begin laying out
the SDRAM-to-FPGA net list today (that part is real and complete), but
could not close the schematic without resolving items 15 above.
@@ -41,19 +41,21 @@
% --- compact block diagram on the title page --- % --- compact block diagram on the title page ---
\begin{center} \begin{center}
\resizebox{\textwidth}{!}{%
\begin{tikzpicture}[node distance=7mm and 10mm] \begin{tikzpicture}[node distance=7mm and 10mm]
\node[fnblockD,minimum width=26mm] (host) {HOST\\{\scriptsize graph loader}}; \node[fnblockD,minimum width=26mm] (host) {HOST\\{\scriptsize graph loader}};
\node[fnblockT,right=14mm of host,minimum width=30mm] (dm) {Dependency\\Manager}; \node[fnblockT,right=14mm of host,minimum width=30mm] (dm) {Dependency\\Manager};
\node[fnblockT,right=14mm of dm,minimum width=28mm] (dir) {Neural\\Director}; \node[fnblockT,right=14mm of dm,minimum width=28mm] (dir) {Neural\\Director};
\node[fnblock,right=14mm of dir,minimum width=34mm] (slots) {N\_SLOTS $\times$ (Memory\\Manager $+$ Neural Proc.)}; \node[fnblock,right=14mm of dir,minimum width=34mm] (slots) {N\_SLOTS $\times$ (Memory\\Manager $+$ Neural Proc.)};
\node[fnblock,right=14mm of slots,minimum width=24mm] (ram) {PSRAM\\{\scriptsize 8\,MB, real V1 chain}}; \node[fnblock,right=10mm of slots,minimum width=20mm] (ram) {SDRAM\\{\scriptsize 64\,MB}};
\draw[fnbus] (host) -- (dm); \draw[fnbus] (host) -- (dm);
\draw[fnbus] (dm) -- (dir); \draw[fnbus] (dm) -- (dir);
\draw[fnbus] (dir) -- (slots); \draw[fnbus] (dir) -- (slots);
\draw[fnbus] (slots) -- node[fnlbl,above]{16-bit word} (ram); \draw[fnbus] (slots) -- node[fnlbl,above]{16-bit word} (ram);
\node[below=1mm of slots,font=\scriptsize\itshape,text=fnGrey] \node[below=1mm of slots,font=\scriptsize\itshape,text=fnGrey]
{computation entirely on-chip, dependency graph resolved autonomously}; {computation entirely on-chip, dependency graph resolved autonomously};
\end{tikzpicture} \end{tikzpicture}%
}
\end{center} \end{center}
\vfill \vfill
@@ -63,14 +65,18 @@
\footnotesize \footnotesize
\textbf{\color{fnDark}Reference target device:} Lattice ECP5 \code{LFE5U-45F-8BG381C} \textbf{\color{fnDark}Reference target device:} Lattice ECP5 \code{LFE5U-45F-8BG381C}
(speed grade $-8$, CABGA381) --- identical device and board as V1.\\[2pt] (speed grade $-8$, CABGA381) --- identical device and board as V1.\\[2pt]
\textbf{\color{fnDark}Recommended configuration:} INT8/INT32, \code{P\_IN}=8, \textbf{\color{fnDark}Production configuration:} INT8/INT32, \code{P\_IN}=8,
\code{N\_SLOTS}=2 (real, measured net win --- see ch.~\ref{ch:impl2}), same real, \code{N\_SLOTS}=4 (real 8/8-seed timing closure at 64\,MHz --- see
unmodified V1 PSRAM backend, ISSI \code{IS66WVE4M16EBLL-70BLI}.\\[2pt] ch.~\ref{ch:hw}), single unified SDR SDRAM (Alliance Memory
\code{AS4C32M16SB-7BIN}, 64\,MB), real board-level pinout and KiCad
schematic/BOM.\\[2pt]
\textbf{\color{fnDark}Status:} RTL verified in real Verilator simulation and real \textbf{\color{fnDark}Status:} RTL verified in real Verilator simulation and real
synthesis + place\&route (Yosys + nextpnr-ecp5). Full benchmark campaign, two synthesis + place\&route (Yosys + nextpnr-ecp5). Full benchmark campaign, two
post-campaign memory optimizations, and a full alternative memory-subsystem post-campaign memory optimizations, an alternative memory-subsystem
redesign (the Neural Memory System, ch.~\ref{ch:nms}) complete and measured. redesign that became the current architecture (the Neural Memory System,
Document describing the project as of \datasheetdate. ch.~\ref{ch:nms}), and a real, board-level schematic/BOM verification pass
(ch.~\ref{ch:hw}) all complete and measured. Document describing the
project as of \datasheetdate.
}; };
\end{tikzpicture} \end{tikzpicture}
\end{center} \end{center}
@@ -79,8 +85,10 @@ Document describing the project as of \datasheetdate.
Project author: Michele Bigi \textbullet{} MIKILAB / manvalan.\\ Project author: Michele Bigi \textbullet{} MIKILAB / manvalan.\\
This datasheet documents V2 of the RTL code, documentation and benchmarks This datasheet documents V2 of the RTL code, documentation and benchmarks
present in the repository \texttt{github.com/manvalan/FPGA-Neural}. V1 remains present in the repository \texttt{github.com/manvalan/FPGA-Neural}. V1 remains
frozen and unmodified as the project's golden functional/performance reference; frozen and unmodified as the project's golden functional/performance
it is documented in a separate datasheet.\par} reference; its own datasheet previously lived alongside this one in this
repository and was consolidated out of the working tree as part of a
2026-09-09 documentation cleanup (recoverable from git history).\par}
\end{titlepage} \end{titlepage}
% ====================================================================== % ======================================================================
@@ -89,7 +97,7 @@ it is documented in a separate datasheet.\par}
\input{chapters/00-features} \input{chapters/00-features}
% ====================================================================== % ======================================================================
% PINOUT SUMMARY (honesty note -- no real ball assignment for V2 yet) % PINOUT SUMMARY (real, board-verified ball assignment)
% ====================================================================== % ======================================================================
\newpage \newpage
\input{chapters/00b-pinout} \input{chapters/00b-pinout}
@@ -14,10 +14,11 @@ and unmodified as the project's golden reference). Where V1 executes one
neuron at a time under host-driven SPI control, V2 registers a neuron at a time under host-driven SPI control, V2 registers a
\textbf{dependency graph of neurons} and keeps \code{N\_SLOTS} independent \textbf{dependency graph of neurons} and keeps \code{N\_SLOTS} independent
Neural Processor $+$ Memory Manager pairs busy concurrently, resolving data Neural Processor $+$ Memory Manager pairs busy concurrently, resolving data
dependencies and hiding PSRAM latency in hardware, without host dependencies and hiding memory latency in hardware, without host
intervention once a graph is loaded. Computation (INT8 MAC, ReLU, intervention once a graph is loaded. Computation (INT8 MAC, ReLU,
saturation) is bit-exact identical to V1's own datapath; what changed is saturation) is bit-exact identical to V1's own datapath; what changed is
everything \emph{around} it.} everything \emph{around} it, including, mid-project, the external memory
device itself (\S\ref{sec:sdram-mem-addendum}).}
\vspace{8pt} \vspace{8pt}
\begin{multicols}{2} \begin{multicols}{2}
@@ -29,22 +30,21 @@ everything \emph{around} it.}
once every producer it depends on has genuinely completed --- verified once every producer it depends on has genuinely completed --- verified
for 1-hop shared-producer/multi-consumer graphs and 2-hop transitive for 1-hop shared-producer/multi-consumer graphs and 2-hop transitive
(diamond) graphs. (diamond) graphs.
\item \code{N\_SLOTS} independent \textbf{Neural Processor + Memory Manager} \item \code{N\_SLOTS}=4 independent \textbf{Neural Processor + Memory
pairs (default recommended: \textbf{2}), each running the identical Manager} pairs (production baseline), each running the identical
8-stage INT8 pipeline inherited from V1. 8-stage INT8 pipeline inherited from V1.
\item \textbf{Word-level burst memory backend}: fetches move a full 16-bit \item \textbf{Single unified SDRAM}: one external SDR SDRAM device serves
PSRAM word per transaction instead of one byte, reusing weights, activations, AND results through one arbitrated backend
\code{memory\_interface.v}/\code{psram\_controller.v} directly and its (\code{sdram\_unified\_backend.v}) --- no PSRAM, no second physical
already-implemented page-mode support --- \textbf{2.24--2.37$\times$} memory device, in the current, frozen hardware path.
real wall-clock speedup, measured. \item \textbf{Real physical host transport}: a placed, ball-assigned SPI
\item \textbf{Shared on-chip activation cache}: a vector of activations Mode~0 slave (\code{spi\_host\_bridge.v}) plus a real
shared by many neurons of the same layer is fetched from PSRAM \code{FPGA\_DATA\_READY} completion pin --- both verified on real
\emph{once}, not once per neuron --- a further real \code{nextpnr-ecp5} place\&route, not just in simulation.
\textbf{1.66--2.00$\times$} cycle reduction on shared-input workloads. \item \textbf{Real, board-level verification}: a real KiCad schematic
\item Same \textbf{real, unmodified V1 PSRAM backend} throughout capture, a real exported BOM, and real component selections
(\code{memory\_interface.v}, \code{psram\_controller.v}) --- V1 remains (regulators, oscillator, configuration flash) all cross-checked
the frozen golden reference and was never altered to make V2 look against this datasheet --- not merely a simulated design.
faster.
\item \textbf{Real, measured} characterization at every step: Verilator \item \textbf{Real, measured} characterization at every step: Verilator
RTL simulation, Yosys synthesis, real \code{nextpnr-ecp5} RTL simulation, Yosys synthesis, real \code{nextpnr-ecp5}
place\&route --- no theoretical number reported without a matching place\&route --- no theoretical number reported without a matching
@@ -56,15 +56,17 @@ everything \emph{around} it.}
{\color{fnDark}\large\bfseries Honest, measured limitations}\\[2pt] {\color{fnDark}\large\bfseries Honest, measured limitations}\\[2pt]
{\footnotesize {\footnotesize
\begin{itemize}[leftmargin=1.1em] \begin{itemize}[leftmargin=1.1em]
\item The system is \textbf{memory-bound}, not compute-bound: real compute- \item \code{N\_SLOTS}=8 is \textbf{functionally correct but not
to-memory-wait ratio on the order of 1:170--1:220. A single shared timing-closed}: only 3/8 tested placement seeds pass 64\,MHz ---
PSRAM port saturates at $\approx$90\% utilization regardless of deferred, not production-frozen (\S\ref{sec:clock-closure-current}).
\code{N\_SLOTS}$\ge$2 --- real parallel scaling beyond 2 slots is \item \textbf{Hold-time closure is a genuine, disclosed tool-chain
essentially flat for large workloads. limitation}: no \code{pytrellis}/vendor static-timing-analysis path
\item \code{N\_SLOTS=4} is \textbf{not recommended}: it delivers no is available in this environment to check min-delay/hold, only
additional real throughput once the shared PSRAM port saturates, setup (\S\ref{sec:clock-closure-current}).
and with the activation cache active it \textbf{fails the 80\,MHz \item \textbf{FPGA dynamic power/current draw is not measured}: no ECP5
timing target outright} (65.01\,MHz measured). power estimator is available in this toolchain; regulator sizing
uses datasheet-based engineering margin, not a computed budget
(\S\ref{sec:power-addendum}).
\item Fixed, lowest-index-priority arbitration (Director and memory \item Fixed, lowest-index-priority arbitration (Director and memory
arbiter alike) is not fairness-balanced --- a real, measured arbiter alike) is not fairness-balanced --- a real, measured
per-slot workload imbalance exists under sustained contention. per-slot workload imbalance exists under sustained contention.
@@ -74,22 +76,23 @@ everything \emph{around} it.}
{\color{fnDark}\large\bfseries Target \& toolchain}\\[2pt] {\color{fnDark}\large\bfseries Target \& toolchain}\\[2pt]
{\footnotesize {\footnotesize
\begin{itemize}[leftmargin=1.1em] \begin{itemize}[leftmargin=1.1em]
\item FPGA: Lattice ECP5 \code{LFE5U-45F-8BG381C} ($-8$, CABGA381) --- same \item FPGA: Lattice ECP5 \code{LFE5U-45F-8BG381C} ($-8$, commercial grade,
target device as V1. 381-ball caBGA, 0.8\,mm pitch) --- same target device as V1.
\item Synthesis: Yosys; place\&route: real \code{nextpnr-ecp5}. \item SDRAM: Alliance Memory \code{AS4C32M16SB-7BIN} (512\,Mbit/64\,MB,
4M$\times$16, 54-ball FBGA).
\item Synthesis: Yosys; place\&route: real \code{nextpnr-ecp5} 0.11.1.
\item Simulation: Verilator 5.050 (\code{--binary --timing}) --- adopted \item Simulation: Verilator 5.050 (\code{--binary --timing}) --- adopted
for V2 after two independent Icarus Verilog v13.0 scheduling for V2 after two independent Icarus Verilog v13.0 scheduling
defects were found and reproduced on minimal repros (V1's own defects were found and reproduced on minimal repros (V1's own
certification, performed separately, was unaffected). certification, performed separately, was unaffected).
\item PSRAM: ISSI \code{IS66WVE4M16EBLL-70BLI} (64\,Mb, 4M$\times$16),
real chain reused byte-for-byte from V1.
\end{itemize}} \end{itemize}}
\end{multicols} \end{multicols}
\vspace{2pt} \vspace{2pt}
% --- key parameter table --- % --- key parameter table ---
\noindent \noindent
{\small\color{fnDark}\bfseries Key parameters (recommended configuration, real measured data)} {\small\color{fnDark}\bfseries Key parameters (production configuration,
real measured data)}
\vspace{2pt} \vspace{2pt}
\noindent \noindent
@@ -100,11 +103,12 @@ everything \emph{around} it.}
Data precision & INT8 (signed) & \code{DATA\_WIDTH}=8, identical to V1 \\ Data precision & INT8 (signed) & \code{DATA\_WIDTH}=8, identical to V1 \\
\rowa Accumulator & INT32 (signed) & \code{ACC\_WIDTH}=32 \\ \rowa Accumulator & INT32 (signed) & \code{ACC\_WIDTH}=32 \\
Dot-product width & 8 & \code{P\_IN}=8 parallel MAC lanes per neuron \\ Dot-product width & 8 & \code{P\_IN}=8 parallel MAC lanes per neuron \\
\rowa Recommended concurrency & \code{N\_SLOTS}=2 & real, measured net win; see ch.~\ref{ch:impl2} \\ \rowa Production concurrency & \code{N\_SLOTS}=4 & real, 8/8-seed timing closure; see \S\ref{sec:clock-closure-current} \\
Fmax, full system (\code{N\_SLOTS}=2) & 87.72~MHz & real place\&route, word-burst + activation cache active \\ System clock & 64\,MHz & 16\,MHz oscillator $\to$ \code{EHXPLLL} PLL; 80\,MHz confirmed NO-GO (genuine regenerated PLL, 0/8 seeds) \\
\rowa cycles/neuron (1 neuron, 8 inputs, real PSRAM) & 166 (V1: 209) & \textbf{2.6$\times$} real wall-clock speedup vs V1 \\ \rowa Fmax, \code{N\_SLOTS}=4 (real P\&R, 8 seeds) & worst 64.55\,MHz / best 72.37\,MHz & production baseline, 8/8 PASS \\
Combined real speedup vs baseline (\code{N\_SLOTS}=2) & \textbf{2.45$\times$} & word-burst $+$ activation cache, D-Stress workload \\ D-Stress regression (256 neurons) & 49,927 cycles, 256/256 bit-exact & 780\,\textmu s wall-clock @ 64\,MHz \\
\rowa Address space & 23~bit (byte) & \code{ADDR\_WIDTH}=23, unchanged from V1 \\ \rowa SPI host clock, verified & 12\,MHz recommended (12.8\,MHz hard CDC edge) & simulation-verified, real margin below the deterministic edge \\
Address space & 26~bit (byte), single SDRAM & \code{ADDR\_WIDTH}=26 \\
\bottomrule \bottomrule
\end{tabularx} \end{tabularx}
@@ -112,28 +116,29 @@ Combined real speedup vs baseline (\code{N\_SLOTS}=2) & \textbf{2.45$\times$} &
\noindent \noindent
{\small\color{fnDark}\bfseries System block diagram} {\small\color{fnDark}\bfseries System block diagram}
\begin{center} \begin{center}
\resizebox{\textwidth}{!}{%
\begin{tikzpicture}[node distance=6mm and 9mm,font=\footnotesize] \begin{tikzpicture}[node distance=6mm and 9mm,font=\footnotesize]
\node[fnblockD,minimum width=24mm,minimum height=15mm] (host){HOST\\{\scriptsize registers a node graph}}; \node[fnblockD,minimum width=24mm,minimum height=15mm] (host){HOST (SPI)\\{\scriptsize registers a node graph}};
\node[fnblockT,right=14mm of host,minimum width=30mm,minimum height=13mm] (dm){Dependency\\Manager}; \node[fnblockT,right=14mm of host,minimum width=30mm,minimum height=13mm] (dm){Dependency\\Manager};
\node[fnblockT,right=14mm of dm,minimum width=28mm,minimum height=13mm] (dir){Neural\\Director}; \node[fnblockT,right=14mm of dm,minimum width=28mm,minimum height=13mm] (dir){Neural\\Director};
\node[fnreg,fill=white,right=14mm of dir,minimum width=30mm,minimum height=20mm] (slots){ \node[fnreg,fill=white,right=14mm of dir,minimum width=30mm,minimum height=20mm] (slots){
\begin{tabular}{c} \begin{tabular}{c}
N\_SLOTS $\times$ \\ N\_SLOTS=4 $\times$ \\
Memory Manager \\ Memory Manager \\
$+$ Neural Processor $+$ Neural Processor
\end{tabular}}; \end{tabular}};
\node[fnblockA,below=9mm of dir,minimum width=28mm,minimum height=11mm] (cache){Activation\\Cache}; \node[fnblock,right=14mm of slots,minimum width=26mm,minimum height=15mm] (ram){SDRAM 64\,MB\\{\scriptsize unified backend}};
\node[fnblock,right=14mm of slots,minimum width=22mm,minimum height=15mm] (ram){PSRAM 8\,MB\\{\scriptsize real V1 backend}};
\draw[fnbus] (host) -- (dm); \draw[fnbus] (host) -- (dm);
\draw[fnbus] (dm) -- node[fnlbl,above]{ready node} (dir); \draw[fnbus] (dm) -- node[fnlbl,above]{ready node} (dir);
\draw[fnbus] (dir) -- (slots); \draw[fnbus] (dir) -- (slots);
\draw[fnarrowT] (slots.south) |- (cache.east); \draw[fnbus] (slots) -- node[fnlbl,above]{W / AR ports} (ram);
\draw[fnarrowT] (cache.north) |- node[fnlbl,above]{producer done} (dm.south); \draw[fnarrowT] (slots.south) |- ++(0,-4mm) -| node[fnlbl,below]{producer done} (dm.south);
\draw[fnbus] (slots) -- node[fnlbl,above]{16-bit word} (ram); \end{tikzpicture}%
\draw[fnbus] (cache.east) -- ++(6mm,0) |- ([yshift=-2mm]ram.south); }
\end{tikzpicture}
\end{center} \end{center}
\begin{center}\footnotesize\itshape\color{fnGrey} \begin{center}\footnotesize\itshape\color{fnGrey}
A slot's completion feeds back to the Director (frees the slot) and to the A slot's completion feeds back to the Director (frees the slot) and to the
Dependency Manager (wakes up any node waiting on it) --- closing the Dependency Manager (wakes up any node waiting on it) --- closing the
dataflow loop entirely on-chip.\end{center} dataflow loop entirely on-chip. \code{FPGA\_DATA\_READY} (ball G3) goes high
once every registered node has both resolved and dispatched
(\S\ref{sec:host-addendum}).\end{center}
@@ -0,0 +1,67 @@
\thispagestyle{plain}
\noindent
\begin{tikzpicture}
\node[fill=fnDark,text=white,rounded corners=2pt,inner sep=6pt,
minimum width=\textwidth,anchor=west]
{\large\bfseries Pinout summary --- real, board-verified};
\end{tikzpicture}
\vspace{6pt}
\noindent
{\footnotesize
V2's top-level module, \code{fpga\_neural\_v2\_top.v}, has a complete,
real ball assignment: every signal --- SDRAM bus, SPI host transport,
clock/reset, \code{FPGA\_DATA\_READY}, JTAG, configuration mode straps,
and the boot flash's dedicated MSPI pins --- carries a real CABGA381 ball
site, sourced from the official Lattice pinout CSV (rev 3.0) and
cross-checked against Project Trellis's own \code{iodb.json}. This
supersedes an earlier V2 milestone in which the board-level top had been
placed only \textbf{unconstrained}; a full, constrained \code{.lpf} now
exists (\code{hardware/v2/constraints/v2\_board\_top.lpf}) and every
Fmax number in this datasheet (\S\ref{sec:clock-closure-current}) is
measured against it.
}
\vspace{6pt}
\begin{fnnote}[What is real]
Every ball in the summary table below is placed, P\&R-confirmed, and
cross-checked against a real, exported KiCad schematic and BOM
(\S\ref{sec:schematic-capture}--\ref{sec:bom}) --- not a simulation-only
placeholder. No PSRAM signals exist anywhere in this revision: the
single external memory is SDR SDRAM (\S\ref{sec:sdram-mem-addendum}).
\end{fnnote}
\begin{fnwarn}[What remains open]
FPGA dynamic power/current draw has not been measured post-implementation
(no ECP5 power estimator is available in this toolchain), so exact
decoupling/regulator sizing uses datasheet-based engineering margin, not
a computed budget. Hold-time closure is a genuine tool-chain limitation
(no min-delay analysis path available) --- setup timing is fully
verified. See ch.~\ref{ch:hw} for the complete, disclosed list.
\end{fnwarn}
\vspace{6pt}
\noindent
{\small\color{fnDark}\bfseries Ball summary (see ch.~\ref{ch:hw} for the
complete, per-signal table)}
\vspace{2pt}
\noindent
\begin{tabularx}{\textwidth}{L{3.4cm}L{2.4cm}Y}
\toprule
\rowh \thd{Interface} & \thd{Ball count} & \thd{Notes} \\
\midrule
SDRAM bus (A[0:12], BA[0:1], DQ[0:15], DQM[0:1], CKE/CS\#/RAS\#/CAS\#/WE\#) & 35 & Bank 6/7, real, P\&R-confirmed \\
\rowa SPI host transport (\code{sclk}/\code{mosi}/\code{miso}/\code{cs\_n}) & 4 & Bank 6/7, plain GPIO \\
\code{FPGA\_DATA\_READY}, \code{osc\_clk}, \code{ext\_rst\_n}, \code{sdram\_clk}, \code{pll\_locked} & 5 & Bank 6/7 \\
\rowa JTAG (TCK/TMS/TDI/TDO) & 4 & Bank 40, to ESP32 \\
Config control (PROGRAMN/INITN/DONE) + CFG[2:0] straps & 6 & Bank 8 \\
\rowa Boot-flash dedicated MSPI (CSSPIN/MCLK/D0/D1) & 4 & Bank 8, dual-function \\
\bottomrule
\end{tabularx}
\vspace{4pt}
\noindent
{\footnotesize\color{fnGrey}
Complete per-signal ball tables and the real KiCad schematic/BOM: ch.~\ref{ch:hw}.
Logical (not physical) register-level port list: ch.~\ref{ch:regs}.\par}
@@ -26,11 +26,15 @@ runs autonomously --- no per-neuron host intervention.
ReLU/linear activation with saturation --- \code{neural\_processor.v} ReLU/linear activation with saturation --- \code{neural\_processor.v}
is a direct, bit-exact-verified port of V1's own is a direct, bit-exact-verified port of V1's own
\code{neuron\_parallel.v}/\code{mac8.v}/\code{mac\_unit.v}. \code{neuron\_parallel.v}/\code{mac8.v}/\code{mac\_unit.v}.
\item The real PSRAM backend: \code{memory\_interface.v} and \item V1's own PSRAM backend files (\code{memory\_interface.v},
\code{psram\_controller.v} are reused \textbf{byte-for-byte, \code{psram\_controller.v}) remain byte-for-byte, unmodified
unmodified} from V1 throughout every V2 milestone --- including copies throughout the repository --- V1 itself, as a tree
the two post-campaign optimizations (ch.~\ref{ch:mem}). V1 itself, (\code{hardware/v1/}), is frozen and was never touched.
as a tree (\code{hardware/v1/}), is frozen and was never touched. \textbf{Not currently part of V2's physical board}, however: the
project has since replaced external memory with a single SDR
SDRAM device (\S\ref{sec:sdram-mem-addendum}); the PSRAM-era
chapters that follow document real, correctly-measured work for
the architecture it was measured on, not the current board.
\item The target device (Lattice ECP5 \code{LFE5U-45F-8BG381C}) and the \item The target device (Lattice ECP5 \code{LFE5U-45F-8BG381C}) and the
real-toolchain-only measurement discipline: every number in this real-toolchain-only measurement discipline: every number in this
datasheet is labelled \textsc{Theoretical}, \textsc{Simulated}, datasheet is labelled \textsc{Theoretical}, \textsc{Simulated},
@@ -70,8 +74,21 @@ also accounted for, \code{N\_SLOTS}=4 measures as \emph{slower} in real
wall-clock time than \code{N\_SLOTS}=1 for the largest workload tested wall-clock time than \code{N\_SLOTS}=1 for the largest workload tested
--- more hardware parallelism made that specific configuration worse, --- more hardware parallelism made that specific configuration worse,
not better, because the bottleneck was never compute. This finding not better, because the bottleneck was never compute. This finding
directly shaped both post-campaign optimizations in ch.~\ref{ch:mem} and directly shaped both post-campaign optimizations in ch.~\ref{ch:mem}.
the \code{N\_SLOTS}=2 recommendation carried throughout this datasheet.
\begin{fnwarn}[Architecture changed since this finding: SDRAM, not PSRAM]
This memory-bound finding was measured on the PSRAM-era architecture
described above. The project has since replaced PSRAM with a single
SDR SDRAM device (\S\ref{sec:sdram-mem-addendum}) and closed on
\textbf{\code{N\_SLOTS}=4 as the production configuration} --- chosen
primarily because it is the largest slot count that reliably closes
real timing (8/8 seeds @ 64\,MHz, ch.~\ref{ch:hw}
\S\ref{sec:clock-closure-current}), not from a re-run of this specific
utilization/scaling study. Whether the SDRAM backend's own
utilization/saturation ratio matches the PSRAM-era $\approx$90\% figure
above has \textbf{not been independently re-measured} --- disclosed as
an open item, not assumed to carry over.
\end{fnwarn}
\begin{fnnote}[Reproducibility] \begin{fnnote}[Reproducibility]
Every real number in this datasheet traces to a specific, append-only Every real number in this datasheet traces to a specific, append-only
@@ -1,6 +1,18 @@
\chapter{Architecture} \chapter{Architecture}
\label{ch:arch} \label{ch:arch}
\begin{fnnote}[Scheduling core unchanged; memory backend and slot count
have]
\code{dependency\_manager.v} and \code{neural\_director.v} (this
chapter's own subject) are identical between the PSRAM-era milestone
described below and the current, real SDRAM board --- the scheduling
logic itself did not change. What changed since is the memory backend
(single SDR SDRAM, not PSRAM, \S\ref{sec:sdram-mem-addendum}), the
absence of the shared \textbf{Activation Cache} module from the current
physical top (ch.~\ref{ch:toplevel}), and the production slot count
(\code{N\_SLOTS}=4, not 2).
\end{fnnote}
\section{Module map} \section{Module map}
\begin{center} \begin{center}
\begin{tikzpicture}[node distance=7mm and 11mm,font=\footnotesize] \begin{tikzpicture}[node distance=7mm and 11mm,font=\footnotesize]
@@ -26,9 +38,11 @@
\end{tikzpicture} \end{tikzpicture}
\end{center} \end{center}
\begin{center}\footnotesize\itshape\color{fnGrey} \begin{center}\footnotesize\itshape\color{fnGrey}
N\_SLOTS=2 shown (the recommended configuration); the architecture is PSRAM-era diagram, N\_SLOTS=2 shown; the architecture is parametric in
parametric in N\_SLOTS. Every arrow is a real signal path verified in N\_SLOTS. Every arrow is a real signal path verified in Verilator
Verilator simulation and real Yosys/nextpnr-ecp5 synthesis.\end{center} simulation and real Yosys/nextpnr-ecp5 synthesis. The current, real
board (N\_SLOTS=4, single SDRAM, no Activation Cache module) is shown
in ch.~\ref{ch:toplevel}'s own hierarchy listing.\end{center}
\section{Dependency Manager} \section{Dependency Manager}
Holds a table of \code{N\_NODES} job descriptors, each tracking: node Holds a table of \code{N\_NODES} job descriptors, each tracking: node
@@ -2,33 +2,38 @@
\label{ch:param} \label{ch:param}
\section{Build parameters (synthesis-time)} \section{Build parameters (synthesis-time)}
\begin{fnwarn}[Current, real board parameters (\code{fpga\_neural\_v2\_top.v})]
The table below reflects the real, current SDRAM-architecture top
level. The PSRAM-era \S\S\ref{ch:mem} chapters below this one describe
an earlier, real, correctly-measured milestone with different defaults
(notably \code{ADDR\_WIDTH}=23 and a PSRAM data-bus parameter) ---
superseded, not deleted, since that data remains accurate for the
architecture it was measured on.
\end{fnwarn}
\begin{tabularx}{\textwidth}{L{3.0cm} C{1.8cm} Y} \begin{tabularx}{\textwidth}{L{3.0cm} C{1.8cm} Y}
\toprule \toprule
\rowh \thd{Parameter} & \thd{Default} & \thd{Meaning} \\ \rowh \thd{Parameter} & \thd{Default} & \thd{Meaning} \\
\midrule \midrule
\code{DATA\_WIDTH} & 8 & Data width (INT8), unchanged from V1. \\ \code{DATA\_WIDTH} & 8 & Data width (INT8), unchanged from V1. \\
\rowa \code{ACC\_WIDTH} & 32 & Accumulator width; \textbf{24 recommended} for new P\_IN=8 configurations (ch.~\ref{ch:datapath}). \\ \rowa \code{ACC\_WIDTH} & 32 & Accumulator width. \\
\code{P\_IN} & 8 & Parallel MAC lanes per neuron per tile; must be even (word-level burst constraint, ch.~\ref{ch:mem}). \\ \code{P\_IN} & 8 & Parallel MAC lanes per neuron per tile; must be even (word-level burst constraint, ch.~\ref{ch:mem}). \\
\rowa \code{ADDR\_WIDTH} & 23 & Byte-address width (8~MB), unchanged from V1. \\ \rowa \code{ADDR\_WIDTH} & 26 & Byte-address width (widened 23$\to$26 for the 64\,MB SDRAM device, DEC-0039). \\
\code{N\_SLOTS} & 4 (RTL default) & Concurrent Memory Manager$+$Neural Processor pairs. \textbf{2 recommended} --- see the honesty note below. \\ \code{N\_SLOTS} & 4 & Concurrent Memory Manager$+$Neural Processor pairs. \textbf{Production configuration} --- real 8/8-seed timing closure at 64\,MHz (ch.~\ref{ch:hw} \S\ref{sec:clock-closure-current}). \\
\rowa \code{N\_NODES} & 16 & Dependency Manager node-table depth. Sized to the largest node-id range a graph will ever use; never reclaimed (ch.~\ref{ch:arch}). \\ \rowa \code{N\_NODES} & 16 & Dependency Manager node-table depth. Sized to the largest node-id range a graph will ever use; never reclaimed (ch.~\ref{ch:arch}). \\
\code{MAX\_DEPS} & 4 & Maximum producers a single node can list. \\ \code{MAX\_DEPS} & 4 & Maximum producers a single node can list. \\
\rowa \code{QUEUE\_DEPTH} & 8 & Neural Director's own ready-job FIFO depth. \\ \rowa \code{QUEUE\_DEPTH} & 8 & Neural Director's own ready-job FIFO depth. \\
\code{MAX\_TILES} & 16 (internal, activation\_cache.v) & Longest activation vector the shared cache can hold; not yet exposed as a top-level parameter. \\ \code{MAX\_TILES} & 16 & Longest activation/weight tile run a job can request. \\
\rowa \code{PSRAM\_DATA\_WIDTH} & 16 & Physical PSRAM data bus width, unchanged from V1. \\ \rowa \code{CLK\_FREQ\_MHZ} & 64 & Real system clock, generated by \code{ecp5\_pll\_sys\_clk.v} from the 16\,MHz oscillator; 80\,MHz confirmed NO-GO (ch.~\ref{ch:hw} \S\ref{sec:clock-closure-current}). \\
\code{CLK\_FREQ\_MHZ} & 80 & Frequency used in \code{psram\_controller.v}'s own timing formulas (unmodified V1 module). \\
\bottomrule \bottomrule
\end{tabularx} \end{tabularx}
\begin{fnwarn}[\texttt{N\_SLOTS} is a real, measured trade-off, not a free parameter] \begin{fnwarn}[\texttt{N\_SLOTS} is a real, measured trade-off, not a free parameter]
Unlike V1's \code{PARALLEL} (a pure resource/frequency trade-off), \code{N\_SLOTS}=4 is the production default: the largest slot count
\code{N\_SLOTS} interacts with a real, measured system bottleneck (the that reliably closes real timing at 64\,MHz on every tested placement
one physical PSRAM port). \code{N\_SLOTS}=1 and \code{N\_SLOTS}=2 both seed (8/8). \code{N\_SLOTS}=8 is functionally correct (bit-exact) but
show a real net wall-clock win over the pre-optimization baseline; only 3/8 seeds close timing --- deferred, not production-frozen. Do
\code{N\_SLOTS}=4 shows \emph{no} additional real throughput and, with not simply raise \code{N\_SLOTS} without re-running the real 8-seed
the activation cache active, \textbf{fails the 80\,MHz timing target \code{nextpnr-ecp5} matrix.
outright} (ch.~\ref{ch:impl2}). Do not simply raise \code{N\_SLOTS} for
more perceived parallelism without re-running the real benchmark suite.
\end{fnwarn} \end{fnwarn}
\section{Word-alignment constraint (post word-burst rewrite)} \section{Word-alignment constraint (post word-burst rewrite)}
@@ -40,7 +45,9 @@ land on an even byte address. \code{P\_IN} even and \code{x\_base}/
every job --- true of every address this project's own testbenches use, every job --- true of every address this project's own testbenches use,
and a trivial constraint for any real loader/host to satisfy. and a trivial constraint for any real loader/host to satisfy.
\section{Characterized configurations} \section{Characterized configurations (PSRAM-era; see ch.~\ref{ch:hw}
\S\ref{sec:clock-closure-current} for the current SDRAM-architecture
numbers)}
\begin{tabularx}{\textwidth}{C{1.6cm} C{2.6cm} C{2.6cm} Y} \begin{tabularx}{\textwidth}{C{1.6cm} C{2.6cm} C{2.6cm} Y}
\toprule \toprule
\rowh \thd{N\_SLOTS} & \thd{Fmax, word-burst only} & \thd{Fmax, $+$activation cache} & \thd{Notes} \\ \rowh \thd{N\_SLOTS} & \thd{Fmax, word-burst only} & \thd{Fmax, $+$activation cache} & \thd{Notes} \\
@@ -169,17 +169,21 @@ two copies of the same real data).
\subsection{Real, measured clock closure} \subsection{Real, measured clock closure}
\textbf{N\_SLOTS=4 @ 64\,MHz is the frozen production configuration}: \textbf{N\_SLOTS=4 @ 64\,MHz is the frozen production configuration}:
real \code{nextpnr-ecp5} P\&R, 8/8 tested seeds PASS (worst 66.58\,MHz, real \code{nextpnr-ecp5} P\&R, 8/8 tested seeds PASS. \textbf{N\_SLOTS=8
worst WNS $+0.605$\,ns). \textbf{N\_SLOTS=8 @ 64\,MHz remains an open @ 64\,MHz is deferred}, not production-frozen: 3/8 seeds PASS in the
item}: 5/8 seeds PASS (worst 60.12\,MHz, worst WNS $-1.009$\,ns) after final, current RTL state. 80\,MHz was tested with a genuinely
a real critical-path optimization (\code{sdram\_unified\_backend.v}'s
weight-cache hit-index encoder, rewritten from a serially-dependent
priority scan to a flat, parallel one-hot compare --- real errors.log
ERR-0029/decisions.log DEC-0040). 80\,MHz was tested with a genuinely
regenerated PLL (not merely a \code{--freq} flag) and is \textbf{not regenerated PLL (not merely a \code{--freq} flag) and is \textbf{not
achievable} at either processor count (0/8 seeds pass, both before and achievable} at either processor count --- the achievable Fmax is a
after the ERR-0029 optimization) --- the achievable Fmax is a property property of the routed fabric, confirmed identical between the
of the routed fabric, confirmed identical between the 64\,MHz- and 64\,MHz- and 80\,MHz-targeted netlists. Bit-exact functional
80\,MHz-targeted netlists. Bit-exact functional correctness (D-Stress, correctness (D-Stress, 256/256 neurons vs.\ golden model) is unaffected
256/256 neurons vs.\ golden model) is unaffected at every configuration at every configuration tested.
tested, including through this optimization.
\begin{fnnote}[Single source of truth for exact numbers]
The exact per-seed Fmax/WNS table, its full revision history (three
successive real critical-path fixes: ERR-0027, ERR-0028, ERR-0029, plus
a later fan-out fix, DEC-0042), and the SDRAM directed boundary-test
result (21/21 PASS, both 64\,MHz and 166\,MHz) are kept in one place to
avoid two copies of the same real data --- see ch.~\ref{ch:hw}
\S\ref{sec:clock-closure-current} and \S\ref{sec:sdram-addendum}.
\end{fnnote}
@@ -71,7 +71,10 @@ real ball assignments (\code{spi\_sclk}/\code{spi\_mosi}/
ch.~\ref{ch:hw}). WRITE\_JOB carries the full table from ch.~\ref{ch:hw}). WRITE\_JOB carries the full table from
\S\ref{ch:host} above as an 18-byte payload (grew from 15 after the \S\ref{ch:host} above as an 18-byte payload (grew from 15 after the
64MB memory upgrade widened every address field from 3 to 4 bytes --- 64MB memory upgrade widened every address field from 3 to 4 bytes ---
\code{decisions.log} DEC-0039). \code{decisions.log} DEC-0039). \textbf{Maximum verified operating
clock: 12\,MHz recommended} (exact deterministic CDC edge at
12.8\,MHz $=$ 64\,MHz/5, triple-flop synchronizer) --- see
ch.~\ref{ch:hw} \S\ref{sec:spi-max-verified} for the full sweep.
\textbf{Completion notification}: \code{FPGA\_DATA\_READY}, a real \textbf{Completion notification}: \code{FPGA\_DATA\_READY}, a real
output pin (ball \code{G3}, bank~7), closes the exact gap this output pin (ball \code{G3}, bank~7), closes the exact gap this
@@ -0,0 +1,64 @@
\chapter{Top-level module}
\label{ch:toplevel}
\begin{fnwarn}[Real, board-level top --- not the PSRAM-era compute core]
This chapter describes \code{fpga\_neural\_v2\_top.v}, the module that
is actually placed\&routed against real balls
(\code{hardware/v2/constraints/v2\_board\_top.lpf}) and whose Fmax
numbers appear throughout this datasheet. It supersedes an earlier
milestone's \code{neural\_multiprocessor.v} top level, which drove
V1's own PSRAM chain directly and is retained in the repository for
regression purposes (\code{tb\_nms\_dstress\_sdram\_unified.v}'s own
wrapper, \S\ref{sec:sdram-mem-addendum}) but is not the physical top.
\end{fnwarn}
\section{\texttt{fpga\_neural\_v2\_top.v}}
The real, board-level top: a PLL/reset front-end, a real SPI host
bridge, the compute/scheduling core, and a single unified SDRAM
backend --- 18 physical ports, every one ball-assigned.
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.2cm} C{1.6cm} Y}
\toprule
\rowh \thd{Port} & \thd{Dir} & \thd{Width} & \thd{Function} \\
\midrule
\code{osc\_clk} & IN & 1 & 16\,MHz board oscillator (ball H5). \\
\rowa \code{ext\_rst\_n} & IN & 1 & External POR/supervisor, active-low (ball B4). \\
\code{spi\_sclk}, \code{spi\_mosi}, \code{spi\_cs\_n} & IN & 1 each & Physical SPI host transport (ch.~\ref{ch:host}). \\
\rowa \code{spi\_miso} & OUT & 1 & SPI host transport, response direction. \\
\code{sdram\_clk} & OUT & 1 & SDRAM chip's own \code{CLK} pin --- a real board-level output, not internal-only routing (found missing during this session's own schematic review; ball J4). \\
\rowa \code{sdram\_cke}, \code{sdram\_cs\_n}, \code{sdram\_ras\_n}, \code{sdram\_cas\_n}, \code{sdram\_we\_n} & OUT & 1 each & SDRAM control lines. \\
\code{sdram\_ba} & OUT & 2 & SDRAM bank address. \\
\rowa \code{sdram\_a} & OUT & 13 & SDRAM row/column address (widened 12$\to$13 bits for the 64\,MB device, DEC-0039). \\
\code{sdram\_dq} & INOUT & 16 & SDRAM bidirectional data bus. \\
\rowa \code{sdram\_dqm} & OUT & 2 & SDRAM byte mask. \\
\code{data\_ready} & OUT & 1 & \code{FPGA\_DATA\_READY}, system-idle completion flag (ball G3, \S\ref{sec:host-addendum}). \\
\rowa \code{pll\_locked} & OUT & 1 & PLL lock status, bring-up/debug (ball L1). \\
\bottomrule
\end{tabularx}
\section{Internal hierarchy}
\noindent\code{fpga\_neural\_v2\_top.v}
\begin{itemize}[leftmargin=2.4em]
\footnotesize
\item \code{u\_pll} : \code{ecp5\_pll\_sys\_clk.v} (real \code{EHXPLLL} primitive, 16$\to$64\,MHz)
\item \code{u\_reset\_sync} : \code{reset\_sync.v} (async assert, sync deassert, gated by \code{ext\_rst\_n} AND \code{pll\_locked})
\item \code{u\_spi\_bridge} : \code{spi\_host\_bridge.v} (real SPI Mode~0 slave, triple-flop CDC)
\item \code{u\_dataflow\_core} : \code{nms\_dataflow\_core\_sdram.v}
\begin{itemize}
\item \code{u\_dep\_mgr} : \code{dependency\_manager.v}
\item \code{u\_director} : \code{neural\_director.v}
\item \code{GEN\_SLOT[0..N\_SLOTS-1]}: \code{nms\_memory\_manager\_stream\_wide.v} $+$ \code{neural\_processor.v}
\end{itemize}
\item \code{u\_arbiter\_w}, \code{u\_arbiter\_ar} : \code{slot\_mem\_arbiter.v} (one per logical SDRAM port, W and AR)
\item \code{u\_sdram\_backend} : \code{sdram\_unified\_backend.v} $\to$ \code{sdram\_controller.v} (single physical SDRAM)
\end{itemize}
\begin{fnnote}[No shared activation cache in this datapath]
The PSRAM-era shared activation cache (\code{activation\_cache.v},
ch.~\ref{ch:mem} \S\ref{sec:cache}) is not part of the current SDRAM
top-level's instantiation tree --- \code{nms\_memory\_manager\_stream\_wide.v}
handles per-slot activation/weight/result streaming directly against
the unified SDRAM backend. The PSRAM-era module remains real, correct,
and documented for the architecture it was measured on
(ch.~\ref{ch:mem}), but is not reused here.
\end{fnnote}
@@ -169,7 +169,7 @@ DSP/LUT/FF availability & Not the bottleneck at N\_SLOTS$\le$4 --- all well unde
\bottomrule \bottomrule
\end{tabularx} \end{tabularx}
\section{Limitations, honestly stated} \section{Limitations, honestly stated (PSRAM-era campaign above)}
\begin{itemize} \begin{itemize}
\item V1's own memory-utilization/stall figures were not re-measured \item V1's own memory-utilization/stall figures were not re-measured
this session (V1 is frozen); only its already-certified numbers this session (V1 is frozen); only its already-certified numbers
@@ -185,3 +185,73 @@ DSP/LUT/FF availability & Not the bottleneck at N\_SLOTS$\le$4 --- all well unde
chain) after either memory optimization --- only chain) after either memory optimization --- only
\code{dataflow\_core.v} alone, pre-optimization (92.63~MHz). \code{dataflow\_core.v} alone, pre-optimization (92.63~MHz).
\end{itemize} \end{itemize}
\section{SDRAM-era benchmark addendum (2026-09-07) --- current,
authoritative results}
\label{sec:impl-sdram-addendum}
\begin{fnwarn}[Supersedes the PSRAM/\code{N\_SLOTS}$\le$2-era campaign
above for the current hardware baseline]
Every section above (V1 vs.\ V2 comparison, \code{N\_SLOTS} sweep,
parallel scaling, memory optimizations \#1/\#2, bottleneck analysis)
describes an earlier V2 milestone built on V1's own PSRAM chain,
recommending \code{N\_SLOTS}=2. The project has since replaced external
memory with a single SDR SDRAM device (ch.~\ref{ch:mem}
\S\ref{sec:sdram-mem-addendum}) and closed on
\textbf{\code{N\_SLOTS}=4 as the production configuration}. This
section is the current, real, measured state; the PSRAM-era numbers
above remain real and correctly measured for the architecture they
describe, but do not apply to the current board.
\end{fnwarn}
\subsection{Real resource utilization (\code{N\_SLOTS}=4, SDRAM
architecture, post real critical-path fixes)}
\begin{tabularx}{\textwidth}{L{4.2cm} C{2.4cm} Y}
\toprule
\rowh \thd{Resource} & \thd{Count} & \thd{Notes} \\
\midrule
TRELLIS\_COMB (LUT4-equiv) & 7,175 / 43,848 (16.4\%) & Real Yosys synthesis, most recent measurement (post-ERR-0029) \\
\rowa MULT18X18D & 32 / 72 (44.4\%) & Exactly $4\times8$ (\code{N\_SLOTS}$\times$\code{P\_IN}), confirmed --- the ERR-0027 fix removed a spurious 33rd multiplier \\
DP16KD (block RAM) & 0 / 108 & All small SRAMs synthesize to distributed RAM \\
\rowa EHXPLLL & 1 & Real \code{EHXPLLL} primitive, \code{ecppll}-derived parameters \\
TRELLIS\_FF & $\ge$6,322 (last individually re-quoted figure) & Real, same SDRAM architecture, pre-dates the ERR-0027/0028/0029 restructuring; not independently re-synthesized standalone since --- disclosed as a lower-bound reference, not re-invented as exact \\
\bottomrule
\end{tabularx}
\subsection{Real, current clock closure and functional regression}
See ch.~\ref{ch:hw} \S\ref{sec:clock-closure-current} for the complete
per-seed Fmax/WNS table (single source of truth, not duplicated here):
\textbf{\code{N\_SLOTS}=4 @ 64\,MHz, 8/8 seeds PASS} (worst 64.55\,MHz,
best 72.37\,MHz); \code{N\_SLOTS}=8 deferred (3/8); 80\,MHz confirmed
NO-GO at either processor count with a genuinely regenerated PLL.
D-Stress functional regression (256 neurons, 256/256 bit-exact vs.\
golden model): \textbf{49,927 cycles} at \code{N\_SLOTS}=4 ---
\textbf{780\,\textmu s} real wall-clock at the P\&R-verified 64\,MHz
system clock ($49{,}927 / 64{,}000{,}000$, \textsc{Derived}). SDRAM
directed boundary verification (ch.~\ref{ch:hw}
\S\ref{sec:sdram-addendum}): 21/21 PASS, zero bugs found, both 64\,MHz
and 166\,MHz.
\subsection{Real SPI host protocol throughput}
Board-level smoke test (\code{tb\_fpga\_neural\_v2\_top\_smoke.v}, 11/11
PASS): single job 99--100 cycles/job; back-to-back 88--100 cycles/job;
steady-state throughput unaffected by inter-job gap (100\,ns/5\,\textmu
s/50\,\textmu s tested). Maximum verified SPI host clock: \textbf{12\,MHz
recommended} (exact deterministic CDC edge at 12.8\,MHz $=$ 64\,MHz/5)
--- see ch.~\ref{ch:hw} \S\ref{sec:spi-max-verified} for the full sweep.
\subsection{Limitations, honestly stated (current SDRAM architecture)}
\begin{itemize}
\item Power/energy: \textbf{NOT MEASURED} --- no ECP5 power estimator
is available in this toolchain (unchanged from the PSRAM-era
disclosure above).
\item Hold-time closure: \textbf{OPEN --- tool-chain limitation}, not a
real defect; see ch.~\ref{ch:hw} \S\ref{sec:hw-open-items} for
the complete, consolidated open-items list.
\item \code{N\_SLOTS}=8 is functionally correct but not
timing-closed on every tested seed --- deferred by explicit
project direction, not attempted further this pass.
\item No embedded-host (ESP32-class) physical baseline exists; all
host-side numbers above are protocol-level simulation, not
measured on real silicon.
\end{itemize}
@@ -1,64 +1,39 @@
\chapter{Hardware and board} \chapter{Hardware and board}
\label{ch:hw} \label{ch:hw}
\section{Unchanged from V1} \section{Board summary}
V2 targets the identical board and component set as V1: Lattice ECP5 V2 targets Lattice ECP5 \code{LFE5U-45F-8BG381C} ($-8$, commercial
\code{LFE5U-45F-8BG381C} ($-8$, CABGA381), ISSI grade, 381-ball caBGA, 0.8\,mm pitch, real package geometry
\code{IS66WVE4M16EBLL-70BLI} PSRAM (64\,Mb, 4M$\times$16), same 16\,MHz 17$\times$17$\times$1.76\,mm) --- the same die/package family as V1,
reference oscillator. The real PSRAM controller but the board around it has diverged substantially: V2 replaces V1's
(\code{psram\_controller.v}) and its byte$\leftrightarrow$word adapter PSRAM with a single external SDR SDRAM device (\S\ref{sec:sdram-addendum}),
(\code{memory\_interface.v}) are reused byte-for-byte, unmodified, from adds a real, placed SPI host transport and \code{FPGA\_DATA\_READY}
\code{hardware/v1/} throughout every V2 milestone --- their real, completion pin (ch.~\ref{ch:host}), and has a real, exported KiCad
already-verified electrical/timing requirements and page-mode behavior schematic capture and BOM (\S\ref{sec:schematic-capture}--\ref{sec:bom}).
are unchanged, because the controller itself was never touched. Every top-level signal of \code{fpga\_neural\_v2\_top.v} carries a real
ball assignment in \code{hardware/v2/constraints/v2\_board\_top.lpf} ---
no unconstrained/placeholder pins remain in this revision.
\begin{fnnote}[Real ball assignment: defer to V1's own chapter] \begin{fnnote}[V1's own PSRAM chain: retained in RTL, not on this board]
V1's own hardware chapter documents a real, \code{iodb.json}-verified, \code{psram\_controller.v}/\code{memory\_interface.v} remain byte-for-byte
place\&route-confirmed ball assignment for every PSRAM signal identical to V1's own copies in the repository (frozen golden reference),
(\code{psram\_a}, \code{psram\_dq}, \code{psram\_ce\_n/oe\_n/we\_n/ but are \textbf{not instantiated anywhere in V2's real physical top}
lb\_n/ub\_n/zz\_n}). Since V2's own \code{neural\_multiprocessor.v} --- confirmed by inspection (\code{grep -ri psram hardware/v2/} returns
drives these signals through the identical, unmodified controller, that nothing outside historical commentary). V1's own PSRAM ball assignment
same real ball assignment applies unchanged if V2 is deployed on the therefore does not apply to this board.
same physical board --- it is not repeated here to avoid maintaining two
copies of the same real data; see the V1 datasheet directly.
\end{fnnote} \end{fnnote}
\section{What V2 has not yet placed on real hardware}
As stated in ch.~\ref{ch:host}, V2's own node-registration bus has no
physical pin assignment in this revision --- every V2 characterization
to date used either a Verilator testbench or an unconstrained
(\code{--lpf-allow-unconstrained}) synthesis top-level. A real deployment
would need:
\begin{itemize}
\item A physical host transport for the registration bus (ch.~\ref{ch:host}).
\item A real, constrained \code{nextpnr-ecp5} place\&route run
producing a genuine \code{.lpf}/ball assignment for
\code{neural\_multiprocessor.v}'s own top-level pins, analogous to
V1's own \code{tools/pinout/gen\_lpf.py} flow.
\item Re-verification that the real Fmax numbers in ch.~\ref{ch:impl2}
(obtained unconstrained) hold once real pin locations are fixed ---
pin placement can itself affect routing and therefore Fmax.
\end{itemize}
\section{Power supply, oscillator, configuration}
Unchanged from V1: same board-level power sequencing, same oscillator,
same JTAG/config-SPI boot path (fixed-function dedicated pins, outside
RTL scope). No V2-specific hardware change was made or is required
beyond the (not yet placed) registration-bus transport above.
\section{SDRAM upgrade addendum (2026-09-07) --- current, authoritative \section{SDRAM upgrade addendum (2026-09-07) --- current, authoritative
board state} board state}
\label{sec:sdram-addendum} \label{sec:sdram-addendum}
\begin{fnwarn}[This section supersedes the PSRAM description above for \begin{fnwarn}[Real, closed architectural decision]
the current hardware baseline] An earlier V2 milestone reused V1's own PSRAM chain, placed
The sections above describe an earlier V2 milestone that still reused unconstrained. The project has since made a closed architectural
V1's own PSRAM chain unconstrained. The project has since made a decision (real \code{decisions.log} DEC-0034) to replace external
closed architectural decision (real \code{decisions.log} DEC-0034) to memory with a single SDR SDRAM device, and has since upgraded that
replace external memory with a single SDR SDRAM device, and has since device's capacity (8\,MB $\to$ 64\,MB) and re-verified real,
upgraded that device's capacity and re-verified real, constrained constrained place\&route timing end to end. This section is the
place\&route timing. This section is the current, real, measured state current, real, measured state.
--- see \code{hardware/v2/docs/MEMORY\_UPGRADE\_64MB\_N8.md} in the
repository for the full investigation.
\end{fnwarn} \end{fnwarn}
\subsection{Memory device} \subsection{Memory device}
@@ -141,15 +116,48 @@ See \code{decisions.log} DEC-0042 for full detail. A further
pipelining fix on the same broadcast path is a real, identified, pipelining fix on the same broadcast path is a real, identified,
not-yet-attempted option if more margin is ever needed. not-yet-attempted option if more margin is ever needed.
\subsection{Directed SDRAM boundary verification}
A dedicated directed testbench (\code{tb\_sdram\_boundary.v}, 21 checks)
covers every address/row/bank boundary the randomized D-Stress
regression does not directly target: exact first/last address
(\code{0x000000}/\code{0x3FFFFF}), the row-10/row-11 column boundary,
all three inter-bank crossings, the real V2 memory-map boundaries
(weights/activations/results base and last-word-before-next-region),
and all four byte-mask combinations with distinct deterministic
patterns. All 21 addresses are written first, then read back in
\textbf{reversed} order with address-derived patterns, proving no
write corrupts any neighbouring address. \textbf{Result: 21/21 PASS at
both 64\,MHz and 166\,MHz --- no bug found}, closing the one directed
boundary-test gap disclosed earlier in the project's own verification
history.
\subsection{Verified SPI host operating clock}
\label{sec:spi-max-verified}
A dedicated sweep testbench (\code{tb\_spi\_freq\_sweep.v}) drives the
real \code{fpga\_neural\_v2\_top} (not \code{spi\_host\_bridge} in
isolation) at the real 64\,MHz system clock and sweeps the SPI bit
rate across single-job, back-to-back, gapped, and raw
\code{WRITE\_MEM}/\code{READ\_MEM} traffic. The breakpoint is
\textbf{exact and deterministic}: PASS at every rate up to
\textbf{12.8\,MHz (precisely 64\,MHz/5)}, FAIL (data corruption, then
protocol FSM hang) at every rate at or above it --- the triple-flop CDC
synchronizer plus edge-detect/FSM reaction in \code{spi\_host\_bridge.v}
requires at least 5 full system-clock cycles per SPI bit period to
reliably track \code{sclk}/\code{mosi}/\code{cs\_n} transitions, a real
property of the CDC design (correct, standard practice), not a bug.
\textbf{SPI\_MAX\_VERIFIED = 12\,MHz} is the recommended host operating
point (real margin below the hard 12.8\,MHz edge, $\approx$6.7\%
headroom). Board-level electrical limits (trace length, driver
rise/fall time, ground bounce, real metastability risk) are
\textbf{not} modeled by this deterministic simulation and remain to be
confirmed empirically at bring-up.
\section{Power supply design (2026-09-07) --- verified against the real \section{Power supply design (2026-09-07) --- verified against the real
Lattice hardware checklist} Lattice hardware checklist}
\label{sec:power-addendum} \label{sec:power-addendum}
\begin{fnwarn}[Supersedes the generic \S3 stub above] \begin{fnwarn}[Real design data, not estimated]
The ``Power supply, oscillator, configuration'' section earlier in The actual rail topology, sized against the real, primary-source
this chapter only said ``unchanged from V1'' without real design data. Lattice and TI documents below.
This section replaces that stub with the actual rail topology, sized
against the real, primary-source Lattice and TI documents below --- not
estimated.
\end{fnwarn} \end{fnwarn}
\subsection{Rail topology} \subsection{Rail topology}
@@ -236,11 +244,18 @@ without external caps at the regulator itself); the 10\,\textmu F$+$
100\,nF on \code{VCCAUX} above are the FPGA-side filter from 100\,nF on \code{VCCAUX} above are the FPGA-side filter from
FPGA-TN-02038, not regulator-stability caps, and are still required. FPGA-TN-02038, not regulator-stability caps, and are still required.
\begin{fnnote}[Open item carried from \S3 above] \begin{fnnote}[16\,MHz oscillator: frozen]
The 16\,MHz reference oscillator's exact manufacturer part number is \textbf{ECS Inc. International \code{ECS-3225MV-160-BN-TR}} --- a
not yet specified in this document (only ``16\,MHz'' as a frequency quartz crystal oscillator (XO, not a bare crystal; direct digital clock
requirement) --- flagged, not invented, pending finalization of the output, no external oscillator circuit needed), 3225 SMD package
schematic capture. (3.2$\times$2.5\,mm, 4-pad, matching the real KiCad footprint for U5),
3.3\,V supply (matches \code{osc\_clk}'s real \code{IO\_TYPE=LVCMOS33}
ball H5 exactly, no level-shifting needed), $\pm$50\,ppm stability,
$-40$ to $+85^{\circ}$C. One 100\,nF decoupling capacitor across
\code{VDD}/\code{GND}, placed close to the supply pin. The exact
terminal order-code suffix (stability/output-enable option letters)
should be cross-checked against ECS's current published datasheet at
BOM lock --- normal due diligence, not an open architectural question.
\end{fnnote} \end{fnnote}
\subsection{Power tree} \subsection{Power tree}
@@ -290,9 +305,11 @@ timing recovery that followed) for the complete history.
\end{fnwarn} \end{fnwarn}
\subsection{One physical flash chip: boot bitstream only} \subsection{One physical flash chip: boot bitstream only}
Connects exclusively to the ECP5's own dedicated sysCONFIG pins, \textbf{Winbond \code{W25Q128JVPIM}} (128\,Mbit, WSON-8, 6$\times$5\,mm
Master SPI mode, auto-boots every power-up, zero ESP32 involvement in --- real BOM entry U9, \S\ref{sec:bom}). Connects exclusively to the
normal operation. No second flash device, no on-board neural-network ECP5's own dedicated sysCONFIG pins, Master SPI mode, auto-boots every
power-up, zero ESP32 involvement in normal operation. No second flash
device, no on-board neural-network
weight persistence in the current design --- the host (ESP32) is weight persistence in the current design --- the host (ESP32) is
responsible for pushing weight/activation data into SDRAM fresh each responsible for pushing weight/activation data into SDRAM fresh each
session via the real SPI application protocol session via the real SPI application protocol
@@ -527,3 +544,21 @@ Target: a castellated-edge SMD module, approximately
dimensions and pin-out placeholder, real layout pending. This section dimensions and pin-out placeholder, real layout pending. This section
will be filled in with the actual module outline, castellation pin will be filled in with the actual module outline, castellation pin
map, and mechanical drawing once available. map, and mechanical drawing once available.
\section{Verification status --- real, disclosed open items}
\label{sec:hw-open-items}
Everything above is real (simulated, synthesized, and/or place\&route
measured); this section lists what is genuinely \textbf{not yet}
verified, honestly, rather than silently omitted.
\begin{tabularx}{\textwidth}{L{4.4cm} Y}
\toprule
\rowh \thd{Item} & \thd{Status} \\
\midrule
Hold-time closure & \textbf{OPEN --- tool-chain limitation.} \code{nextpnr-ecp5}'s own timing report contains setup-side (posedge$\to$posedge max-delay) data only; no hold/min-delay analysis. No \code{pytrellis}-based min-delay pass or vendor (Lattice Diamond/Radiant) static timing analysis is available in this environment. Setup timing is fully verified (\S\ref{sec:clock-closure-current}). \\
\rowa FPGA dynamic power/current draw & \textbf{OPEN --- not computable without post-implementation tools.} No ECP5 power estimator (\code{ecppower} or equivalent) is available in this toolchain. Regulator current ratings (\S\ref{sec:power-addendum}) are real, datasheet-supported engineering margin against this unknown, not a computed budget. \\
N\_SLOTS=8 @ 64\,MHz & \textbf{Deferred, not production-frozen} --- functionally correct (bit-exact), 3/8 seeds pass timing closure. See \S\ref{sec:clock-closure-current}. \\
\rowa Board-level SPI electrical limit & \textbf{OPEN --- requires real hardware.} \S\ref{sec:spi-max-verified}'s 12\,MHz recommendation is a simulation-verified logical limit; real trace length, driver rise/fall time, and metastability risk are not modeled by simulation. \\
Embedded-host (ESP32-class) benchmark baseline & \textbf{OPEN --- no hardware available.} No comparison against a real ESP32 host exists; all host-side timing is protocol-level (ch.~\ref{ch:host}), not measured on real silicon. \\
\bottomrule
\end{tabularx}
@@ -1,13 +1,22 @@
\chapter{Register-level interface \& internal state encodings} \chapter{Register-level interface \& internal state encodings}
\label{ch:regs} \label{ch:regs}
\begin{fnwarn}[No SPI register map in this revision] \begin{fnwarn}[Real SPI opcode map exists; state encodings below are
V1's own quick-reference chapter documents a real SPI opcode/register per-module reference]
map (\code{STATUS}, \code{SET\_BASE}, \code{READ\_CONFIG}, \ldots). V2 has Ch.~\ref{ch:host} now documents V2's real, physical SPI opcode map
no equivalent yet (ch.~\ref{ch:host}) --- this chapter instead documents (\code{WRITE\_JOB}/\code{WRITE\_MEM}/\code{READ\_MEM}/\code{STATUS}/
the \textbf{node-registration field layout} (repeated here for quick \code{RESET}) --- this chapter's own node-registration field layout
reference) and the \textbf{internal FSM state encodings} exposed by each below remains the logical field reference (repeated here for quick
module, useful for simulation-level debug and for a future host driver. reference). The \textbf{internal FSM state encodings} below are useful
for simulation-level debug; \S\S\ref{ch:regs}'s Dependency
Manager/Neural Director tables are shared by every V2 architecture
(unchanged between the PSRAM-era and current SDRAM boards). The Memory
Manager and Neural Processor tables were captured from the PSRAM-era
\code{memory\_manager.v}/\code{neural\_processor.v} pairing (ch.~\ref{ch:arch})
--- the current SDRAM board's \code{nms\_memory\_manager\_stream\_wide.v}
implements the same functional handshake (prefetch $\to$ stream $\to$
write-back $\to$ done) against the SDRAM backend instead of PSRAM, but
its own internal state encoding was not re-transcribed into this table.
\end{fnwarn} \end{fnwarn}
\section{Node registration fields (quick reference)} \section{Node registration fields (quick reference)}
@@ -17,6 +17,26 @@ data shows the choice between them is configuration-dependent, not a
strict win for either. strict win for either.
\end{fnnote} \end{fnnote}
\begin{fnwarn}[This is the direct ancestor of the current, real board
--- read this before the rest of the chapter]
The \code{nms\_*}-prefixed modules introduced in this chapter
(\code{nms\_dataflow\_core.v}, \code{nms\_neural\_multiprocessor.v},
\ldots) are the \textbf{direct code ancestors} of the real, current
board-level RTL documented in ch.~\ref{ch:hw}/\ref{ch:toplevel}
(\code{nms\_dataflow\_core\_sdram.v}, \code{fpga\_neural\_v2\_top.v}).
The project's own path was: Current V2 (PSRAM, ch.~\ref{ch:arch}) $\to$
NMS (this chapter, still PSRAM, replicated on-chip SRAM) $\to$
\textbf{single unified SDRAM} (ch.~\ref{ch:hw}
\S\ref{sec:sdram-mem-addendum}, the current, real, shipped board). This
chapter's own STEP9/10 recommendation below (``adopt NMS at
\code{N\_SLOTS}$\le$2'') was itself superseded by that final SDRAM
step, which changed the backing memory device and re-closed timing at
\code{N\_SLOTS}=4 (ch.~\ref{ch:hw} \S\ref{sec:clock-closure-current}).
Read this chapter as \textbf{real history explaining how the current
architecture was reached}, not as a currently-open choice between three
systems.
\end{fnwarn}
\section{Design goal} \section{Design goal}
Current V2's own memory path is fundamentally an on-demand, Current V2's own memory path is fundamentally an on-demand,
per-request architecture: every tile fetch is a fresh transaction, per-request architecture: every tile fetch is a fresh transaction,
@@ -1,119 +0,0 @@
% ======================================================================
% FPGA-Neural -- INT8 Neural Network Engine
% Datasheet / Manuale di riferimento tecnico
% Repository: github.com/manvalan/FPGA-Neural
% ======================================================================
\documentclass[11pt,a4paper,openany]{report}
\newcommand{\datasheetrev}{A1}
\newcommand{\datasheetdate}{Settembre 2026}
\input{preamble}
\begin{document}
\sloppy
% ======================================================================
% FRONTESPIZIO
% ======================================================================
\begin{titlepage}
\thispagestyle{empty}
\begin{tikzpicture}[remember picture,overlay]
\fill[fnDark] (current page.north west) rectangle
([yshift=-4.3cm]current page.north east);
\fill[fnTeal] ([yshift=-4.3cm]current page.north west) rectangle
([yshift=-4.55cm]current page.north east);
\node[anchor=north west,text=white,font=\Huge\bfseries]
at ([xshift=2.2cm,yshift=-1.15cm]current page.north west)
{FPGA\,--\,Neural};
\node[anchor=north west,text=fnLight,font=\large]
at ([xshift=2.25cm,yshift=-2.15cm]current page.north west)
{INT8 Neural Network Engine per FPGA};
\node[anchor=north west,text=fnLight2,font=\normalsize]
at ([xshift=2.25cm,yshift=-2.85cm]current page.north west)
{Acceleratore hardware parametrico -- Datasheet e manuale di riferimento};
\node[anchor=north east,text=white,font=\ttfamily\small]
at ([xshift=-2.2cm,yshift=-3.55cm]current page.north east)
{Rev.~\datasheetrev~~\textbullet~~\datasheetdate};
\end{tikzpicture}
\vspace*{5.0cm}
% --- diagramma a blocchi sintetico sul frontespizio ---
\begin{center}
\begin{tikzpicture}[node distance=7mm and 12mm]
\node[fnblockD,minimum width=30mm] (host) {HOST\\{\scriptsize Linux / ESP32 / MCU / PC}};
\node[fnblockT,right=18mm of host,minimum width=34mm] (fpga)
{FPGA\\{\scriptsize Neural Network Engine}};
\node[fnblock,right=18mm of fpga,minimum width=26mm] (ram)
{PSRAM\\{\scriptsize 8\,MB dedicata}};
\draw[fnbus] (host) -- node[fnlbl,above]{SPI Mode 0} (fpga);
\draw[fnbus] (fpga) -- node[fnlbl,above]{async 16-bit} (ram);
\node[below=1mm of fpga,font=\scriptsize\itshape,text=fnGrey]
{calcolo interamente on-chip};
\end{tikzpicture}
\end{center}
\vfill
\begin{center}
\begin{tikzpicture}
\node[draw=fnRule,rounded corners=3pt,inner sep=10pt,fill=fnLight,text width=15.5cm]{
\footnotesize
\textbf{\color{fnDark}Dispositivo target di riferimento:} Lattice ECP5 \code{LFE5U-45F-8BG381C}
(speed grade $-8$, CABGA381, 72$\times$MULT18X18D, $\approx$44k LUT).\\[2pt]
\textbf{\color{fnDark}Configurazione baseline:} INT8/INT32, \code{N\_INPUTS}=256, \code{N\_NEURONS}=4,
\code{PARALLEL} parametrico, memoria di lavoro PSRAM ISSI \code{IS66WVE4M16EBLL-70BLI}.\\[2pt]
\textbf{\color{fnDark}Stato:} RTL verificato in simulazione (Icarus) e sintesi reale
(Yosys + nextpnr-ecp5). Documento descrittivo del progetto allo stato del \datasheetdate.
};
\end{tikzpicture}
\end{center}
\vspace{0.6cm}
{\footnotesize\color{fnGrey}\raggedright
Autore del progetto: Michele Bigi \textbullet{} MIKILAB / manvalan.\\
Questo datasheet documenta il codice RTL, la documentazione e i benchmark
presenti nella repository \texttt{github.com/manvalan/FPGA-Neural}.\par}
\end{titlepage}
% ======================================================================
% PAGINA "FEATURES" (stile datasheet)
% ======================================================================
\input{chapters/00-features}
% ======================================================================
% SUNTO PINOUT (pagine 2-3, pin per pin -- non a bus)
% ======================================================================
\newpage
\input{chapters/00b-pinout}
% ======================================================================
% INDICE
% ======================================================================
\newpage
\pagenumbering{roman}
{\color{fnDark}\tableofcontents}
\newpage
\pagenumbering{arabic}
% ======================================================================
% CAPITOLI
% ======================================================================
\include{chapters/01-overview}
\include{chapters/02-architettura}
\include{chapters/03-datapath}
\include{chapters/04-parametri}
\include{chapters/05-memoria}
\include{chapters/06-sequencer}
\include{chapters/06b-grafo}
\include{chapters/07-spi}
\include{chapters/07b-programmazione}
\include{chapters/08-toplevel}
\include{chapters/09-implementazione}
\include{chapters/10-hardware}
\include{chapters/11-registri}
\include{chapters/12-roadmap}
\appendix
\include{chapters/A-moduli}
\end{document}
@@ -1,45 +0,0 @@
# FPGA-Neural — Datasheet
Datasheet tecnico multicapitolo dell'engine FPGA-Neural, in italiano e inglese.
Ricostruito a partire dal codice RTL, dalla documentazione e dai benchmark presenti
nella repository (revisione A1, settembre 2026).
## Struttura
```
docs/datasheet/
├── FPGA-Neural-Datasheet.pdf ← PDF italiano (36 pagine)
├── FPGA-Neural-Datasheet.tex ← sorgente principale (IT)
├── preamble.tex ← stili, palette, box, TikZ
├── chapters/ ← 14 capitoli (IT)
└── en/
├── FPGA-Neural-Datasheet-EN.pdf ← PDF inglese (36 pagine)
├── FPGA-Neural-Datasheet-EN.tex ← sorgente principale (EN)
├── preamble.tex ← stili (EN)
└── chapters/ ← 14 capitoli (EN)
```
## Compilazione
Serve una distribuzione LaTeX con `pgfplots`, `tikz-timing`, `tcolorbox`,
`ltablex`, `listings`, `babel`.
```sh
# Italiano
cd docs/datasheet
pdflatex FPGA-Neural-Datasheet.tex
pdflatex FPGA-Neural-Datasheet.tex # 2ª passata per indice e riferimenti
# Inglese
cd docs/datasheet/en
pdflatex FPGA-Neural-Datasheet-EN.tex
pdflatex FPGA-Neural-Datasheet-EN.tex
```
## Nota sul pinout
Il capitolo *Progetto hardware e mappa dei segnali* riporta l'analisi completa
segnale-per-segnale del top-level `spi_neuron_top`, con la colonna **Ball**
compilata con assegnazioni CABGA381 reali (53 segnali, `.lpf` reale in
`synth/`) e verificata da un place\&route reale (`nextpnr-ecp5`, 0 errori di
vincolo, `Program finished normally`) — non più auto-piazzate.
@@ -1,119 +0,0 @@
\thispagestyle{plain}
\noindent
\begin{tikzpicture}
\node[fill=fnDark,text=white,rounded corners=2pt,inner sep=6pt,
minimum width=\textwidth,anchor=west]
{\large\bfseries FPGA-Neural --- Descrizione generale e caratteristiche};
\end{tikzpicture}
\vspace{6pt}
\noindent
{\small FPGA-Neural è un \textbf{acceleratore hardware parametrico per reti neurali}
feed-forward completamente contenuto nell'FPGA. Il calcolo (moltiplicazione,
accumulo, bias, attivazione, saturazione) avviene interamente on-chip in aritmetica
intera INT8/INT32; il sistema host fornisce solo configurazione, pesi, dati di
ingresso e controllo attraverso una semplice interfaccia SPI, senza mai far parte
del datapath computazionale. Un unico bitstream serve qualunque topologia fino al
massimo di build.}
\vspace{8pt}
\begin{multicols}{2}
{\color{fnDark}\large\bfseries Caratteristiche}\\[2pt]
{\footnotesize
\begin{itemize}[leftmargin=1.1em]
\item Datapath \textbf{INT8 $\times$ INT8 $\to$ INT16 $\to$ INT32}, accumulo a 32~bit
con estensione di segno.
\item \textbf{Balanced binary adder tree} ($O(\log_2 \text{PARALLEL})$) al posto della
riduzione lineare.
\item MAC parallelo configurabile: \code{PARALLEL} MAC hardware simultanei per neurone,
mappati su DSP \code{MULT18X18D}.
\item Architettura completamente \textbf{parametrica}: \code{N\_INPUTS}, \code{N\_NEURONS},
\code{PARALLEL}, \code{DATA\_WIDTH}, \code{ACC\_WIDTH}, \code{N\_LAYERS}.
\item \textbf{Larghezza di rete a runtime}: \code{n\_inputs\_real}/\code{n\_neurons\_real}
per-layer, un solo bitstream per ogni topologia fino al massimo.
\item Attivazioni configurabili: \code{ACT\_RELU} (default) e \code{ACT\_NONE} (lineare
con saturazione bilaterale), con saturazione INT8.
\item \textbf{Due tipi di rete}: classica multi-layer dense (\code{layer\_sequencer},
buffer ping-pong) e \textbf{grafo arbitrario sparse} (\code{graph\_engine} +
buffer di attivazione in block RAM \code{DP16KD}), selezionabili a runtime.
\item Sottosistema di \textbf{memoria dedicata}: interfaccia byte$\leftrightarrow$word,
controller PSRAM parallelo asincrono con \textbf{page mode} (70~ns accesso
casuale, 20~ns burst di pagina), 8~MB indirizzabili (23~bit).
\item Interfaccia host \textbf{SPI Mode 0} MSB-first, \code{SET\_NET\_TYPE}+dispatch, \code{STATUS.done}
sticky/clear-on-read, \code{READ\_CONFIG} runtime.
\item \textbf{Sottosistema flash} boot/persistenza: accesso esclusivo della FPGA a una
\code{W25Q128JV} SPI NOR (16~MB) via SPI master dedicato, copy engine
flash$\leftrightarrow$PSRAM e catalogo a 16 slot con CRC32, 8 opcode host.
\item Verificato in \textbf{simulazione} (Icarus Verilog) e \textbf{sintesi reale}
(Yosys + nextpnr-ecp5 + ecppack).
\end{itemize}}
\columnbreak
{\color{fnDark}\large\bfseries Applicazioni}\\[2pt]
{\footnotesize
\begin{itemize}[leftmargin=1.1em]
\item Inferenza a bassa latenza deterministica come periferica di
SoC Linux, Raspberry-Pi-like, ESP32, microcontrollori.
\item Blocco hardware riusabile integrabile in progetti eterogenei
(piattaforma, non singola rete).
\item Edge AI su reti dense compatte quantizzate INT8.
\item Off-loading del carico neurale dalla CPU host verso hardware
dedicato con throughput prevedibile.
\end{itemize}}
\vspace{4pt}
{\color{fnDark}\large\bfseries Target \& toolchain}\\[2pt]
{\footnotesize
\begin{itemize}[leftmargin=1.1em]
\item FPGA: Lattice ECP5 \code{LFE5U-45F-8BG381C} ($-8$, CABGA381).
\item Sintesi: Yosys; place\&route: nextpnr-ecp5; bitstream: Project~Trellis
(\code{ecppack}).
\item Simulazione: Icarus Verilog (\code{-g2012}).
\item PSRAM: ISSI \code{IS66WVE4M16EBLL-70BLI} (64\,Mb, 4M$\times$16).
\end{itemize}}
\end{multicols}
\vspace{2pt}
% --- tabella parametri chiave ---
\noindent
{\small\color{fnDark}\bfseries Parametri chiave (configurazione baseline caratterizzata)}
\vspace{2pt}
\noindent
\begin{tabularx}{\textwidth}{L{3.2cm}L{3.6cm}Y}
\toprule
\rowh \thd{Grandezza} & \thd{Valore} & \thd{Note} \\
\midrule
Precisione dati & INT8 (signed) & \code{DATA\_WIDTH}=8 \\
\rowa Accumulatore & INT32 (signed) & \code{ACC\_WIDTH}=32 \\
Ingressi / neuroni & 256 / 4 & baseline benchmark datapath \\
\rowa MAC simultanei & $2\ldots64$ & $=$\code{PARALLEL}$\times$\code{N\_NEURONS} \\
Attivazioni & ReLU, lineare & \code{ACT\_RELU} / \code{ACT\_NONE} \\
\rowa Fmax (P=2, datapath) & 87.88~MHz & benchmark datapath isolato \\
Fmax (P=2, sistema integrato) & 67.91~MHz & sistema completo incl. sottosistema flash, place\&route reale \\
Throughput MAC (P=16) & $\approx$3.34~G\,MAC/s & teorico, solo datapath \\
\rowa Memoria di lavoro & 8~MB PSRAM & bus parallelo 16-bit, 70~ns / 20~ns page mode \\
Spazio indirizzi & 23~bit (byte) & \code{ADDR\_WIDTH}=23 \\
\bottomrule
\end{tabularx}
\vspace{8pt}
\noindent
{\small\color{fnDark}\bfseries Diagramma a blocchi del sistema}
\begin{center}
\begin{tikzpicture}[node distance=6mm and 10mm,font=\footnotesize]
\node[fnblockD,minimum width=26mm,minimum height=13mm] (host){HOST\\{\scriptsize configura / addestra / controlla}};
\node[fnblockT,right=16mm of host,minimum width=52mm,minimum height=22mm] (eng){};
\node[anchor=north,font=\footnotesize\bfseries,text=fnDark] at (eng.north){FPGA -- Neural Network Engine};
\node[fnreg,fill=white] (spi) at ([yshift=-2mm]eng.center){\code{spi\_slave} + \code{spi\_engine}};
\node[fnreg,fill=white,below=2.5mm of spi] (arb){\code{mem\_arbiter} + \code{layer\_sequencer}};
\node[fnreg,fill=white,above=2.5mm of spi] (core){\code{neuron\_memory} $\to$ \code{neuron\_parallel} $\to$ \code{mac8}};
\node[fnblock,right=16mm of eng,minimum width=24mm,minimum height=13mm] (ram){PSRAM 8\,MB\\{\scriptsize \code{psram\_controller}}};
\draw[fnbus] (host) -- node[fnlbl,above]{SPI} (eng.west|-host);
\draw[fnbus] (eng.east|-ram) -- node[fnlbl,above]{16-bit async} (ram);
\end{tikzpicture}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
Il datapath neurale è interamente nell'FPGA; l'host non partecipa alle singole
operazioni MAC.\end{center}
@@ -1,104 +0,0 @@
\thispagestyle{plain}
\noindent
\begin{tikzpicture}
\node[fill=fnDark,text=white,rounded corners=2pt,inner sep=6pt,
minimum width=\textwidth,anchor=west]
{\large\bfseries Sunto del pinout --- collegamento pin per pin};
\end{tikzpicture}
\vspace{6pt}
\noindent
{\footnotesize
Tabella di riferimento rapido: i \textbf{57 segnali reali} del top-level
\code{spi\_neuron\_top}, ciascuno con la propria ball \code{CABGA381}
individuale (\textbf{non} un intervallo di bus) --- dati reali dal database
di dispositivo di Project~Trellis (\code{iodb.json}), \textbf{verificati da
un place\&route \code{nextpnr-ecp5} completo a 0 errori} (non un pinout
pianificato). Descrizione completa, razionale di collocazione per banco e
schema di collegamento pin-per-pin verso la PSRAM ISSI: cap.~\ref{ch:hw}.
}
\vspace{4pt}
\noindent
\renewcommand{\arraystretch}{1.08}
\begin{tabularx}{\textwidth}{L{2.7cm} C{1.0cm} C{1.0cm} C{0.9cm} Y}
\toprule
\rowh \thd{Segnale} & \thd{Ball} & \thd{Banco} & \thd{Dir} & \thd{Pin corrispondente / funzione} \\
\midrule
\multicolumn{5}{l}{\textit{\color{fnDark}Clock e reset}}\\
\code{clk} & H5 & 7 & IN & Clock di sistema, pad \code{GR\_PCLK7\_0} (clock globale dedicato). \\
\rowa \code{rst} & B4 & 7 & IN & Reset globale sincrono, attivo alto. \\
\multicolumn{5}{l}{\textit{\color{fnDark}SPI applicativo (host $\leftrightarrow$ FPGA, Mode~0)}}\\
\code{sclk} & B5 & 7 & IN & SPI clock (CPOL=0, CPHA=0). \\
\rowa \code{mosi} & C5 & 7 & IN & Master-Out Slave-In. \\
\code{miso} & A3 & 7 & OUT & Master-In Slave-Out. \\
\rowa \code{cs\_n} & B3 & 7 & IN & Chip-select, attivo basso. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Attenzione host (attivi bassi, di livello)}}\\
\code{data\_ready\_n} & C3 & 7 & OUT & Basso finché un risultato attende lettura. \\
\rowa \code{irq\_n} & C4 & 7 & OUT & Basso se il guard load-time del grafo è scattato. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Flash subsystem --- SPI verso W25Q128JV (boot/persistenza)}}\\
\code{flash\_sclk} & E3 & 7 & OUT & SPI clock verso la flash --- GPIO ordinario, indipendente (Fase F7, cap.~\ref{ch:hw}). \\
\rowa \code{flash\_mosi} & D3 & 7 & OUT & Master-Out Slave-In verso la flash NOR onboard. \\
\code{flash\_miso} & D5 & 7 & IN & Master-In Slave-Out dalla flash. \\
\rowa \code{flash\_cs\_n} & E4 & 7 & OUT & Chip-select flash, attivo basso. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Bus indirizzi PSRAM --- \code{psram\_a[21:0]} (22 linee reali)}}\\
\code{psram\_a[0]} & E16 & 2 & OUT & PSRAM A0 \\
\rowa \code{psram\_a[1]} & F16 & 2 & OUT & PSRAM A1 \\
\code{psram\_a[2]} & D18 & 2 & OUT & PSRAM A2 \\
\rowa \code{psram\_a[3]} & E17 & 2 & OUT & PSRAM A3 \\
\code{psram\_a[4]} & E18 & 2 & OUT & PSRAM A4 \\
\rowa \code{psram\_a[5]} & F18 & 2 & OUT & PSRAM A5 \\
\code{psram\_a[6]} & F17 & 2 & OUT & PSRAM A6 \\
\rowa \code{psram\_a[7]} & G16 & 2 & OUT & PSRAM A7 \\
\code{psram\_a[8]} & G18 & 2 & OUT & PSRAM A8 \\
\rowa \code{psram\_a[9]} & H16 & 2 & OUT & PSRAM A9 \\
\code{psram\_a[10]} & H17 & 2 & OUT & PSRAM A10 \\
\rowa \code{psram\_a[11]} & H18 & 2 & OUT & PSRAM A11 \\
\code{psram\_a[12]} & J16 & 2 & OUT & PSRAM A12 \\
\rowa \code{psram\_a[13]} & J17 & 2 & OUT & PSRAM A13 \\
\code{psram\_a[14]} & C20 & 2 & OUT & PSRAM A14 \\
\rowa \code{psram\_a[15]} & D19 & 2 & OUT & PSRAM A15 \\
\code{psram\_a[16]} & E19 & 2 & OUT & PSRAM A16 \\
\rowa \code{psram\_a[17]} & E20 & 2 & OUT & PSRAM A17 \\
\code{psram\_a[18]} & F19 & 2 & OUT & PSRAM A18 \\
\rowa \code{psram\_a[19]} & F20 & 2 & OUT & PSRAM A19 \\
\code{psram\_a[20]} & G20 & 2 & OUT & PSRAM A20 \\
\rowa \code{psram\_a[21]} & H20 & 2 & OUT & PSRAM A21 \\
\code{psram\_a[22]} & P18 & 3 & OUT & Sempre 0 (shift byte$\to$word) --- NC su scheda. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Bus dati PSRAM --- \code{psram\_dq[15:0]} (bidirezionale)}}\\
\rowa \code{psram\_dq[0]} & K18 & 2 & IO & PSRAM DQ0 \\
\code{psram\_dq[1]} & C18 & 2 & IO & PSRAM DQ1 \\
\rowa \code{psram\_dq[2]} & D17 & 2 & IO & PSRAM DQ2 \\
\code{psram\_dq[3]} & D20 & 2 & IO & PSRAM DQ3 \\
\rowa \code{psram\_dq[4]} & G19 & 2 & IO & PSRAM DQ4 \\
\code{psram\_dq[5]} & J18 & 2 & IO & PSRAM DQ5 \\
\rowa \code{psram\_dq[6]} & J19 & 2 & IO & PSRAM DQ6 \\
\code{psram\_dq[7]} & J20 & 2 & IO & PSRAM DQ7 \\
\rowa \code{psram\_dq[8]} & K19 & 2 & IO & PSRAM DQ8 \\
\code{psram\_dq[9]} & K20 & 2 & IO & PSRAM DQ9 \\
\rowa \code{psram\_dq[10]} & L17 & 3 & IO & PSRAM DQ10 \\
\code{psram\_dq[11]} & M18 & 3 & IO & PSRAM DQ11 \\
\rowa \code{psram\_dq[12]} & M17 & 3 & IO & PSRAM DQ12 \\
\code{psram\_dq[13]} & N16 & 3 & IO & PSRAM DQ13 \\
\rowa \code{psram\_dq[14]} & N18 & 3 & IO & PSRAM DQ14 \\
\code{psram\_dq[15]} & P17 & 3 & IO & PSRAM DQ15 \\
\multicolumn{5}{l}{\textit{\color{fnDark}Controllo PSRAM}}\\
\rowa \code{psram\_ce\_n} & N17 & 3 & OUT & PSRAM CE\# --- chip enable, attivo basso. \\
\code{psram\_oe\_n} & R16 & 3 & OUT & PSRAM OE\# --- output enable (lettura). \\
\rowa \code{psram\_we\_n} & R17 & 3 & OUT & PSRAM WE\# --- write enable (scrittura). \\
\code{psram\_lb\_n} & T16 & 3 & OUT & PSRAM LB\# --- lower-byte enable (DQ[7:0]). \\
\rowa \code{psram\_ub\_n} & N19 & 3 & OUT & PSRAM UB\# --- upper-byte enable (DQ[15:8]). \\
\code{psram\_zz\_n} & N20 & 3 & OUT & PSRAM ZZ\# --- sleep/snooze (alto in funzionamento). \\
\bottomrule
\end{tabularx}
\renewcommand{\arraystretch}{1.25}
\vspace{4pt}
\noindent
{\footnotesize\color{fnGrey}
Standard I/O: LVCMOS33 su tutti i 57 segnali. Ball di JTAG e config-SPI di
boot (pin dedicati a funzione fissa, senza porta RTL) non compaiono in
questa tabella --- vedi cap.~\ref{ch:hw} §``Configurazione e
programmazione''. Sorgente: \code{synth/ecp5/spi\_neuron\_top.lpf},
generato da \code{tools/pinout/gen\_lpf.py} contro
\code{iodb.json} di Project~Trellis.\par}
@@ -1,94 +0,0 @@
\chapter{Panoramica del sistema}
\label{ch:overview}
\section{Obiettivo del progetto}
FPGA-Neural implementa un \textbf{Neural Network Engine riusabile in hardware FPGA}.
L'insieme è composto da tre elementi: l'FPGA, che è il vero acceleratore; una RAM
dedicata fisicamente associata all'FPGA e non condivisa con l'host; e un'interfaccia
host indipendente dal sistema operativo, inizialmente SPI (con possibile estensione
futura a Dual~SPI).
Il principio fondante è la separazione fra chi \emph{esegue} il calcolo e chi lo
\emph{usa}: il calcolo della rete neurale avviene interamente dentro l'FPGA, mentre
il sistema host fornisce solo configurazione, parametri di rete, dati di ingresso,
controllo e lettura dei risultati. L'host non fa parte del datapath computazionale.
Sistemi host possibili includono SoC Linux, sistemi tipo Raspberry~Pi, ESP32,
microcontrollori e PC di sviluppo: la stessa architettura di engine deve poter essere
usata in sistemi completamente diversi.
\begin{center}
\begin{tikzpicture}[font=\footnotesize,node distance=8mm]
\node[fnblockD,minimum width=42mm,minimum height=20mm] (host){\textbf{HOST}\\[2pt]
{\scriptsize Configurazione}\\{\scriptsize Addestramento}\\{\scriptsize Controllo}};
\node[fnblockT,below=14mm of host,minimum width=42mm,minimum height=20mm] (fpga)
{\textbf{FPGA}\\[2pt]{\scriptsize Neural Network Engine}\\{\scriptsize Compute / Control}};
\node[fnblock,below=14mm of fpga,minimum width=42mm,minimum height=13mm] (ram)
{\textbf{RAM dedicata}\\{\scriptsize pesi / bias / buffer}};
\draw[fnbus] (host) -- node[fnlbl,right]{SPI / Dual SPI} (fpga);
\draw[fnbus] (fpga) -- node[fnlbl,right]{bus parallelo} (ram);
\end{tikzpicture}
\end{center}
\section{Configurazione hardware contro configurazione di rete}
Il progetto distingue con precisione fra l'\textbf{architettura hardware}
dell'acceleratore e i \textbf{parametri della rete neurale}.
L'architettura fisica dell'engine è definita al momento della sintesi e
dell'implementazione dell'FPGA. I parametri hardware tipici sono \code{N\_INPUTS},
\code{N\_NEURONS}, \code{N\_LAYERS}, \code{PARALLEL}, \code{DATA\_WIDTH},
\code{ACC\_WIDTH}: sono parametri Verilog risolti in fase di sintesi e determinano il
datapath contenuto nel bitstream. I parametri della rete --- pesi, bias, parametri di
attivazione e di quantizzazione, costanti specifiche --- vengono invece caricati a
runtime attraverso l'interfaccia host e memorizzati nella RAM associata all'FPGA.
\begin{fnnote}[Principio architetturale centrale]
Una build fissa il \emph{soffitto} della macchina (numero massimo di layer, larghezza
massima, \code{PARALLEL}); l'host configura la rete \emph{reale} --- numero di layer,
larghezza ingressi/uscite per-layer, attivazione per-layer e parametri addestrati ---
interamente a runtime, via SPI, nella memoria locale dell'FPGA. Un solo bitstream
serve qualunque topologia fino a quel soffitto.
\end{fnnote}
\section{Boot e inizializzazione}
L'FPGA viene configurato all'accensione tramite il consueto meccanismo di
configurazione (caricamento del bitstream da flash SPI). Il bitstream definisce
l'architettura hardware dell'engine; l'host non costruisce dinamicamente il datapath
durante il funzionamento normale, ma configura i dati di rete su cui il datapath già
esistente opera.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=4.5mm,start chain=going below,
every node/.style={on chain}]
\node[fnblockA,minimum width=60mm](p){Power-on};
\node[fnblock,minimum width=60mm]{Configurazione FPGA (bitstream da flash)};
\node[fnblockT,minimum width=60mm]{Neural Network Engine disponibile};
\node[fnblock,minimum width=60mm]{Inizializzazione host (SPI)};
\node[fnblock,minimum width=60mm]{Caricamento parametri di rete / pesi / bias};
\node[fnblockD,minimum width=60mm]{Engine pronto};
\begin{scope}[every path/.style={fnarrow}]
\foreach \a/\b in {1/2,2/3,3/4,4/5,5/6}{}
\end{scope}
\foreach \i [count=\j from 2] in {1,...,5}{
\draw[fnarrow] (chain-\i) -- (chain-\j);}
\end{tikzpicture}
\end{center}
\section{Addestramento e inferenza}
Addestramento e inferenza sono concettualmente separati. La prima implementazione non
richiede che l'FPGA esegua l'addestramento: i pesi possono essere calcolati
esternamente (PC/Linux/altro host) e trasferiti via SPI nella RAM dell'FPGA, che poi
esegue l'inferenza. Questo riduce drasticamente la complessità dell'hardware iniziale,
senza precludere una futura implementazione di training assistito o interamente
hardware (Fase~8 della roadmap, cap.~\ref{ch:roadmap}). Durante l'inferenza l'host
fornisce solo i dati di ingresso e recupera il risultato, ottenendo calcolo
deterministico, carico ridotto sull'host, parallelismo hardware, latenza prevedibile e
indipendenza dall'architettura della CPU host.
\section{Filosofia di progetto e riuso}
Il progetto va inteso come una \emph{piattaforma di accelerazione neurale FPGA
riusabile} più che come una singola rete. L'applicazione determina dimensione degli
ingressi, topologia, numero di layer e neuroni, parallelismo, precisione numerica,
funzioni di attivazione, requisiti di memoria e prestazioni; il processo di
generazione hardware produce l'implementazione FPGA corrispondente. La stessa
architettura HDL rimane concettualmente invariata mentre i parametri di sintesi
generano implementazioni appropriate ai diversi target applicativi.
@@ -1,80 +0,0 @@
\chapter[Architettura RTL]{Architettura RTL e gerarchia dei moduli}
\label{ch:arch}
\section{Organizzazione gerarchica}
Il design è organizzato per livelli, dal moltiplicatore-accumulatore elementare fino
al top-level integrato con interfaccia SPI e PSRAM. Ogni livello incapsula il
precedente e ne astrae i dettagli: il datapath validato (\code{mac\_unit},
\code{mac8}, \code{neuron\_parallel}) non viene mai modificato dai livelli di
orchestrazione superiori.
\begin{center}
\begin{tikzpicture}[font=\footnotesize,every node/.style={fnblock,minimum width=40mm},
level distance=13mm,sibling distance=0mm]
\node[fnblockD,minimum width=62mm](top){\code{spi\_neuron\_top} \\ {\scriptsize top-level integrato}};
\node[fnblockT,minimum width=62mm,below=8mm of top](arb){\code{mem\_arbiter} \;/\; \code{layer\_sequencer} \\ {\scriptsize arbitraggio 3 porte + sequenza layer}};
\node[fnblock,minimum width=62mm,below=8mm of arb](nm){\code{neuron\_memory} \\ {\scriptsize ponte memoria $\leftrightarrow$ neurone, loop neuroni}};
\node[fnblock,minimum width=62mm,below=8mm of nm](np){\code{neuron\_parallel} \\ {\scriptsize FSM neurone: gruppi, bias, attivazione, saturazione}};
\node[fnblockT,minimum width=62mm,below=8mm of np](m8){\code{mac8} \\ {\scriptsize \code{PARALLEL} MAC + balanced adder tree}};
\node[fnblock,minimum width=62mm,below=8mm of m8](mu){\code{mac\_unit} \\ {\scriptsize $x\cdot w$ + estensione segno + accumulo}};
\foreach \a/\b in {top/arb,arb/nm,nm/np,np/m8,m8/mu}
\draw[fnarrow] (\a) -- (\b);
% rami memoria a destra
\node[fnblockA,minimum width=34mm,right=14mm of nm](ma){\code{int8\_memory\_access}\\{\scriptsize byte $\leftrightarrow$ word 16-bit}};
\node[fnblockA,minimum width=34mm,below=6mm of ma](mi){\code{memory\_interface}\\{\scriptsize handshake req/ready}};
\node[fnblockA,minimum width=34mm,below=6mm of mi](pc){\code{psram\_controller}\\{\scriptsize bus fisico PSRAM}};
\draw[fnarrowT] (ma)--(mi); \draw[fnarrowT] (mi)--(pc);
\draw[fnarrowT,dashed] (nm.east) -- (ma.west);
% rami SPI a sinistra
\node[fnblockA,minimum width=30mm,left=14mm of arb,yshift=6mm](ss){\code{spi\_slave}\\{\scriptsize layer fisico Mode 0}};
\node[fnblockA,minimum width=30mm,below=6mm of ss](se){\code{spi\_engine}\\{\scriptsize FSM opcode + registri}};
\draw[fnarrowT] (ss)--(se);
\draw[fnarrowT,dashed] (se.east) -- (arb.west);
\end{tikzpicture}
\end{center}
\section{Ruolo di ciascun modulo}
\begin{tabularx}{\textwidth}{L{3.4cm}Y}
\toprule
\rowh \thd{Modulo} & \thd{Funzione} \\
\midrule
\code{mac\_unit} & Singolo prodotto-accumulatore: $\mathrm{acc\_out}=\mathrm{acc\_in}+(x\cdot w)$, con estensione di segno del prodotto ad \code{ACC\_WIDTH}. Parametrico su \code{DATA\_WIDTH}/\code{ACC\_WIDTH}. \\
\rowa \code{mac8} & \code{PARALLEL} istanze di \code{mac\_unit} i cui prodotti vengono sommati da un \emph{balanced binary adder tree} di profondità $\log_2(\text{PARALLEL})$; il risultato è aggiunto all'accumulatore in ingresso. \\
\code{neuron\_parallel} & FSM di un singolo neurone: elabora \code{N\_INPUTS} ingressi in gruppi di \code{PARALLEL}, accumula tra i gruppi, somma il bias, applica l'attivazione e satura a INT8. Include il guard di elaborazione su \code{N\_INPUTS \% PARALLEL} e la larghezza runtime \code{n\_inputs\_real}. \\
\rowa \code{layer} & Istanzia \code{N\_NEURONS} neuroni \emph{in parallelo} sullo stesso vettore di ingresso; \code{busy}=OR, \code{done}=AND dei neuroni. Percorso puramente combinatorio-di-dati usato nei benchmark del datapath. \\
\code{neuron\_memory} & Integra il calcolo con la memoria: legge $X$ (condiviso) una volta, poi per ogni neurone rilegge $W$ e bias dalla RAM e riusa una singola istanza \code{neuron\_parallel} (memory-bound, un neurone per volta). Uscita \code{y\_bus} packed neuron-major. \\
\rowa \code{layer\_sequencer} & Concatena fino a \code{N\_LAYERS} esecuzioni di \code{neuron\_memory} leggendo una tabella descrittori scritta dall'host e alternando i buffer ping-pong in RAM (Fase~5). \\
\code{act\_buffer} & Buffer di attivazione globale in block RAM \code{DP16KD}, indicizzato per id di segnale (Tipo \#2). \\
\rowa \code{graph\_engine} & Motore della rete a grafo (Tipo \#2): gather da \code{act\_buffer}, riusa \code{neuron\_parallel}, scrive le uscite per id (cap.~\ref{ch:grafo}). \\
\code{int8\_memory\_access} & Converte l'interfaccia byte/INT8 (indirizzo di byte) nell'interfaccia a parola 16-bit, selezionando il byte basso/alto tramite \code{lb\_n}/\code{ub\_n} e \code{addr>>1}. \\
\rowa \code{memory\_interface} & FSM di handshake a 2 stati (IDLE/WAIT) che serializza la singola transazione verso il controller. \\
\code{psram\_controller} & Controller del bus PSRAM parallelo asincrono con \textbf{page mode} di lettura: accesso casuale a 70~ns (\code{tAA}), burst nella stessa pagina a 20~ns (\code{tAPA}) con CE\#/OE\# tenuti attivi; abilita il page mode sul chip all'avvio via registro di configurazione (cap.~\ref{ch:mem}, \S~5.5). Pilota \code{ce\_n/oe\_n/we\_n/lb\_n/ub\_n/zz\_n} e il bus dati tri-state. \\
\rowa \code{mem\_arbiter} & Arbitro a priorità fissa (B$>$C$>$A) fra tre master byte-level: \code{spi\_engine} (A), \code{neuron\_memory} (B), \code{layer\_sequencer} (C). \\
\code{spi\_slave} & Layer fisico SPI Mode 0, MSB-first, sincronizzatore CDC a 3 stadi su SCLK/MOSI/CS\_N, shift-register e framing di CS. \\
\rowa \code{spi\_engine} & FSM di protocollo/opcode e banco registri (\code{x\_base}, \code{w\_base}, \code{bias\_addr}, base ping-pong, attivazione, larghezze runtime\ldots), con \code{STATUS.done} sticky/clear-on-read. \\
\code{spi\_neuron\_top} & Top-level: collega SPI, arbitro, sequencer, \code{neuron\_memory} e catena PSRAM; multiplexa il controllo di \code{neuron\_memory} fra sequencer e percorso diretto single-layer. \\
\bottomrule
\end{tabularx}
\vspace{6pt}
\begin{fnnote}[Modelli di simulazione]
\code{psram\_model.v} (in \code{sim/}) e \code{memory\_model.v} sono modelli
comportamentali della memoria usati nei testbench; non fanno parte del design
sintetizzabile ma riproducono la latenza reale per la verifica end-to-end.
\end{fnnote}
\section{Due percorsi di esecuzione}
Il top-level espone due modalità mutuamente esclusive verso lo stesso motore di
calcolo \code{neuron\_memory}:
\begin{itemize}
\item \textbf{Percorso single-layer / manuale}: l'host imposta le basi con
\op{SET\_BASE}, avvia con \op{START} e legge con \op{READ\_OUTPUT}. \code{spi\_engine}
pilota direttamente \code{neuron\_memory}.
\item \textbf{Percorso multi-layer}: l'host scrive la tabella descrittori e avvia con
\op{RUN\_NETWORK}; \code{layer\_sequencer} prende possesso del controllo di
\code{neuron\_memory} (mentre \code{seq\_busy} è alto) e concatena i layer.
\end{itemize}
Il multiplexer del top-level commuta le linee di controllo di \code{neuron\_memory}
in base a \code{seq\_busy}, restituendo il motore al percorso diretto a fine sequenza.
@@ -1,165 +0,0 @@
\chapter{Datapath di calcolo}
\label{ch:datapath}
\section{Catena aritmetica INT8/INT32}
Il datapath elementare implementa la sequenza tipica di un neurone quantizzato:
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=3mm,start chain=going right,
every node/.style={fnblock,minimum width=15mm,minimum height=8mm,on chain}]
\node[fnblockT]{INT8\\$\times$\,INT8};
\node{INT16\\prodotto};
\node{sign-ext\\INT32};
\node[fnblockD]{accumulo\\INT32};
\node{$+$ bias};
\node[fnblockA]{attivazione};
\node[fnblockT]{sat. INT8};
\foreach \i [count=\j from 2] in {1,...,6}
\draw[fnarrow] (chain-\i) -- (chain-\j);
\end{tikzpicture}
\end{center}
Ogni prodotto INT8$\times$INT8 sta in 16~bit; viene esteso con segno a 32~bit prima
dell'accumulo, così l'accumulatore non trabocca su vettori lunghi. Bias e attivazione
operano a 32~bit; solo l'uscita finale viene saturata a INT8.
\section{\texttt{mac\_unit} --- moltiplicatore-accumulatore}
Il modulo \code{mac\_unit} è puramente combinatorio e parametrico su \code{DATA\_WIDTH}
e \code{ACC\_WIDTH}. Calcola:
\[
\mathrm{acc\_out} = \mathrm{acc\_in} + \mathrm{signext}_{ACC}(x \cdot w)
\]
Il prodotto ha larghezza $2\times$\code{DATA\_WIDTH} e viene esteso con segno replicando
il bit più significativo. Su ECP5 la moltiplicazione mappa su un blocco DSP
\code{MULT18X18D}.
\begin{lstlisting}[caption={\texttt{rtl/mac\_unit.v} --- nucleo aritmetico},label={lst:macunit}]
localparam PROD_WIDTH = 2 * DATA_WIDTH;
wire signed [PROD_WIDTH-1:0] product = x * w;
wire signed [ACC_WIDTH-1:0] product_ext =
{{(ACC_WIDTH-PROD_WIDTH){product[PROD_WIDTH-1]}}, product};
assign acc_out = acc_in + product_ext;
\end{lstlisting}
\section{\texttt{mac8} --- MAC parallelo e balanced adder tree}
\code{mac8} istanzia \code{PARALLEL} unità \code{mac\_unit} che generano
\code{PARALLEL} prodotti indipendenti, poi li somma con un \emph{albero di addizione
binario bilanciato}. Rispetto alla riduzione lineare
$((((p_0{+}p_1){+}p_2){+}p_3){+}\dots)$, di profondità $O(\text{PARALLEL})$, l'albero
ha profondità $O(\log_2 \text{PARALLEL})$, riducendo drasticamente il percorso
combinatorio.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,level distance=11mm,
every node/.style={fnreg,minimum width=8mm},
level 1/.style={sibling distance=30mm},
level 2/.style={sibling distance=15mm},
level 3/.style={sibling distance=8mm},
edge from parent/.style={fnarrowT,draw}]
\node[fnblockD]{sum}
child {node[fnblockT]{$+$}
child {node[fnblockT]{$+$}
child {node{$p_0$}} child {node{$p_1$}}}
child {node[fnblockT]{$+$}
child {node{$p_2$}} child {node{$p_3$}}}}
child {node[fnblockT]{$+$}
child {node[fnblockT]{$+$}
child {node{$p_4$}} child {node{$p_5$}}}
child {node[fnblockT]{$+$}
child {node{$p_6$}} child {node{$p_7$}}}};
\end{tikzpicture}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
Esempio con PARALLEL=8: 3 livelli. PARALLEL=16 $\to$ 4 livelli; PARALLEL=32 $\to$ 5
livelli.\end{center}
\begin{fnnote}[PARALLEL potenza di due]
L'albero è pensato per \code{PARALLEL} potenza di due (8, 16, 32\ldots). Questo è anche
il valore usato in tutte le configurazioni del progetto.
\end{fnnote}
\section{\texttt{neuron\_parallel} --- FSM del neurone}
\code{neuron\_parallel} elabora \code{N\_INPUTS} ingressi in gruppi di \code{PARALLEL},
mantenendo l'accumulatore tra un gruppo e il successivo. Alla fine somma il bias,
applica l'attivazione e satura a INT8. Il numero di gruppi è
$\text{GROUPS}=\text{N\_INPUTS}/\text{PARALLEL}$.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=4mm,start chain=going below,
every node/.style={on chain,fnblock,minimum width=46mm}]
\node[fnblockA]{\code{start}};
\node{gruppo 0 $\to$ accumulo};
\node{gruppo 1 $\to$ accumulo};
\node[draw=none,fill=none]{\vdots};
\node{gruppo GROUPS$-$1 $\to$ accumulo};
\node{$+$ bias};
\node[fnblockA]{attivazione (ACT\_RELU / ACT\_NONE)};
\node[fnblockT]{saturazione INT8};
\node[fnblockD]{\code{done}, \code{y}};
\foreach \i [count=\j from 2] in {1,...,8}
\draw[fnarrow] (chain-\i) -- (chain-\j);
\end{tikzpicture}
\end{center}
\subsection{Guard di parametri (elaboration-time)}
Se \code{PARALLEL} non divide esattamente \code{N\_INPUTS} si verificano due guasti,
entrambi confermati empiricamente in \code{sim/parameter\_sweep\_tb.v}:
\begin{itemize}
\item la divisione intera tronca \code{GROUPS} e gli ingressi in eccesso non vengono
mai letti $\to$ risultato \textbf{errato}, senza errore né avviso;
\item se \code{PARALLEL > N\_INPUTS}, \code{GROUPS=0} e la condizione terminale non è mai
soddisfatta $\to$ il neurone \textbf{si blocca} (busy alto, done mai asserito).
\end{itemize}
La soluzione non modifica il datapath validato: un blocco \code{generate} istanzia un
modulo deliberatamente indefinito quando $\text{N\_INPUTS} \bmod \text{PARALLEL}\neq0$,
forzando un errore in \emph{elaborazione} sia in simulazione sia in sintesi. Per le
configurazioni valide il ramo non viene mai elaborato.
\begin{lstlisting}[caption={\texttt{rtl/neuron\_parallel.v} --- guard di parametri}]
generate
if (N_INPUTS == 0 || N_INPUTS % PARALLEL != 0) begin : PARAMETER_ERROR
neuron_parallel_requires_N_INPUTS_multiple_of_PARALLEL
invalid_parameter_combination();
end
endgenerate
\end{lstlisting}
\begin{fnnote}[Caso limite \texttt{N\_INPUTS=0} (corretto 2026-09-04)]
La condizione originale (\code{N\_INPUTS \% PARALLEL != 0}) non intercetta
\code{N\_INPUTS=0}, poiché $0 \bmod \text{PARALLEL}=0$ per ogni \code{PARALLEL}: il modulo
elaborava con successo (sia in simulazione sia in sintesi reale Yosys) lasciando
\code{x\_bus}/\code{w\_bus} non pilotati e \code{start} silenziosamente inefficace. Trovato
durante la campagna di ri-certificazione (\code{docs/validation/bugs.md}, BUG-002) e
corretto estendendo il guard come sopra --- \code{N\_INPUTS=0} ora fallisce l'elaborazione
esattamente come gli altri casi degeneri.
\end{fnnote}
\section{Funzioni di attivazione}
\code{neuron\_parallel} accetta una porta \code{activation} a 2~bit. Il default è
\code{ACT\_RELU}, l'unico comportamento esistente prima dell'introduzione della porta,
così ogni chiamante preesistente resta invariato.
\begin{tabularx}{\textwidth}{L{2.6cm} C{1.4cm} Y}
\toprule
\rowh \thd{Codifica} & \thd{Valore} & \thd{Comportamento} \\
\midrule
\code{ACT\_NONE} & \code{2'd0} & Lineare: nessun clamp a zero, saturazione bilaterale al range INT8 $[-128,+127]$. \\
\rowa \code{ACT\_RELU} & \code{2'd1} & $\max(0,x)$, poi saturazione positiva a $+127$ (default; fallback anche per codifiche riservate). \\
\bottomrule
\end{tabularx}
\section{Saturazione INT8}
Dopo bias e attivazione, l'accumulatore a 32~bit viene ridotto a INT8:
\[
y=\begin{cases}
+127 & \text{se } \mathrm{final\_acc} > 127\\
-128 & \text{se } \mathrm{final\_acc} < -128 \ \text{(solo ACT\_NONE)}\\
0 & \text{se } \mathrm{final\_acc}\le 0 \ \text{(solo ACT\_RELU)}\\
\mathrm{final\_acc}[7:0] & \text{altrimenti}
\end{cases}
\]
\section{\texttt{layer} --- neuroni in parallelo}
\code{layer} istanzia \code{N\_NEURONS} neuroni che condividono il vettore di ingresso
\code{x\_bus} ma hanno pesi e bias distinti; \code{busy} è l'OR e \code{done} l'AND dei
segnali dei neuroni. È il modulo usato nei benchmark del datapath (cap.~\ref{ch:impl}),
dove tutti i neuroni lavorano simultaneamente. La convenzione di indirizzamento è
neuron-major: i pesi del neurone $n$ occupano \code{weights\_bus[n*N\_INPUTS*DATA\_WIDTH +: N\_INPUTS*DATA\_WIDTH]}.
@@ -1,87 +0,0 @@
\chapter{Parametri e configurabilità}
\label{ch:param}
\section{Parametri di build (synthesis-time)}
L'architettura hardware è fissata alla sintesi tramite i parametri Verilog seguenti.
Determinano il datapath contenuto nel bitstream e il suo \emph{soffitto} di capacità.
\begin{tabularx}{\textwidth}{L{3.0cm} C{1.8cm} Y}
\toprule
\rowh \thd{Parametro} & \thd{Default} & \thd{Significato} \\
\midrule
\code{DATA\_WIDTH} & 8 & Larghezza dei dati (INT8). \\
\rowa \code{ACC\_WIDTH} & 32 & Larghezza dell'accumulatore (INT32). \\
\code{N\_INPUTS} & 32 / 256 & Numero massimo di ingressi per neurone (baseline benchmark: 256). \\
\rowa \code{N\_NEURONS} & 1 / 4 & Numero massimo di neuroni per layer. \\
\code{PARALLEL} & 8 & MAC hardware simultanei per neurone; deve dividere \code{N\_INPUTS} e conviene sia potenza di due. \\
\rowa \code{N\_LAYERS} & 4 & Numero massimo di layer concatenabili da \code{layer\_sequencer}. \\
\code{ADDR\_WIDTH} & 23 & Larghezza dell'indirizzo di byte (8~MB). \\
\rowa \code{MEM\_DATA\_WIDTH} & 16 & Larghezza del bus dati fisico PSRAM. \\
\code{CLK\_FREQ\_MHZ} & 80 & Frequenza usata per le formule di temporizzazione PSRAM (va allineata all'oscillatore reale). \\
\bottomrule
\end{tabularx}
\begin{fnwarn}[Vincolo \texttt{N\_INPUTS} \% \texttt{PARALLEL}]
\code{PARALLEL} deve dividere esattamente \code{N\_INPUTS}, altrimenti scatta il guard
di elaborazione (§\ref{ch:datapath}). Lo stesso vincolo vale a runtime su
\code{n\_inputs\_real}.
\end{fnwarn}
\section{Larghezza di rete a runtime}
Un singolo bitstream serve qualunque topologia \emph{fino} al massimo di build. La
larghezza reale di ciascuna esecuzione è un valore separato, impostato dall'host:
\begin{itemize}
\item \code{n\_inputs\_real} --- ingressi realmente usati in questa esecuzione (deve
essere multiplo di \code{PARALLEL});
\item \code{n\_neurons\_real} --- neuroni realmente calcolati in questa esecuzione.
\end{itemize}
Entrambi hanno default pari al massimo di build, così ogni chiamante che li lascia
scollegati elabora l'intera larghezza come prima dell'introduzione delle porte.
\begin{fnnote}[Terminazione anticipata reale]
Non si tratta di semplice contabilità di indirizzi: i due valori limitano
direttamente i loop hardware (letture X/W di \code{neuron\_memory}, conteggio gruppi
MAC di \code{neuron\_parallel} e lunghezza della copia ping-pong per \code{RUN\_NETWORK}).
Un layer più stretto \emph{calcola} e \emph{copia} davvero più in fretta e non richiede
zero-padding della RAM per la coda non usata: i dati oltre
\code{n\_inputs\_real}/\code{n\_neurons\_real} non vengono mai letti.
\end{fnnote}
Questo permette a una rete di rastremarsi dentro una sola esecuzione concatenata, ad
esempio $256\to64\to16\to4$, con ogni layer che dichiara la propria larghezza reale
nella tabella descrittori (cap.~\ref{ch:seq}).
\subsection{Risparmio misurato}
La terminazione anticipata è stata misurata end-to-end:
\begin{tabularx}{\textwidth}{L{5.5cm} C{3.0cm} Y}
\toprule
\rowh \thd{Test} & \thd{Cicli} & \thd{Confronto} \\
\midrule
\code{neuron\_parallel\_tb.v} (T7) & 3 vs 6 & ridotto vs pieno, con dati ``spazzatura'' nelle corsie saltate (prova che non vengono lette). \\
\rowa \code{neuron\_memory\_tb.v} (T5) & 209 vs 788 & 8-di-32 vs 32 pieni, attraverso lo stack PSRAM reale. \\
\bottomrule
\end{tabularx}
\section{Configurazioni caratterizzate}
Alcune combinazioni convalidate in simulazione e/o sintesi:
\begin{tabularx}{\textwidth}{C{2.0cm} C{2.0cm} C{2.0cm} Y}
\toprule
\rowh \thd{N\_INPUTS} & \thd{N\_NEURONS} & \thd{PARALLEL} & \thd{Note} \\
\midrule
32 & 4 & 8 & Primo test parametrico funzionale (Fase~1). \\
\rowa 256 & 4 & 2/4/8/16 & Sweep di benchmark del datapath (Fase~7). \\
32 & 1..3 & 8 & Integrazione memoria mono/multi-neurone (Fase~3). \\
\rowa 4 & 4 & 2 & Test end-to-end \code{RUN\_NETWORK} a 2 layer su SPI reale. \\
\bottomrule
\end{tabularx}
\section{Riepilogo build contro runtime}
\begin{center}
\begin{tikzpicture}[font=\footnotesize,node distance=6mm]
\node[fnblockD,minimum width=54mm,minimum height=15mm](b){\textbf{BUILD (sintesi)}\\[2pt]
{\scriptsize N\_INPUTS, N\_NEURONS, N\_LAYERS,}\\{\scriptsize PARALLEL, DATA\_WIDTH, ACC\_WIDTH}\\{\scriptsize $\Rightarrow$ soffitto della macchina}};
\node[fnblockT,right=16mm of b,minimum width=54mm,minimum height=15mm](r){\textbf{RUNTIME (host, SPI)}\\[2pt]
{\scriptsize n\_inputs\_real, n\_neurons\_real,}\\{\scriptsize attivazione, num\_layers, pesi/bias}\\{\scriptsize $\Rightarrow$ rete effettiva}};
\draw[fnbus] (b) -- node[fnlbl,above]{$\le$} (r);
\end{tikzpicture}
\end{center}
@@ -1,193 +0,0 @@
\chapter{Sottosistema di memoria}
\label{ch:mem}
\section{Catena di memoria}
Il motore di calcolo lavora con indirizzi e dati a livello di \emph{byte} (INT8), mentre
la PSRAM è un dispositivo a parola da 16~bit. Tre moduli in cascata realizzano la
conversione e l'accesso fisico:
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=8mm]
\node[fnblockD,minimum width=30mm,minimum height=12mm](nm){master byte-level\\{\scriptsize \code{neuron\_memory} / \code{spi\_engine} / \code{layer\_sequencer}}};
\node[fnblockT,right=10mm of nm,minimum width=28mm,minimum height=12mm](ia){\code{int8\_memory\_access}\\{\scriptsize byte $\leftrightarrow$ word 16-bit}};
\node[fnblock,right=10mm of ia,minimum width=26mm,minimum height=12mm](mi){\code{memory\_interface}\\{\scriptsize FSM IDLE/WAIT}};
\node[fnblockA,below=9mm of mi,minimum width=26mm,minimum height=12mm](pc){\code{psram\_controller}\\{\scriptsize bus fisico async 70\,ns}};
\node[fnblock,left=10mm of pc,minimum width=26mm,minimum height=12mm](ps){PSRAM\\{\scriptsize 8\,MB 4M$\times$16}};
\draw[fnbus] (nm)--node[fnlbl,above]{req/wr/addr}(ia);
\draw[fnbus] (ia)--node[fnlbl,above]{16-bit}(mi);
\draw[fnbus] (mi)--(pc);
\draw[fnbus] (pc)--node[fnlbl,above]{DQ/A/ctrl}(ps);
\end{tikzpicture}
\end{center}
\section{\texttt{int8\_memory\_access} --- conversione byte/word}
Converte l'interfaccia INT8 (indirizzo di byte) nell'interfaccia a parola. L'indirizzo
di byte viene diviso per due (\code{addr>>1}) per ottenere l'indirizzo di parola; il
bit meno significativo seleziona il byte:
\begin{itemize}
\item \code{addr[0]=0} $\to$ byte basso: \code{lb\_n=0}, \code{ub\_n=1}, dato su DQ[7:0];
\item \code{addr[0]=1} $\to$ byte alto: \code{lb\_n=1}, \code{ub\_n=0}, dato su DQ[15:8].
\end{itemize}
In lettura estrae il byte corretto da \code{mem\_rdata}. La FSM ha due stati (IDLE,
WAIT) e restituisce \code{ready} come impulso di un ciclo.
\section{\texttt{memory\_interface} --- handshake}
FSM a due stati che serializza una singola transazione: in IDLE, alla richiesta
\code{req}, latcha \code{wr/addr/wdata/lb\_n/ub\_n} ed emette un impulso \code{mem\_req}
di un ciclo verso il controller; in WAIT attende \code{mem\_ready}, cattura
\code{rdata} in lettura e asserisce \code{ready}. Garantisce il contratto ``una
transazione per volta''.
\section{\texttt{psram\_controller} --- bus fisico}
Controller del bus PSRAM parallelo asincrono, con supporto al \textbf{page mode}
di lettura del chip (\S~\ref{sec:pagemode}). La macchina a stati principale è:
\begin{center}
\begin{tikzpicture}[font=\scriptsize]
\node[fnstate](init) at (0,0){INIT};
\node[fnstate](idle) at (3.2,0){IDLE};
\node[fnstate](read) at (7,2.7){READ};
\node[fnstate](popen) at (11,2.7){PAGE\\OPEN};
\node[fnstate](write) at (7,-2.7){WRITE};
\node[fnstate](ww) at (11,-2.7){WRITE\\WAIT};
\draw[fnarrow] (init)--node[fnlbl,above]{INIT\_CYCLES + CR load}(idle);
\draw[fnarrow] (idle)--node[fnlbl,above,sloped]{req \& !wr}(read);
\draw[fnarrow] (idle)--node[fnlbl,below,sloped]{req \& wr}(write);
\draw[fnarrow] (read)--node[fnlbl,above]{ready}(popen);
\draw[fnarrowT] (popen) to[bend left=25] node[fnlbl,below]{req \& !wr}(read);
\draw[fnarrow] (popen) to[bend right=20] node[fnlbl,above,sloped]{req \& wr}(write);
\draw[fnarrow] (popen) to[out=-100,in=15,looseness=1.15] node[fnlbl,pos=0.55]{timeout tCEM}(idle);
\draw[fnarrow] (write)--node[fnlbl,above]{ACCESS\_CYCLES}(ww);
\draw[fnarrow] (ww) to[out=160,in=-70] node[fnlbl,pos=0.5,left]{ready}(idle);
\end{tikzpicture}
\end{center}
Da INIT il controller passa automaticamente per una sotto-sequenza di caricamento del
registro di configurazione (\code{STATE\_CR\_INIT}, 4 passi) prima di raggiungere IDLE
per la prima volta --- vedi \S~\ref{sec:pagemode}. La transizione PAGE~OPEN
$\to$~WRITE (freccia in basso a destra) passa internamente per due micro-stati di
transito, \code{STATE\_PAGE\_CLOSE} e \code{STATE\_PAGE\_REOPEN} (un ciclo ciascuno):
il primo forza CE\#/OE\# alti per almeno un ciclo prima che il controller inizi a
pilotare il bus dati, evitando contesa con l'uscita ancora attiva della PSRAM
($\geq t_{HZ}$); il secondo riavvia la transazione già latchata esattamente come
farebbe IDLE. Non sono disegnati come nodi separati per non appesantire la figura.
\subsection{Temporizzazione}
\begin{fnspec}[Formule di temporizzazione]
$\text{ACCESS\_CYCLES}=\lceil (70\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(latenza di accesso casuale, $t_{AA}$/$t_{RC}$ = 70~ns)\\[3pt]
$\text{PAGE\_CYCLES}=\lceil (20\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(continuazione nella stessa pagina, $t_{APA}$/$t_{PC}$ = 20~ns)\\[3pt]
$\text{INIT\_CYCLES}=150\times \text{CLK\_FREQ\_MHZ}$ \quad
(inizializzazione di power-up, $t_{PU}$ = 150~\textmu s)\\[3pt]
$\text{PAGE\_TIMEOUT\_CYCLES}=\lceil (6000\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(chiusura automatica della pagina, margine di sicurezza sotto $t_{CEM}$ = 8~\textmu s)
\end{fnspec}
Il bus dati è pilotato in tri-state: \code{psram\_dq = dq\_oe ? dq\_out : Z}. In lettura
\code{dq\_oe=0}; in scrittura \code{dq\_oe=1} durante l'impulso di \code{we\_n}. Uno
stato di WRITE\_WAIT mantiene attivi \code{ce\_n/lb\_n/ub\_n} per l'hold finale prima
del rilascio.
\begin{fnwarn}[Non è QSPI]
Questa è un'interfaccia SRAM-asincrona classica, \textbf{non} QSPI: la maggior parte
delle ``PSRAM'' serie/QSPI in commercio non è compatibile con questo controller senza
riscrittura. Vedere il cap.~\ref{ch:hw} per la parte raccomandata (ISSI parallela).
\end{fnwarn}
\section{Page mode di lettura}
\label{sec:pagemode}
Il chip raccomandato (cap.~\ref{ch:hw}) è ``asynchronous/\textbf{page mode}'': una
volta fatto un primo accesso casuale a $t_{AA}$~=~70~ns, letture successive
all'interno della stessa pagina da 16 word (bit di indirizzo sopra \code{A[3]}
invariati) costano solo $t_{APA}$/$t_{PC}$~=~20~ns, perché CE\#/OE\# restano attivi
e cambia solo il bus indirizzi. Il page mode è \textbf{disabilitato di default}
all'accensione (bit~7 del registro di configurazione, CR~=~\texttt{0x0070} di
default) e va abilitato esplicitamente.
\begin{itemize}
\item \textbf{Abilitazione all'avvio}: subito dopo INIT, il controller esegue la
``software-access sequence'' del datasheet (2 letture dummy + 2 scritture,
\texttt{0x0000} di sblocco poi CR reale \texttt{0x00F0} = default con il bit
Page attivo) all'indirizzo più alto del chip --- riusa esattamente la stessa
logica READ/WRITE di ogni altra transazione, quindi passa dagli stessi controlli
di temporizzazione.
\item \textbf{Burst di pagina}: dopo una READ il controller non chiude più
CE\#/OE\# (stato PAGE~OPEN). Una READ successiva nella stessa pagina aspetta solo
PAGE\_CYCLES; una READ che attraversa pagina resta comunque senza toggle di CE\#
ma paga un ACCESS\_CYCLES pieno per quella parola (qualunque cambio a
\code{A[4]} o superiore richiede un nuovo $t_{AA}$). Un contatore chiude la
pagina prima del limite $t_{CEM}$ con margine di sicurezza.
\item \textbf{Solo una WRITE chiude la pagina.} I cambi di \code{lb\_n}/\code{ub\_n}
\emph{non} la chiudono: \code{int8\_memory\_access} alterna questi segnali a quasi
ogni accesso (accesso byte-granulare su bus a 16~bit), quindi trattarli come
condizione di chiusura --- primo tentativo di implementazione --- rendeva il
workload reale \emph{più lento}, non più veloce (misurato: 53.25$\to$61.25
cicli/edge sul gather di \code{graph\_engine}); rimosso, corretto a
53.25$\to$37.53 cicli/edge (banda +42\%, \S~\ref{sec:bandwidth}).
\end{itemize}
\begin{fnwarn}[Nessun beneficio senza pattern sequenziale]
Il page mode accelera solo accessi che restano nella stessa pagina (o quasi) mentre
il controller resta in attesa di una nuova richiesta con la pagina ancora aperta.
Accessi isolati e sparsi (indirizzo casuale ogni volta) pagano comunque
ACCESS\_CYCLES pieno, più un piccolo overhead di chiusura/riapertura se preceduti
da una WRITE o da un timeout $t_{CEM}$: non è un guadagno universale, dipende dal
pattern di accesso del chiamante.
\end{fnwarn}
Fmax reale (\code{nextpnr-ecp5}, cap.~\ref{ch:impl}) sul sistema integrato
\code{spi\_neuron\_top} con Tipo~\#2 abilitato: \textbf{75.73~MHz} a
\code{PARALLEL}=2 (era 55.59~MHz prima dell'aggiunta del page mode) e
\textbf{65.13~MHz} a \code{PARALLEL}=8, entrambe ancora FAIL all'obiettivo di
80~MHz ma non regredite. Il percorso critico resta, in entrambi i casi,
interamente dentro \code{u\_graph\_engine.u\_neuron} (catena di accumulo
\code{mac8}/\code{neuron\_parallel}, cap.~\ref{ch:impl}) --- \code{psram\_controller}
non compare mai nel percorso critico nonostante la crescita di risorse del page
mode.
\section{Mappa degli indirizzi e convenzioni}
Lo spazio di indirizzamento è di \code{ADDR\_WIDTH}=23~bit (indirizzo di \emph{byte}),
per 8~MB pieni. Le regioni non hanno indirizzi cablati: le loro basi sono registri
impostati dall'host via \op{SET\_BASE} (percorso single-layer) o lette dalla tabella
descrittori (percorso multi-layer).
\begin{tabularx}{\textwidth}{L{3.2cm} L{3.4cm} Y}
\toprule
\rowh \thd{Regione} & \thd{Base} & \thd{Contenuto / convenzione} \\
\midrule
Ingresso $X$ & \code{x\_base} & Vettore di ingresso condiviso, letto una volta per invocazione. \\
\rowa Pesi $W$ & \code{w\_base} & Neuron-major: i pesi del neurone $n$ a \code{w\_base + n*N\_INPUTS} byte. \\
Bias & \code{bias\_addr} & Un byte per neurone: bias del neurone $n$ a \code{bias\_addr + n}. \\
\rowa Tabella descrittori & \code{table\_base} & \code{N\_LAYERS} voci da 11 byte (cap.~\ref{ch:seq}). \\
Buffer ping-pong A/B & \code{buf\_a\_base} / \code{buf\_b\_base} & Uscite intermedie tra layer. \\
\bottomrule
\end{tabularx}
\subsection{Indirizzamento fisico della PSRAM}
La PSRAM raccomandata è 4M$\times$16 (8~MB), che richiede un indirizzo di parola a
22~bit (A0--A21). \code{int8\_memory\_access} calcola \code{addr>>1} portando l'indirizzo
di byte a 23~bit in un indirizzo di parola a 22~bit che mappa esattamente su A0--A21; il
bit~22 di \code{psram\_a} è quindi sempre 0 e sul PCB restano 22 linee di indirizzo
reali.
\section{Larghezza di banda}
\label{sec:bandwidth}
Misurata sul gather della lista di edge di \code{graph\_engine} (cap.~\ref{ch:grafo}),
per differenza tra due dimensioni di grafo per isolare il costo per-edge dall'overhead
fisso per-neurone (\code{sim/graph\_engine\_bandwidth\_tb.v}):
\begin{tabularx}{\textwidth}{L{5.2cm} Y Y Y}
\toprule
\rowh \thd{} & \thd{Prima (no page mode)} & \thd{Dopo (page mode)} & \thd{$\Delta$} \\
\midrule
Cicli/edge & 53.25 & 37.53 & $-29.5\%$ \\
\rowa Banda @80\,MHz & 6.01\,MB/s & 8.53\,MB/s & $+41.9\%$ \\
Banda @16\,MHz\textsuperscript{*} & 1.20\,MB/s & 1.71\,MB/s & $+41.9\%$ \\
\bottomrule
\end{tabularx}
\textsuperscript{*}oscillatore reale raccomandato (cap.~\ref{ch:hw}).
Il modello resta comunque memory-bound per costruzione: \code{neuron\_memory} legge
$X$ una volta e rilegge $W$/bias per ciascun neurone (cap.~\ref{ch:seq}), un neurone
per volta; il page mode riduce il costo per-byte dell'accesso sequenziale, non elimina
il pattern di accesso stesso.
@@ -1,111 +0,0 @@
\chapter[Memoria, multi-neurone e multi-layer]{Integrazione memoria, multi-neurone e multi-layer}
\label{ch:seq}
\section{\texttt{neuron\_memory} --- ponte memoria/neurone}
\code{neuron\_memory} collega il datapath di calcolo alla memoria e gestisce il loop sui
neuroni. Legge il vettore $X$ una sola volta (ingresso condiviso), poi per ciascun
neurone rilegge $W$ e bias dalla RAM e li invia a una singola istanza riusata di
\code{neuron\_parallel}: il progetto è memory-bound, un neurone calcolato per volta,
senza duplicare il datapath. L'uscita è \code{y\_bus}, packed neuron-major
(\code{DATA\_WIDTH*N\_NEURONS} bit).
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=13mm]
\node[fnstate](idle){IDLE};
\node[fnstate,right=of idle](rx){READ\_X};
\node[fnstate,right=of rx](rw){READ\_W};
\node[fnstate,below=10mm of rw](rb){READ\_BIAS};
\node[fnstate,left=of rb](sn){START\_N};
\node[fnstate,left=of sn](wn){WAIT\_N};
\draw[fnarrow] (idle)--node[fnlbl,above]{start}(rx);
\draw[fnarrow] (rx)--node[fnlbl,above]{X letto}(rw);
\draw[fnarrow] (rw)--(rb);
\draw[fnarrow] (rb)--(sn);
\draw[fnarrow] (sn)--(wn);
\draw[fnarrow] (wn) to[bend left=18] node[fnlbl,above]{neurone succ.}(rw);
\draw[fnarrow] (wn) to[bend right=28] node[fnlbl,below]{ultimo neurone: done}(idle);
\end{tikzpicture}
\end{center}
Gli stati sono IDLE, READ\_X, READ\_W, READ\_BIAS, START\_N, WAIT\_N. Dopo l'ultimo
neurone la FSM torna in IDLE e asserisce \code{done}. Il conteggio di neuroni e ingressi
realmente elaborati è dato da \code{n\_neurons\_real}/\code{n\_inputs\_real}
(cap.~\ref{ch:param}).
\section{\texttt{layer\_sequencer} --- rete multi-layer}
\code{layer\_sequencer} concatena fino a \code{N\_LAYERS} esecuzioni della stessa
istanza \code{neuron\_memory}, realizzando una rete densa feed-forward \emph{senza}
toccare il core di calcolo validato. Legge una tabella descrittori scritta dall'host e
alterna i due buffer di uscita in RAM (ping-pong).
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=13mm]
\node[fnstate](i){IDLE};
\node[fnstate,right=of i](rd){READ\\DESC};
\node[fnstate,right=of rd](rw){READ\\WAIT};
\node[fnstate,below=10mm of rw](sl){START\\LAYER};
\node[fnstate,left=of sl](wl){WAIT\\LAYER};
\node[fnstate,left=of wl](ci){COPY\\ISSUE};
\node[fnstate,below=9mm of ci](cw){COPY\\WAIT};
\draw[fnarrow] (i)--node[fnlbl,above]{run\_start}(rd);
\draw[fnarrow] (rd)--(rw);
\draw[fnarrow] (rw)--(sl);
\draw[fnarrow] (sl)--(wl);
\draw[fnarrow] (wl)--(ci);
\draw[fnarrow] (ci)--(cw);
\draw[fnarrow] (cw) to[bend left=15] node[fnlbl,left]{layer succ.}(rd);
\draw[fnarrow] (cw) to[bend right=12] node[fnlbl,below]{ultimo: seq\_done}(i);
\end{tikzpicture}
\end{center}
\subsection{Buffer ping-pong}
Il layer~0 legge l'ingresso esterno \code{x\_base}. Il layer $k>0$ legge dal buffer
scritto dal layer $k-1$; l'uscita di ciascun layer viene copiata nell'altro buffer,
alternando A e B. L'uscita finale resta sia in \code{y\_bus} (leggibile con
\op{READ\_OUTPUT}) sia nel buffer ping-pong su cui è stata copiata.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=7mm]
\node[fnblockA,minimum width=18mm](x){X\\\code{x\_base}};
\node[fnblockD,right=10mm of x,minimum width=20mm](l0){Layer 0};
\node[fnblock,right=10mm of l0,minimum width=18mm](ba){buf A};
\node[fnblockD,right=10mm of ba,minimum width=20mm](l1){Layer 1};
\node[fnblock,right=10mm of l1,minimum width=18mm](bb){buf B};
\node[fnblockD,right=10mm of bb,minimum width=20mm](l2){Layer 2};
\draw[fnarrow] (x)--(l0); \draw[fnarrow] (l0)--(ba);
\draw[fnarrow] (ba)--(l1); \draw[fnarrow] (l1)--(bb);
\draw[fnarrow] (bb)--(l2);
\draw[fnarrowT,dashed] (l2.south) to[bend left=25] node[fnlbl,below]{copia in buf A} (ba.south);
\end{tikzpicture}
\end{center}
\subsection{Tabella descrittori}
Scritta dall'host in RAM a \code{table\_base} con \op{WRITE\_RAM}; \code{N\_LAYERS} voci
da 11 byte ciascuna, MSB-first:
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.6cm} Y}
\toprule
\rowh \thd{Campo} & \thd{Byte} & \thd{Significato} \\
\midrule
\code{w\_base} & 3 & Base dei pesi del layer. \\
\rowa \code{bias\_addr} & 3 & Base dei bias del layer. \\
\code{activation} & 1 & Attivazione del layer (2 bit bassi, cfr. \code{ACT\_*}). \\
\rowa \code{n\_inputs\_real} & 2 & Ingressi reali del layer (multiplo di \code{PARALLEL}). \\
\code{n\_neurons\_real} & 2 & Neuroni reali del layer. \\
\midrule
\rowh \thd{Totale} & \thd{11} & per voce/layer \\
\bottomrule
\end{tabularx}
\begin{fnnote}[Copia proporzionale alla larghezza reale]
Il sequencer copia esattamente \code{n\_neurons\_real} byte di \code{y\_bus} nel buffer
ping-pong (non l'intera larghezza di build): un layer più stretto viene copiato più in
fretta, senza zero-padding in RAM. Ogni attivazione è letta per-layer dalla tabella,
indipendente dal registro \code{activation} del percorso single-layer.
\end{fnnote}
\section{Gerarchia dei segnali \texttt{busy}/\texttt{done}}
Nel percorso multi-layer, \code{STATUS.busy} è l'OR dei busy single-layer e sequencer,
mentre \code{STATUS.done} latcha solo al completamento dell'\emph{ultimo} layer, non a
ogni layer intermedio (cap.~\ref{ch:spi}). Il top-level restituisce il controllo di
\code{neuron\_memory} al percorso diretto \op{START} al termine della sequenza.
@@ -1,191 +0,0 @@
\chapter[Rete a grafo (Tipo \#2)]{Configurazione a due livelli: rete a grafo (Tipo \#2)}
\label{ch:grafo}
\section{Due tipi di rete}
L'engine espone due \emph{tipi di rete} selezionabili dall'host, con lo stesso comando di
avvio che instrada verso il motore corretto:
\begin{itemize}
\item \textbf{Tipo \#1 --- rete classica (dense).} Layer con neuroni per layer, fully
connected tra layer consecutivi. È il percorso di \code{layer\_sequencer}
(cap.~\ref{ch:seq}), avviato da \op{RUN\_NETWORK}. Connessioni \emph{implicite per
posizione}: non si enumera nulla, si definiscono solo i pesi indirizzati come
\code{w\_base + k*n\_inputs + j}.
\item \textbf{Tipo \#2 --- grafo arbitrario (sparse).} A partire dagli id dei neuroni di
ingresso si definiscono le connessioni di ogni neurone fino all'uscita, tramite una
\emph{edge-list sparsa} per-neurone. Connessioni \emph{esplicite per enumerazione}: ogni
connessione è un edge \code{(src\_id, peso)}; se non è nella lista, non esiste.
\end{itemize}
\begin{fnnote}[La differenza in una riga]
Dense: definisci i \emph{pesi} per posizione in una matrice. Graph: definisci ogni
\emph{connessione} come edge \code{(src\_id, peso)} in una lista per-neurone. Le due
tabelle descrittori hanno lo stesso formato di 11~byte ma campi diversi; il registro
\code{net\_type} dice al motore quale interpretazione usare.
\end{fnnote}
\section{Buffer di attivazione globale}
Il Tipo \#2 introduce un \textbf{buffer di attivazione} indicizzato per \emph{id di
segnale}, un byte INT8 per id, realizzato in \textbf{block RAM on-chip \code{DP16KD}}
(\code{rtl/act\_buffer.v}). Gli id \code{0..N\_in-1} sono gli ingressi; ogni neurone
scrive la propria uscita nel proprio id. Il gather delle sorgenti legge da qui a
\emph{accesso random a un ciclo}: è ciò che rende economico il grafo, perché è l'accesso
che la PSRAM (70~ns, sequenziale) non potrebbe accelerare.
\begin{fnspec}[Dimensionamento V1]
\code{N\_TOTAL}=4096 segnali, id a 16~bit (spazio fino a 65.536 senza cambiare formato).
Buffer = 4~KB, cioè 2 blocchi \code{DP16KD} su 108. Il vincolo reale diventa la capacità
PSRAM per gli edge ($\approx$2\,M edge a 4~B), non la block RAM.
\end{fnspec}
\section{DAG feed-forward e vincolo \texttt{src\_id < out\_id}}
Il grafo è un DAG feed-forward: ogni connessione punta a un id \textbf{già calcolato}
(\code{src\_id < out\_id}). I neuroni si elaborano in ordine di id crescente, così quando
si calcola un neurone tutte le sue sorgenti sono pronte nel buffer. Cicli e ricorrenza
sono fuori scope per la V1. Il vincolo è verificato a due livelli: dall'assemblatore host
(a compile time) e da un guard a runtime in \code{graph\_engine} (\code{STATUS.err}),
nella stessa filosofia del guard di elaborazione su \code{N\_INPUTS \% PARALLEL}.
\section{Formati dati}
Entrambi i descrittori sono da 11~byte/voce, MSB-first, a \code{table\_base}.
\subsection{Descrittore Tipo \#2 (grafo)}
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.6cm} Y}
\toprule
\rowh \thd{Campo} & \thd{Byte} & \thd{Significato} \\
\midrule
\code{conn\_ptr} & 3 & Indirizzo byte in PSRAM del blocco edge del neurone. \\
\rowa \code{n\_conn} & 2 & Connessioni reali (pre-padding). \\
\code{out\_id} & 2 & Id in cui scrivere l'uscita del neurone. \\
\rowa \code{activation} & 1 & \code{ACT\_RELU} / \code{ACT\_NONE} (2 bit bassi). \\
\code{bias} & 1 & Bias del neurone (INT8). \\
\rowa \code{reserved} & 2 & 0. \\
\midrule
\rowh \thd{Totale} & \thd{11} & voci in ordine di \code{out\_id} crescente \\
\bottomrule
\end{tabularx}
\subsection{Edge del grafo (4~byte, allineato)}
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.6cm} Y}
\toprule
\rowh \thd{Campo} & \thd{Byte} & \thd{Significato} \\
\midrule
\code{src\_id} & 2 & Id sorgente (uint16 BE). \\
\rowa \code{weight} & 1 & Peso (INT8). \\
\code{reserved} & 1 & 0 (allineamento a 4~byte). \\
\bottomrule
\end{tabularx}
\begin{fnnote}[Padding a \texttt{PARALLEL}]
\code{n\_conn} arbitrario non è multiplo di \code{PARALLEL}: la edge-list del neurone è
riempita fino al multiplo con edge a \textbf{peso zero} (spreco $\le$\code{PARALLEL}$-1$
per neurone). Così il datapath e il suo guard restano intatti.
\end{fnnote}
\section{\texttt{graph\_engine} --- motore del grafo}
\code{rtl/graph\_engine.v} orchestra il Tipo \#2 \textbf{riusando \code{neuron\_parallel}
senza modificarlo}, come fa \code{neuron\_memory} per il caso denso. Differenza chiave: tra
i due modi cambia \emph{solo l'indirizzamento di X}. In Tipo \#1 l'input è contiguo
(\code{x\_base + i}); in Tipo \#2 è un gather (\code{act\_buf[src\_id]}). Il core aritmetico
non si tocca.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=4mm,start chain=going below,
every node/.style={on chain,fnblock,minimum width=52mm}]
\node[fnblockA]{\code{COPY\_INPUTS}: PSRAM \code{x\_base} $\to$ \code{act\_buf[0..N\_in-1]}};
\node{\code{READ\_DESC}: descrittore del neurone k};
\node{\code{READ\_EDGES}: stream edge + gather \code{act\_buf[src\_id]}};
\node{\code{START\_N} / \code{WAIT\_N}: gruppo da \code{PARALLEL} $\to$ \code{neuron\_parallel}};
\node[fnblockT]{\code{WRITE\_ACT}: y $\to$ \code{act\_buf[out\_id]}};
\node{neurone successivo (ordine di id)};
\node[fnblockD]{\code{WRITE\_OUTPUTS}: ultimi \code{n\_out} $\to$ PSRAM \code{out\_base}};
\foreach \i [count=\j from 2] in {1,...,6}
\draw[fnarrow] (chain-\i) -- (chain-\j);
\end{tikzpicture}
\end{center}
Le uscite sono gli \textbf{ultimi \code{n\_out}} id: nei DAG con l'ordinamento
\code{src\_id < out\_id} i neuroni di uscita (sink, non riusati come sorgente) finiscono
naturalmente con gli id più alti. A fine esecuzione \code{graph\_engine} copia questi
\code{n\_out} byte in una regione PSRAM a \code{out\_base}, che l'host rilegge con
\op{READ\_RAM}.
\section{Opcode e registri del Tipo \#2}
La selezione del tipo avviene con un nuovo opcode; \op{RUN\_NETWORK} fa il dispatch sul
registro \code{net\_type} (dettagli in cap.~\ref{ch:spi}).
\begin{tabularx}{\textwidth}{L{2.6cm} L{3.4cm} Y}
\toprule
\rowh \thd{Opcode / sel} & \thd{Nome} & \thd{Funzione} \\
\midrule
\op{0x11} & SET\_NET\_TYPE & \code{type(1B)}: \code{0x01}=dense (\#1), \code{0x02}=graph (\#2). Default dopo \op{RESET}=dense. \\
\rowa \code{SET\_BASE sel 9} & num\_neurons\_graph & Numero di neuroni del grafo (uint16). \\
\code{SET\_BASE sel 10} & n\_out & Numero di id di uscita (uint16). \\
\bottomrule
\end{tabularx}
\begin{fnnote}[Zero regressioni sul Tipo \#1]
Con \code{net\_type=dense} (valore di default dopo \op{RESET}) il percorso \#1 è
bit-identico a prima: \op{RUN\_NETWORK} mantiene il payload \code{num\_layers(1B)} e il
framing degli opcode esistenti non cambia.
\end{fnnote}
\section{Occupazione (Tipo \#2 abilitato)}
Sintesi Yosys del sistema completo \code{spi\_neuron\_top} con Tipo \#2 abilitato
(\code{PARALLEL}=2):
\begin{tabularx}{\textwidth}{L{4.6cm} Y}
\toprule
\rowh \thd{Risorsa} & \thd{Uso} \\
\midrule
\code{DP16KD} (block RAM) & 2 (buffer di attivazione) \\
\rowa \code{MULT18X18D} (DSP) & 4 (2 \code{neuron\_memory} + 2 \code{graph\_engine}) \\
LUT4 & 2619 \\
\rowa TRELLIS\_FF & 2467 \\
\code{\$\_TBUF\_} (bus PSRAM) & 16 \\
\bottomrule
\end{tabularx}
Il device (108 \code{DP16KD}, 72 DSP, $\approx$44k LUT/FF) resta ampiamente sotto la
saturazione: il Tipo \#2 aggiunge una modalità completa a costo di risorse contenuto.
LUT4/TRELLIS\_FF sono cresciuti rispetto a una misura precedente (2367/2406) per via del
page mode PSRAM aggiunto al controller (cap.~\ref{ch:mem}, \S~5.5) --- sotto il 6\% di
utilizzo, nessun impatto pratico.
\section{Banda del gather (misurata)}
Il costo per-edge del gather è stato \textbf{isolato} costruendo due grafi identici per
struttura ma con conteggio edge diverso e differenziando i cicli: la sottrazione cancella
l'overhead fisso per-neurone e lascia il solo costo dell'edge.
\begin{fnspec}[Costo per-edge]
\textbf{37.53 cicli/edge} con il page mode PSRAM abilitato (cap.~\ref{ch:mem},
\S~5.5) --- \textbf{53.25 cicli/edge} senza (baseline pre-page-mode, coerente con la
teoria: 4~byte/edge $\times$ $\approx$13 cicli/byte via PSRAM asincrona
$\approx$52). A 80~MHz: $\approx$2.13\,M edge/s ($\approx$8.5~MB/s, +42\% vs
baseline); al clock reale di 16~MHz: $\approx$426\,k edge/s ($\approx$1.71~MB/s).
\end{fnspec}
Il page-mode read (roadmap G7, cap.~\ref{ch:roadmap}) è stato implementato e misurato:
l'accesso sequenziale del gather ne beneficia direttamente, riducendo il costo per-edge
del 29.5\% (53.25$\to$37.53 cicli/edge). Ogni edge continua comunque a pagare l'accesso
byte-granulare di \code{int8\_memory\_access} (4 byte/edge); il page mode riduce il costo
di ciascun byte sequenziale, non il numero di accessi.
\section{Assemblatore host \texttt{netasm}}
La configurazione leggibile della rete non richiede logica dedicata in FPGA: uno
pseudo-assembly viene compilato \emph{sull'host} (\code{tools/netasm/}) nei byte esatti
delle tabelle e degli edge, poi caricati con \op{WRITE\_RAM}. L'assemblatore valida a
compile time (\code{src\_id < out\_id}, limiti \code{N\_TOTAL}, padding a \code{PARALLEL}),
complementando il guard runtime.
\begin{lstlisting}[language=,caption={Esempio di pseudo-assembly (grafo)},basicstyle=\ttfamily\scriptsize]
NET graph
INPUTS 4 ; id 0..3
NEURON n4 relu bias=2
CONN 0 w=5
CONN 1 w=-3
NEURON n5 none bias=0
CONN n4 w=2 ; riferimento simbolico all'uscita di n4
CONN 2 w=7
OUTPUT n5
END
\end{lstlisting}
@@ -1,297 +0,0 @@
\chapter{Interfaccia host SPI}
\label{ch:spi}
\section{Livello fisico}
L'FPGA è sempre \textbf{slave} SPI. Il protocollo v1 usa SPI \textbf{Mode~0}
(CPOL=0, CPHA=0), MSB-first, single-SPI. Un comando per periodo di CS basso; il byte~0
di ogni transazione è l'opcode. I campi multi-byte sono big-endian.
\begin{fnspec}[Campionamento Mode 0]
\code{mosi} è campionato sul fronte di \textbf{salita} di \code{sclk}; \code{miso} è
pilotato sul fronte di \textbf{discesa} (stabile prima del successivo campionamento del
master). \code{spi\_slave} sincronizza \code{sclk/mosi/cs\_n} con un doppio flip-flop
(CDC a 3 stadi) prima di ogni rilevazione di fronte.
\end{fnspec}
\begin{center}
\begin{tikztimingtable}[timing/dslope=0.1,timing/.style={x=3.4ex,y=2.2ex},
xscale=1.0,font=\scriptsize]
\sig{CS\_N} & H 1L 16L 1H \\
\sig{SCLK} & L 1L {2C(2)}8{2C(2)} 6L \\
\sig{MOSI} & U 1U 2D{b7} 2D{b6} 2D{b5} 2D{b4} 2D{b3} 2D{b2} 2D{b1} 2D{b0} 2U \\
\sig{MISO} & Z 1Z 16D{dato} 1Z \\
\end{tikztimingtable}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
Framing di un byte: CS scende, 8 colpi di SCLK, MSB per primo; MISO in tri-state fuori
transazione.\end{center}
\begin{fnnote}[Contratto \texttt{tx\_byte\_req}]
\code{tx\_byte\_req} è un \emph{prefetch hint}, non un evento ``byte consumato'': un
consumatore deve avanzare i puntatori (indirizzo RAM, indice byte di risposta) su
\code{rx\_valid}, che pulsa esattamente una volta per byte reale trasferito.
\end{fnnote}
\section{Framing e lunghezza esplicita}
La lunghezza dei trasferimenti RAM è \textbf{esplicita}, non delimitata dal fronte di
CS: \op{WRITE\_RAM}/\op{READ\_RAM} portano un campo lunghezza a 2~byte, così il
controller SPI necessita solo di un contatore di byte. Gli indirizzi di byte sono a
23~bit, trasportati in un campo di 3~byte con il bit più alto riservato a 0.
\section{Tabella degli opcode}
\renewcommand{\arraystretch}{1.16}
\begin{longtable}{C{1.1cm} L{2.4cm} L{3.9cm} L{2.4cm} L{4.0cm}}
\toprule
\rowh \thd{Op} & \thd{Nome} & \thd{Payload (host$\to$FPGA)} & \thd{Risposta} & \thd{Funzione} \\
\midrule
\endfirsthead
\rowh \thd{Op} & \thd{Nome} & \thd{Payload} & \thd{Risposta} & \thd{Funzione} \\ \midrule
\endhead
\bottomrule
\endfoot
\op{0x00} & NOP & --- & --- & Nessuna operazione (idle/dummy clocking). \\
\rowa \op{0x01} & WRITE\_RAM & addr(3B)+len(2B)+dati & --- & Scrive un blocco in PSRAM (X, pesi, bias, parametri). \\
\op{0x02} & READ\_RAM & addr(3B)+len(2B) & \code{len} byte & Rilegge un blocco da PSRAM. \\
\rowa \op{0x0F} & RESET & --- & --- & Reset sincrono del motore e azzeramento del latch STATUS; non cancella la PSRAM. \\
\op{0x10} & SET\_BASE & sel(1B)+addr(3B) & --- & Imposta le basi/registri (vedi §\ref{sec:setbase}). \\
\rowa \op{0x11} & SET\_NET\_TYPE & type(1B) & --- & Tipo di rete: \code{0x01}=dense (\#1), \code{0x02}=graph (\#2). Default dopo RESET=dense. \\
\rowa \op{0x20} & START & --- & --- & Avvia \code{neuron\_memory} (percorso single-layer); ignorato se busy. \\
\op{0x21} & STATUS & --- & 1 byte & bit0=\code{busy} (live), bit1=\code{done} (sticky, clear-on-read), bit2=\code{err} (guard grafo), bit3=\code{flash\_err} (sticky, clear-on-read), bit4=\code{flash\_busy} (live); bit7:5=0. \\
\rowa \op{0x22} & READ\_OUTPUT & --- & \code{N\_NEURONS} byte & \code{y\_bus} neuron-major (byte~0 = neurone~0); solo percorso dense (Tipo \#1). \\
\op{0x23} & RUN\_NETWORK & num\_layers(1B) & --- & Avvia l'esecuzione: dispatch su \code{net\_type} verso \code{layer\_sequencer} (\#1) o \code{graph\_engine} (\#2); ignorato se busy. \\
\rowa \op{0x30} & READ\_CONFIG & --- & 11 byte & Record di configurazione hardware (§\ref{sec:readcfg}). \\
\op{0x40} & FLASH\_READ\_BLOCK & flash\_addr(3B)+psram\_addr(3B)+len(3B) & --- & Lettura raw flash$\to$PSRAM, bypassa il catalogo. \\
\rowa \op{0x41} & FLASH\_WRITE\_BLOCK & psram\_addr(3B)+flash\_addr(3B)+len(3B) & --- & Scrittura raw PSRAM$\to$flash (erase-before-write interno + loop Page Program $\leq$256B + poll WIP, trasparente all'host), bypassa il catalogo. \\
\op{0x42} & FLASH\_ERASE & sector\_addr(3B) & --- & Erase di un settore da 4~KB (deve essere sector-aligned), bypassa il catalogo. \\
\rowa \op{0x43} & CAT\_READ & --- & --- & Ricarica il catalogo a 16 slot (registri on-chip) dal settore riservato in flash. \\
\op{0x44} & CAT\_WRITE\_SLOT & slot\_id(1B)+offset(3B)+len(3B)+tipo(1B) & --- & Registra/aggiorna (offset, lunghezza, tipo) dello slot nel catalogo on-chip e lo persiste in flash; marca lo slot \emph{non valido} finché \op{SAVE\_SLOT} non lo conferma. \\
\rowa \op{0x45} & LOAD\_SLOT & slot\_id(1B)+psram\_addr(3B) & --- & Flash$\to$PSRAM per lo slot (offset/lunghezza dal catalogo), verifica CRC32 live; \code{STATUS.flash\_err} se lo slot non è valido o il CRC non torna. \\
\op{0x46} & SAVE\_SLOT & slot\_id(1B)+psram\_addr(3B)+len(3B) & --- & PSRAM$\to$flash all'offset già registrato dello slot, calcola il CRC32 live; a esito positivo aggiorna e persiste la entry di catalogo (lunghezza, CRC, valid=1). \\
\rowa \op{0x47} & CAT\_INSPECT & slot\_id(1B) & 16 byte & Lettura sincrona di una entry di catalogo già caricata: offset[3]+len[3]+tipo[1]+valid[1]+CRC32[4]+riservato[4], MSB-first. \\
\end{longtable}
Tutti gli opcode flash sono \emph{fire-and-forget}: l'host fa polling su \op{STATUS}
(bit4=\code{flash\_busy}, bit3=\code{flash\_err}) o sui pin \code{irq\_n}/\code{data\_ready\_n}
per l'esito, eccetto \op{CAT\_INSPECT} che risponde in modo sincrono.
Gli 8 opcode flash (\op{0x40}--\op{0x47}) sono descritti in dettaglio, con
razionale di progetto e latenze reali misurate, in §\ref{sec:flashspi} sotto.
\section{Selettori \texttt{SET\_BASE}}
\label{sec:setbase}
\begin{tabularx}{\textwidth}{C{1.2cm} L{3.2cm} Y}
\toprule
\rowh \thd{sel} & \thd{Registro} & \thd{Uso} \\
\midrule
0 & \code{x\_base} & Base ingresso $X$. \\
\rowa 1 & \code{w\_base} & Base pesi. \\
2 & \code{bias\_addr} & Base bias. \\
\rowa 3 & \code{table\_base} & Base tabella descrittori (multi-layer). \\
4 & \code{buf\_a\_base} & Buffer ping-pong A. \\
\rowa 5 & \code{buf\_b\_base} & Buffer ping-pong B. \\
6 & \code{activation} & Attivazione (2 bit bassi) --- solo percorso single-layer. \\
\rowa 7 & \code{n\_inputs\_real} & Larghezza ingressi runtime (16-bit BE) --- single-layer. \\
8 & \code{n\_neurons\_real} & Larghezza neuroni runtime (16-bit BE) --- single-layer. \\
\rowa 9 & \code{num\_neurons\_graph} & Numero neuroni del grafo (16-bit BE) --- Tipo \#2. \\
10 & \code{n\_out} & Numero id di uscita (16-bit BE) --- Tipo \#2. \\
\bottomrule
\end{tabularx}
I selettori 6--8 riguardano solo il percorso single-layer/manuale; con \op{RUN\_NETWORK}
i valori equivalenti sono letti per-layer dalla tabella descrittori.
\begin{fnwarn}[Casi limite ``reale=0'' corretti (2026-09-04)]
La campagna di ri-certificazione (\code{docs/validation/bugs.md}) ha trovato che diversi
valori runtime pari a zero non erano protetti da alcun guard, con esiti che andavano da un
risultato silenziosamente ignorato fino a hang o scritture PSRAM a indirizzi arbitrari.
Tutti e cinque i casi seguenti sono ora no-op sicuri, verificati indipendentemente:
\begin{itemize}
\item \code{n\_inputs\_real=0} (selettore 7): completa in 1 ciclo con
$y=\text{activation}(\text{bias})$ (BUG-003).
\item \code{n\_neurons\_real=0} (selettore 8): completa senza eseguire alcun calcolo
per-neurone, molto più rapido di un run a piena larghezza (BUG-004).
\item \code{num\_neurons\_graph=0} (selettore 9): completa immediatamente dopo la copia
degli ingressi, senza mai entrare nel loop dei descrittori (BUG-006).
\item \op{RUN\_NETWORK} con \code{num\_layers=0} (percorso dense): no-op immediato ---
\textbf{prima del fix eseguiva 256 layer fasulli leggendo dati PSRAM arbitrari come
descrittori} (BUG-005, CRITICO, vedi \S\ref{sec:run-network} sotto).
\item \op{SET\_NET\_TYPE} ricevuto mentre un run è in corso: ora rifiutato silenziosamente
(nessun effetto, nessun errore SPI) invece di rimappare il multiplexer dell'arbitro a metà
esecuzione --- \textbf{prima del fix causava un hang permanente del motore in corso}
(BUG-007, CRITICO).
\end{itemize}
Dettagli, evidenza e verifica di ciascun fix in \code{docs/validation/bugs.md}.
\end{fnwarn}
\section{\texttt{STATUS.done} sticky / clear-on-read}
In \code{neuron\_memory} il segnale \code{done} è un impulso di un solo ciclo. Un host
che effettua polling via SPI (molto più lento del clock FPGA) mancherebbe quasi
certamente un impulso grezzo di un ciclo. Il banco registri SPI latcha quindi
\code{done} in un bit sticky sull'impulso e lo azzera quando l'host legge \op{STATUS}
(o \op{RESET}). Il bit \code{busy} è invece mantenuto a livello per tutta la
computazione e si legge live.
\begin{fnwarn}[Race corretto (2026-09-02)]
Una race reale nel meccanismo sticky (presente dalla Fase~4) è stata corretta latchando
uno \code{status\_snapshot} all'accettazione dell'opcode \op{STATUS} e condizionando la
pulizia del bit sticky a \code{status\_snapshot[1]} (si azzera solo se il byte
effettivamente trasmesso mostrava \code{done=1}). Un \code{done} che arriva troppo tardi
per uno snapshot viene riportato al polling successivo invece di essere perso.
\end{fnwarn}
\section{Pin di attenzione host (\texttt{data\_ready\_n}, \texttt{irq\_n})}
Oltre al polling di \op{STATUS}, il top-level espone due pin fisici attivi bassi (banco 7,
cap.~\ref{ch:hw}) che rispecchiano i bit sticky senza richiedere una transazione SPI,
utili per pilotare un GPIO/IRQ dell'host:
\begin{itemize}
\item \code{data\_ready\_n} = $\sim$\code{STATUS.done} (sticky): basso quando un risultato è
pronto da leggere, torna alto alla lettura di \op{STATUS} (clear-on-read).
\item \code{irq\_n} = $\sim$\code{STATUS.err} (guard grafo): basso quando il guard load-time
di \code{graph\_engine} è scattato. \textbf{Non} è clear-on-read: si azzera solo con
\op{RESET} o un nuovo avvio di grafo, così un errore non passa inosservato tra un polling e
l'altro.
\end{itemize}
Sono porte aggiuntive: non toccano gli opcode né i registri esistenti.
\begin{fnwarn}[\code{flash\_err} non ha un pin dedicato]
\code{STATUS.flash\_err} (bit3) è riportato \textbf{solo} nel byte \op{STATUS}, per scelta
di progetto: riusare \code{irq\_n} lo avrebbe confuso con gli errori del guard grafo (due
domini di errore indipendenti sullo stesso pin), mentre un'operazione flash è sempre
avviata dall'host con un opcode appena emesso, quindi il polling di \op{STATUS} subito dopo
--- già implicito nella convenzione ``fire-and-forget, poi polling \op{STATUS}/
\code{data\_ready\_n}'' --- è già naturale, senza bisogno di un pin asincrono in più.
\code{data\_ready\_n} invece \emph{si azzera anche al termine di un'operazione flash}: lo
specchia \code{STATUS.done} (bit1), che ora latcha anche sul completamento di un op flash,
non solo su \op{RUN\_NETWORK}/\op{START}.
\end{fnwarn}
\section{\texttt{READ\_CONFIG}}
\label{sec:readcfg}
Payload fisso di \textbf{11 byte}: permette a un unico firmware host di funzionare con
bitstream diversi senza ricompilare. I valori \code{N\_INPUTS}/\code{N\_NEURONS} riportano
il \emph{massimo} di build (il soffitto), non necessariamente la rete correntemente
caricata.
\begin{tabularx}{\textwidth}{C{1.6cm} L{3.6cm} Y}
\toprule
\rowh \thd{Byte} & \thd{Campo} & \thd{Sorgente} \\
\midrule
0 & \code{ADDR\_WIDTH} (bit) & \code{neuron\_memory.ADDR\_WIDTH} \\
\rowa 1--2 & \code{N\_INPUTS} (16-bit BE) & massimo di build \\
3 & \code{N\_NEURONS} & massimo di build \\
\rowa 4 & \code{PARALLEL} & parametro di build \\
5 & \code{DATA\_WIDTH} (bit) & parametro di build \\
\rowa 6--7 & versione protocollo (BE) & \code{0x0001} \\
8--9 & \code{N\_TOTAL} (16-bit BE) & massimo segnali grafo (Tipo \#2) \\
\rowa 10 & flag di capacità & bit0=\code{GRAPH\_SUPPORTED}=1 \\
\bottomrule
\end{tabularx}
\section{Sottosistema flash (opcode 0x40--0x47, completato 2026-09-04)}
\label{sec:flashspi}
La FPGA ha accesso \textbf{esclusivo} alla flash di boot/persistenza onboard (Winbond
\code{W25Q128JV}, 16~MB SPI NOR, cap.~\ref{ch:hw} §6/§7) tramite un SPI master dedicato e
fisicamente separato (\code{rtl/spi\_flash\_master.v}), mai per accesso diretto dell'host ai
pin della flash. \textbf{Non} è un filesystem: un catalogo a dimensione fissa (16 slot,
\code{rtl/flash\_slot\_manager.v}) mappa \code{slot\_id}~$\to$~(offset, lunghezza, tipo,
valid, CRC32) in un settore riservato della flash (settore 0) --- nessuna allocazione
dinamica, nessun garbage collection.
\begin{fnnote}[Stratificazione (ogni livello testabile a sé)]
\begin{itemize}
\item \code{rtl/spi\_flash\_master.v} --- SPI master grezzo verso il chip flash
(RDID/READ/WREN/PP/SE/RDSR-1). Bus a 4 fili completamente indipendente
(\code{sclk}/\code{mosi}/\code{miso}/\code{cs\_n}, tutti GPIO ordinario ---
Fase F7, 2026-09-04): una versione precedente riusava il pad \code{CCLK} di boot
via la primitiva ECP5 \code{USRMCLK} per risparmiare un pin, abbandonato perché
rendeva fuorviante l'affermazione di ``bus esclusivo'' (elettricamente dipendeva
comunque dal motore di configurazione) e comportava un gap di verifica mai chiuso
(timing di \code{USRMCLKTS} mai verificato contro la guida Lattice primaria).
\item \code{rtl/flash\_copy\_engine.v} --- motore di streaming a blocchi: flash$\to$PSRAM
(\code{DIR\_LOAD}), PSRAM$\to$flash con erase-before-write interno + loop Page
Program $\leq$256B + poll WIP (\code{DIR\_SAVE}), erase di settore standalone
(\code{DIR\_ERASE}). Master a bassa priorità (Porta D) su \code{rtl/mem\_arbiter.v}:
le operazioni flash sono su scala dei ms e non bloccano mai l'inferenza.
\item \code{rtl/flash\_slot\_manager.v} --- il catalogo a slot sopra, più un CRC32
(\code{rtl/crc32.v}, IEEE~802.3/zlib) calcolato live sul flusso di byte reale durante
\op{LOAD\_SLOT}/\op{SAVE\_SLOT}, così uno slot corrotto o scritto a metà (es.
alimentazione persa durante l'erase) è rilevato anche quando l'operazione flash
sottostante ha riportato successo.
\end{itemize}
\end{fnnote}
\begin{fnwarn}[Allineamento a settore obbligatorio]
\op{SAVE\_SLOT} (e i raw \op{FLASH\_WRITE\_BLOCK}/\op{FLASH\_ERASE}) richiedono che
l'indirizzo flash target sia allineato a settore da 4~KB --- rifiutato come errore
altrimenti, invece di un silenzioso read-modify-erase-write parziale del settore (non
esiste un buffer di scratch abbastanza grande per farlo, e ogni \op{SAVE\_SLOT} reale scrive
già uno slot intero e allineato per costruzione).
\end{fnwarn}
Razionale completo, ogni citazione da datasheet, ogni test avversariale (CRC non
corrispondente, slot mai salvato, attraversamento di confine pagina, simulazione di perdita
di alimentazione, contesa sull'arbitro) e i due bug reali trovati e corretti durante il
bring-up (uno pre-esistente in \code{psram\_controller.v}, uno nel nuovo handshake di
richiesta dell'arbitro) sono in \code{WORKLOG.md} (voci Fasi F1-F6) e
\code{docs/FPGA-Neural-Flash-Subsystem-Verification.md} (sunto di copertura per modulo, non
ripetuto qui).
\begin{tabularx}{\textwidth}{L{3.4cm}Y}
\toprule
\rowh \thd{Operazione} & \thd{Latenza reale misurata} \\
\midrule
ERASE (settore 4~KB) & $\approx$400~ms (dominata dal tSE interno del chip flash, indipendente dal clock host) \\
\rowa SAVE (pagina 256~B, incl. erase interno) & $\approx$403~ms (idem, tSE+tPP) \\
LOAD (4096~B) & 1.74~ms (2.35~MB/s) @80~MHz; 8.71~ms (0.47~MB/s) @16~MHz (solo SPI-clock-bound) \\
\bottomrule
\end{tabularx}
Metodologia di misura completa in \code{docs/FPGA-Neural-Flash-Subsystem-Verification.md}.
\section{Sequenze di sessione}
\subsection{Percorso single-layer}
\begin{lstlisting}[language=,caption={Sessione single-layer},basicstyle=\ttfamily\scriptsize]
RESET -> 0x0F
READ_CONFIG -> 0x30 (l'host apprende N_INPUTS/N_NEURONS/...)
WRITE_RAM (pesi) -> 0x01 ...
WRITE_RAM (bias) -> 0x01 ...
SET_BASE (X/W/BIAS) -> 0x10 x3
WRITE_RAM (input X) -> 0x01 ...
START -> 0x20
poll STATUS -> 0x21 (finche' done=1; si azzera a questa lettura)
READ_OUTPUT -> 0x22
\end{lstlisting}
\subsection{Percorso multi-layer (RUN\_NETWORK)}
\label{sec:run-network}
\begin{lstlisting}[language=,caption={Sessione multi-layer},basicstyle=\ttfamily\scriptsize]
WRITE_RAM (tabella descrittori) -> 0x01 ...
WRITE_RAM (pesi/bias per layer, X layer0)-> 0x01 ...
SET_BASE (X/TABLE/BUF_A/BUF_B) -> 0x10 x4
RUN_NETWORK(num_layers) -> 0x23 <num_layers>
poll STATUS -> 0x21 (finche' done=1)
READ_OUTPUT -> 0x22 (y_bus del layer finale)
\end{lstlisting}
\begin{fnnote}[Fuori ambito per v1]
Dual~SPI e CRC/checksum sui trasferimenti host (SPI assunto affidabile su traccia di
scheda --- da non confondere con il CRC32 del catalogo flash, §\ref{sec:flashspi}, che
protegge un dominio diverso: la persistenza flash$\leftrightarrow$PSRAM, non il link SPI
host).
\end{fnnote}
\begin{fnwarn}[\op{WRITE\_RAM}/\op{READ\_RAM} senza backpressure verso l'host --- rischio reale, non teorico]
Ogni byte ricevuto/prodotto deve essere completamente processato da \code{spi\_engine}
prima che arrivi il successivo confine di byte scandito da SCLK --- ragionevole per il
bulk-loading iniziale di pesi/ingressi, non un percorso real-time. Il rischio concreto: se
un host emette \op{WRITE\_RAM}/\op{READ\_RAM} prima che la sequenza di power-up di
\code{psram\_controller.v} sia completata ($\sim$150~\textmu s dopo il reset,
\code{STATE\_INIT}+\code{STATE\_CR\_INIT}), \code{spi\_engine} si blocca in attesa che il
primo accesso PSRAM completi, mentre l'host --- non rallentato da alcun handshake ---
continua a scandire byte. I byte ricevuti durante quello stallo vengono \textbf{scartati
silenziosamente}, senza errore e senza hang: solo dati sbagliati in PSRAM. Trovato durante
il lavoro sul sottosistema flash (\code{WORKLOG.md}, Fase~F5) con una riproduzione minimale
solo-\op{WRITE\_RAM}, senza alcun opcode flash coinvolto: è un rischio generale per
qualunque host, non specifico agli opcode flash. \textbf{Mitigazione attuale: l'host deve
attendere il power-up della PSRAM (o assicurarsi che la FPGA sia fuori reset da
$>$150~\textmu s) prima del suo primo \op{WRITE\_RAM}/\op{READ\_RAM}.} Non risolto a livello
di protocollo (richiederebbe una vera backpressure, una modifica più ampia) ---
dichiarato qui come rischio aperto, non aggirato silenziosamente.
\end{fnwarn}
@@ -1,233 +0,0 @@
\chapter[Programmazione della rete]{Programmazione della rete neurale}
\label{ch:prog}
Questo capitolo è la guida pratica alla codifica di una rete per FPGA-Neural: come si
dispone in memoria, quali registri si impostano e come si avvia, per entrambe le
topologie. Presuppone gli opcode SPI (cap.~\ref{ch:spi}) e i formati descrittore
(cap.~\ref{ch:seq}, \ref{ch:grafo}).
\section{Flusso generale}
Qualunque sia il tipo, il ciclo è lo stesso: l'host \emph{costruisce le strutture dati in
RAM}, imposta i \emph{registri base}, dichiara il \emph{tipo di rete}, \emph{avvia} e
\emph{rilegge} il risultato.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=3mm,start chain=going below,
every node/.style={on chain,fnblock,minimum width=64mm}]
\node[fnblockA]{1. \op{RESET} --- azzera il motore e il latch STATUS};
\node{2. \op{SET\_NET\_TYPE} --- dense (\#1) o graph (\#2)};
\node{3. \op{WRITE\_RAM} --- tabelle, pesi/edge, bias, input X};
\node{4. \op{SET\_BASE} --- registri base (x, table, \ldots)};
\node[fnblockT]{5. \op{RUN\_NETWORK} --- dispatch su \code{net\_type}};
\node{6. \op{STATUS} in polling --- attende \code{done}};
\node[fnblockD]{7. \op{READ\_OUTPUT} / \op{READ\_RAM} --- risultato};
\foreach \i [count=\j from 2] in {1,...,6} \draw[fnarrow] (chain-\i)--(chain-\j);
\end{tikzpicture}
\end{center}
\section{Registri e opcode coinvolti}
Tutti i valori base si impostano con \op{SET\_BASE} \code{sel(1B)+addr(3B)}. Selettori:
\begin{tabularx}{\textwidth}{C{1.0cm} L{3.4cm} C{1.4cm} C{1.4cm} Y}
\toprule
\rowh \thd{sel} & \thd{Registro} & \thd{Tipo \#1} & \thd{Tipo \#2} & \thd{Uso} \\
\midrule
0 & \code{x\_base} & \checkmark & \checkmark & Base input $X$. \\
\rowa 3 & \code{table\_base} & \checkmark & \checkmark & Tabella descrittori. \\
4 & \code{buf\_a\_base} & \checkmark & \checkmark\textsuperscript{$\ast$} & Ping-pong A (\#1) / \code{out\_base} riuso (\#2). \\
\rowa 5 & \code{buf\_b\_base} & \checkmark & --- & Ping-pong B (\#1). \\
9 & \code{num\_neurons\_graph} & --- & \checkmark & Numero neuroni del grafo. \\
\rowa 10 & \code{n\_out} & --- & \checkmark & Numero id di uscita. \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
$\ast$ In Tipo \#2 i buffer ping-pong non servono: il selettore 4 è riusato come
\code{out\_base} (regione dove copiare le uscite). I selettori 1/2/6/7/8 riguardano solo
il percorso single-layer manuale (\op{START}), non \op{RUN\_NETWORK}.\end{center}
Per il Tipo \#1, i \code{w\_base}/\code{bias\_addr} \emph{per-layer} \textbf{non} si
impostano con \op{SET\_BASE}: sono campi della tabella descrittori. \op{SET\_NET\_TYPE}
default dopo \op{RESET} è \emph{dense}, quindi una rete \#1 funziona anche senza emetterlo.
% ======================================================================
\section{Tipo \#1 --- rete densa}
\subsection{Layout in memoria}
\begin{tabularx}{\textwidth}{L{3.4cm} Y}
\toprule
\rowh \thd{Struttura} & \thd{Formato} \\
\midrule
Input $X$ & \code{n\_inputs\_real} byte INT8 a \code{x\_base}. \\
\rowa Pesi (per layer) & Neuron-major: neurone $k$ a \code{w\_base + k*n\_inputs\_real}, \code{n\_neurons*n\_inputs} byte. \\
Bias (per layer) & Un byte INT8 per neurone a \code{bias\_addr}. \\
\rowa Tabella descrittori & \code{num\_layers} voci da 11 byte a \code{table\_base}. \\
Buffer A/B & Uscite intermedie ping-pong. \\
\bottomrule
\end{tabularx}
Descrittore (11 byte, MSB-first): \code{w\_base}(3) $|$ \code{bias\_addr}(3) $|$
\code{activation}(1) $|$ \code{n\_inputs\_real}(2) $|$ \code{n\_neurons\_real}(2).
\subsection{Esempio completo: rete $4\to4\to2$}
Layer~0: 4 input, 4 neuroni, ReLU. Layer~1: 4 input, 2 neuroni, lineare
(\code{PARALLEL}=2, quindi ogni \code{n\_inputs\_real} è multiplo di 2). Indirizzi scelti:
\code{table\_base}=\code{0x000000}, \code{x\_base}=\code{0x001000}, pesi/bias L0 a
\code{0x002000}/\code{0x002100}, L1 a \code{0x002200}/\code{0x002300}, buffer a
\code{0x003000}/\code{0x003100}.
\begin{lstlisting}[language=,caption={Tabella descrittori dense (22 byte)},basicstyle=\ttfamily\scriptsize]
Layer 0: 00 20 00 | 00 21 00 | 01 | 00 04 | 00 04
w_base bias_addr ReLU n_in=4 n_neu=4
Layer 1: 00 22 00 | 00 23 00 | 00 | 00 04 | 00 02
w_base bias_addr NONE n_in=4 n_neu=2
\end{lstlisting}
\begin{lstlisting}[language=,caption={Sessione SPI (dense)},basicstyle=\ttfamily\scriptsize]
0x0F RESET
0x11 01 SET_NET_TYPE = dense
0x01 000000 0016 <22 byte tabella> WRITE_RAM tabella
0x01 002000 0010 <16 byte pesi L0> WRITE_RAM pesi L0 (neuron-major)
0x01 002100 0004 <4 byte bias L0>
0x01 002200 0008 <8 byte pesi L1>
0x01 002300 0002 <2 byte bias L1>
0x01 001000 0004 <x0 x1 x2 x3> WRITE_RAM input X
0x10 00 001000 SET_BASE x_base
0x10 03 000000 SET_BASE table_base
0x10 04 003000 SET_BASE buf_a
0x10 05 003100 SET_BASE buf_b
0x23 02 RUN_NETWORK num_layers=2
0x21 ... poll STATUS finche' done=1
0x22 READ_OUTPUT -> 2 byte (layer finale)
\end{lstlisting}
\subsection{Pseudocodice host (dense)}
\begin{lstlisting}[language=,caption={Codifica e caricamento di una rete densa},basicstyle=\ttfamily\scriptsize]
def load_dense(layers, X): # layers in ordine di esecuzione
spi(RESET); spi(SET_NET_TYPE, DENSE)
table = b""
for L in layers: # L: pesi[n][k], bias[n], act, n_in, n_out
assert L.n_in % PARALLEL == 0
w = alloc(L.weights_neuron_major) # k lento, input veloce
b = alloc(L.bias)
table += u24(w)+u24(b)+u8(L.act)+u16(L.n_in)+u16(L.n_out)
write_ram(TABLE_BASE, table)
write_ram(X_BASE, X)
set_base(0, X_BASE); set_base(3, TABLE_BASE)
set_base(4, BUF_A); set_base(5, BUF_B)
spi(RUN_NETWORK, len(layers))
wait_status_done()
return read_output(layers[-1].n_out)
\end{lstlisting}
% ======================================================================
\section{Tipo \#2 --- rete a grafo}
\subsection{Layout in memoria}
\begin{tabularx}{\textwidth}{L{3.4cm} Y}
\toprule
\rowh \thd{Struttura} & \thd{Formato} \\
\midrule
Input $X$ & \code{N\_in} byte a \code{x\_base}; copiati in \code{act\_buf[0..N\_in-1]} all'avvio. \\
\rowa Tabella descrittori & \code{num\_neurons\_graph} voci da 11 byte a \code{table\_base}, in ordine di \code{out\_id} crescente. \\
Blocchi edge & Per neurone: \code{n\_conn} edge da 4 byte a \code{conn\_ptr}, con padding a multiplo di \code{PARALLEL} (edge peso 0). \\
\rowa Uscite & \code{n\_out} byte scritti a \code{out\_base} (=selettore 4). \\
\bottomrule
\end{tabularx}
Descrittore graph (11 byte): \code{conn\_ptr}(3) $|$ \code{n\_conn}(2) $|$ \code{out\_id}(2)
$|$ \code{activation}(1) $|$ \code{bias}(1) $|$ \code{reserved}(2). \quad
Edge (4 byte): \code{src\_id}(2) $|$ \code{weight}(1) $|$ \code{reserved}(1). \quad
Vincolo: \code{src\_id < out\_id} (DAG feed-forward).
\subsection{Esempio completo}
4 ingressi (id 0--3). Neurone n4 (\code{out\_id}=4, ReLU, bias=2) connesso agli id 0 e 1;
neurone n5 (\code{out\_id}=5, lineare, bias=0) connesso a n4 (id~4) e all'id~2; uscita = n5
(\code{n\_out}=1). \code{PARALLEL}=2, entrambi hanno 2 connessioni (nessun padding).
Indirizzi: \code{table\_base}=\code{0x000000}, edge a \code{0x000100}, \code{x\_base}=
\code{0x001000}, \code{out\_base}=\code{0x002000}.
\begin{lstlisting}[language=,caption={Descrittori + edge grafo},basicstyle=\ttfamily\scriptsize]
Descrittori (a 0x000000, 22 byte):
n4: 00 01 00 | 00 02 | 00 04 | 01 | 02 | 00 00
conn_ptr n_conn out_id ReLU bias rsv
n5: 00 01 08 | 00 02 | 00 05 | 00 | 00 | 00 00
conn_ptr n_conn out_id NONE bias rsv
Blocchi edge (a 0x000100, 4 byte/edge: src_id, weight, rsv):
n4 @0x000100: 00 00 05 00 (src=0, w=+5)
00 01 FD 00 (src=1, w=-3) ; -3 = 0xFD
n5 @0x000108: 00 04 02 00 (src=4, w=+2) ; id4 = uscita di n4
00 02 07 00 (src=2, w=+7)
\end{lstlisting}
\begin{lstlisting}[language=,caption={Sessione SPI (graph)},basicstyle=\ttfamily\scriptsize]
0x0F RESET
0x11 02 SET_NET_TYPE = graph
0x01 000000 0016 <22 byte tabella> WRITE_RAM descrittori
0x01 000100 0010 <16 byte edge> WRITE_RAM blocchi edge
0x01 001000 0004 <x0 x1 x2 x3> WRITE_RAM input X
0x10 00 001000 SET_BASE x_base
0x10 03 000000 SET_BASE table_base
0x10 04 002000 SET_BASE out_base (riuso sel 4)
0x10 09 000002 SET_BASE num_neurons_graph = 2
0x10 0A 000001 SET_BASE n_out = 1
0x23 00 RUN_NETWORK (dispatch a graph_engine)
0x21 ... poll STATUS (bit2=err se src_id>=out_id)
0x02 002000 0001 READ_RAM out_base -> 1 byte (uscita n5)
\end{lstlisting}
\subsection{Pseudocodice host (graph)}
\begin{lstlisting}[language=,caption={Codifica e caricamento di un grafo},basicstyle=\ttfamily\scriptsize]
def load_graph(neurons, X, n_out): # neurons ordinati per out_id crescente
spi(RESET); spi(SET_NET_TYPE, GRAPH)
edges = b""; table = b""
for N in neurons: # N: out_id, conns=[(src_id,w)...], act, bias
for (src,_) in N.conns:
assert src < N.out_id and src < N_TOTAL # regola DAG
conn_ptr = EDGE_BASE + len(edges)
padded = pad(N.conns, PARALLEL, fill=(0,0)) # edge peso 0
for (src,w) in padded:
edges += u16(src)+i8(w)+u8(0)
table += u24(conn_ptr)+u16(len(N.conns))+u16(N.out_id) \
+ u8(N.act)+i8(N.bias)+u16(0)
write_ram(TABLE_BASE, table); write_ram(EDGE_BASE, edges)
write_ram(X_BASE, X)
set_base(0, X_BASE); set_base(3, TABLE_BASE); set_base(4, OUT_BASE)
set_base(9, len(neurons)); set_base(10, n_out)
spi(RUN_NETWORK, 0) # payload ignorato in graph
wait_status_done()
return read_ram(OUT_BASE, n_out)
\end{lstlisting}
\subsection{Pseudo-assembly \texttt{netasm}}
La descrizione leggibile viene compilata dall'assemblatore host (\code{tools/netasm/})
esattamente nei byte delle tabelle e degli edge sopra. Esempio equivalente al grafo
dell'esempio:
\begin{lstlisting}[language=,caption={netasm: sorgente e byte generati},basicstyle=\ttfamily\scriptsize]
; --- sorgente ---
NET graph
INPUTS 4 ; id 0..3
NEURON n4 relu bias=2
CONN 0 w=5
CONN 1 w=-3
NEURON n5 none bias=0
CONN n4 w=2 ; riferimento simbolico -> id 4
CONN 2 w=7
OUTPUT n5
END
; --- l'assemblatore emette ---
; id assegnati: n4=4, n5=5 (garantito src_id < out_id)
; descrittori: 00 01 00 00 02 00 04 01 02 00 00
; 00 01 08 00 02 00 05 00 00 00 00
; edge: 00 00 05 00 00 01 FD 00 (n4)
; 00 04 02 00 00 02 07 00 (n5)
; registri: table_base, x_base, out_base, num_neurons=2, n_out=1
; validato a compile-time: src_id<out_id, N_TOTAL, padding a PARALLEL
\end{lstlisting}
\begin{fnnote}[Perche' due livelli di codifica]
Lo pseudocodice host e il \code{netasm} producono gli \emph{stessi byte}. Il primo è utile
quando la rete è generata a runtime (es. pesi da training); il secondo quando la topologia
è scritta a mano o versionata come sorgente. In entrambi i casi l'FPGA riceve solo tabelle
e dati via \op{WRITE\_RAM}: nessun interprete a bordo.
\end{fnnote}
@@ -1,78 +0,0 @@
\chapter[Arbitraggio e top-level]{Arbitraggio e integrazione top-level}
\label{ch:top}
\section{\texttt{mem\_arbiter} --- arbitro a tre porte}
Un unico master di memoria byte-level (che alimenta la catena condivisa
\code{int8\_memory\_access} $\to$ \code{memory\_interface} $\to$ \code{psram\_controller})
è arbitrato tra tre richiedenti:
\begin{tabularx}{\textwidth}{C{1.3cm} L{3.4cm} Y}
\toprule
\rowh \thd{Porta} & \thd{Master} & \thd{Accessi} \\
\midrule
A & \code{spi\_engine} & \op{WRITE\_RAM} / \op{READ\_RAM}. \\
\rowa B & \code{neuron\_memory} & Letture X/W/bias durante un'esecuzione. \\
C & \code{layer\_sequencer} & Letture descrittori + scritture buffer tra layer. \\
\bottomrule
\end{tabularx}
Priorità fissa \textbf{B $>$ C $>$ A}: un'inferenza in corso è più critica della
contabilità del sequencer, che a sua volta è più critica di un accesso SPI manuale
appena arrivato. In funzionamento normale B e C sono comunque temporalmente disgiunti
(\code{neuron\_memory} richiede solo durante un'esecuzione, \code{layer\_sequencer} solo
nelle pause tra layer), quindi la priorità conta soprattutto per il caso limite di un
\op{WRITE\_RAM}/\op{READ\_RAM} manuale che arriva durante un'esecuzione multi-layer.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=6mm]
\node[fnblock,minimum width=30mm](a){Port A --- \code{spi\_engine}};
\node[fnblock,below=4mm of a,minimum width=30mm](b){Port B --- \code{neuron\_memory}};
\node[fnblock,below=4mm of b,minimum width=30mm](c){Port C --- \code{layer\_sequencer}};
\node[fnblockD,right=16mm of b,minimum width=26mm,minimum height=16mm](arb){\code{mem\_arbiter}\\{\scriptsize B$>$C$>$A}};
\node[fnblockT,right=14mm of arb,minimum width=26mm](m){catena memoria\\{\scriptsize condivisa}};
\draw[fnarrow] (a)-|(arb.west|-a); \draw[fnarrow] (b)--(arb.west);
\draw[fnarrow] (c)-|(arb.west|-c);
\draw[fnbus] (arb)--(m);
\end{tikzpicture}
\end{center}
Concesso l'accesso, l'arbitro mantiene la proprietà fino all'impulso \code{m\_ready}
della singola transazione, poi rilascia: tutti e tre i master emettono \code{req} come
impulso pulito di un ciclo, quindi è sufficiente un design grant-and-forward senza code.
\section{\texttt{spi\_neuron\_top} --- integrazione completa}
Il top-level collega SPI (\code{spi\_slave}+\code{spi\_engine}), l'arbitro, il sequencer,
\code{neuron\_memory} e la catena PSRAM. Il reset di \code{neuron\_memory} è l'OR del
reset globale con l'impulso di soft-reset dell'opcode \op{RESET}, così l'host può
recuperare il motore via SPI senza reset fisico (la RAM resta intatta).
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=7mm]
\node[fnblockA,minimum width=22mm](ss){\code{spi\_slave}};
\node[fnblockA,right=8mm of ss,minimum width=22mm](se){\code{spi\_engine}};
\node[fnblockT,below=8mm of se,minimum width=26mm](sq){\code{layer\_sequencer}};
\node[fnblockD,right=10mm of se,minimum width=24mm](mux){MUX ctrl\\{\scriptsize su \code{seq\_busy}}};
\node[fnblock,below=8mm of mux,minimum width=26mm](nm){\code{neuron\_memory}};
\node[fnblockD,right=10mm of mux,minimum width=22mm](arb){\code{mem\_arbiter}};
\node[fnblockA,right=8mm of arb,minimum width=26mm](mem){catena PSRAM};
\draw[fnarrow] (ss)--(se);
\draw[fnarrow] (se)--(mux);
\draw[fnarrow] (sq)--(mux);
\draw[fnarrow] (mux)--(nm);
\draw[fnarrow] (se.south) to[bend right=10] (arb.north west);
\draw[fnarrow] (nm)--(arb);
\draw[fnarrow] (sq.east) to[bend right=20] (arb.south west);
\draw[fnbus] (arb)--(mem);
\end{tikzpicture}
\end{center}
Il multiplexer commuta le linee di controllo di \code{neuron\_memory} tra il sequencer
(mentre \code{seq\_busy} è alto) e il percorso diretto di \code{spi\_engine} (modalità
single-layer legacy), restituendo il motore al percorso diretto a fine sequenza.
\begin{fnnote}[Verifica end-to-end]
\code{spi\_neuron\_top} è verificato in simulazione con PSRAM reale
(\code{psram\_model.v}, nessun mock): RESET/READ\_CONFIG/WRITE\_RAM/READ\_RAM/SET\_BASE/
START/STATUS/READ\_OUTPUT e \op{RUN\_NETWORK} sono esercitati puramente su SPI simulato
(cap.~\ref{ch:impl}).
\end{fnnote}
@@ -1,172 +0,0 @@
\chapter[Implementazione ECP5]{Implementazione e caratterizzazione ECP5}
\label{ch:impl}
\section{Flusso e verifica}
Il progetto è verificato su due piani complementari: \textbf{simulazione} funzionale con
Icarus Verilog (algebra signed, prodotti, accumulo, gruppi, bias, ReLU, saturazione,
segnali busy/done) e \textbf{implementazione} reale con Yosys (sintesi) $+$
nextpnr-ecp5 (place\&route, timing) $+$ Project~Trellis (\code{ecppack}).
\begin{tabularx}{\textwidth}{L{5.0cm} C{3.0cm} Y}
\toprule
\rowh \thd{Fase di verifica} & \thd{Esito} & \thd{Copre} \\
\midrule
RTL funzionale & \PASS & correttezza del datapath \\
\rowa Simulazione parametrica & \PASS & sweep di configurazioni \\
Sintesi ECP5 & \PASS & sintetizzabilità, mapping \\
\rowa Placement / Routing & \PASS & LUT/FF/DSP, timing \\
Bitstream (\code{ecppack}) & \PASS & flusso completo, 0 errori (P2 e P8) \\
\bottomrule
\end{tabularx}
\begin{fnnote}[Toolchain end-to-end fino al bitstream]
L'intero flusso RTL $\to$ Yosys $\to$ nextpnr-ecp5 $\to$ \code{ecppack} produce un
bitstream valido per P2 e P8, \textbf{0 errori in ogni stadio}. Header verificato
byte-per-byte: \code{Part: LFE5U-45F-8CABGA381}, il part number reale del target, non un
placeholder. Verificata la sola \emph{generazione}: nessun test su hardware fisico in
questa sessione.
\end{fnnote}
\section{Benchmark del datapath (256$\times$4)}
Configurazione: INT8/INT32, \code{N\_INPUTS}=256, \code{N\_NEURONS}=4, \code{PARALLEL}
variabile, target 80~MHz, dispositivo \code{LFE5U-45F-8BG381C} ($-8$). I bus di test
sono generati \emph{dentro} il wrapper di benchmark per non esporre migliaia di I/O; il
top-level espone solo \code{clk/rst/start/y\_bus/busy/done}.
\begin{tabularx}{\textwidth}{C{1.4cm} C{1.8cm} C{1.4cm} C{1.6cm} C{1.6cm} C{1.5cm} C{1.4cm}}
\toprule
\rowh \thd{PAR} & \thd{MAC tot} & \thd{DSP} & \thd{Fmax} & \thd{Tcrit} & \thd{80\,MHz} & \thd{LUT4} \\
\midrule
16 & 64 & 64/72 & 52.13 & 19.18 & \FAIL & $\approx$2531 \\
\rowa 8 & 32 & 32/72 & 61.71 & 16.20 & \FAIL & --- \\
4 & 16 & 16/72 & 75.01 & 13.33 & \FAIL & 804 \\
\rowa 2 & 8 & 8/72 & 87.88 & 11.38 & \PASS & 481 \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
Fmax e Tcrit in MHz e ns. MAC totali $=$ PARALLEL$\times$4 neuroni.\end{center}
\subsection{Fmax e throughput contro parallelismo}
\begin{center}
\begin{tikzpicture}
\begin{axis}[
width=0.62\textwidth,height=6.0cm,
axis y line*=left, axis x line=bottom,
xlabel={\footnotesize PARALLEL}, ylabel={\footnotesize Fmax [MHz]},
xtick={2,4,8,16}, xmode=log, log basis x=2,
ymin=40,ymax=95, ytick={40,55,70,85},
tick label style={font=\scriptsize}, label style={font=\footnotesize},
grid=major, grid style={fnRule!40},
legend style={font=\scriptsize,at={(0.5,-0.28)},anchor=north,legend columns=2}]
\addplot[fnTeal,mark=*,thick,mark options={fill=fnTeal}]
coordinates {(2,87.88)(4,75.01)(8,61.71)(16,52.13)};
\addlegendentry{Fmax}
\draw[fnAmber,dashed,thick] (axis cs:2,80)--(axis cs:16,80);
\node[font=\scriptsize,text=fnAmber] at (axis cs:11,82.5){target 80 MHz};
\end{axis}
\begin{axis}[
width=0.62\textwidth,height=6.0cm,
axis y line*=right, axis x line=none,
xmode=log, log basis x=2, xmin=2,xmax=16,
ylabel={\footnotesize throughput [G\,MAC/s]},
ymin=0,ymax=3.6, ytick={0,1,2,3},
tick label style={font=\scriptsize}, label style={font=\footnotesize}]
\addplot[fnBlue,mark=square*,thick,mark options={fill=fnBlue}]
coordinates {(2,0.703)(4,1.20)(8,1.97)(16,3.34)};
\label{plt:tp}
\end{axis}
\end{tikzpicture}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
Trade-off fondamentale: al crescere di PARALLEL la Fmax cala (routing/albero più
profondi) ma il throughput teorico sale. La linea blu (quadrati) è il throughput
$\approx$MAC/ciclo$\times$Fmax.\end{center}
\subsection{Interpretazione}
Riducendo \code{PARALLEL} calano MAC simultanei, DSP, profondità dell'adder tree e
congestione di routing, quindi la Fmax sale; ma aumenta il numero di gruppi e quindi la
latenza. La sola frequenza non basta a scegliere: conta il throughput complessivo
$\approx$MAC/ciclo$\times$frequenza.
\begin{fnnote}[Scelte architetturali]
\code{PARALLEL=8} è il candidato per la V1 orientata al throughput: esattamente 32~MAC
simultanei con 4 neuroni, DSP al $\approx$44\%, lasciando risorse per controller,
buffer, SPI e pipeline future. \code{PARALLEL=2} è il riferimento orientato alla
frequenza: 87.88~MHz, unico a superare il target 80~MHz, ma richiede 128 gruppi per un
neurone da 256 ingressi.
\end{fnnote}
\subsection{Percorso critico e limite a 100~MHz}
Il target 100~MHz non è raggiunto (miglior risultato 87.88~MHz con P2). Il limite è
\emph{temporale}, non di occupazione: con P2 l'FPGA è usato pochissimo (DSP $\approx$11\%,
LUT $\approx$1\%). Il percorso critico attraversa FF pesi $\to$ \code{MULT18X18D} $\to$
prodotti $\to$ adder/carry $\to$ \code{acc\_next} $\to$ ReLU/saturazione $\to$ FF uscita.
Superare 100~MHz richiederà una o più pipeline interne, non ancora necessarie per
proseguire.
\section{Sistema integrato completo}
Sintesi reale di \code{spi\_neuron\_top} (SPI + arbitro + \code{neuron\_memory} +
\code{graph\_engine} + catena PSRAM), speed grade $-8$. Prima della timing closure il
sistema integrato mancava il target 80~MHz (P2 $\approx$55~MHz, P8 $\approx$45~MHz), con
un percorso critico interamente interno a \code{neuron\_parallel}.
\subsection{Causa: catena di saturazione/ReLU}
L'utilizzo di risorse non è la causa (device sotto il 10\% ovunque). Il percorso critico
del sistema integrato è la \textbf{catena di riporto \code{CCU2C} del comparatore di
saturazione/ReLU} in \code{neuron\_parallel.v} --- \emph{non} lo SPI, l'arbitro, la PSRAM
né i moduli del Tipo \#2. La saturazione era scritta come confronto aritmetico
(\code{acc > 127}, \code{acc < -128}), mappato dal sintetizzatore su un sottrattore a
32~bit con carry chain lunga.
\subsection{Timing closure (2026-09-03)}
Deroga esplicita al vincolo ``datapath intoccabile'' per un task separato di timing
closure, con l'unico vincolo dell'\textbf{equivalenza bit-esatta} su tutta la regressione.
Due passi:
\begin{itemize}
\item \textbf{Passo 1 --- saturazione/ReLU come bit-test.} Un valore signed a 32~bit sta
in INT8 se e solo se \code{acc[31:7]} sono tutti uguali: riduzione AND/OR su una fetta di
bit invece di 32~bit di riporto. Semplificazione corretta e verificata bit-esatta, guadagno
di logica reale ma da solo sommerso dal rumore di piazzamento.
\item \textbf{Passo 2 --- registro di pipeline} tra accumulo e attivazione (\code{+1}
ciclo di latenza per neurone, assorbito dall'handshake \code{start}/\code{done}, trasparente
per i chiamanti). È il passo decisivo.
\end{itemize}
\begin{tabularx}{\textwidth}{L{4.6cm} C{2.6cm} C{2.4cm} Y}
\toprule
\rowh \thd{Config} & \thd{Prima} & \thd{Dopo} & \thd{$\Delta$} \\
\midrule
P2, \code{.lpf} reale & 54.58 & \textbf{75.30} & $+38\%$ \\
\rowa P2, sweep 5 seed & 55.59 & 73.38--75.55 & robusto \\
P8, unconstrained & 45.47 & \textbf{60.26} & $+33\%$ \\
\rowa P8, sweep 5 seed & 43.15--50.48 & 60.26--68.87 & non sovrapposto \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
Fmax in MHz, place\&route reale (\code{nextpnr-ecp5}). Guadagno robusto su 5 seed, non
attribuibile a fortuna di placement.\end{center}
\begin{fnnote}[Criterio di stop e margine reale]
80~MHz non è raggiunto (75.30~MHz a P2, 94\% del target) ma il guadagno è enorme e reale
($+38\%$/$+33\%$). Il passo successivo (registro di uscita del \code{MULT18X18D}, che
toccherebbe \code{mac\_unit.v}) è stato lasciato: gli 80~MHz sono \emph{headroom} in vista
del \code{.lpf} reale, non un requisito operativo. Con l'oscillatore previsto a 16~MHz,
anche il numero peggiore misurato ($\approx$45~MHz a P8) ha $2.8\times$ di margine.
\textbf{Superato 2026-09-04}: dopo l'aggiunta del sottosistema flash (cap.~\ref{ch:spi}
§\ref{sec:flashspi}, cap.~\ref{ch:roadmap}) la Fmax del sistema completo (P2, stesso
pinout reale + 3 nuovi segnali flash) era 66.68~MHz, percorso critico ancora sulla
stessa catena di accumulo di \code{neuron\_parallel} identificata qui sopra --- non un
nuovo collo di bottiglia, la differenza rispetto a 75.30~MHz rumore di piazzamento/routing
dovuto ai pin/logica aggiuntivi. \textbf{Aggiornato di nuovo lo stesso giorno (Fase F7)}:
reso il bus SPI della flash genuinamente indipendente (rimosso il riuso del pad \code{CCLK}
via \code{USRMCLK}, aggiunto un 4°~pin \code{flash\_sclk} ordinario), Fmax ri-misurata
\textbf{67.91~MHz} (leggero miglioramento, percorso critico confermato ancora identico).
Margine sull'oscillatore 16~MHz: $4.2\times$.
\end{fnnote}
\begin{fnnote}[Ottimizzazione futura separata]
Indipendente dalla timing closure: gli array \code{x\_mem}/\code{w\_mem} di
\code{neuron\_memory} sono ancora inferiti come RAM distribuita su LUT anziché su
\code{DP16KD}. Spostarli su block RAM libererebbe LUT ed è un candidato per la Fase~7 ---
non era però sul percorso critico risolto qui.
\end{fnnote}
@@ -1,325 +0,0 @@
\chapter[Progetto hardware e pinout]{Progetto hardware e mappa dei segnali}
\label{ch:hw}
\begin{fnnote}[Stato del pinout --- assegnato e verificato]
Esiste ora un \code{.lpf} reale (\code{synth/ecp5/spi\_neuron\_top.lpf}) con i \textbf{57
segnali} del top-level assegnati a ball CABGA381 concrete, \textbf{verificato
da un place\&route \code{nextpnr-ecp5} completo a 0 errori} (non più
\code{-{}-lpf-allow-unconstrained}). Le ball derivano dal database di dispositivo di
Project~Trellis (\code{iodb.json}, lo stesso che usa nextpnr) e sono state validate in modo
indipendente contro la §4.3.2 del datasheet Lattice ufficiale (conteggi GPIO per banco:
coincidenza esatta su 6 banchi su 7, scostamento di 1 ball sul banco 3, irrilevante perché
nessun segnale assegnato lo usa). \code{TRELLIS\_IO}: 57/245 (23\%). Fmax del build
corrente (sistema completo incl. sottosistema flash con bus SPI indipendente, Fase F7,
2026-09-04) \textbf{67.91~MHz}, percorso critico confermato ancora sulla catena di accumulo
di \code{neuron\_parallel}, invariato rispetto alle build precedenti (cap.~\ref{ch:impl}).
Sunto pin-per-pin a inizio documento (pagg.~2--3). Le ball di config-SPI di boot e JTAG non
compaiono qui perché sono pin dedicati a funzione fissa, senza porta RTL corrispondente:
nextpnr non le richiede mai (0 errori), contano solo per lo schematic PCB.
\end{fnnote}
\section{Dispositivo target}
\begin{tabularx}{\textwidth}{L{4.2cm}Y}
\toprule
\rowh \thd{Parametro} & \thd{Valore} \\
\midrule
Dispositivo & Lattice ECP5 \code{LFE5U-45F-8BG381C} \\
\rowa Package & CABGA381 (381 ball) \\
Speed grade & $-8$ (il più veloce della famiglia ECP5) \\
\rowa Risorse & $\approx$44k LUT/FF, 72$\times$\code{MULT18X18D}, block RAM \code{DP16KD} \\
I/O utilizzabili & $\approx$232 ball su 381 (resto: alimentazione/massa/NC) \\
\bottomrule
\end{tabularx}
\section{Budget dei pin}
Il progetto richiede circa 60 segnali su $\approx$232 I/O utilizzabili: ampio margine
($>$170 pin liberi), quindi la scheda non è pin-constrained.
\begin{tabularx}{\textwidth}{Y C{2.2cm}}
\toprule
\rowh \thd{Funzione} & \thd{Pin} \\
\midrule
PSRAM (indirizzi 22, dati 16, controllo 6) & fino a 44 \\
\rowa SPI applicativo (\code{sclk/mosi/miso/cs\_n}) & 4 \\
Clock, reset & 2 \\
\rowa Pin attenzione host (\code{irq\_n}, \code{data\_ready\_n}) & 2 \\
Bus SPI flash runtime (\code{flash\_sclk/flash\_mosi/flash\_miso/flash\_cs\_n}, GPIO ordinario, bus indipendente --- Fase F7) & 4 \\
\rowa JTAG (bring-up / debug, consigliato) & 4 \\
\midrule
\rowh \thd{Totale} & \thd{$\approx$60} \\
\bottomrule
\end{tabularx}
\section{Mappa dei segnali (top-level \texttt{spi\_neuron\_top}) --- ball reali}
Assegnazione reale dei 57 segnali del top-level, verificata da place\&route, \textbf{ball
individuale per ogni bit} (mai un intervallo di bus). Standard I/O: LVCMOS33
(alimentazione I/O a 3.3~V). Le ball provengono dal \code{.lpf} reale
place\&route-verified. Sunto compatto della stessa tabella anche a inizio documento
(pagg.~2--3).
\renewcommand{\arraystretch}{1.1}
\begin{tabularx}{\textwidth}{L{3.0cm} C{1.0cm} C{1.9cm} C{1.0cm} Y}
\toprule
\rowh \thd{Segnale} & \thd{Dir} & \thd{Ball} & \thd{Banco} & \thd{Funzione} \\
\midrule
\multicolumn{5}{l}{\textit{\color{fnDark}Clock e reset (banco 7, lato sinistro)}}\\
\code{clk} & IN & H5 & 7 & Clock di sistema su pad \code{GR\_PCLK7\_0} (clock globale dedicato). \\
\rowa \code{rst} & IN & B4 & 7 & Reset globale sincrono, attivo alto. \\
\multicolumn{5}{l}{\textit{\color{fnDark}SPI applicativo (banco 7, opposto al bus PSRAM)}}\\
\code{sclk} & IN & B5 & 7 & SPI clock (CPOL=0, CPHA=0). \\
\rowa \code{mosi} & IN & C5 & 7 & Master-Out Slave-In. \\
\code{miso} & OUT & A3 & 7 & Master-In Slave-Out (pilotato sul fronte di discesa). \\
\rowa \code{cs\_n} & IN & B3 & 7 & Chip-select attivo basso. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Pin di attenzione host (banco 7, attivi bassi, di livello)}}\\
\code{data\_ready\_n} & OUT & C3 & 7 & Basso finché un risultato attende lettura (specchio di \code{STATUS.done}, clear su lettura STATUS). \\
\rowa \code{irq\_n} & OUT & C4 & 7 & Basso se il guard load-time del grafo è scattato (specchio di \code{STATUS.err}); si azzera solo su \code{RESET} o nuovo \code{run\_start}, \emph{non} su lettura STATUS. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Flash subsystem --- SPI verso W25Q128JV onboard, bus indipendente (banco 7, Fasi F1-F7)}}\\
\code{flash\_sclk} & OUT & E3 & 7 & SPI clock verso la flash --- GPIO ordinario, nessuna primitiva di config coinvolta (Fase F7). \\
\rowa \code{flash\_mosi} & OUT & D3 & 7 & Master-Out Slave-In verso la flash. \\
\code{flash\_miso} & IN & D5 & 7 & Master-In Slave-Out dalla flash. \\
\rowa \code{flash\_cs\_n} & OUT & E4 & 7 & Chip-select flash, attivo basso. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Bus PSRAM indirizzi \code{psram\_a[21:0]} --- 22 ball individuali (banco 2)}}\\
\code{psram\_a[0]} & OUT & E16 & 2 & PSRAM A0 \\
\rowa \code{psram\_a[1]} & OUT & F16 & 2 & PSRAM A1 \\
\code{psram\_a[2]} & OUT & D18 & 2 & PSRAM A2 \\
\rowa \code{psram\_a[3]} & OUT & E17 & 2 & PSRAM A3 \\
\code{psram\_a[4]} & OUT & E18 & 2 & PSRAM A4 \\
\rowa \code{psram\_a[5]} & OUT & F18 & 2 & PSRAM A5 \\
\code{psram\_a[6]} & OUT & F17 & 2 & PSRAM A6 \\
\rowa \code{psram\_a[7]} & OUT & G16 & 2 & PSRAM A7 \\
\code{psram\_a[8]} & OUT & G18 & 2 & PSRAM A8 \\
\rowa \code{psram\_a[9]} & OUT & H16 & 2 & PSRAM A9 \\
\code{psram\_a[10]} & OUT & H17 & 2 & PSRAM A10 \\
\rowa \code{psram\_a[11]} & OUT & H18 & 2 & PSRAM A11 \\
\code{psram\_a[12]} & OUT & J16 & 2 & PSRAM A12 \\
\rowa \code{psram\_a[13]} & OUT & J17 & 2 & PSRAM A13 \\
\code{psram\_a[14]} & OUT & C20 & 2 & PSRAM A14 \\
\rowa \code{psram\_a[15]} & OUT & D19 & 2 & PSRAM A15 \\
\code{psram\_a[16]} & OUT & E19 & 2 & PSRAM A16 \\
\rowa \code{psram\_a[17]} & OUT & E20 & 2 & PSRAM A17 \\
\code{psram\_a[18]} & OUT & F19 & 2 & PSRAM A18 \\
\rowa \code{psram\_a[19]} & OUT & F20 & 2 & PSRAM A19 \\
\code{psram\_a[20]} & OUT & G20 & 2 & PSRAM A20 \\
\rowa \code{psram\_a[21]} & OUT & H20 & 2 & PSRAM A21 \\
\code{psram\_a[22]} & OUT & P18 & 3 & Sempre 0 (shift byte$\to$word): NC sulla scheda. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Bus PSRAM dati \code{psram\_dq[15:0]} --- 16 ball individuali (banchi 2 e 3)}}\\
\rowa \code{psram\_dq[0]} & IO & K18 & 2 & PSRAM DQ0 \\
\code{psram\_dq[1]} & IO & C18 & 2 & PSRAM DQ1 (dual-function, usata come GPIO ordinario). \\
\rowa \code{psram\_dq[2]} & IO & D17 & 2 & PSRAM DQ2 \\
\code{psram\_dq[3]} & IO & D20 & 2 & PSRAM DQ3 \\
\rowa \code{psram\_dq[4]} & IO & G19 & 2 & PSRAM DQ4 \\
\code{psram\_dq[5]} & IO & J18 & 2 & PSRAM DQ5 \\
\rowa \code{psram\_dq[6]} & IO & J19 & 2 & PSRAM DQ6 \\
\code{psram\_dq[7]} & IO & J20 & 2 & PSRAM DQ7 \\
\rowa \code{psram\_dq[8]} & IO & K19 & 2 & PSRAM DQ8 \\
\code{psram\_dq[9]} & IO & K20 & 2 & PSRAM DQ9 \\
\rowa \code{psram\_dq[10]} & IO & L17 & 3 & PSRAM DQ10 \\
\code{psram\_dq[11]} & IO & M18 & 3 & PSRAM DQ11 \\
\rowa \code{psram\_dq[12]} & IO & M17 & 3 & PSRAM DQ12 \\
\code{psram\_dq[13]} & IO & N16 & 3 & PSRAM DQ13 \\
\rowa \code{psram\_dq[14]} & IO & N18 & 3 & PSRAM DQ14 \\
\code{psram\_dq[15]} & IO & P17 & 3 & PSRAM DQ15 (bus dati bidirezionale tri-state, \code{dq\_oe} = direzione). \\
\multicolumn{5}{l}{\textit{\color{fnDark}Controllo PSRAM (banco 3)}}\\
\rowa \code{psram\_ce\_n} & OUT & N17 & 3 & Chip enable, attivo basso. \\
\code{psram\_oe\_n} & OUT & R16 & 3 & Output enable (lettura). \\
\rowa \code{psram\_we\_n} & OUT & R17 & 3 & Write enable (scrittura). \\
\code{psram\_lb\_n} & OUT & T16 & 3 & Lower-byte enable (DQ[7:0]). \\
\rowa \code{psram\_ub\_n} & OUT & N19 & 3 & Upper-byte enable (DQ[15:8]). \\
\code{psram\_zz\_n} & OUT & N20 & 3 & Sleep/snooze (inattivo=alto in funzionamento). \\
\bottomrule
\end{tabularx}
\renewcommand{\arraystretch}{1.25}
\begin{fnnote}[Segnali di scheda non esposti come porte RTL]
Non sono porte di \code{spi\_neuron\_top} ma vanno previsti a livello di scheda: le linee
di \textbf{SPI di configurazione} verso la flash NOR onboard (\code{PROGRAMN}/\code{INITN}/
\code{DONE}/\code{CCLK}\ldots, i ``Miscellaneous Dedicated Pins'' del datasheet) e le 4
linee \textbf{JTAG} (\code{TCK}/\code{TMS}/\code{TDI}/\code{TDO}), l'\textbf{oscillatore}
sul pad \code{PCLK}, le \textbf{alimentazioni}. I loro numeri di ball non sono nel datasheet
Lattice (file separato) ma non servono qui: sono pin dedicati senza porta RTL, nextpnr non
li richiede mai (0 errori), contano solo per lo schematic PCB.
\end{fnnote}
\begin{fnwarn}[SPI applicativo separato dallo SPI di configurazione]
L'SPI applicativo (\code{sclk/mosi/miso/cs\_n}) deve cadere su I/O ordinarie,
\textbf{mai} sui pin dell'SPI di configurazione: il pin di clock della config-SPI non è
riutilizzabile come ingresso generico dopo la configurazione senza workaround a livello
di scheda. Tenerli fisicamente separati evita quel problema.
\end{fnwarn}
\section{Allocazione per banchi (geometria reale del die)}
La collocazione segue la geometria dei bordi del die (da \code{globals.json} di Trellis,
ball~$\to$~(col,row)~$\to$~banco): i banchi \textbf{2 e 3} sono contigui lungo il bordo
\textbf{destro} del chip e ospitano insieme l'intero bus PSRAM (44+1 segnali) --- esattamente
gli ``uno o due banchi adiacenti'' raccomandati. Il banco \textbf{7} (bordo \textbf{sinistro},
fisicamente opposto al bus PSRAM) ospita SPI applicativo e clock/reset, deliberatamente sul
lato opposto per non far incrociare i due bus. \code{clk} è sul pad dedicato \code{H5}
(\code{GR\_PCLK7\_0}). Dove un banco esauriva le ball ``plain'' (parte di \code{psram\_dq}),
è stata usata la ball dual-function successiva come GPIO ordinario, confermata utilizzabile
dal place\&route reale.
\begin{tabularx}{\textwidth}{Y C{1.6cm} L{4.4cm}}
\toprule
\rowh \thd{Gruppo di segnali} & \thd{N. pin} & \thd{Banco (reale)} \\
\midrule
Indirizzi PSRAM \code{psram\_a[21:0]} & 22 & banco 2 (bordo destro) \\
\rowa Dati PSRAM \code{psram\_dq[15:0]} & 16 & banchi 2 + 3 (adiacenti) \\
Controllo PSRAM (ce/oe/we/lb/ub/zz) & 6 & banco 3 \\
\rowa SPI applicativo & 4 & banco 7 (bordo sinistro) \\
Pin attenzione host (\code{irq\_n}, \code{data\_ready\_n}) & 2 & banco 7 \\
\rowa Bus SPI flash indipendente (\code{flash\_sclk/flash\_mosi/flash\_miso/flash\_cs\_n}) & 4 & banco 7 \\
Clock / reset & 2 & banco 7, \code{clk} su \code{GR\_PCLK7\_0} \\
\rowa Config SPI boot / JTAG & --- & pin dedicati (fuori RTL, solo PCB) \\
\bottomrule
\end{tabularx}
\section{Sottosistema PSRAM}
Il controller \code{psram\_controller.v} implementa un'interfaccia \textbf{parallela
asincrona} (bus indirizzi, dati 16-bit, \code{ce\_n/oe\_n/we\_n} e byte-lane
\code{lb\_n/ub\_n}, più \code{zz\_n}) con latenza di accesso \textbf{70~ns} cablata come
$\lceil 70\,\text{ns}\times f_{clk}\rceil$. È un bus in stile SRAM asincrona, non QSPI.
\begin{tabularx}{\textwidth}{L{3.0cm}Y}
\toprule
\rowh \thd{Ruolo} & \thd{Componente} \\
\midrule
Memoria di lavoro & ISSI \code{IS66WVE4M16EBLL-70BLI} --- PSRAM parallela 64\,Mbit (4M$\times$16, 8~MB), async, 70~ns, corrispondente esatto alla temporizzazione del controller. \\
\rowa Fallback & ISSI \code{IS61WV6416DBLL} / \code{IS61WV102416BLL} (SRAM async vera, drop-in sugli stessi segnali, \code{zz\_n} inattivo, $\sim$10~ns, densità minore). \\
Storage persistente & Winbond \code{W25Q128JV} --- flash NOR SPI 16~MB per bitstream, pesi, bias, metadati di rete. \\
\bottomrule
\end{tabularx}
\subsection{Collegamento PSRAM (esclusivo della FPGA)}
La PSRAM è pilotata \textbf{esclusivamente dalla FPGA} tramite \code{psram\_controller.v}:
nessun master esterno accede al bus. L'host esterno (RPi/ESP32/MCU) parla solo SPI con la
FPGA e non tocca mai queste linee. Collegamento pin-per-pin FPGA~$\leftrightarrow$~ISSI
\code{IS66WVE4M16EBLL-70BLI}:
\begin{tabularx}{\textwidth}{L{3.6cm} L{3.0cm} Y}
\toprule
\rowh \thd{Segnale FPGA} & \thd{Pin PSRAM} & \thd{Funzione} \\
\midrule
\code{psram\_a[21:0]} & A0--A21 & Bus indirizzi (22 linee, 8~MB word address). \\
\rowa \code{psram\_dq[15:0]} & DQ0--DQ15 & Bus dati bidirezionale (tri-state, \code{dq\_oe}=direzione). \\
\code{psram\_ce\_n} & CE\# & Chip enable (attivo basso). \\
\rowa \code{psram\_oe\_n} & OE\# & Output enable (lettura). \\
\code{psram\_we\_n} & WE\# & Write enable (scrittura). \\
\rowa \code{psram\_lb\_n} & LB\# & Lower-byte enable (DQ[7:0]). \\
\code{psram\_ub\_n} & UB\# & Upper-byte enable (DQ[15:8]). \\
\rowa \code{psram\_zz\_n} & ZZ\# & Sleep/snooze (tenuto alto in funzionamento). \\
\bottomrule
\end{tabularx}
Alimentazione PSRAM: \textbf{3.3~V} (variante BLL), sullo stesso rail I/O dei banchi 2/3
a cui è cablata (cap.~\ref{ch:hw}, ball reali). Disaccoppiamento per pin di alimentazione
secondo il datasheet ISSI.
\section{Clock}
\label{sec:clock}
Non esiste ancora alcun PLL nell'RTL: \code{CLK\_FREQ\_MHZ} è un \emph{parametro di
temporizzazione} (alimenta le formule di accesso PSRAM), non un generatore di clock.
L'oscillatore montato pilota \code{clk} direttamente. Raccomandazione: oscillatore MEMS
16~MHz (famiglia SiTime SiT2001B), ben al di sotto dei 67.91~MHz di Fmax del sistema
integrato completo (incl. sottosistema flash, cap.~\ref{ch:impl}). \code{CLK\_FREQ\_MHZ} deve essere impostato al valore reale
dell'oscillatore montato, altrimenti la temporizzazione PSRAM risulta errata.
\section{Alimentazione}
Albero a \textbf{tre rail} (la sezione SERDES dell'eval board Lattice non serve e va
omessa: niente \code{VCCA}/\code{VCCHTX} a 1.2~V):
\begin{tabularx}{\textwidth}{L{3.4cm} C{2.0cm} Y}
\toprule
\rowh \thd{Rail} & \thd{Tensione} & \thd{Alimenta / regolatore} \\
\midrule
\code{VCC} (core) & 1.1~V & Core logico FPGA. Buck \code{TLV62568}, $\geq$600~mA. \\
\rowa \code{VCCIO0/2/3/6/7} & 3.3~V & I/O di tutti i banchi usati + PSRAM. Buck \code{TLV62568}, 1~A. \\
\code{VCCAUX} & 2.5~V & Ausiliario FPGA. LDO \code{TLV73325}, 10~mA. \\
\bottomrule
\end{tabularx}
Disaccoppiamento: almeno un condensatore per pin di alimentazione + bulk per rail, secondo
la checklist hardware ECP5 Lattice. Ingresso: 12~V esterno (o adatta i buck alla sorgente).
\section{Configurazione e programmazione}
\label{sec:config}
La ``scrittura della mappa'' dell'FPGA (bitstream) avviene tramite pin dedicati del
silicio, \textbf{non} porte del top-level RTL. Modo di default: \textbf{MSPI} --- boot
automatico dalla flash NOR all'accensione (prodotto standalone); JTAG disponibile per lo
sviluppo.
\subsection{JTAG (sviluppo / debug)}
\begin{tabularx}{\textwidth}{L{3.0cm} C{2.2cm} Y}
\toprule
\rowh \thd{Segnale} & \thd{Ball\textsuperscript{$\dagger$}} & \thd{Funzione} \\
\midrule
\code{TCK} & T5 & Test clock. \\
\rowa \code{TDI} & R5 & Test data in. \\
\code{TDO} & V4 & Test data out. \\
\rowa \code{TMS} & U5 & Test mode select. \\
\bottomrule
\end{tabularx}
\subsection{Config-SPI verso boot flash}
La FPGA carica il bitstream dalla \textbf{Winbond \code{W25Q128JV}} (128~Mbit SPI NOR,
Quad read) all'accensione. Il sottosistema flash (\code{rtl/flash\_slot\_manager.v}, Fasi
F1-F7, cap.~\ref{ch:impl}) usa la \textbf{stessa flash fisica} per pesi/bias/metadati di
rete a runtime, accesso esclusivo della FPGA: dopo la configurazione, la FPGA riprende il
controllo del chip via un bus SPI a 4 fili completamente indipendente,
\code{flash\_sclk/flash\_mosi/flash\_miso/flash\_cs\_n} (tutti GPIO ordinario, pagg.~2--3 e
§``Mappa dei segnali'' --- nessuna primitiva di configurazione ECP5 coinvolta, Fase F7) ---
implica comunque un doppio collegamento a livello di scheda (DI/DO/CS/CLK della flash
cablati sia ai pin dedicati di boot sotto sia a queste 4 ball ordinarie, poiché è lo stesso
chip fisico a svolgere entrambi i ruoli), non ancora riportato in uno schematico (nessuno
esiste ancora, vedi checklist sotto).
\begin{tabularx}{\textwidth}{L{3.4cm} C{2.2cm} Y}
\toprule
\rowh \thd{Segnale} & \thd{Ball\textsuperscript{$\dagger$}} & \thd{Funzione} \\
\midrule
\code{CCLK/MCLK/SCK} & U3 & Clock di configurazione. \\
\rowa \code{DQ0\_MOSI} & W2 & Dato config (MOSI). \\
\code{DQ1\_MISO} & V2 & Dato config (MISO). \\
\rowa \code{BUSY\_CSSPIN} & R2 & Chip-select flash. \\
\code{DQ2 / DQ3} & Y2 / W1 & Linee per Quad read. \\
\rowa \code{PROGRAMN} & W3 & Avvia riconfigurazione (pulsante, attivo basso). \\
\code{INITN} & V3 & Init / errore di configurazione (LED). \\
\rowa \code{DONE} & Y3 & Configurazione completata (LED). \\
\code{CFGMDN[2:0]} & R4/T4/U4 & Selezione modo (vedi sotto). \\
\bottomrule
\end{tabularx}
\subsection{Modi di configurazione (\texttt{CFGMDN})}
\begin{tabularx}{\textwidth}{L{4.0cm} C{4.0cm} Y}
\toprule
\rowh \thd{Modo} & \thd{CFGMDN[2:0]} & \thd{Uso} \\
\midrule
MSPI (boot da flash) & \code{010} & \textbf{Default} --- standalone. \\
\rowa SSPI (slave SPI) & \code{001} & Config da host esterno. \\
SCM (slave serial) & \code{101} & Config seriale. \\
\rowa SPCM (slave parallel) & \code{111} & Config parallela 8-bit. \\
\bottomrule
\end{tabularx}
\begin{fnwarn}[Ball di configurazione da verificare sul 45F]
\textsuperscript{$\dagger$}Le ball di JTAG e config-SPI qui riportate sono il
\emph{riferimento} dell'eval board Lattice (device 85F). JTAG e config-SPI sono pin
dedicati e in gran parte fissi nella famiglia ECP5, ma le posizioni esatte sul target
\code{LFE5U-45F-8BG381C} vanno confermate sul file pinout Lattice del 45F (Diamond/Radiant
o database Trellis) prima di committarle nello schematico, come già fatto per i segnali
applicativi (cap.~\ref{ch:hw}).
\end{fnwarn}
\section{Attività aperte prima della cattura schematica}
\begin{itemize}
\item[\OK] \code{ADDR\_WIDTH}=23 (8~MB pieni) su tutti i moduli e testbench.
\item[\OK] \code{.lpf} reale con l'assegnazione ball CABGA381, place\&route-verified a
0 errori (\code{synth/ecp5/spi\_neuron\_top.lpf}, 57 segnali incl. sottosistema flash).
\item[\OK] Sottosistema flash boot/persistenza (Fasi F1-F7): SPI master, copy engine,
catalogo a slot con CRC32, bus SPI a 4 fili indipendente (nessuna primitiva di
configurazione condivisa), sintesi reale a 0 errori, Fmax 67.91~MHz (\code{WORKLOG.md}).
\item[$\square$] Confermare signal integrity PSRAM/SPI al clock effettivamente montato.
\item[$\square$] Schema di doppio collegamento DI/DO/CS/CLK della flash (pin dedicati di
boot + le 4 ball ordinarie del sottosistema flash) --- non ancora catturato a
schematico.
\item[$\square$] Scelta del footprint del connettore JTAG.
\item[$\square$] Cattura schematica (KiCad o altro): nessuno schema esiste ancora per
questa combinazione dispositivo/package.
\end{itemize}
@@ -1,81 +0,0 @@
\chapter{Riferimento rapido}
\label{ch:ref}
\section{Opcode SPI}
\begin{tabularx}{\textwidth}{C{1.4cm} L{3.2cm} C{2.4cm} Y}
\toprule
\rowh \thd{Valore} & \thd{Nome} & \thd{Risposta} & \thd{Sintesi} \\
\midrule
\op{0x00} & NOP & --- & idle \\
\rowa \op{0x01} & WRITE\_RAM & --- & scrittura blocco PSRAM \\
\op{0x02} & READ\_RAM & \code{len} B & lettura blocco PSRAM \\
\rowa \op{0x0F} & RESET & --- & reset motore + latch STATUS \\
\op{0x10} & SET\_BASE & --- & imposta base/registro (sel 0..10) \\
\rowa \op{0x11} & SET\_NET\_TYPE & --- & tipo rete \#1/\#2 \\
\rowa \op{0x20} & START & --- & avvio single-layer \\
\op{0x21} & STATUS & 1 B & busy(live)/done(sticky) \\
\rowa \op{0x22} & READ\_OUTPUT & N\_NEURONS B & \code{y\_bus} \\
\op{0x23} & RUN\_NETWORK & --- & avvio multi-layer \\
\rowa \op{0x30} & READ\_CONFIG & 11 B & record configurazione \\
\bottomrule
\end{tabularx}
\section{Byte di STATUS}
\begin{center}
\begin{tikzpicture}[font=\scriptsize]
\foreach \i/\lbl [count=\x from 0] in {7/0,6/0,5/0,4/0,3/0,2/0,1/{done},0/{busy}}{
\node[fnreg,minimum width=13mm,minimum height=9mm] (b\x) at (\x*13mm,0) {\lbl};
\node[font=\tiny,text=fnGrey,above=0.5mm of b\x] {bit \i};
}
\node[fill=fnAmber,text=white,rounded corners=1pt,inner sep=1.5pt,font=\tiny]
at (b7.center){riservati = 0};
\node[fill=fnTeal,text=white,rounded corners=1pt,inner sep=1.5pt,font=\tiny]
at (b6.center){};
\end{tikzpicture}
\end{center}
\code{done} è sticky, clear-on-read; \code{busy} è live; \code{bit2=err} (guard grafo).
\section{Selettori SET\_BASE}
\begin{multicols}{2}\footnotesize
\begin{itemize}
\item 0 --- \code{x\_base}
\item 1 --- \code{w\_base}
\item 2 --- \code{bias\_addr}
\item 3 --- \code{table\_base}
\item 4 --- \code{buf\_a\_base}
\columnbreak
\item 5 --- \code{buf\_b\_base}
\item 6 --- \code{activation} (single-layer)
\item 7 --- \code{n\_inputs\_real} (single-layer)
\item 8 --- \code{n\_neurons\_real} (single-layer)
\item 9 --- \code{num\_neurons\_graph} (Tipo \#2)
\item 10 --- \code{n\_out} (Tipo \#2)
\end{itemize}
\end{multicols}
\section{Tabella descrittori (11 byte/layer, MSB-first)}
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=0mm]
\node[fnreg,minimum width=20mm,minimum height=8mm](a){\code{w\_base}\\3B};
\node[fnreg,minimum width=20mm,minimum height=8mm,right=0mm of a](b){\code{bias\_addr}\\3B};
\node[fnreg,minimum width=14mm,minimum height=8mm,right=0mm of b](c){\code{act}\\1B};
\node[fnreg,minimum width=22mm,minimum height=8mm,right=0mm of c](d){\code{n\_inputs\_real}\\2B};
\node[fnreg,minimum width=22mm,minimum height=8mm,right=0mm of d](e){\code{n\_neurons\_real}\\2B};
\end{tikzpicture}
\end{center}
\section{Parametri di build}
\begin{multicols}{2}\footnotesize
\begin{itemize}
\item \code{DATA\_WIDTH} --- 8 (INT8)
\item \code{ACC\_WIDTH} --- 32 (INT32)
\item \code{N\_INPUTS} --- max ingressi
\item \code{N\_NEURONS} --- max neuroni
\item \code{PARALLEL} --- MAC simultanei
\columnbreak
\item \code{N\_LAYERS} --- max layer
\item \code{ADDR\_WIDTH} --- 23 (8 MB)
\item \code{MEM\_DATA\_WIDTH} --- 16
\item \code{CLK\_FREQ\_MHZ} --- timing PSRAM
\end{itemize}
\end{multicols}
@@ -1,65 +0,0 @@
\chapter{Roadmap e stato di sviluppo}
\label{ch:roadmap}
\section{Fasi di sviluppo}
\begin{tabularx}{\textwidth}{C{1.2cm} L{4.6cm} C{1.8cm} Y}
\toprule
\rowh \thd{Fase} & \thd{Titolo} & \thd{Stato} & \thd{Contenuto} \\
\midrule
1 & Layer parametrico & \OK & ingressi/neuroni/parallelismo, accumulo, bias, ReLU; test 32$\times$4/P=8. \\
\rowa 2 & Parameter sweep & \OK & configurazioni multiple incl. non-multiple e degeneri; guard di elaborazione aggiunto. \\
3 & Architettura di memoria & \OK & \code{neuron\_memory} mono/multi-neurone, PSRAM reale testata; buffer multi-layer $\to$ Fase~5. \\
\rowa 4 & Interfaccia SPI & \OK & \code{spi\_slave}+\code{spi\_engine}, 17 opcode incl. sottosistema flash, Fmax controllata a livello di sistema completo. \\
5 & Rete multi-layer & \OK$^\dagger$ & \code{layer\_sequencer}, attivazioni configurabili, larghezza runtime; toolchain reale controllata. \\
\rowa 6 & Software host & pianificata & driver Linux ed ESP32 sullo stesso protocollo. \\
7 & Ottimizzazione & in corso & timing closure fatta (55$\to$75~MHz); page-mode PSRAM fatto (banda gather +42\%); resta block RAM per $x$/$w$. \\
\rowa 8 & Training hardware (opz.) & futura & backprop, gradienti, aggiornamento pesi. \\
9 & Sottosistema flash (F1-F7) & \OK & SPI master dedicato, copy engine flash$\leftrightarrow$PSRAM, catalogo a 16 slot con CRC32, bus SPI a 4 fili indipendente (F7), 8 opcode (\op{0x40}--\op{0x47}, cap.~\ref{ch:spi} §\ref{sec:flashspi}); sintesi reale 0 errori. \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
$\dagger$ RTL, unit test ed end-to-end su SPI simulato completi; timing closure eseguita:
75.30~MHz (P2) / 60.26~MHz (P8) al tempo della Fase~5, bit-esatta su tutta la regressione;
Fmax del sistema completo dopo Fase~9 (incl. sottosistema flash indipendente): \textbf{67.91~MHz}
(cap.~\ref{ch:impl}).\end{center}
\section{Stato dei componenti}
\begin{tabularx}{\textwidth}{Y C{4.2cm}}
\toprule
\rowh \thd{Componente} & \thd{Stato} \\
\midrule
Layer neurale parametrico & \OK{} funzionante \\
\rowa Ingressi/neuroni/parallelismo parametrici & \OK \\
Accumulo, bias, ReLU & \OK \\
\rowa Validazione 32$\times$4 / P=8 & \OK \\
RAM dedicata (interfaccia + controller + accesso INT8) & \OK{} testata su PSRAM reale \\
\rowa Interfaccia SPI (17 opcode incl. RUN\_NETWORK + flash) & \OK{} Fmax a livello di sistema completo \\
Dual SPI & futura \\
\rowa Motore multi-layer & \OK{} timing closure 75.30~MHz (P2) al tempo della Fase~5 \\
Attivazioni configurabili (ACT\_NONE/ACT\_RELU) & \OK \\
\rowa Larghezza rete runtime (un bitstream, ogni topologia) & \OK{} risparmio misurato \\
Rete a grafo Tipo \#2 (act\_buffer, graph\_engine, netasm) & \OK{} RTL + test + sintesi \\
\rowa Pinout CABGA381 (\code{.lpf} reale, 57 segnali incl. flash) & \OK{} place\&route-verified 0 errori \\
Page-mode PSRAM (G7) & \OK{} fatto (37.53 cicli/edge, banda +42\%) \\
\rowa Sottosistema flash (SPI master, copy engine, catalogo CRC32, bus indipendente F7) & \OK{} sintesi reale 0 errori, Fmax 67.91~MHz \\
Bitstream reale (\code{ecppack}, P2/P8) & \OK{} 0 errori, part LFE5U-45F-8CABGA381 \\
\rowa Driver host Linux / ESP32 & pianificato \\
Training hardware & futuro \\
\bottomrule
\end{tabularx}
\section{Principio architetturale (sintesi)}
\begin{fnspec}[Fondamento del progetto]
L'FPGA implementa la macchina neurale e possiede la propria RAM; l'host configura e usa
la macchina. Una build fissa il \emph{soffitto} (max layer, max larghezza, PARALLEL);
l'host configura la rete \emph{reale} --- numero di layer, larghezza per-layer,
attivazione per-layer, parametri addestrati --- interamente a runtime, via SPI, nella
memoria locale dell'FPGA. Un solo bitstream serve qualunque topologia fino a quel
soffitto.
\end{fnspec}
\section{Visione a lungo termine}
L'obiettivo finale è un blocco hardware riusabile integrabile in progetti futuri
diversi: la piattaforma host può cambiare (Linux, ESP32, MCU, PC) senza cambiare
l'architettura fondamentale dell'engine. L'FPGA diventa una periferica di computazione
neurale dedicata, ottimizzata per la topologia richiesta da ciascuna applicazione.
@@ -1,97 +0,0 @@
\chapter[Moduli e toolchain]{Moduli, porte e toolchain}
\label{ch:appmod}
\section{Elenco dei moduli RTL}
\begin{tabularx}{\textwidth}{L{3.4cm} C{2.0cm} Y}
\toprule
\rowh \thd{File} & \thd{Tipo} & \thd{Ruolo} \\
\midrule
\code{rtl/mac\_unit.v} & combinatorio & prodotto-accumulatore singolo \\
\rowa \code{rtl/mac8.v} & combinatorio & MAC parallelo + adder tree bilanciato \\
\code{rtl/neuron\_parallel.v} & FSM & neurone: gruppi, bias, attivazione, saturazione \\
\rowa \code{rtl/layer.v} & strutturale & N\_NEURONS neuroni in parallelo \\
\code{rtl/neuron\_memory.v} & FSM & ponte memoria/neurone, loop neuroni \\
\rowa \code{rtl/layer\_sequencer.v} & FSM & sequenza multi-layer, ping-pong \\
\code{rtl/int8\_memory\_access.v} & FSM & conversione byte $\leftrightarrow$ word \\
\rowa \code{rtl/memory\_interface.v} & FSM & handshake req/ready \\
\code{rtl/psram\_controller.v} & FSM & bus fisico PSRAM async, page mode 70/20~ns \\
\rowa \code{rtl/mem\_arbiter.v} & arbitro & 3 porte, priorità B$>$C$>$A \\
\code{rtl/spi\_slave.v} & FSM & layer fisico SPI Mode 0 + CDC \\
\rowa \code{rtl/spi\_engine.v} & FSM & opcode + banco registri \\
\code{rtl/act\_buffer.v} & block RAM & buffer di attivazione DP16KD (Tipo \#2) \\
\rowa \code{rtl/graph\_engine.v} & FSM & motore rete a grafo (Tipo \#2) \\
\code{rtl/spi\_neuron\_top.v} & top & integrazione completa \\
\rowa \code{rtl/memory\_model.v} & modello & RAM comportamentale (sim) \\
\bottomrule
\end{tabularx}
\section{Porte del top-level \texttt{spi\_neuron\_top}}
Vedere la tabella segnale-per-segnale completa nel cap.~\ref{ch:hw}. In sintesi: clock
e reset (\code{clk}, \code{rst}); SPI applicativo (\code{sclk}, \code{mosi},
\code{miso}, \code{cs\_n}); bus PSRAM (\code{psram\_a[22:0]}, \code{psram\_dq[15:0]},
\code{psram\_ce\_n/oe\_n/we\_n/lb\_n/ub\_n/zz\_n}).
\section{Toolchain}
\begin{tabularx}{\textwidth}{L{3.6cm} L{3.4cm} Y}
\toprule
\rowh \thd{Strumento} & \thd{Versione} & \thd{Uso} \\
\midrule
Yosys & 0.68+post & sintesi RTL $\to$ netlist JSON, mapping ECP5 \\
\rowa nextpnr-ecp5 & 0.11.1-19-g8dbcee5 & placement, routing, timing \\
Project Trellis & install & \code{ecppack}/\code{ecppll}/\code{ecpbram} \\
\rowa Icarus Verilog & \code{-g2012} & simulazione funzionale \\
\bottomrule
\end{tabularx}
\subsection{Parametri nextpnr principali}
\begin{lstlisting}[language=,basicstyle=\ttfamily\scriptsize]
--45k seleziona LFE5U-45F
--package CABGA381 package
--speed 8 speed grade -8
--json <netlist> netlist da Yosys
--lpf <vincoli> vincoli di pin (attualmente vuoti)
--lpf-allow-unconstrained permette I/O non vincolate (benchmark)
--freq 80 timing target 80 MHz
\end{lstlisting}
\subsection{Esempio di simulazione}
\begin{lstlisting}[language=,basicstyle=\ttfamily\scriptsize]
iverilog -g2012 -Ptb.PARALLEL=16 -o sim/parametric_256x4_p16 \
sim/parametric_tb.v rtl/mac_unit.v rtl/mac8.v \
rtl/neuron_parallel.v rtl/layer.v
vvp sim/parametric_256x4_p16
\end{lstlisting}
\section{Testbench principali}
\begin{tabularx}{\textwidth}{L{5.4cm} Y}
\toprule
\rowh \thd{Testbench} & \thd{Copertura} \\
\midrule
\code{parametric\_tb.v} & datapath 256$\times$4, casi accumulo/bias/ReLU/saturazione \\
\rowa \code{parameter\_sweep\_tb.v} & sweep configurazioni valide \\
\code{neuron\_parallel\_tb.v} & attivazioni, larghezza runtime (T7) \\
\rowa \code{neuron\_memory\_tb.v} / \code{\_multi\_tb.v} & integrazione memoria mono/multi-neurone, PSRAM reale (T5) \\
\code{psram\_controller\_tb.v} & controller PSRAM \\
\rowa \code{psram\_page\_mode\_tb.v} & burst di pagina, attraversamento pagina, chiusura su WRITE/timeout $t_{CEM}$, cambi di byte-enable (§~5.5) \\
\code{spi\_slave\_tb.v} & layer fisico SPI (4 test) \\
\rowa \code{spi\_engine\_tb.v} & opcode, registri (10+ test) \\
\code{spi\_neuron\_top\_tb.v} & end-to-end, PSRAM reale su SPI simulato \\
\rowa \code{spi\_neuron\_top\_runnetwork\_tb.v} & RUN\_NETWORK 2 layer end-to-end \\
\code{layer\_sequencer\_tb.v} & sequenza 2 layer, ping-pong, copia byte-exact \\
\bottomrule
\end{tabularx}
\vfill
\begin{center}
\begin{tikzpicture}
\node[draw=fnRule,rounded corners=3pt,inner sep=8pt,fill=fnLight,text width=15.5cm]{
\footnotesize\color{fnGrey}
Questo datasheet è generato a partire dal codice RTL, dalla documentazione e dai
benchmark presenti nella repository \texttt{github.com/manvalan/FPGA-Neural} allo stato
del \datasheetdate. I valori di Fmax, utilizzo risorse e throughput sono quelli
riportati nelle misure della repository (\texttt{.lpf} reale già assegnato e
verificato da place\&route, cap.~\ref{ch:hw}) e vanno riverificati ad ogni
variazione sostanziale dell'RTL o della chiusura del timing di Fase~7, tuttora in
corso (cap.~\ref{ch:roadmap}).};
\end{tikzpicture}
\end{center}
@@ -1,119 +0,0 @@
% ======================================================================
% FPGA-Neural -- INT8 Neural Network Engine
% Datasheet / Technical reference manual
% Repository: github.com/manvalan/FPGA-Neural
% ======================================================================
\documentclass[11pt,a4paper,openany]{report}
\newcommand{\datasheetrev}{A1}
\newcommand{\datasheetdate}{September 2026}
\input{preamble}
\begin{document}
\sloppy
% ======================================================================
% TITLE PAGE
% ======================================================================
\begin{titlepage}
\thispagestyle{empty}
\begin{tikzpicture}[remember picture,overlay]
\fill[fnDark] (current page.north west) rectangle
([yshift=-4.3cm]current page.north east);
\fill[fnTeal] ([yshift=-4.3cm]current page.north west) rectangle
([yshift=-4.55cm]current page.north east);
\node[anchor=north west,text=white,font=\Huge\bfseries]
at ([xshift=2.2cm,yshift=-1.15cm]current page.north west)
{FPGA\,--\,Neural};
\node[anchor=north west,text=fnLight,font=\large]
at ([xshift=2.25cm,yshift=-2.15cm]current page.north west)
{INT8 Neural Network Engine for FPGA};
\node[anchor=north west,text=fnLight2,font=\normalsize]
at ([xshift=2.25cm,yshift=-2.85cm]current page.north west)
{Parametric hardware accelerator -- Datasheet and reference manual};
\node[anchor=north east,text=white,font=\ttfamily\small]
at ([xshift=-2.2cm,yshift=-3.55cm]current page.north east)
{Rev.~\datasheetrev~~\textbullet~~\datasheetdate};
\end{tikzpicture}
\vspace*{5.0cm}
% --- compact block diagram on the title page ---
\begin{center}
\begin{tikzpicture}[node distance=7mm and 12mm]
\node[fnblockD,minimum width=30mm] (host) {HOST\\{\scriptsize Linux / ESP32 / MCU / PC}};
\node[fnblockT,right=18mm of host,minimum width=34mm] (fpga)
{FPGA\\{\scriptsize Neural Network Engine}};
\node[fnblock,right=18mm of fpga,minimum width=26mm] (ram)
{PSRAM\\{\scriptsize 8\,MB dedicated}};
\draw[fnbus] (host) -- node[fnlbl,above]{SPI Mode 0} (fpga);
\draw[fnbus] (fpga) -- node[fnlbl,above]{async 16-bit} (ram);
\node[below=1mm of fpga,font=\scriptsize\itshape,text=fnGrey]
{computation entirely on-chip};
\end{tikzpicture}
\end{center}
\vfill
\begin{center}
\begin{tikzpicture}
\node[draw=fnRule,rounded corners=3pt,inner sep=10pt,fill=fnLight,text width=15.5cm]{
\footnotesize
\textbf{\color{fnDark}Reference target device:} Lattice ECP5 \code{LFE5U-45F-8BG381C}
(speed grade $-8$, CABGA381, 72$\times$MULT18X18D, $\approx$44k LUT).\\[2pt]
\textbf{\color{fnDark}Baseline configuration:} INT8/INT32, \code{N\_INPUTS}=256, \code{N\_NEURONS}=4,
parametric \code{PARALLEL}, PSRAM working memory ISSI \code{IS66WVE4M16EBLL-70BLI}.\\[2pt]
\textbf{\color{fnDark}Status:} RTL verified in simulation (Icarus) and real synthesis
(Yosys + nextpnr-ecp5). Document describing the project as of \datasheetdate.
};
\end{tikzpicture}
\end{center}
\vspace{0.6cm}
{\footnotesize\color{fnGrey}\raggedright
Project author: Michele Bigi \textbullet{} MIKILAB / manvalan.\\
This datasheet documents the RTL code, documentation and benchmarks
present in the repository \texttt{github.com/manvalan/FPGA-Neural}.\par}
\end{titlepage}
% ======================================================================
% "FEATURES" PAGE (datasheet style)
% ======================================================================
\input{chapters/00-features}
% ======================================================================
% PINOUT SUMMARY (pages 2-3, pin-by-pin -- not bus ranges)
% ======================================================================
\newpage
\input{chapters/00b-pinout}
% ======================================================================
% TABLE OF CONTENTS
% ======================================================================
\newpage
\pagenumbering{roman}
{\color{fnDark}\tableofcontents}
\newpage
\pagenumbering{arabic}
% ======================================================================
% CHAPTERS
% ======================================================================
\include{chapters/01-overview}
\include{chapters/02-architettura}
\include{chapters/03-datapath}
\include{chapters/04-parametri}
\include{chapters/05-memoria}
\include{chapters/06-sequencer}
\include{chapters/06b-grafo}
\include{chapters/07-spi}
\include{chapters/07b-programmazione}
\include{chapters/08-toplevel}
\include{chapters/09-implementazione}
\include{chapters/10-hardware}
\include{chapters/11-registri}
\include{chapters/12-roadmap}
\appendix
\include{chapters/A-moduli}
\end{document}
@@ -1,120 +0,0 @@
\thispagestyle{plain}
\noindent
\begin{tikzpicture}
\node[fill=fnDark,text=white,rounded corners=2pt,inner sep=6pt,
minimum width=\textwidth,anchor=west]
{\large\bfseries FPGA-Neural --- General description and features};
\end{tikzpicture}
\vspace{6pt}
\noindent
{\small FPGA-Neural is a \textbf{parametric hardware accelerator for feed-forward
neural networks} contained entirely within the FPGA. Computation (multiplication,
accumulation, bias, activation, saturation) takes place entirely on-chip in INT8/INT32
integer arithmetic; the host system only provides configuration, weights, input data
and control through a simple SPI interface, without ever being part of the
computational datapath. A single bitstream serves any topology up to the build
maximum.}
\vspace{8pt}
\begin{multicols}{2}
{\color{fnDark}\large\bfseries Features}\\[2pt]
{\footnotesize
\begin{itemize}[leftmargin=1.1em]
\item \textbf{INT8 $\times$ INT8 $\to$ INT16 $\to$ INT32} datapath, 32-bit accumulation
with sign extension.
\item \textbf{Balanced binary adder tree} ($O(\log_2 \text{PARALLEL})$) instead of
linear reduction.
\item Configurable parallel MAC: \code{PARALLEL} simultaneous hardware MACs per neuron,
mapped onto \code{MULT18X18D} DSPs.
\item Fully \textbf{parametric} architecture: \code{N\_INPUTS}, \code{N\_NEURONS},
\code{PARALLEL}, \code{DATA\_WIDTH}, \code{ACC\_WIDTH}, \code{N\_LAYERS}.
\item \textbf{Runtime network width}: per-layer \code{n\_inputs\_real}/\code{n\_neurons\_real},
a single bitstream for every topology up to the maximum.
\item Configurable activations: \code{ACT\_RELU} (default) and \code{ACT\_NONE} (linear
with bilateral saturation), with INT8 saturation.
\item \textbf{Two network types}: classic multi-layer dense (\code{layer\_sequencer},
ping-pong buffers) and \textbf{arbitrary sparse graph} (\code{graph\_engine} +
activation buffer in \code{DP16KD} block RAM), selectable at runtime.
\item \textbf{Dedicated memory} subsystem: byte$\leftrightarrow$word interface,
asynchronous parallel PSRAM controller with \textbf{page mode} (70~ns random
access, 20~ns page burst), 8~MB addressable (23~bit).
\item \textbf{SPI Mode 0} MSB-first host interface, \code{SET\_NET\_TYPE}+dispatch, \code{STATUS.done}
sticky/clear-on-read, runtime \code{READ\_CONFIG}.
\item \textbf{Flash subsystem} for boot/persistence: FPGA-exclusive access to a
\code{W25Q128JV} SPI NOR (16~MB) via a dedicated SPI master, a
flash$\leftrightarrow$PSRAM copy engine, and a 16-slot catalog with CRC32,
8 host opcodes.
\item Verified in \textbf{simulation} (Icarus Verilog) and \textbf{real synthesis}
(Yosys + nextpnr-ecp5 + ecppack).
\end{itemize}}
\columnbreak
{\color{fnDark}\large\bfseries Applications}\\[2pt]
{\footnotesize
\begin{itemize}[leftmargin=1.1em]
\item Deterministic low-latency inference as a peripheral of a
Linux SoC, Raspberry-Pi-like board, ESP32, microcontrollers.
\item Reusable hardware block integrable into heterogeneous projects
(a platform, not a single network).
\item Edge AI on compact dense INT8-quantized networks.
\item Off-loading the neural workload from the host CPU to dedicated
hardware with predictable throughput.
\end{itemize}}
\vspace{4pt}
{\color{fnDark}\large\bfseries Target \& toolchain}\\[2pt]
{\footnotesize
\begin{itemize}[leftmargin=1.1em]
\item FPGA: Lattice ECP5 \code{LFE5U-45F-8BG381C} ($-8$, CABGA381).
\item Synthesis: Yosys; place\&route: nextpnr-ecp5; bitstream: Project~Trellis
(\code{ecppack}).
\item Simulation: Icarus Verilog (\code{-g2012}).
\item PSRAM: ISSI \code{IS66WVE4M16EBLL-70BLI} (64\,Mb, 4M$\times$16).
\end{itemize}}
\end{multicols}
\vspace{2pt}
% --- key parameter table ---
\noindent
{\small\color{fnDark}\bfseries Key parameters (characterized baseline configuration)}
\vspace{2pt}
\noindent
\begin{tabularx}{\textwidth}{L{3.2cm}L{3.6cm}Y}
\toprule
\rowh \thd{Quantity} & \thd{Value} & \thd{Notes} \\
\midrule
Data precision & INT8 (signed) & \code{DATA\_WIDTH}=8 \\
\rowa Accumulator & INT32 (signed) & \code{ACC\_WIDTH}=32 \\
Inputs / neurons & 256 / 4 & datapath benchmark baseline \\
\rowa Simultaneous MACs & $2\ldots64$ & $=$\code{PARALLEL}$\times$\code{N\_NEURONS} \\
Activations & ReLU, linear & \code{ACT\_RELU} / \code{ACT\_NONE} \\
\rowa Fmax (P=2, datapath) & 87.88~MHz & isolated datapath benchmark \\
Fmax (P=2, integrated system) & 67.91~MHz & full system incl. flash subsystem, real place\&route \\
MAC throughput (P=16) & $\approx$3.34~G\,MAC/s & theoretical, datapath only \\
\rowa Working memory & 8~MB PSRAM & 16-bit parallel bus, 70~ns / 20~ns page mode \\
Address space & 23~bit (byte) & \code{ADDR\_WIDTH}=23 \\
\bottomrule
\end{tabularx}
\vspace{8pt}
\noindent
{\small\color{fnDark}\bfseries System block diagram}
\begin{center}
\begin{tikzpicture}[node distance=6mm and 10mm,font=\footnotesize]
\node[fnblockD,minimum width=26mm,minimum height=13mm] (host){HOST\\{\scriptsize configures / trains / controls}};
\node[fnblockT,right=16mm of host,minimum width=52mm,minimum height=22mm] (eng){};
\node[anchor=north,font=\footnotesize\bfseries,text=fnDark] at (eng.north){FPGA -- Neural Network Engine};
\node[fnreg,fill=white] (spi) at ([yshift=-2mm]eng.center){\code{spi\_slave} + \code{spi\_engine}};
\node[fnreg,fill=white,below=2.5mm of spi] (arb){\code{mem\_arbiter} + \code{layer\_sequencer}};
\node[fnreg,fill=white,above=2.5mm of spi] (core){\code{neuron\_memory} $\to$ \code{neuron\_parallel} $\to$ \code{mac8}};
\node[fnblock,right=16mm of eng,minimum width=24mm,minimum height=13mm] (ram){PSRAM 8\,MB\\{\scriptsize \code{psram\_controller}}};
\draw[fnbus] (host) -- node[fnlbl,above]{SPI} (eng.west|-host);
\draw[fnbus] (eng.east|-ram) -- node[fnlbl,above]{16-bit async} (ram);
\end{tikzpicture}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
The neural datapath is entirely inside the FPGA; the host does not take part in the
individual MAC operations.\end{center}
@@ -1,103 +0,0 @@
\thispagestyle{plain}
\noindent
\begin{tikzpicture}
\node[fill=fnDark,text=white,rounded corners=2pt,inner sep=6pt,
minimum width=\textwidth,anchor=west]
{\large\bfseries Pinout summary --- pin-by-pin connection};
\end{tikzpicture}
\vspace{6pt}
\noindent
{\footnotesize
Quick-reference table: the \textbf{57 real signals} of the top-level
\code{spi\_neuron\_top}, each with its own individual \code{CABGA381} ball
(\textbf{not} a bus range) --- real data from Project~Trellis's device
database (\code{iodb.json}), \textbf{verified by a complete
\code{nextpnr-ecp5} place\&route run at 0 errors} (not a planned pinout).
Full description, per-bank placement rationale and the pin-by-pin
connection to the ISSI PSRAM: ch.~\ref{ch:hw}.
}
\vspace{4pt}
\noindent
\renewcommand{\arraystretch}{1.08}
\begin{tabularx}{\textwidth}{L{2.7cm} C{1.0cm} C{1.0cm} C{0.9cm} Y}
\toprule
\rowh \thd{Signal} & \thd{Ball} & \thd{Bank} & \thd{Dir} & \thd{Corresponding pin / function} \\
\midrule
\multicolumn{5}{l}{\textit{\color{fnDark}Clock and reset}}\\
\code{clk} & H5 & 7 & IN & System clock, pad \code{GR\_PCLK7\_0} (dedicated global clock). \\
\rowa \code{rst} & B4 & 7 & IN & Global synchronous reset, active high. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Application SPI (host $\leftrightarrow$ FPGA, Mode~0)}}\\
\code{sclk} & B5 & 7 & IN & SPI clock (CPOL=0, CPHA=0). \\
\rowa \code{mosi} & C5 & 7 & IN & Master-Out Slave-In. \\
\code{miso} & A3 & 7 & OUT & Master-In Slave-Out. \\
\rowa \code{cs\_n} & B3 & 7 & IN & Chip-select, active low. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Host attention (active-low, level)}}\\
\code{data\_ready\_n} & C3 & 7 & OUT & Low while a result is waiting to be read. \\
\rowa \code{irq\_n} & C4 & 7 & OUT & Low while the graph engine's load-time guard has tripped. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Flash subsystem --- SPI toward W25Q128JV (boot/persistence)}}\\
\code{flash\_sclk} & E3 & 7 & OUT & SPI clock toward the flash --- ordinary GPIO, independent (Phase F7, ch.~\ref{ch:hw}). \\
\rowa \code{flash\_mosi} & D3 & 7 & OUT & Master-Out Slave-In toward the onboard flash. \\
\code{flash\_miso} & D5 & 7 & IN & Master-In Slave-Out from the flash. \\
\rowa \code{flash\_cs\_n} & E4 & 7 & OUT & Flash chip-select, active low. \\
\multicolumn{5}{l}{\textit{\color{fnDark}PSRAM address bus \code{psram\_a[21:0]} --- 22 individual balls (bank 2)}}\\
\code{psram\_a[0]} & E16 & 2 & OUT & PSRAM A0 \\
\rowa \code{psram\_a[1]} & F16 & 2 & OUT & PSRAM A1 \\
\code{psram\_a[2]} & D18 & 2 & OUT & PSRAM A2 \\
\rowa \code{psram\_a[3]} & E17 & 2 & OUT & PSRAM A3 \\
\code{psram\_a[4]} & E18 & 2 & OUT & PSRAM A4 \\
\rowa \code{psram\_a[5]} & F18 & 2 & OUT & PSRAM A5 \\
\code{psram\_a[6]} & F17 & 2 & OUT & PSRAM A6 \\
\rowa \code{psram\_a[7]} & G16 & 2 & OUT & PSRAM A7 \\
\code{psram\_a[8]} & G18 & 2 & OUT & PSRAM A8 \\
\rowa \code{psram\_a[9]} & H16 & 2 & OUT & PSRAM A9 \\
\code{psram\_a[10]} & H17 & 2 & OUT & PSRAM A10 \\
\rowa \code{psram\_a[11]} & H18 & 2 & OUT & PSRAM A11 \\
\code{psram\_a[12]} & J16 & 2 & OUT & PSRAM A12 \\
\rowa \code{psram\_a[13]} & J17 & 2 & OUT & PSRAM A13 \\
\code{psram\_a[14]} & C20 & 2 & OUT & PSRAM A14 \\
\rowa \code{psram\_a[15]} & D19 & 2 & OUT & PSRAM A15 \\
\code{psram\_a[16]} & E19 & 2 & OUT & PSRAM A16 \\
\rowa \code{psram\_a[17]} & E20 & 2 & OUT & PSRAM A17 \\
\code{psram\_a[18]} & F19 & 2 & OUT & PSRAM A18 \\
\rowa \code{psram\_a[19]} & F20 & 2 & OUT & PSRAM A19 \\
\code{psram\_a[20]} & G20 & 2 & OUT & PSRAM A20 \\
\rowa \code{psram\_a[21]} & H20 & 2 & OUT & PSRAM A21 \\
\code{psram\_a[22]} & P18 & 3 & OUT & Always 0 (byte$\to$word shift): NC on the board. \\
\multicolumn{5}{l}{\textit{\color{fnDark}PSRAM data bus \code{psram\_dq[15:0]} --- 16 individual balls (banks 2 and 3)}}\\
\rowa \code{psram\_dq[0]} & K18 & 2 & IO & PSRAM DQ0 \\
\code{psram\_dq[1]} & C18 & 2 & IO & PSRAM DQ1 (dual-function ball, used as ordinary GPIO). \\
\rowa \code{psram\_dq[2]} & D17 & 2 & IO & PSRAM DQ2 \\
\code{psram\_dq[3]} & D20 & 2 & IO & PSRAM DQ3 \\
\rowa \code{psram\_dq[4]} & G19 & 2 & IO & PSRAM DQ4 \\
\code{psram\_dq[5]} & J18 & 2 & IO & PSRAM DQ5 \\
\rowa \code{psram\_dq[6]} & J19 & 2 & IO & PSRAM DQ6 \\
\code{psram\_dq[7]} & J20 & 2 & IO & PSRAM DQ7 \\
\rowa \code{psram\_dq[8]} & K19 & 2 & IO & PSRAM DQ8 \\
\code{psram\_dq[9]} & K20 & 2 & IO & PSRAM DQ9 \\
\rowa \code{psram\_dq[10]} & L17 & 3 & IO & PSRAM DQ10 \\
\code{psram\_dq[11]} & M18 & 3 & IO & PSRAM DQ11 \\
\rowa \code{psram\_dq[12]} & M17 & 3 & IO & PSRAM DQ12 \\
\code{psram\_dq[13]} & N16 & 3 & IO & PSRAM DQ13 \\
\rowa \code{psram\_dq[14]} & N18 & 3 & IO & PSRAM DQ14 \\
\code{psram\_dq[15]} & P17 & 3 & IO & PSRAM DQ15 \\
\multicolumn{5}{l}{\textit{\color{fnDark}PSRAM control}}\\
\rowa \code{psram\_ce\_n} & N17 & 3 & OUT & PSRAM CE\# --- chip enable, active low. \\
\code{psram\_oe\_n} & R16 & 3 & OUT & PSRAM OE\# --- output enable (read). \\
\rowa \code{psram\_we\_n} & R17 & 3 & OUT & PSRAM WE\# --- write enable. \\
\code{psram\_lb\_n} & T16 & 3 & OUT & PSRAM LB\# --- lower-byte enable (DQ[7:0]). \\
\rowa \code{psram\_ub\_n} & N19 & 3 & OUT & PSRAM UB\# --- upper-byte enable (DQ[15:8]). \\
\code{psram\_zz\_n} & N20 & 3 & OUT & PSRAM ZZ\# --- sleep/snooze (high during normal operation). \\
\bottomrule
\end{tabularx}
\renewcommand{\arraystretch}{1.25}
\vspace{4pt}
\noindent
{\footnotesize\color{fnGrey}
Standard I/O: LVCMOS33 on all 57 signals. Boot config-SPI and JTAG balls (fixed-function
dedicated pins, no RTL port) do not appear in this table --- see ch.~\ref{ch:hw}
§``Configuration and programming''. Source: \code{synth/ecp5/spi\_neuron\_top.lpf},
generated by \code{tools/pinout/gen\_lpf.py} against Project~Trellis's
\code{iodb.json}.\par}
@@ -1,93 +0,0 @@
\chapter{System overview}
\label{ch:overview}
\section{Project goal}
FPGA-Neural implements a \textbf{reusable Neural Network Engine in FPGA hardware}.
The whole is made of three elements: the FPGA, which is the actual accelerator; a
dedicated RAM physically associated with the FPGA and not shared with the host; and a
host interface independent of the operating system, initially SPI (with possible
future extension to Dual~SPI).
The founding principle is the separation between who \emph{executes} the computation
and who \emph{uses} it: the neural network computation happens entirely inside the
FPGA, while the host system only provides configuration, network parameters, input
data, control and result readback. The host is not part of the computational datapath.
Possible host systems include Linux SoCs, Raspberry~Pi-like systems, ESP32,
microcontrollers and development PCs: the same engine architecture must be usable in
completely different systems.
\begin{center}
\begin{tikzpicture}[font=\footnotesize,node distance=8mm]
\node[fnblockD,minimum width=42mm,minimum height=20mm] (host){\textbf{HOST}\\[2pt]
{\scriptsize Configuration}\\{\scriptsize Training}\\{\scriptsize Control}};
\node[fnblockT,below=14mm of host,minimum width=42mm,minimum height=20mm] (fpga)
{\textbf{FPGA}\\[2pt]{\scriptsize Neural Network Engine}\\{\scriptsize Compute / Control}};
\node[fnblock,below=14mm of fpga,minimum width=42mm,minimum height=13mm] (ram)
{\textbf{Dedicated RAM}\\{\scriptsize weights / bias / buffers}};
\draw[fnbus] (host) -- node[fnlbl,right]{SPI / Dual SPI} (fpga);
\draw[fnbus] (fpga) -- node[fnlbl,right]{parallel bus} (ram);
\end{tikzpicture}
\end{center}
\section{Hardware configuration versus network configuration}
The project draws a precise distinction between the accelerator's \textbf{hardware
architecture} and the \textbf{neural network parameters}.
The physical architecture of the engine is defined at FPGA synthesis and
implementation time. Typical hardware parameters are \code{N\_INPUTS},
\code{N\_NEURONS}, \code{N\_LAYERS}, \code{PARALLEL}, \code{DATA\_WIDTH},
\code{ACC\_WIDTH}: they are Verilog parameters resolved at synthesis and they
determine the datapath contained in the bitstream. The network parameters --- weights,
bias, activation and quantization parameters, specific constants --- are instead loaded
at runtime through the host interface and stored in the RAM associated with the FPGA.
\begin{fnnote}[Central architectural principle]
A build fixes the \emph{ceiling} of the machine (maximum number of layers, maximum
width, \code{PARALLEL}); the host configures the \emph{actual} network --- number of
layers, per-layer input/output width, per-layer activation and trained parameters ---
entirely at runtime, over SPI, into the FPGA's local memory. A single bitstream serves
any topology up to that ceiling.
\end{fnnote}
\section{Boot and initialization}
The FPGA is configured at power-on through the usual configuration mechanism (bitstream
loading from SPI flash). The bitstream defines the hardware architecture of the engine;
the host does not dynamically build the datapath during normal operation, but rather
configures the network data on which the already existing datapath operates.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=4.5mm,start chain=going below,
every node/.style={on chain}]
\node[fnblockA,minimum width=60mm](p){Power-on};
\node[fnblock,minimum width=60mm]{FPGA configuration (bitstream from flash)};
\node[fnblockT,minimum width=60mm]{Neural Network Engine available};
\node[fnblock,minimum width=60mm]{Host initialization (SPI)};
\node[fnblock,minimum width=60mm]{Loading network parameters / weights / bias};
\node[fnblockD,minimum width=60mm]{Engine ready};
\begin{scope}[every path/.style={fnarrow}]
\foreach \a/\b in {1/2,2/3,3/4,4/5,5/6}{}
\end{scope}
\foreach \i [count=\j from 2] in {1,...,5}{
\draw[fnarrow] (chain-\i) -- (chain-\j);}
\end{tikzpicture}
\end{center}
\section{Training and inference}
Training and inference are conceptually separate. The first implementation does not
require the FPGA to perform training: weights can be computed externally
(PC/Linux/other host) and transferred over SPI into the FPGA's RAM, which then performs
inference. This drastically reduces the complexity of the initial hardware, without
precluding a future implementation of assisted or fully hardware training (roadmap
Phase~8, ch.~\ref{ch:roadmap}). During inference the host only provides the input data
and retrieves the result, obtaining deterministic computation, reduced host load,
hardware parallelism, predictable latency and independence from the host CPU
architecture.
\section{Design philosophy and reuse}
The project should be understood as a \emph{reusable FPGA neural acceleration platform}
rather than a single network. The application determines input size, topology, number
of layers and neurons, parallelism, numeric precision, activation functions, memory and
performance requirements; the hardware generation process produces the corresponding
FPGA implementation. The same HDL architecture remains conceptually unchanged while the
synthesis parameters generate implementations appropriate to the different application
targets.
@@ -1,80 +0,0 @@
\chapter[RTL architecture]{RTL architecture and module hierarchy}
\label{ch:arch}
\section{Hierarchical organization}
The design is organized in layers, from the elementary multiply-accumulator up to the
integrated top-level with SPI interface and PSRAM. Each layer encapsulates the previous
one and abstracts away its details: the validated datapath (\code{mac\_unit},
\code{mac8}, \code{neuron\_parallel}) is never modified by the higher orchestration
layers.
\begin{center}
\begin{tikzpicture}[font=\footnotesize,every node/.style={fnblock,minimum width=40mm},
level distance=13mm,sibling distance=0mm]
\node[fnblockD,minimum width=62mm](top){\code{spi\_neuron\_top} \\ {\scriptsize integrated top-level}};
\node[fnblockT,minimum width=62mm,below=8mm of top](arb){\code{mem\_arbiter} \;/\; \code{layer\_sequencer} \\ {\scriptsize 3-port arbitration + layer sequencing}};
\node[fnblock,minimum width=62mm,below=8mm of arb](nm){\code{neuron\_memory} \\ {\scriptsize memory $\leftrightarrow$ neuron bridge, neuron loop}};
\node[fnblock,minimum width=62mm,below=8mm of nm](np){\code{neuron\_parallel} \\ {\scriptsize neuron FSM: groups, bias, activation, saturation}};
\node[fnblockT,minimum width=62mm,below=8mm of np](m8){\code{mac8} \\ {\scriptsize \code{PARALLEL} MACs + balanced adder tree}};
\node[fnblock,minimum width=62mm,below=8mm of m8](mu){\code{mac\_unit} \\ {\scriptsize $x\cdot w$ + sign extension + accumulate}};
\foreach \a/\b in {top/arb,arb/nm,nm/np,np/m8,m8/mu}
\draw[fnarrow] (\a) -- (\b);
% memory branches on the right
\node[fnblockA,minimum width=34mm,right=14mm of nm](ma){\code{int8\_memory\_access}\\{\scriptsize byte $\leftrightarrow$ 16-bit word}};
\node[fnblockA,minimum width=34mm,below=6mm of ma](mi){\code{memory\_interface}\\{\scriptsize req/ready handshake}};
\node[fnblockA,minimum width=34mm,below=6mm of mi](pc){\code{psram\_controller}\\{\scriptsize physical PSRAM bus}};
\draw[fnarrowT] (ma)--(mi); \draw[fnarrowT] (mi)--(pc);
\draw[fnarrowT,dashed] (nm.east) -- (ma.west);
% SPI branches on the left
\node[fnblockA,minimum width=30mm,left=14mm of arb,yshift=6mm](ss){\code{spi\_slave}\\{\scriptsize Mode 0 physical layer}};
\node[fnblockA,minimum width=30mm,below=6mm of ss](se){\code{spi\_engine}\\{\scriptsize opcode FSM + registers}};
\draw[fnarrowT] (ss)--(se);
\draw[fnarrowT,dashed] (se.east) -- (arb.west);
\end{tikzpicture}
\end{center}
\section{Role of each module}
\begin{tabularx}{\textwidth}{L{3.4cm}Y}
\toprule
\rowh \thd{Module} & \thd{Function} \\
\midrule
\code{mac\_unit} & Single multiply-accumulate: $\mathrm{acc\_out}=\mathrm{acc\_in}+(x\cdot w)$, with sign extension of the product to \code{ACC\_WIDTH}. Parametric on \code{DATA\_WIDTH}/\code{ACC\_WIDTH}. \\
\rowa \code{mac8} & \code{PARALLEL} instances of \code{mac\_unit} whose products are summed by a \emph{balanced binary adder tree} of depth $\log_2(\text{PARALLEL})$; the result is added to the input accumulator. \\
\code{neuron\_parallel} & FSM of a single neuron: processes \code{N\_INPUTS} inputs in groups of \code{PARALLEL}, accumulates across groups, adds the bias, applies the activation and saturates to INT8. Includes the processing guard on \code{N\_INPUTS \% PARALLEL} and the runtime width \code{n\_inputs\_real}. \\
\rowa \code{layer} & Instantiates \code{N\_NEURONS} neurons \emph{in parallel} on the same input vector; \code{busy}=OR, \code{done}=AND of the neurons. A purely data-combinational path used in the datapath benchmarks. \\
\code{neuron\_memory} & Integrates computation with memory: reads $X$ (shared) once, then for each neuron re-reads $W$ and bias from RAM and reuses a single \code{neuron\_parallel} instance (memory-bound, one neuron at a time). Output \code{y\_bus} packed neuron-major. \\
\rowa \code{layer\_sequencer} & Chains up to \code{N\_LAYERS} executions of \code{neuron\_memory} by reading a descriptor table written by the host and alternating the ping-pong buffers in RAM (Phase~5). \\
\code{act\_buffer} & Global activation buffer in \code{DP16KD} block RAM, indexed by signal id (Type \#2). \\
\rowa \code{graph\_engine} & Graph-network engine (Type \#2): gather from \code{act\_buffer}, reuses \code{neuron\_parallel}, writes outputs by id (ch.~\ref{ch:grafo}). \\
\code{int8\_memory\_access} & Converts the byte/INT8 interface (byte address) into the 16-bit word interface, selecting the low/high byte via \code{lb\_n}/\code{ub\_n} and \code{addr>>1}. \\
\rowa \code{memory\_interface} & 2-state handshake FSM (IDLE/WAIT) that serializes the single transaction toward the controller. \\
\code{psram\_controller} & Asynchronous parallel PSRAM bus controller with read \textbf{page mode}: 70~ns random access (\code{tAA}), 20~ns same-page bursts (\code{tAPA}) with CE\#/OE\# held asserted; enables page mode on the chip at boot via the configuration register (ch.~\ref{ch:mem}, \S~5.5). Drives \code{ce\_n/oe\_n/we\_n/lb\_n/ub\_n/zz\_n} and the tri-state data bus. \\
\rowa \code{mem\_arbiter} & Fixed-priority arbiter (B$>$C$>$A) among three byte-level masters: \code{spi\_engine} (A), \code{neuron\_memory} (B), \code{layer\_sequencer} (C). \\
\code{spi\_slave} & SPI Mode 0 physical layer, MSB-first, 3-stage CDC synchronizer on SCLK/MOSI/CS\_N, shift register and CS framing. \\
\rowa \code{spi\_engine} & Protocol/opcode FSM and register bank (\code{x\_base}, \code{w\_base}, \code{bias\_addr}, ping-pong base, activation, runtime widths\ldots), with sticky/clear-on-read \code{STATUS.done}. \\
\code{spi\_neuron\_top} & Top-level: connects SPI, arbiter, sequencer, \code{neuron\_memory} and the PSRAM chain; multiplexes control of \code{neuron\_memory} between the sequencer and the direct single-layer path. \\
\bottomrule
\end{tabularx}
\vspace{6pt}
\begin{fnnote}[Simulation models]
\code{psram\_model.v} (in \code{sim/}) and \code{memory\_model.v} are behavioral memory
models used in the testbenches; they are not part of the synthesizable design but they
reproduce the real latency for end-to-end verification.
\end{fnnote}
\section{Two execution paths}
The top-level exposes two mutually exclusive modes toward the same \code{neuron\_memory}
compute engine:
\begin{itemize}
\item \textbf{Single-layer / manual path}: the host sets the bases with
\op{SET\_BASE}, starts with \op{START} and reads with \op{READ\_OUTPUT}.
\code{spi\_engine} drives \code{neuron\_memory} directly.
\item \textbf{Multi-layer path}: the host writes the descriptor table and starts with
\op{RUN\_NETWORK}; \code{layer\_sequencer} takes over control of \code{neuron\_memory}
(while \code{seq\_busy} is high) and chains the layers.
\end{itemize}
The top-level multiplexer switches the control lines of \code{neuron\_memory} based on
\code{seq\_busy}, returning the engine to the direct path at the end of the sequence.
@@ -1,166 +0,0 @@
\chapter{Compute datapath}
\label{ch:datapath}
\section{INT8/INT32 arithmetic chain}
The elementary datapath implements the typical sequence of a quantized neuron:
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=3mm,start chain=going right,
every node/.style={fnblock,minimum width=15mm,minimum height=8mm,on chain}]
\node[fnblockT]{INT8\\$\times$\,INT8};
\node{INT16\\product};
\node{sign-ext\\INT32};
\node[fnblockD]{accumulate\\INT32};
\node{$+$ bias};
\node[fnblockA]{activation};
\node[fnblockT]{sat. INT8};
\foreach \i [count=\j from 2] in {1,...,6}
\draw[fnarrow] (chain-\i) -- (chain-\j);
\end{tikzpicture}
\end{center}
Each INT8$\times$INT8 product fits in 16~bits; it is sign-extended to 32~bits before
accumulation, so the accumulator does not overflow on long vectors. Bias and activation
operate at 32~bits; only the final output is saturated to INT8.
\section{\texttt{mac\_unit} --- multiply-accumulator}
The \code{mac\_unit} module is purely combinational and parametric on \code{DATA\_WIDTH}
and \code{ACC\_WIDTH}. It computes:
\[
\mathrm{acc\_out} = \mathrm{acc\_in} + \mathrm{signext}_{ACC}(x \cdot w)
\]
The product has width $2\times$\code{DATA\_WIDTH} and is sign-extended by replicating
the most significant bit. On ECP5 the multiplication maps onto a \code{MULT18X18D} DSP
block.
\begin{lstlisting}[caption={\texttt{rtl/mac\_unit.v} --- arithmetic core},label={lst:macunit}]
localparam PROD_WIDTH = 2 * DATA_WIDTH;
wire signed [PROD_WIDTH-1:0] product = x * w;
wire signed [ACC_WIDTH-1:0] product_ext =
{{(ACC_WIDTH-PROD_WIDTH){product[PROD_WIDTH-1]}}, product};
assign acc_out = acc_in + product_ext;
\end{lstlisting}
\section{\texttt{mac8} --- parallel MAC and balanced adder tree}
\code{mac8} instantiates \code{PARALLEL} \code{mac\_unit} units that generate
\code{PARALLEL} independent products, then sums them with a \emph{balanced binary
adder tree}. Compared to the linear reduction
$((((p_0{+}p_1){+}p_2){+}p_3){+}\dots)$, of depth $O(\text{PARALLEL})$, the tree has
depth $O(\log_2 \text{PARALLEL})$, drastically reducing the combinational path.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,level distance=11mm,
every node/.style={fnreg,minimum width=8mm},
level 1/.style={sibling distance=30mm},
level 2/.style={sibling distance=15mm},
level 3/.style={sibling distance=8mm},
edge from parent/.style={fnarrowT,draw}]
\node[fnblockD]{sum}
child {node[fnblockT]{$+$}
child {node[fnblockT]{$+$}
child {node{$p_0$}} child {node{$p_1$}}}
child {node[fnblockT]{$+$}
child {node{$p_2$}} child {node{$p_3$}}}}
child {node[fnblockT]{$+$}
child {node[fnblockT]{$+$}
child {node{$p_4$}} child {node{$p_5$}}}
child {node[fnblockT]{$+$}
child {node{$p_6$}} child {node{$p_7$}}}};
\end{tikzpicture}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
Example with PARALLEL=8: 3 levels. PARALLEL=16 $\to$ 4 levels; PARALLEL=32 $\to$ 5
levels.\end{center}
\begin{fnnote}[PARALLEL as a power of two]
The tree is designed for \code{PARALLEL} as a power of two (8, 16, 32\ldots). This is
also the value used in all project configurations.
\end{fnnote}
\section{\texttt{neuron\_parallel} --- neuron FSM}
\code{neuron\_parallel} processes \code{N\_INPUTS} inputs in groups of \code{PARALLEL},
maintaining the accumulator from one group to the next. At the end it adds the bias,
applies the activation and saturates to INT8. The number of groups is
$\text{GROUPS}=\text{N\_INPUTS}/\text{PARALLEL}$.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=4mm,start chain=going below,
every node/.style={on chain,fnblock,minimum width=46mm}]
\node[fnblockA]{\code{start}};
\node{group 0 $\to$ accumulate};
\node{group 1 $\to$ accumulate};
\node[draw=none,fill=none]{\vdots};
\node{group GROUPS$-$1 $\to$ accumulate};
\node{$+$ bias};
\node[fnblockA]{activation (ACT\_RELU / ACT\_NONE)};
\node[fnblockT]{INT8 saturation};
\node[fnblockD]{\code{done}, \code{y}};
\foreach \i [count=\j from 2] in {1,...,8}
\draw[fnarrow] (chain-\i) -- (chain-\j);
\end{tikzpicture}
\end{center}
\subsection{Parameter guard (elaboration-time)}
If \code{PARALLEL} does not exactly divide \code{N\_INPUTS} two failures occur, both
confirmed empirically in \code{sim/parameter\_sweep\_tb.v}:
\begin{itemize}
\item integer division truncates \code{GROUPS} and the excess inputs are never read
$\to$ \textbf{wrong} result, with no error and no warning;
\item if \code{PARALLEL > N\_INPUTS}, \code{GROUPS=0} and the terminal condition is
never satisfied $\to$ the neuron \textbf{hangs} (busy high, done never asserted).
\end{itemize}
The solution does not modify the validated datapath: a \code{generate} block
instantiates a deliberately undefined module when
$\text{N\_INPUTS} \bmod \text{PARALLEL}\neq0$, forcing an error at \emph{elaboration}
both in simulation and in synthesis. For valid configurations the branch is never
elaborated.
\begin{lstlisting}[caption={\texttt{rtl/neuron\_parallel.v} --- parameter guard}]
generate
if (N_INPUTS == 0 || N_INPUTS % PARALLEL != 0) begin : PARAMETER_ERROR
neuron_parallel_requires_N_INPUTS_multiple_of_PARALLEL
invalid_parameter_combination();
end
endgenerate
\end{lstlisting}
\begin{fnnote}[Edge case \texttt{N\_INPUTS=0} (fixed 2026-09-04)]
The original condition (\code{N\_INPUTS \% PARALLEL != 0}) does not catch
\code{N\_INPUTS=0}, since $0 \bmod \text{PARALLEL}=0$ for any \code{PARALLEL}: the module
elaborated successfully (both in simulation and in real Yosys synthesis) while leaving
\code{x\_bus}/\code{w\_bus} undriven and \code{start} silently ineffective. Found during
the re-certification campaign (\code{docs/validation/bugs.md}, BUG-002) and fixed by
extending the guard as above --- \code{N\_INPUTS=0} now fails elaboration exactly like the
other degenerate cases.
\end{fnnote}
\section{Activation functions}
\code{neuron\_parallel} accepts a 2-bit \code{activation} port. The default is
\code{ACT\_RELU}, the only behavior that existed before the port was introduced, so
every pre-existing caller remains unchanged.
\begin{tabularx}{\textwidth}{L{2.6cm} C{1.4cm} Y}
\toprule
\rowh \thd{Encoding} & \thd{Value} & \thd{Behavior} \\
\midrule
\code{ACT\_NONE} & \code{2'd0} & Linear: no clamp to zero, bilateral saturation to the INT8 range $[-128,+127]$. \\
\rowa \code{ACT\_RELU} & \code{2'd1} & $\max(0,x)$, then positive saturation to $+127$ (default; also the fallback for reserved encodings). \\
\bottomrule
\end{tabularx}
\section{INT8 saturation}
After bias and activation, the 32-bit accumulator is reduced to INT8:
\[
y=\begin{cases}
+127 & \text{if } \mathrm{final\_acc} > 127\\
-128 & \text{if } \mathrm{final\_acc} < -128 \ \text{(ACT\_NONE only)}\\
0 & \text{if } \mathrm{final\_acc}\le 0 \ \text{(ACT\_RELU only)}\\
\mathrm{final\_acc}[7:0] & \text{otherwise}
\end{cases}
\]
\section{\texttt{layer} --- neurons in parallel}
\code{layer} instantiates \code{N\_NEURONS} neurons that share the input vector
\code{x\_bus} but have distinct weights and bias; \code{busy} is the OR and \code{done}
the AND of the neurons' signals. It is the module used in the datapath benchmarks
(ch.~\ref{ch:impl}), where all neurons work simultaneously. The addressing convention
is neuron-major: the weights of neuron $n$ occupy
\code{weights\_bus[n*N\_INPUTS*DATA\_WIDTH +: N\_INPUTS*DATA\_WIDTH]}.
@@ -1,88 +0,0 @@
\chapter{Parameters and configurability}
\label{ch:param}
\section{Build parameters (synthesis-time)}
The hardware architecture is fixed at synthesis through the following Verilog
parameters. They determine the datapath contained in the bitstream and its capacity
\emph{ceiling}.
\begin{tabularx}{\textwidth}{L{3.0cm} C{1.8cm} Y}
\toprule
\rowh \thd{Parameter} & \thd{Default} & \thd{Meaning} \\
\midrule
\code{DATA\_WIDTH} & 8 & Data width (INT8). \\
\rowa \code{ACC\_WIDTH} & 32 & Accumulator width (INT32). \\
\code{N\_INPUTS} & 32 / 256 & Maximum number of inputs per neuron (benchmark baseline: 256). \\
\rowa \code{N\_NEURONS} & 1 / 4 & Maximum number of neurons per layer. \\
\code{PARALLEL} & 8 & Simultaneous hardware MACs per neuron; must divide \code{N\_INPUTS} and should be a power of two. \\
\rowa \code{N\_LAYERS} & 4 & Maximum number of layers chainable by \code{layer\_sequencer}. \\
\code{ADDR\_WIDTH} & 23 & Byte-address width (8~MB). \\
\rowa \code{MEM\_DATA\_WIDTH} & 16 & Width of the physical PSRAM data bus. \\
\code{CLK\_FREQ\_MHZ} & 80 & Frequency used in the PSRAM timing formulas (must be aligned to the real oscillator). \\
\bottomrule
\end{tabularx}
\begin{fnwarn}[\texttt{N\_INPUTS} \% \texttt{PARALLEL} constraint]
\code{PARALLEL} must divide \code{N\_INPUTS} exactly, otherwise the elaboration guard
fires (§\ref{ch:datapath}). The same constraint applies at runtime to
\code{n\_inputs\_real}.
\end{fnwarn}
\section{Runtime network width}
A single bitstream serves any topology \emph{up to} the build maximum. The actual width
of each execution is a separate value, set by the host:
\begin{itemize}
\item \code{n\_inputs\_real} --- inputs actually used in this execution (must be a
multiple of \code{PARALLEL});
\item \code{n\_neurons\_real} --- neurons actually computed in this execution.
\end{itemize}
Both default to the build maximum, so any caller that leaves them unconnected processes
the full width as before the ports were introduced.
\begin{fnnote}[Real early termination]
This is not mere address bookkeeping: the two values directly bound the hardware loops
(X/W reads of \code{neuron\_memory}, MAC group count of \code{neuron\_parallel} and the
length of the ping-pong copy for \code{RUN\_NETWORK}). A narrower layer actually
\emph{computes} and \emph{copies} faster and does not require zero-padding of the RAM
for the unused tail: data beyond \code{n\_inputs\_real}/\code{n\_neurons\_real} is never
read.
\end{fnnote}
This lets a network taper within a single chained execution, for example
$256\to64\to16\to4$, with each layer declaring its own actual width in the descriptor
table (ch.~\ref{ch:seq}).
\subsection{Measured savings}
Early termination was measured end-to-end:
\begin{tabularx}{\textwidth}{L{5.5cm} C{3.0cm} Y}
\toprule
\rowh \thd{Test} & \thd{Cycles} & \thd{Comparison} \\
\midrule
\code{neuron\_parallel\_tb.v} (T7) & 3 vs 6 & reduced vs full, with ``garbage'' data in the skipped lanes (proof that they are not read). \\
\rowa \code{neuron\_memory\_tb.v} (T5) & 209 vs 788 & 8-of-32 vs full 32, through the real PSRAM stack. \\
\bottomrule
\end{tabularx}
\section{Characterized configurations}
Some combinations validated in simulation and/or synthesis:
\begin{tabularx}{\textwidth}{C{2.0cm} C{2.0cm} C{2.0cm} Y}
\toprule
\rowh \thd{N\_INPUTS} & \thd{N\_NEURONS} & \thd{PARALLEL} & \thd{Notes} \\
\midrule
32 & 4 & 8 & First functional parametric test (Phase~1). \\
\rowa 256 & 4 & 2/4/8/16 & Datapath benchmark sweep (Phase~7). \\
32 & 1..3 & 8 & Single/multi-neuron memory integration (Phase~3). \\
\rowa 4 & 4 & 2 & End-to-end 2-layer \code{RUN\_NETWORK} test over real SPI. \\
\bottomrule
\end{tabularx}
\section{Build versus runtime summary}
\begin{center}
\begin{tikzpicture}[font=\footnotesize,node distance=6mm]
\node[fnblockD,minimum width=54mm,minimum height=15mm](b){\textbf{BUILD (synthesis)}\\[2pt]
{\scriptsize N\_INPUTS, N\_NEURONS, N\_LAYERS,}\\{\scriptsize PARALLEL, DATA\_WIDTH, ACC\_WIDTH}\\{\scriptsize $\Rightarrow$ machine ceiling}};
\node[fnblockT,right=16mm of b,minimum width=54mm,minimum height=15mm](r){\textbf{RUNTIME (host, SPI)}\\[2pt]
{\scriptsize n\_inputs\_real, n\_neurons\_real,}\\{\scriptsize activation, num\_layers, weights/bias}\\{\scriptsize $\Rightarrow$ actual network}};
\draw[fnbus] (b) -- node[fnlbl,above]{$\le$} (r);
\end{tikzpicture}
\end{center}
@@ -1,187 +0,0 @@
\chapter{Memory subsystem}
\label{ch:mem}
\section{Memory chain}
The compute engine works with addresses and data at the \emph{byte} level (INT8), while
the PSRAM is a 16-bit word device. Three cascaded modules realize the conversion and
the physical access:
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=8mm]
\node[fnblockD,minimum width=30mm,minimum height=12mm](nm){byte-level master\\{\scriptsize \code{neuron\_memory} / \code{spi\_engine} / \code{layer\_sequencer}}};
\node[fnblockT,right=10mm of nm,minimum width=28mm,minimum height=12mm](ia){\code{int8\_memory\_access}\\{\scriptsize byte $\leftrightarrow$ 16-bit word}};
\node[fnblock,right=10mm of ia,minimum width=26mm,minimum height=12mm](mi){\code{memory\_interface}\\{\scriptsize IDLE/WAIT FSM}};
\node[fnblockA,below=9mm of mi,minimum width=26mm,minimum height=12mm](pc){\code{psram\_controller}\\{\scriptsize async 70\,ns physical bus}};
\node[fnblock,left=10mm of pc,minimum width=26mm,minimum height=12mm](ps){PSRAM\\{\scriptsize 8\,MB 4M$\times$16}};
\draw[fnbus] (nm)--node[fnlbl,above]{req/wr/addr}(ia);
\draw[fnbus] (ia)--node[fnlbl,above]{16-bit}(mi);
\draw[fnbus] (mi)--(pc);
\draw[fnbus] (pc)--node[fnlbl,above]{DQ/A/ctrl}(ps);
\end{tikzpicture}
\end{center}
\section{\texttt{int8\_memory\_access} --- byte/word conversion}
Converts the INT8 interface (byte address) into the word interface. The byte address is
divided by two (\code{addr>>1}) to obtain the word address; the least significant bit
selects the byte:
\begin{itemize}
\item \code{addr[0]=0} $\to$ low byte: \code{lb\_n=0}, \code{ub\_n=1}, data on DQ[7:0];
\item \code{addr[0]=1} $\to$ high byte: \code{lb\_n=1}, \code{ub\_n=0}, data on DQ[15:8].
\end{itemize}
On read it extracts the correct byte from \code{mem\_rdata}. The FSM has two states
(IDLE, WAIT) and returns \code{ready} as a one-cycle pulse.
\section{\texttt{memory\_interface} --- handshake}
Two-state FSM that serializes a single transaction: in IDLE, on the \code{req} request,
it latches \code{wr/addr/wdata/lb\_n/ub\_n} and emits a one-cycle \code{mem\_req} pulse
toward the controller; in WAIT it waits for \code{mem\_ready}, captures \code{rdata} on
read and asserts \code{ready}. It guarantees the ``one transaction at a time'' contract.
\section{\texttt{psram\_controller} --- physical bus}
Asynchronous parallel PSRAM bus controller, with support for the chip's read
\textbf{page mode} (\S~\ref{sec:pagemode}). The main state machine is:
\begin{center}
\begin{tikzpicture}[font=\scriptsize]
\node[fnstate](init) at (0,0){INIT};
\node[fnstate](idle) at (3.2,0){IDLE};
\node[fnstate](read) at (7,2.7){READ};
\node[fnstate](popen) at (11,2.7){PAGE\\OPEN};
\node[fnstate](write) at (7,-2.7){WRITE};
\node[fnstate](ww) at (11,-2.7){WRITE\\WAIT};
\draw[fnarrow] (init)--node[fnlbl,above]{INIT\_CYCLES + CR load}(idle);
\draw[fnarrow] (idle)--node[fnlbl,above,sloped]{req \& !wr}(read);
\draw[fnarrow] (idle)--node[fnlbl,below,sloped]{req \& wr}(write);
\draw[fnarrow] (read)--node[fnlbl,above]{ready}(popen);
\draw[fnarrowT] (popen) to[bend left=25] node[fnlbl,below]{req \& !wr}(read);
\draw[fnarrow] (popen) to[bend right=20] node[fnlbl,above,sloped]{req \& wr}(write);
\draw[fnarrow] (popen) to[out=-100,in=15,looseness=1.15] node[fnlbl,pos=0.55]{tCEM timeout}(idle);
\draw[fnarrow] (write)--node[fnlbl,above]{ACCESS\_CYCLES}(ww);
\draw[fnarrow] (ww) to[out=160,in=-70] node[fnlbl,pos=0.5,left]{ready}(idle);
\end{tikzpicture}
\end{center}
From INIT the controller automatically goes through a configuration-register load
sub-sequence (\code{STATE\_CR\_INIT}, 4 steps) before reaching IDLE for the first
time --- see \S~\ref{sec:pagemode}. The PAGE~OPEN~$\to$~WRITE transition
(bottom-right arrow) internally passes through two transit micro-states,
\code{STATE\_PAGE\_CLOSE} and \code{STATE\_PAGE\_REOPEN} (one cycle each): the
first forces CE\#/OE\# high for at least one cycle before the controller starts
driving the data bus, avoiding contention with the PSRAM's still-active output
($\geq t_{HZ}$); the second restarts the already-latched transaction exactly as
IDLE would. They are not drawn as separate nodes to keep the figure readable.
\subsection{Timing}
\begin{fnspec}[Timing formulas]
$\text{ACCESS\_CYCLES}=\lceil (70\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(random-access latency, $t_{AA}$/$t_{RC}$ = 70~ns)\\[3pt]
$\text{PAGE\_CYCLES}=\lceil (20\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(same-page continuation, $t_{APA}$/$t_{PC}$ = 20~ns)\\[3pt]
$\text{INIT\_CYCLES}=150\times \text{CLK\_FREQ\_MHZ}$ \quad
(power-up initialization, $t_{PU}$ = 150~\textmu s)\\[3pt]
$\text{PAGE\_TIMEOUT\_CYCLES}=\lceil (6000\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(automatic page close, safety margin under $t_{CEM}$ = 8~\textmu s)
\end{fnspec}
The data bus is tri-state driven: \code{psram\_dq = dq\_oe ? dq\_out : Z}. On read
\code{dq\_oe=0}; on write \code{dq\_oe=1} during the \code{we\_n} pulse. A WRITE\_WAIT
state keeps \code{ce\_n/lb\_n/ub\_n} active for the final hold before release.
\begin{fnwarn}[This is not QSPI]
This is a classic asynchronous-SRAM interface, \textbf{not} QSPI: most commercial
serial/QSPI ``PSRAM'' parts are not compatible with this controller without a rewrite.
See ch.~\ref{ch:hw} for the recommended part (parallel ISSI).
\end{fnwarn}
\section{Read page mode}
\label{sec:pagemode}
The recommended chip (ch.~\ref{ch:hw}) is ``asynchronous/\textbf{page mode}'': once
an initial random access at $t_{AA}$~=~70~ns has been done, further reads inside the
same 16-word page (address bits above \code{A[3]} unchanged) only cost
$t_{APA}$/$t_{PC}$~=~20~ns, because CE\#/OE\# stay asserted and only the address bus
changes. Page mode is \textbf{disabled by default} at power-up (bit~7 of the
configuration register, CR~=~\texttt{0x0070} by default) and must be explicitly
enabled.
\begin{itemize}
\item \textbf{Enable at boot}: right after INIT, the controller runs the
datasheet's ``software-access sequence'' (2 dummy reads + 2 writes, \texttt{0x0000}
unlock then real CR \texttt{0x00F0} = default with the Page bit set) at the
chip's highest address --- it reuses exactly the same READ/WRITE logic as every
other transaction, so it goes through the same timing checks.
\item \textbf{Page bursts}: after a READ the controller no longer closes CE\#/OE\#
(PAGE~OPEN state). A following read in the same page only waits PAGE\_CYCLES; a
read crossing into a different page still avoids a CE\# toggle but pays a full
ACCESS\_CYCLES for that one word (any change at \code{A[4]} or above requires a
new $t_{AA}$). A counter closes the page before the $t_{CEM}$ limit with a
safety margin.
\item \textbf{Only a WRITE closes the page.} Changes to \code{lb\_n}/\code{ub\_n}
do \emph{not} close it: \code{int8\_memory\_access} alternates these signals on
nearly every access (byte-granular access over the 16-bit bus), so treating them
as a close condition --- the first implementation attempt --- made the real
workload \emph{slower}, not faster (measured: 53.25$\to$61.25 cycles/edge on
\code{graph\_engine}'s gather); removed, corrected to 53.25$\to$37.53
cycles/edge (bandwidth +42\%, \S~\ref{sec:bandwidth}).
\end{itemize}
\begin{fnwarn}[No benefit without a sequential pattern]
Page mode only speeds up accesses that stay in the same page (or nearly) while the
controller is waiting for a new request with the page still open. Isolated,
scattered accesses (a random address every time) still pay a full ACCESS\_CYCLES,
plus a small close/reopen overhead if preceded by a WRITE or a $t_{CEM}$ timeout:
it is not a universal win, it depends on the caller's access pattern.
\end{fnwarn}
Real Fmax (\code{nextpnr-ecp5}, ch.~\ref{ch:impl}) on the integrated
\code{spi\_neuron\_top} system with Type~\#2 enabled: \textbf{75.73~MHz} at
\code{PARALLEL}=2 (was 55.59~MHz before page mode was added) and
\textbf{65.13~MHz} at \code{PARALLEL}=8, both still FAIL against the 80~MHz
target but not regressed. The critical path stays, in both cases, entirely
inside \code{u\_graph\_engine.u\_neuron} (the \code{mac8}/\code{neuron\_parallel}
accumulate chain, ch.~\ref{ch:impl}) --- \code{psram\_controller} never appears
in the critical path despite page mode's resource growth.
\section{Address map and conventions}
The addressing space is \code{ADDR\_WIDTH}=23~bits (\emph{byte} address), for a full
8~MB. The regions do not have hardwired addresses: their bases are registers set by the
host via \op{SET\_BASE} (single-layer path) or read from the descriptor table
(multi-layer path).
\begin{tabularx}{\textwidth}{L{3.2cm} L{3.4cm} Y}
\toprule
\rowh \thd{Region} & \thd{Base} & \thd{Content / convention} \\
\midrule
Input $X$ & \code{x\_base} & Shared input vector, read once per invocation. \\
\rowa Weights $W$ & \code{w\_base} & Neuron-major: weights of neuron $n$ at \code{w\_base + n*N\_INPUTS} bytes. \\
Bias & \code{bias\_addr} & One byte per neuron: bias of neuron $n$ at \code{bias\_addr + n}. \\
\rowa Descriptor table & \code{table\_base} & \code{N\_LAYERS} 11-byte entries (ch.~\ref{ch:seq}). \\
Ping-pong buffers A/B & \code{buf\_a\_base} / \code{buf\_b\_base} & Intermediate outputs between layers. \\
\bottomrule
\end{tabularx}
\subsection{PSRAM physical addressing}
The recommended PSRAM is 4M$\times$16 (8~MB), which requires a 22-bit word address
(A0--A21). \code{int8\_memory\_access} computes \code{addr>>1}, turning the 23-bit byte
address into a 22-bit word address that maps exactly onto A0--A21; bit~22 of
\code{psram\_a} is therefore always 0 and 22 real address lines remain on the PCB.
\section{Bandwidth}
\label{sec:bandwidth}
Measured on \code{graph\_engine}'s edge-list gather (ch.~\ref{ch:grafo}), by
difference between two graph sizes to isolate the per-edge cost from the fixed
per-neuron overhead (\code{sim/graph\_engine\_bandwidth\_tb.v}):
\begin{tabularx}{\textwidth}{L{5.2cm} Y Y Y}
\toprule
\rowh \thd{} & \thd{Before (no page mode)} & \thd{After (page mode)} & \thd{$\Delta$} \\
\midrule
Cycles/edge & 53.25 & 37.53 & $-29.5\%$ \\
\rowa Bandwidth @80\,MHz & 6.01\,MB/s & 8.53\,MB/s & $+41.9\%$ \\
Bandwidth @16\,MHz\textsuperscript{*} & 1.20\,MB/s & 1.71\,MB/s & $+41.9\%$ \\
\bottomrule
\end{tabularx}
\textsuperscript{*}recommended real oscillator (ch.~\ref{ch:hw}).
The model still remains memory-bound by construction: \code{neuron\_memory} reads
$X$ once and re-reads $W$/bias for each neuron (ch.~\ref{ch:seq}), one neuron at a
time; page mode reduces the per-byte cost of a sequential access, it does not
eliminate the access pattern itself.
@@ -1,111 +0,0 @@
\chapter[Memory, multi-neuron and multi-layer]{Memory integration, multi-neuron and multi-layer}
\label{ch:seq}
\section{\texttt{neuron\_memory} --- memory/neuron bridge}
\code{neuron\_memory} connects the compute datapath to memory and manages the loop over
the neurons. It reads the $X$ vector only once (shared input), then for each neuron
re-reads $W$ and bias from RAM and feeds them to a single reused instance of
\code{neuron\_parallel}: the design is memory-bound, one neuron computed at a time,
without duplicating the datapath. The output is \code{y\_bus}, packed neuron-major
(\code{DATA\_WIDTH*N\_NEURONS} bits).
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=13mm]
\node[fnstate](idle){IDLE};
\node[fnstate,right=of idle](rx){READ\_X};
\node[fnstate,right=of rx](rw){READ\_W};
\node[fnstate,below=10mm of rw](rb){READ\_BIAS};
\node[fnstate,left=of rb](sn){START\_N};
\node[fnstate,left=of sn](wn){WAIT\_N};
\draw[fnarrow] (idle)--node[fnlbl,above]{start}(rx);
\draw[fnarrow] (rx)--node[fnlbl,above]{X read}(rw);
\draw[fnarrow] (rw)--(rb);
\draw[fnarrow] (rb)--(sn);
\draw[fnarrow] (sn)--(wn);
\draw[fnarrow] (wn) to[bend left=18] node[fnlbl,above]{next neuron}(rw);
\draw[fnarrow] (wn) to[bend right=28] node[fnlbl,below]{last neuron: done}(idle);
\end{tikzpicture}
\end{center}
The states are IDLE, READ\_X, READ\_W, READ\_BIAS, START\_N, WAIT\_N. After the last
neuron the FSM returns to IDLE and asserts \code{done}. The count of neurons and inputs
actually processed is given by \code{n\_neurons\_real}/\code{n\_inputs\_real}
(ch.~\ref{ch:param}).
\section{\texttt{layer\_sequencer} --- multi-layer network}
\code{layer\_sequencer} chains up to \code{N\_LAYERS} executions of the same
\code{neuron\_memory} instance, realizing a dense feed-forward network \emph{without}
touching the validated compute core. It reads a descriptor table written by the host and
alternates the two output buffers in RAM (ping-pong).
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=13mm]
\node[fnstate](i){IDLE};
\node[fnstate,right=of i](rd){READ\\DESC};
\node[fnstate,right=of rd](rw){READ\\WAIT};
\node[fnstate,below=10mm of rw](sl){START\\LAYER};
\node[fnstate,left=of sl](wl){WAIT\\LAYER};
\node[fnstate,left=of wl](ci){COPY\\ISSUE};
\node[fnstate,below=9mm of ci](cw){COPY\\WAIT};
\draw[fnarrow] (i)--node[fnlbl,above]{run\_start}(rd);
\draw[fnarrow] (rd)--(rw);
\draw[fnarrow] (rw)--(sl);
\draw[fnarrow] (sl)--(wl);
\draw[fnarrow] (wl)--(ci);
\draw[fnarrow] (ci)--(cw);
\draw[fnarrow] (cw) to[bend left=15] node[fnlbl,left]{next layer}(rd);
\draw[fnarrow] (cw) to[bend right=12] node[fnlbl,below]{last: seq\_done}(i);
\end{tikzpicture}
\end{center}
\subsection{Ping-pong buffers}
Layer~0 reads the external input \code{x\_base}. Layer $k>0$ reads from the buffer
written by layer $k-1$; the output of each layer is copied into the other buffer,
alternating A and B. The final output remains both in \code{y\_bus} (readable with
\op{READ\_OUTPUT}) and in the ping-pong buffer into which it was copied.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=7mm]
\node[fnblockA,minimum width=18mm](x){X\\\code{x\_base}};
\node[fnblockD,right=10mm of x,minimum width=20mm](l0){Layer 0};
\node[fnblock,right=10mm of l0,minimum width=18mm](ba){buf A};
\node[fnblockD,right=10mm of ba,minimum width=20mm](l1){Layer 1};
\node[fnblock,right=10mm of l1,minimum width=18mm](bb){buf B};
\node[fnblockD,right=10mm of bb,minimum width=20mm](l2){Layer 2};
\draw[fnarrow] (x)--(l0); \draw[fnarrow] (l0)--(ba);
\draw[fnarrow] (ba)--(l1); \draw[fnarrow] (l1)--(bb);
\draw[fnarrow] (bb)--(l2);
\draw[fnarrowT,dashed] (l2.south) to[bend left=25] node[fnlbl,below]{copy into buf A} (ba.south);
\end{tikzpicture}
\end{center}
\subsection{Descriptor table}
Written by the host into RAM at \code{table\_base} with \op{WRITE\_RAM}; \code{N\_LAYERS}
entries of 11 bytes each, MSB-first:
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.6cm} Y}
\toprule
\rowh \thd{Field} & \thd{Bytes} & \thd{Meaning} \\
\midrule
\code{w\_base} & 3 & Weight base of the layer. \\
\rowa \code{bias\_addr} & 3 & Bias base of the layer. \\
\code{activation} & 1 & Layer activation (low 2 bits, cf. \code{ACT\_*}). \\
\rowa \code{n\_inputs\_real} & 2 & Actual inputs of the layer (multiple of \code{PARALLEL}). \\
\code{n\_neurons\_real} & 2 & Actual neurons of the layer. \\
\midrule
\rowh \thd{Total} & \thd{11} & per entry/layer \\
\bottomrule
\end{tabularx}
\begin{fnnote}[Copy proportional to the actual width]
The sequencer copies exactly \code{n\_neurons\_real} bytes of \code{y\_bus} into the
ping-pong buffer (not the full build width): a narrower layer is copied faster, without
zero-padding in RAM. Each activation is read per-layer from the table, independent of the
\code{activation} register of the single-layer path.
\end{fnnote}
\section{Hierarchy of the \texttt{busy}/\texttt{done} signals}
In the multi-layer path, \code{STATUS.busy} is the OR of the single-layer and sequencer
busy signals, while \code{STATUS.done} latches only at completion of the \emph{last}
layer, not at each intermediate layer (ch.~\ref{ch:spi}). The top-level returns control
of \code{neuron\_memory} to the direct \op{START} path at the end of the sequence.
@@ -1,190 +0,0 @@
\chapter[Graph network (Type \#2)]{Two-level configuration: graph network (Type \#2)}
\label{ch:grafo}
\section{Two network types}
The engine exposes two \emph{network types} selectable by the host, with the same start
command dispatching to the correct engine:
\begin{itemize}
\item \textbf{Type \#1 --- classic network (dense).} Layers with neurons per layer, fully
connected between consecutive layers. It is the \code{layer\_sequencer} path
(ch.~\ref{ch:seq}), started by \op{RUN\_NETWORK}. Connections are \emph{implicit by
position}: nothing is enumerated, only the weights are defined, addressed as
\code{w\_base + k*n\_inputs + j}.
\item \textbf{Type \#2 --- arbitrary graph (sparse).} Starting from the input neuron ids,
each neuron's connections up to the output are defined through a per-neuron \emph{sparse
edge-list}. Connections are \emph{explicit by enumeration}: each connection is an edge
\code{(src\_id, weight)}; if it is not in the list, it does not exist.
\end{itemize}
\begin{fnnote}[The difference in one line]
Dense: you define the \emph{weights} by position in a matrix. Graph: you define each
\emph{connection} as an edge \code{(src\_id, weight)} in a per-neuron list. The two
descriptor tables share the same 11-byte format but different fields; the
\code{net\_type} register tells the engine which interpretation to use.
\end{fnnote}
\section{Global activation buffer}
Type \#2 introduces an \textbf{activation buffer} indexed by \emph{signal id}, one INT8
byte per id, implemented in \textbf{on-chip \code{DP16KD} block RAM}
(\code{rtl/act\_buffer.v}). Ids \code{0..N\_in-1} are the inputs; each neuron writes its
own output into its own id. The source gather reads from here with \emph{single-cycle
random access}: this is what makes the graph cheap, because it is the access that PSRAM
(70~ns, sequential) could not accelerate.
\begin{fnspec}[V1 sizing]
\code{N\_TOTAL}=4096 signals, 16-bit id (room to 65\,536 without changing the format).
Buffer = 4~KB, i.e. 2 \code{DP16KD} blocks out of 108. The real constraint becomes the
PSRAM edge capacity ($\approx$2\,M edges at 4~B), not block RAM.
\end{fnspec}
\section{Feed-forward DAG and the \texttt{src\_id < out\_id} rule}
The graph is a feed-forward DAG: every connection points to an \textbf{already-computed}
id (\code{src\_id < out\_id}). Neurons are processed in ascending id order, so that when a
neuron is computed all its sources are ready in the buffer. Cycles and recurrence are out
of scope for V1. The rule is checked at two levels: by the host assembler (compile time)
and by a runtime guard in \code{graph\_engine} (\code{STATUS.err}), in the same philosophy
as the elaboration guard on \code{N\_INPUTS \% PARALLEL}.
\section{Data formats}
Both descriptors are 11~bytes/entry, MSB-first, at \code{table\_base}.
\subsection{Type \#2 descriptor (graph)}
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.6cm} Y}
\toprule
\rowh \thd{Field} & \thd{Bytes} & \thd{Meaning} \\
\midrule
\code{conn\_ptr} & 3 & Byte address in PSRAM of the neuron's edge block. \\
\rowa \code{n\_conn} & 2 & Real connections (pre-padding). \\
\code{out\_id} & 2 & Id into which the neuron's output is written. \\
\rowa \code{activation} & 1 & \code{ACT\_RELU} / \code{ACT\_NONE} (low 2 bits). \\
\code{bias} & 1 & Neuron bias (INT8). \\
\rowa \code{reserved} & 2 & 0. \\
\midrule
\rowh \thd{Total} & \thd{11} & entries in ascending \code{out\_id} order \\
\bottomrule
\end{tabularx}
\subsection{Graph edge (4~bytes, aligned)}
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.6cm} Y}
\toprule
\rowh \thd{Field} & \thd{Bytes} & \thd{Meaning} \\
\midrule
\code{src\_id} & 2 & Source id (uint16 BE). \\
\rowa \code{weight} & 1 & Weight (INT8). \\
\code{reserved} & 1 & 0 (4-byte alignment). \\
\bottomrule
\end{tabularx}
\begin{fnnote}[Padding to \texttt{PARALLEL}]
An arbitrary \code{n\_conn} is not a multiple of \code{PARALLEL}: the neuron's edge-list
is padded up to the multiple with \textbf{zero-weight} edges (waste
$\le$\code{PARALLEL}$-1$ per neuron). This keeps the datapath and its guard intact.
\end{fnnote}
\section{\texttt{graph\_engine} --- graph engine}
\code{rtl/graph\_engine.v} orchestrates Type \#2 \textbf{reusing \code{neuron\_parallel}
unmodified}, as \code{neuron\_memory} does for the dense case. Key difference: between the
two modes only the \emph{X addressing} changes. In Type \#1 the input is contiguous
(\code{x\_base + i}); in Type \#2 it is a gather (\code{act\_buf[src\_id]}). The arithmetic
core is untouched.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=4mm,start chain=going below,
every node/.style={on chain,fnblock,minimum width=52mm}]
\node[fnblockA]{\code{COPY\_INPUTS}: PSRAM \code{x\_base} $\to$ \code{act\_buf[0..N\_in-1]}};
\node{\code{READ\_DESC}: descriptor of neuron k};
\node{\code{READ\_EDGES}: stream edges + gather \code{act\_buf[src\_id]}};
\node{\code{START\_N} / \code{WAIT\_N}: group of \code{PARALLEL} $\to$ \code{neuron\_parallel}};
\node[fnblockT]{\code{WRITE\_ACT}: y $\to$ \code{act\_buf[out\_id]}};
\node{next neuron (id order)};
\node[fnblockD]{\code{WRITE\_OUTPUTS}: last \code{n\_out} $\to$ PSRAM \code{out\_base}};
\foreach \i [count=\j from 2] in {1,...,6}
\draw[fnarrow] (chain-\i) -- (chain-\j);
\end{tikzpicture}
\end{center}
The outputs are the \textbf{last \code{n\_out}} ids: in a DAG with the
\code{src\_id < out\_id} ordering the output neurons (sinks, not reused as sources)
naturally end up with the highest ids. At the end \code{graph\_engine} copies these
\code{n\_out} bytes into a PSRAM region at \code{out\_base}, which the host reads back with
\op{READ\_RAM}.
\section{Type \#2 opcodes and registers}
The type is selected with a new opcode; \op{RUN\_NETWORK} dispatches on the
\code{net\_type} register (details in ch.~\ref{ch:spi}).
\begin{tabularx}{\textwidth}{L{2.6cm} L{3.4cm} Y}
\toprule
\rowh \thd{Opcode / sel} & \thd{Name} & \thd{Function} \\
\midrule
\op{0x11} & SET\_NET\_TYPE & \code{type(1B)}: \code{0x01}=dense (\#1), \code{0x02}=graph (\#2). Default after \op{RESET}=dense. \\
\rowa \code{SET\_BASE sel 9} & num\_neurons\_graph & Number of graph neurons (uint16). \\
\code{SET\_BASE sel 10} & n\_out & Number of output ids (uint16). \\
\bottomrule
\end{tabularx}
\begin{fnnote}[Zero regression on Type \#1]
With \code{net\_type=dense} (the default value after \op{RESET}) the \#1 path is
bit-identical to before: \op{RUN\_NETWORK} keeps its \code{num\_layers(1B)} payload and the
framing of the existing opcodes does not change.
\end{fnnote}
\section{Occupancy (Type \#2 enabled)}
Yosys synthesis of the full \code{spi\_neuron\_top} system with Type \#2 enabled
(\code{PARALLEL}=2):
\begin{tabularx}{\textwidth}{L{4.6cm} Y}
\toprule
\rowh \thd{Resource} & \thd{Use} \\
\midrule
\code{DP16KD} (block RAM) & 2 (activation buffer) \\
\rowa \code{MULT18X18D} (DSP) & 4 (2 \code{neuron\_memory} + 2 \code{graph\_engine}) \\
LUT4 & 2619 \\
\rowa TRELLIS\_FF & 2467 \\
\code{\$\_TBUF\_} (PSRAM bus) & 16 \\
\bottomrule
\end{tabularx}
The device (108 \code{DP16KD}, 72 DSP, $\approx$44k LUT/FF) stays well below saturation:
Type \#2 adds a complete mode at a contained resource cost. LUT4/TRELLIS\_FF grew from an
earlier measurement (2367/2406) because of the PSRAM page mode added to the controller
(ch.~\ref{ch:mem}, \S~5.5) --- under 6\% utilization, no practical impact.
\section{Gather bandwidth (measured)}
The per-edge gather cost was \textbf{isolated} by building two structurally-identical
graphs with different edge counts and differencing the cycles: the subtraction cancels the
fixed per-neuron overhead and leaves the edge cost alone.
\begin{fnspec}[Per-edge cost]
\textbf{37.53 cycles/edge} with PSRAM page mode enabled (ch.~\ref{ch:mem}, \S~5.5) ---
\textbf{53.25 cycles/edge} without it (pre-page-mode baseline, consistent with theory:
4~bytes/edge $\times$ $\approx$13 cycles/byte over async PSRAM $\approx$52). At 80~MHz:
$\approx$2.13\,M edges/s ($\approx$8.5~MB/s, +42\% vs. baseline); at the real 16~MHz
clock: $\approx$426\,k edges/s ($\approx$1.71~MB/s).
\end{fnspec}
Page-mode read (roadmap G7, ch.~\ref{ch:roadmap}) has been implemented and measured: the
gather's sequential access benefits directly, cutting the per-edge cost by 29.5\%
(53.25$\to$37.53 cycles/edge). Each edge still pays \code{int8\_memory\_access}'s
byte-granular access (4 bytes/edge); page mode reduces the cost of each sequential byte,
not the number of accesses.
\section{\texttt{netasm} host assembler}
Readable network configuration needs no dedicated FPGA logic: a pseudo-assembly is
compiled \emph{on the host} (\code{tools/netasm/}) into the exact bytes of the tables and
edges, then loaded with \op{WRITE\_RAM}. The assembler validates at compile time
(\code{src\_id < out\_id}, \code{N\_TOTAL} bounds, padding to \code{PARALLEL}),
complementing the runtime guard.
\begin{lstlisting}[language=,caption={Pseudo-assembly example (graph)},basicstyle=\ttfamily\scriptsize]
NET graph
INPUTS 4 ; ids 0..3
NEURON n4 relu bias=2
CONN 0 w=5
CONN 1 w=-3
NEURON n5 none bias=0
CONN n4 w=2 ; symbolic reference to n4's output
CONN 2 w=7
OUTPUT n5
END
\end{lstlisting}
@@ -1,290 +0,0 @@
\chapter{SPI host interface}
\label{ch:spi}
\section{Physical layer}
The FPGA is always an SPI \textbf{slave}. The v1 protocol uses SPI \textbf{Mode~0}
(CPOL=0, CPHA=0), MSB-first, single-SPI. One command per low-CS period; byte~0 of each
transaction is the opcode. Multi-byte fields are big-endian.
\begin{fnspec}[Mode 0 sampling]
\code{mosi} is sampled on the \textbf{rising} edge of \code{sclk}; \code{miso} is driven
on the \textbf{falling} edge (stable before the master's next sampling). \code{spi\_slave}
synchronizes \code{sclk/mosi/cs\_n} with a double flip-flop (3-stage CDC) before every
edge detection.
\end{fnspec}
\begin{center}
\begin{tikztimingtable}[timing/dslope=0.1,timing/.style={x=3.4ex,y=2.2ex},
xscale=1.0,font=\scriptsize]
\sig{CS\_N} & H 1L 16L 1H \\
\sig{SCLK} & L 1L {2C(2)}8{2C(2)} 6L \\
\sig{MOSI} & U 1U 2D{b7} 2D{b6} 2D{b5} 2D{b4} 2D{b3} 2D{b2} 2D{b1} 2D{b0} 2U \\
\sig{MISO} & Z 1Z 16D{data} 1Z \\
\end{tikztimingtable}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
Framing of one byte: CS falls, 8 SCLK pulses, MSB first; MISO in tri-state outside a
transaction.\end{center}
\begin{fnnote}[\texttt{tx\_byte\_req} contract]
\code{tx\_byte\_req} is a \emph{prefetch hint}, not a ``byte consumed'' event: a consumer
must advance its pointers (RAM address, response byte index) on \code{rx\_valid}, which
pulses exactly once per real byte transferred.
\end{fnnote}
\section{Framing and explicit length}
The length of RAM transfers is \textbf{explicit}, not delimited by the CS edge:
\op{WRITE\_RAM}/\op{READ\_RAM} carry a 2-byte length field, so the SPI controller only
needs a byte counter. Byte addresses are 23-bit, carried in a 3-byte field with the most
significant bit reserved to 0.
\section{Opcode table}
\renewcommand{\arraystretch}{1.16}
\begin{longtable}{C{1.1cm} L{2.4cm} L{3.9cm} L{2.4cm} L{4.0cm}}
\toprule
\rowh \thd{Op} & \thd{Name} & \thd{Payload (host$\to$FPGA)} & \thd{Response} & \thd{Function} \\
\midrule
\endfirsthead
\rowh \thd{Op} & \thd{Name} & \thd{Payload} & \thd{Response} & \thd{Function} \\ \midrule
\endhead
\bottomrule
\endfoot
\op{0x00} & NOP & --- & --- & No operation (idle/dummy clocking). \\
\rowa \op{0x01} & WRITE\_RAM & addr(3B)+len(2B)+data & --- & Writes a block into PSRAM (X, weights, bias, parameters). \\
\op{0x02} & READ\_RAM & addr(3B)+len(2B) & \code{len} bytes & Reads a block back from PSRAM. \\
\rowa \op{0x0F} & RESET & --- & --- & Synchronous reset of the engine and clearing of the STATUS latch; does not erase PSRAM. \\
\op{0x10} & SET\_BASE & sel(1B)+addr(3B) & --- & Sets the bases/registers (see §\ref{sec:setbase}). \\
\rowa \op{0x11} & SET\_NET\_TYPE & type(1B) & --- & Network type: \code{0x01}=dense (\#1), \code{0x02}=graph (\#2). Default after RESET=dense. \\
\rowa \op{0x20} & START & --- & --- & Starts \code{neuron\_memory} (single-layer path); ignored if busy. \\
\op{0x21} & STATUS & --- & 1 byte & bit0=\code{busy} (live), bit1=\code{done} (sticky, clear-on-read), bit2=\code{err} (graph guard), bit3=\code{flash\_err} (sticky, clear-on-read), bit4=\code{flash\_busy} (live); bit7:5=0. \\
\rowa \op{0x22} & READ\_OUTPUT & --- & \code{N\_NEURONS} bytes & \code{y\_bus} neuron-major (byte~0 = neuron~0); dense path only (Type \#1). \\
\op{0x23} & RUN\_NETWORK & num\_layers(1B) & --- & Starts execution: dispatches on \code{net\_type} to \code{layer\_sequencer} (\#1) or \code{graph\_engine} (\#2); ignored if busy. \\
\rowa \op{0x30} & READ\_CONFIG & --- & 11 bytes & Hardware configuration record (§\ref{sec:readcfg}). \\
\op{0x40} & FLASH\_READ\_BLOCK & flash\_addr(3B)+psram\_addr(3B)+len(3B) & --- & Raw flash$\to$PSRAM read, bypasses the catalog. \\
\rowa \op{0x41} & FLASH\_WRITE\_BLOCK & psram\_addr(3B)+flash\_addr(3B)+len(3B) & --- & Raw PSRAM$\to$flash write (internal erase-before-write + $\leq$256B Page Program loop + WIP poll, transparent to the host), bypasses the catalog. \\
\op{0x42} & FLASH\_ERASE & sector\_addr(3B) & --- & Standalone 4~KB sector erase (must be sector-aligned), bypasses the catalog. \\
\rowa \op{0x43} & CAT\_READ & --- & --- & Reloads the 16-slot catalog (on-chip registers) from the flash's reserved sector. \\
\op{0x44} & CAT\_WRITE\_SLOT & slot\_id(1B)+offset(3B)+len(3B)+type(1B) & --- & Registers/updates the slot's (offset, length, type) in the on-chip catalog and persists it to flash; marks the slot \emph{invalid} until \op{SAVE\_SLOT} confirms it. \\
\rowa \op{0x45} & LOAD\_SLOT & slot\_id(1B)+psram\_addr(3B) & --- & Flash$\to$PSRAM for the slot (offset/length from the catalog), verifies the CRC32 live; \code{STATUS.flash\_err} if the slot is invalid or the CRC does not match. \\
\op{0x46} & SAVE\_SLOT & slot\_id(1B)+psram\_addr(3B)+len(3B) & --- & PSRAM$\to$flash at the slot's already-registered offset, computes the CRC32 live; on success updates and persists the catalog entry (length, CRC, valid=1). \\
\rowa \op{0x47} & CAT\_INSPECT & slot\_id(1B) & 16 bytes & Synchronous read of an already-loaded catalog entry: offset[3]+len[3]+type[1]+valid[1]+CRC32[4]+reserved[4], MSB-first. \\
\end{longtable}
All flash opcodes are \emph{fire-and-forget}: the host polls \op{STATUS} (bit4=
\code{flash\_busy}, bit3=\code{flash\_err}) or the \code{irq\_n}/\code{data\_ready\_n} pins
for the outcome, except \op{CAT\_INSPECT}, which responds synchronously.
The 8 flash opcodes (\op{0x40}--\op{0x47}) are described in full, with design rationale and
measured real latencies, in §\ref{sec:flashspi} below.
\section{\texttt{SET\_BASE} selectors}
\label{sec:setbase}
\begin{tabularx}{\textwidth}{C{1.2cm} L{3.2cm} Y}
\toprule
\rowh \thd{sel} & \thd{Register} & \thd{Use} \\
\midrule
0 & \code{x\_base} & Input base $X$. \\
\rowa 1 & \code{w\_base} & Weight base. \\
2 & \code{bias\_addr} & Bias base. \\
\rowa 3 & \code{table\_base} & Descriptor table base (multi-layer). \\
4 & \code{buf\_a\_base} & Ping-pong buffer A. \\
\rowa 5 & \code{buf\_b\_base} & Ping-pong buffer B. \\
6 & \code{activation} & Activation (low 2 bits) --- single-layer path only. \\
\rowa 7 & \code{n\_inputs\_real} & Runtime input width (16-bit BE) --- single-layer. \\
8 & \code{n\_neurons\_real} & Runtime neuron width (16-bit BE) --- single-layer. \\
\rowa 9 & \code{num\_neurons\_graph} & Number of graph neurons (16-bit BE) --- Type \#2. \\
10 & \code{n\_out} & Number of output ids (16-bit BE) --- Type \#2. \\
\bottomrule
\end{tabularx}
Selectors 6--8 concern only the single-layer/manual path; with \op{RUN\_NETWORK} the
equivalent values are read per-layer from the descriptor table.
\begin{fnwarn}[``real=0'' edge cases fixed (2026-09-04)]
The re-certification campaign (\code{docs/validation/bugs.md}) found that several
runtime values equal to zero were unguarded, with outcomes ranging from a silently
ignored limit to a hang or arbitrary-address PSRAM writes. All five cases below are now
safe no-ops, independently verified:
\begin{itemize}
\item \code{n\_inputs\_real=0} (selector 7): completes in 1 cycle with
$y=\text{activation}(\text{bias})$ (BUG-003).
\item \code{n\_neurons\_real=0} (selector 8): completes without performing any
per-neuron computation, far faster than a full-width run (BUG-004).
\item \code{num\_neurons\_graph=0} (selector 9): completes immediately after the input
copy, without ever entering the descriptor loop (BUG-006).
\item \op{RUN\_NETWORK} with \code{num\_layers=0} (dense path): an immediate no-op ---
\textbf{before the fix it executed 256 fabricated layers, reading arbitrary PSRAM data as
descriptors} (BUG-005, CRITICAL, see \S\ref{sec:run-network} below).
\item \op{SET\_NET\_TYPE} received while a run is in progress: now silently rejected
(no effect, no SPI error) instead of remapping the arbiter's multiplexer mid-execution
--- \textbf{before the fix it caused a permanent hang of the in-progress engine}
(BUG-007, CRITICAL).
\end{itemize}
Details, evidence, and per-fix verification are in \code{docs/validation/bugs.md}.
\end{fnwarn}
\section{\texttt{STATUS.done} sticky / clear-on-read}
In \code{neuron\_memory} the \code{done} signal is a single-cycle pulse. A host polling
over SPI (much slower than the FPGA clock) would almost certainly miss a raw one-cycle
pulse. The SPI register bank therefore latches \code{done} into a sticky bit on the pulse
and clears it when the host reads \op{STATUS} (or \op{RESET}). The \code{busy} bit is
instead held at level for the whole computation and is read live.
\begin{fnwarn}[Race corrected (2026-09-02)]
A real race in the sticky mechanism (present since Phase~4) was corrected by latching a
\code{status\_snapshot} on acceptance of the \op{STATUS} opcode and conditioning the
clearing of the sticky bit on \code{status\_snapshot[1]} (it clears only if the byte
actually transmitted showed \code{done=1}). A \code{done} that arrives too late for a
snapshot is reported on the next poll instead of being lost.
\end{fnwarn}
\section{Host attention pins (\texttt{data\_ready\_n}, \texttt{irq\_n})}
Besides \op{STATUS} polling, the top-level exposes two active-low physical pins (bank 7,
ch.~\ref{ch:hw}) that mirror the sticky bits without an SPI transaction, handy for driving
a host GPIO/IRQ:
\begin{itemize}
\item \code{data\_ready\_n} = $\sim$\code{STATUS.done} (sticky): low when a result is ready
to read, returns high on the \op{STATUS} read (clear-on-read).
\item \code{irq\_n} = $\sim$\code{STATUS.err} (graph guard): low when \code{graph\_engine}'s
load-time guard has tripped. It is \textbf{not} clear-on-read: it clears only on \op{RESET}
or a fresh graph start, so an error is not missed between polls.
\end{itemize}
These are additive ports: they touch neither the existing opcodes nor the registers.
\begin{fnwarn}[\code{flash\_err} has no dedicated pin]
\code{STATUS.flash\_err} (bit3) is reported \textbf{only} in the \op{STATUS} byte, by
design: reusing \code{irq\_n} would have conflated it with graph-guard errors (two
independent error domains on one pin), while a flash operation is always host-initiated
with an opcode just issued, so polling \op{STATUS} right after --- already implicit in the
``fire-and-forget, then poll \op{STATUS}/\code{data\_ready\_n}'' convention --- is already a
natural fit, no extra async pin needed. \code{data\_ready\_n}, on the other hand,
\emph{also clears at the end of a flash operation}: it mirrors \code{STATUS.done} (bit1),
which now latches on a completed flash op too, not only on \op{RUN\_NETWORK}/\op{START}.
\end{fnwarn}
\section{\texttt{READ\_CONFIG}}
\label{sec:readcfg}
Fixed \textbf{11-byte} payload: it lets a single host firmware work with different
bitstreams without recompiling. The \code{N\_INPUTS}/\code{N\_NEURONS} values report the
build \emph{maximum} (the ceiling), not necessarily the currently loaded network.
\begin{tabularx}{\textwidth}{C{1.6cm} L{3.6cm} Y}
\toprule
\rowh \thd{Byte} & \thd{Field} & \thd{Source} \\
\midrule
0 & \code{ADDR\_WIDTH} (bit) & \code{neuron\_memory.ADDR\_WIDTH} \\
\rowa 1--2 & \code{N\_INPUTS} (16-bit BE) & build maximum \\
3 & \code{N\_NEURONS} & build maximum \\
\rowa 4 & \code{PARALLEL} & build parameter \\
5 & \code{DATA\_WIDTH} (bit) & build parameter \\
\rowa 6--7 & protocol version (BE) & \code{0x0001} \\
8--9 & \code{N\_TOTAL} (16-bit BE) & max graph signals (Type \#2) \\
\rowa 10 & capability flag & bit0=\code{GRAPH\_SUPPORTED}=1 \\
\bottomrule
\end{tabularx}
\section{Flash subsystem (opcodes 0x40--0x47, completed 2026-09-04)}
\label{sec:flashspi}
The FPGA has \textbf{exclusive} access to the onboard boot/persistence flash (Winbond
\code{W25Q128JV}, 16~MB SPI NOR, ch.~\ref{ch:hw} §6/§7) through a dedicated, physically
separate SPI master (\code{rtl/spi\_flash\_master.v}), never through direct host access to
the flash pins. This is \textbf{not} a filesystem: a fixed-size catalog (16 slots,
\code{rtl/flash\_slot\_manager.v}) maps \code{slot\_id}~$\to$~(offset, length, type, valid,
CRC32) in a reserved flash sector (sector 0) --- no dynamic allocation, no garbage
collection.
\begin{fnnote}[Layering (each level independently testable)]
\begin{itemize}
\item \code{rtl/spi\_flash\_master.v} --- raw SPI master toward the flash chip
(RDID/READ/WREN/PP/SE/RDSR-1). Fully independent 4-wire bus (\code{sclk}/\code{mosi}/
\code{miso}/\code{cs\_n}, all ordinary GPIO --- Phase F7, 2026-09-04): an earlier
version reused the boot \code{CCLK} pad via the ECP5 \code{USRMCLK} primitive to save
one pin, dropped because it made the ``exclusive flash bus'' claim electrically
misleading (SCLK still depended on the same pad as the config engine) and carried an
unresolved verification gap (\code{USRMCLKTS} timing never checked against the
primary Lattice sysCONFIG Usage Guide).
\item \code{rtl/flash\_copy\_engine.v} --- block-streaming engine on top: flash$\to$PSRAM
(\code{DIR\_LOAD}), PSRAM$\to$flash with internal erase-before-write + $\leq$256B
Page Program loop + WIP polling (\code{DIR\_SAVE}), standalone sector erase
(\code{DIR\_ERASE}). A low-priority master (Port D) on \code{rtl/mem\_arbiter.v}:
flash operations are ms-scale and never block inference.
\item \code{rtl/flash\_slot\_manager.v} --- the slot catalog on top of that, plus a CRC32
(\code{rtl/crc32.v}, IEEE~802.3/zlib) computed live over the real byte stream during
\op{LOAD\_SLOT}/\op{SAVE\_SLOT}, so a corrupted or partially-written slot (e.g. power
lost mid-erase) is detected even when the underlying flash operation itself reported
success.
\end{itemize}
\end{fnnote}
\begin{fnwarn}[Sector alignment is mandatory]
\op{SAVE\_SLOT} (and the raw \op{FLASH\_WRITE\_BLOCK}/\op{FLASH\_ERASE}) require the target
flash address to be 4~KB-sector-aligned --- rejected as an error otherwise, rather than a
silent partial-sector read-modify-erase-write (no scratch buffer large enough exists for
that, and every real \op{SAVE\_SLOT} already writes a whole, sector-aligned slot by
construction).
\end{fnwarn}
Full rationale, every datasheet citation, every adversarial test (CRC mismatch, never-saved
slot, page-boundary crossing, simulated power loss, arbiter contention), and the two real
bugs found and fixed during bring-up (one pre-existing in \code{psram\_controller.v}, one in
the new arbiter request handshake) are in \code{WORKLOG.md} (Phases F1-F6 entries) and
\code{docs/FPGA-Neural-Flash-Subsystem-Verification.md} (per-module coverage summary, not
repeated here).
\begin{tabularx}{\textwidth}{L{3.4cm}Y}
\toprule
\rowh \thd{Operation} & \thd{Measured real latency} \\
\midrule
ERASE (4~KB sector) & $\approx$400~ms (dominated by the flash chip's own internal tSE, independent of the host clock) \\
\rowa SAVE (256~B page, incl. its own erase) & $\approx$403~ms (same, tSE+tPP) \\
LOAD (4096~B) & 1.74~ms (2.35~MB/s) @80~MHz; 8.71~ms (0.47~MB/s) @16~MHz (purely SPI-clock-bound) \\
\bottomrule
\end{tabularx}
Full measurement methodology in \code{docs/FPGA-Neural-Flash-Subsystem-Verification.md}.
\section{Session sequences}
\subsection{Single-layer path}
\begin{lstlisting}[language=,caption={Single-layer session},basicstyle=\ttfamily\scriptsize]
RESET -> 0x0F
READ_CONFIG -> 0x30 (host learns N_INPUTS/N_NEURONS/...)
WRITE_RAM (weights) -> 0x01 ...
WRITE_RAM (bias) -> 0x01 ...
SET_BASE (X/W/BIAS) -> 0x10 x3
WRITE_RAM (input X) -> 0x01 ...
START -> 0x20
poll STATUS -> 0x21 (until done=1; cleared by this read)
READ_OUTPUT -> 0x22
\end{lstlisting}
\subsection{Multi-layer path (RUN\_NETWORK)}
\label{sec:run-network}
\begin{lstlisting}[language=,caption={Multi-layer session},basicstyle=\ttfamily\scriptsize]
WRITE_RAM (descriptor table) -> 0x01 ...
WRITE_RAM (weights/bias per layer, X L0) -> 0x01 ...
SET_BASE (X/TABLE/BUF_A/BUF_B) -> 0x10 x4
RUN_NETWORK(num_layers) -> 0x23 <num_layers>
poll STATUS -> 0x21 (until done=1)
READ_OUTPUT -> 0x22 (y_bus of the final layer)
\end{lstlisting}
\begin{fnnote}[Out of scope for v1]
Dual~SPI and CRC/checksum on host transfers (SPI assumed reliable on a board trace --- not
to be confused with the flash catalog's CRC32, §\ref{sec:flashspi}, which protects a
different domain: flash$\leftrightarrow$PSRAM persistence, not the host SPI link).
\end{fnnote}
\begin{fnwarn}[\op{WRITE\_RAM}/\op{READ\_RAM} have no backpressure to the host --- a real risk, not a theoretical one]
Every received/produced byte must be fully processed by \code{spi\_engine} before the next
SCLK-driven byte boundary arrives --- reasonable for the initial bulk-loading of weights/
inputs, not a real-time path. The concrete risk: if a host issues \op{WRITE\_RAM}/
\op{READ\_RAM} before \code{psram\_controller.v}'s power-up sequence has completed
($\sim$150~\textmu s after reset, \code{STATE\_INIT}+\code{STATE\_CR\_INIT}),
\code{spi\_engine} stalls waiting for the very first PSRAM access to complete, while the
host --- not slowed by any handshake --- keeps clocking bytes. Bytes received during that
stall are \textbf{silently dropped}, with no error and no hang: just wrong data in PSRAM.
Found during the flash-subsystem work (\code{WORKLOG.md}, Phase~F5) via a minimal
\op{WRITE\_RAM}-only reproduction with no flash opcodes involved at all: it is a general
hazard for any host, not specific to the flash opcodes. \textbf{Current mitigation: a host
must wait for PSRAM power-up (or otherwise ensure the FPGA has been out of reset for
$>$150~\textmu s) before its first \op{WRITE\_RAM}/\op{READ\_RAM}.} Not fixed at the
protocol level (would need real backpressure, a larger change) --- declared here as an open
risk, not silently worked around.
\end{fnwarn}
@@ -1,232 +0,0 @@
\chapter[Network programming]{Neural network programming}
\label{ch:prog}
This chapter is the practical guide to encoding a network for FPGA-Neural: how it is laid
out in memory, which registers are set and how it is started, for both topologies. It
assumes the SPI opcodes (ch.~\ref{ch:spi}) and the descriptor formats (ch.~\ref{ch:seq},
\ref{ch:grafo}).
\section{General flow}
Whatever the type, the cycle is the same: the host \emph{builds the data structures in
RAM}, sets the \emph{base registers}, declares the \emph{network type}, \emph{starts} and
\emph{reads back} the result.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=3mm,start chain=going below,
every node/.style={on chain,fnblock,minimum width=64mm}]
\node[fnblockA]{1. \op{RESET} --- clears the engine and the STATUS latch};
\node{2. \op{SET\_NET\_TYPE} --- dense (\#1) or graph (\#2)};
\node{3. \op{WRITE\_RAM} --- tables, weights/edges, bias, input X};
\node{4. \op{SET\_BASE} --- base registers (x, table, \ldots)};
\node[fnblockT]{5. \op{RUN\_NETWORK} --- dispatch on \code{net\_type}};
\node{6. \op{STATUS} polling --- waits for \code{done}};
\node[fnblockD]{7. \op{READ\_OUTPUT} / \op{READ\_RAM} --- result};
\foreach \i [count=\j from 2] in {1,...,6} \draw[fnarrow] (chain-\i)--(chain-\j);
\end{tikzpicture}
\end{center}
\section{Registers and opcodes involved}
All base values are set with \op{SET\_BASE} \code{sel(1B)+addr(3B)}. Selectors:
\begin{tabularx}{\textwidth}{C{1.0cm} L{3.4cm} C{1.4cm} C{1.4cm} Y}
\toprule
\rowh \thd{sel} & \thd{Register} & \thd{Type \#1} & \thd{Type \#2} & \thd{Use} \\
\midrule
0 & \code{x\_base} & \checkmark & \checkmark & Input base $X$. \\
\rowa 3 & \code{table\_base} & \checkmark & \checkmark & Descriptor table. \\
4 & \code{buf\_a\_base} & \checkmark & \checkmark\textsuperscript{$\ast$} & Ping-pong A (\#1) / \code{out\_base} reuse (\#2). \\
\rowa 5 & \code{buf\_b\_base} & \checkmark & --- & Ping-pong B (\#1). \\
9 & \code{num\_neurons\_graph} & --- & \checkmark & Number of graph neurons. \\
\rowa 10 & \code{n\_out} & --- & \checkmark & Number of output ids. \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
$\ast$ In Type \#2 the ping-pong buffers are unused: selector 4 is reused as
\code{out\_base} (region into which outputs are copied). Selectors 1/2/6/7/8 concern only
the manual single-layer path (\op{START}), not \op{RUN\_NETWORK}.\end{center}
For Type \#1, the \emph{per-layer} \code{w\_base}/\code{bias\_addr} are \textbf{not} set
with \op{SET\_BASE}: they are fields of the descriptor table. \op{SET\_NET\_TYPE} defaults
to \emph{dense} after \op{RESET}, so a \#1 network works even without issuing it.
% ======================================================================
\section{Type \#1 --- dense network}
\subsection{Memory layout}
\begin{tabularx}{\textwidth}{L{3.4cm} Y}
\toprule
\rowh \thd{Structure} & \thd{Format} \\
\midrule
Input $X$ & \code{n\_inputs\_real} INT8 bytes at \code{x\_base}. \\
\rowa Weights (per layer) & Neuron-major: neuron $k$ at \code{w\_base + k*n\_inputs\_real}, \code{n\_neurons*n\_inputs} bytes. \\
Bias (per layer) & One INT8 byte per neuron at \code{bias\_addr}. \\
\rowa Descriptor table & \code{num\_layers} 11-byte entries at \code{table\_base}. \\
Buffers A/B & Ping-pong intermediate outputs. \\
\bottomrule
\end{tabularx}
Descriptor (11 bytes, MSB-first): \code{w\_base}(3) $|$ \code{bias\_addr}(3) $|$
\code{activation}(1) $|$ \code{n\_inputs\_real}(2) $|$ \code{n\_neurons\_real}(2).
\subsection{Worked example: a $4\to4\to2$ network}
Layer~0: 4 inputs, 4 neurons, ReLU. Layer~1: 4 inputs, 2 neurons, linear
(\code{PARALLEL}=2, so each \code{n\_inputs\_real} is a multiple of 2). Chosen addresses:
\code{table\_base}=\code{0x000000}, \code{x\_base}=\code{0x001000}, L0 weights/bias at
\code{0x002000}/\code{0x002100}, L1 at \code{0x002200}/\code{0x002300}, buffers at
\code{0x003000}/\code{0x003100}.
\begin{lstlisting}[language=,caption={Dense descriptor table (22 bytes)},basicstyle=\ttfamily\scriptsize]
Layer 0: 00 20 00 | 00 21 00 | 01 | 00 04 | 00 04
w_base bias_addr ReLU n_in=4 n_neu=4
Layer 1: 00 22 00 | 00 23 00 | 00 | 00 04 | 00 02
w_base bias_addr NONE n_in=4 n_neu=2
\end{lstlisting}
\begin{lstlisting}[language=,caption={SPI session (dense)},basicstyle=\ttfamily\scriptsize]
0x0F RESET
0x11 01 SET_NET_TYPE = dense
0x01 000000 0016 <22-byte table> WRITE_RAM table
0x01 002000 0010 <16-byte L0 wts> WRITE_RAM L0 weights (neuron-major)
0x01 002100 0004 <4-byte L0 bias>
0x01 002200 0008 <8-byte L1 wts>
0x01 002300 0002 <2-byte L1 bias>
0x01 001000 0004 <x0 x1 x2 x3> WRITE_RAM input X
0x10 00 001000 SET_BASE x_base
0x10 03 000000 SET_BASE table_base
0x10 04 003000 SET_BASE buf_a
0x10 05 003100 SET_BASE buf_b
0x23 02 RUN_NETWORK num_layers=2
0x21 ... poll STATUS until done=1
0x22 READ_OUTPUT -> 2 bytes (final layer)
\end{lstlisting}
\subsection{Host pseudocode (dense)}
\begin{lstlisting}[language=,caption={Encoding and loading a dense network},basicstyle=\ttfamily\scriptsize]
def load_dense(layers, X): # layers in execution order
spi(RESET); spi(SET_NET_TYPE, DENSE)
table = b""
for L in layers: # L: weights[n][k], bias[n], act, n_in, n_out
assert L.n_in % PARALLEL == 0
w = alloc(L.weights_neuron_major) # k slow, input fast
b = alloc(L.bias)
table += u24(w)+u24(b)+u8(L.act)+u16(L.n_in)+u16(L.n_out)
write_ram(TABLE_BASE, table)
write_ram(X_BASE, X)
set_base(0, X_BASE); set_base(3, TABLE_BASE)
set_base(4, BUF_A); set_base(5, BUF_B)
spi(RUN_NETWORK, len(layers))
wait_status_done()
return read_output(layers[-1].n_out)
\end{lstlisting}
% ======================================================================
\section{Type \#2 --- graph network}
\subsection{Memory layout}
\begin{tabularx}{\textwidth}{L{3.4cm} Y}
\toprule
\rowh \thd{Structure} & \thd{Format} \\
\midrule
Input $X$ & \code{N\_in} bytes at \code{x\_base}; copied into \code{act\_buf[0..N\_in-1]} at start. \\
\rowa Descriptor table & \code{num\_neurons\_graph} 11-byte entries at \code{table\_base}, in ascending \code{out\_id} order. \\
Edge blocks & Per neuron: \code{n\_conn} 4-byte edges at \code{conn\_ptr}, padded to a multiple of \code{PARALLEL} (zero-weight edges). \\
\rowa Outputs & \code{n\_out} bytes written to \code{out\_base} (=selector 4). \\
\bottomrule
\end{tabularx}
Graph descriptor (11 bytes): \code{conn\_ptr}(3) $|$ \code{n\_conn}(2) $|$ \code{out\_id}(2)
$|$ \code{activation}(1) $|$ \code{bias}(1) $|$ \code{reserved}(2). \quad
Edge (4 bytes): \code{src\_id}(2) $|$ \code{weight}(1) $|$ \code{reserved}(1). \quad
Rule: \code{src\_id < out\_id} (feed-forward DAG).
\subsection{Worked example}
4 inputs (ids 0--3). Neuron n4 (\code{out\_id}=4, ReLU, bias=2) connected to ids 0 and 1;
neuron n5 (\code{out\_id}=5, linear, bias=0) connected to n4 (id~4) and id~2; output = n5
(\code{n\_out}=1). \code{PARALLEL}=2, both have 2 connections (no padding). Addresses:
\code{table\_base}=\code{0x000000}, edges at \code{0x000100}, \code{x\_base}=
\code{0x001000}, \code{out\_base}=\code{0x002000}.
\begin{lstlisting}[language=,caption={Graph descriptors + edges},basicstyle=\ttfamily\scriptsize]
Descriptors (at 0x000000, 22 bytes):
n4: 00 01 00 | 00 02 | 00 04 | 01 | 02 | 00 00
conn_ptr n_conn out_id ReLU bias rsv
n5: 00 01 08 | 00 02 | 00 05 | 00 | 00 | 00 00
conn_ptr n_conn out_id NONE bias rsv
Edge blocks (at 0x000100, 4 bytes/edge: src_id, weight, rsv):
n4 @0x000100: 00 00 05 00 (src=0, w=+5)
00 01 FD 00 (src=1, w=-3) ; -3 = 0xFD
n5 @0x000108: 00 04 02 00 (src=4, w=+2) ; id4 = n4's output
00 02 07 00 (src=2, w=+7)
\end{lstlisting}
\begin{lstlisting}[language=,caption={SPI session (graph)},basicstyle=\ttfamily\scriptsize]
0x0F RESET
0x11 02 SET_NET_TYPE = graph
0x01 000000 0016 <22-byte table> WRITE_RAM descriptors
0x01 000100 0010 <16-byte edges> WRITE_RAM edge blocks
0x01 001000 0004 <x0 x1 x2 x3> WRITE_RAM input X
0x10 00 001000 SET_BASE x_base
0x10 03 000000 SET_BASE table_base
0x10 04 002000 SET_BASE out_base (sel 4 reuse)
0x10 09 000002 SET_BASE num_neurons_graph = 2
0x10 0A 000001 SET_BASE n_out = 1
0x23 00 RUN_NETWORK (dispatch to graph_engine)
0x21 ... poll STATUS (bit2=err if src_id>=out_id)
0x02 002000 0001 READ_RAM out_base -> 1 byte (n5 output)
\end{lstlisting}
\subsection{Host pseudocode (graph)}
\begin{lstlisting}[language=,caption={Encoding and loading a graph},basicstyle=\ttfamily\scriptsize]
def load_graph(neurons, X, n_out): # neurons sorted by ascending out_id
spi(RESET); spi(SET_NET_TYPE, GRAPH)
edges = b""; table = b""
for N in neurons: # N: out_id, conns=[(src_id,w)...], act, bias
for (src,_) in N.conns:
assert src < N.out_id and src < N_TOTAL # DAG rule
conn_ptr = EDGE_BASE + len(edges)
padded = pad(N.conns, PARALLEL, fill=(0,0)) # zero-weight edges
for (src,w) in padded:
edges += u16(src)+i8(w)+u8(0)
table += u24(conn_ptr)+u16(len(N.conns))+u16(N.out_id) \
+ u8(N.act)+i8(N.bias)+u16(0)
write_ram(TABLE_BASE, table); write_ram(EDGE_BASE, edges)
write_ram(X_BASE, X)
set_base(0, X_BASE); set_base(3, TABLE_BASE); set_base(4, OUT_BASE)
set_base(9, len(neurons)); set_base(10, n_out)
spi(RUN_NETWORK, 0) # payload ignored in graph
wait_status_done()
return read_ram(OUT_BASE, n_out)
\end{lstlisting}
\subsection{\texttt{netasm} pseudo-assembly}
The readable description is compiled by the host assembler (\code{tools/netasm/}) into
exactly the table and edge bytes above. Example equivalent to the worked graph:
\begin{lstlisting}[language=,caption={netasm: source and generated bytes},basicstyle=\ttfamily\scriptsize]
; --- source ---
NET graph
INPUTS 4 ; ids 0..3
NEURON n4 relu bias=2
CONN 0 w=5
CONN 1 w=-3
NEURON n5 none bias=0
CONN n4 w=2 ; symbolic reference -> id 4
CONN 2 w=7
OUTPUT n5
END
; --- the assembler emits ---
; assigned ids: n4=4, n5=5 (guarantees src_id < out_id)
; descriptors: 00 01 00 00 02 00 04 01 02 00 00
; 00 01 08 00 02 00 05 00 00 00 00
; edges: 00 00 05 00 00 01 FD 00 (n4)
; 00 04 02 00 00 02 07 00 (n5)
; registers: table_base, x_base, out_base, num_neurons=2, n_out=1
; compile-time checks: src_id<out_id, N_TOTAL, padding to PARALLEL
\end{lstlisting}
\begin{fnnote}[Why two encoding levels]
The host pseudocode and \code{netasm} produce the \emph{same bytes}. The former is useful
when the network is generated at runtime (e.g. trained weights); the latter when the
topology is hand-written or version-controlled as source. In both cases the FPGA receives
only tables and data via \op{WRITE\_RAM}: no on-board interpreter.
\end{fnnote}
@@ -1,78 +0,0 @@
\chapter[Arbitration and top-level]{Arbitration and top-level integration}
\label{ch:top}
\section{\texttt{mem\_arbiter} --- three-port arbiter}
A single byte-level memory master (which feeds the shared chain
\code{int8\_memory\_access} $\to$ \code{memory\_interface} $\to$ \code{psram\_controller})
is arbitrated among three requesters:
\begin{tabularx}{\textwidth}{C{1.3cm} L{3.4cm} Y}
\toprule
\rowh \thd{Port} & \thd{Master} & \thd{Accesses} \\
\midrule
A & \code{spi\_engine} & \op{WRITE\_RAM} / \op{READ\_RAM}. \\
\rowa B & \code{neuron\_memory} & X/W/bias reads during an execution. \\
C & \code{layer\_sequencer} & Descriptor reads + buffer writes between layers. \\
\bottomrule
\end{tabularx}
Fixed priority \textbf{B $>$ C $>$ A}: an inference in progress is more critical than the
sequencer's bookkeeping, which in turn is more critical than a manual SPI access that has
just arrived. In normal operation B and C are anyway temporally disjoint
(\code{neuron\_memory} requests only during an execution, \code{layer\_sequencer} only in
the pauses between layers), so the priority matters mostly for the corner case of a
manual \op{WRITE\_RAM}/\op{READ\_RAM} arriving during a multi-layer execution.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=6mm]
\node[fnblock,minimum width=30mm](a){Port A --- \code{spi\_engine}};
\node[fnblock,below=4mm of a,minimum width=30mm](b){Port B --- \code{neuron\_memory}};
\node[fnblock,below=4mm of b,minimum width=30mm](c){Port C --- \code{layer\_sequencer}};
\node[fnblockD,right=16mm of b,minimum width=26mm,minimum height=16mm](arb){\code{mem\_arbiter}\\{\scriptsize B$>$C$>$A}};
\node[fnblockT,right=14mm of arb,minimum width=26mm](m){shared memory\\{\scriptsize chain}};
\draw[fnarrow] (a)-|(arb.west|-a); \draw[fnarrow] (b)--(arb.west);
\draw[fnarrow] (c)-|(arb.west|-c);
\draw[fnbus] (arb)--(m);
\end{tikzpicture}
\end{center}
Once access is granted, the arbiter retains ownership until the single transaction's
\code{m\_ready} pulse, then releases: all three masters emit \code{req} as a clean
one-cycle pulse, so a queue-less grant-and-forward design suffices.
\section{\texttt{spi\_neuron\_top} --- full integration}
The top-level connects SPI (\code{spi\_slave}+\code{spi\_engine}), the arbiter, the
sequencer, \code{neuron\_memory} and the PSRAM chain. The reset of \code{neuron\_memory}
is the OR of the global reset with the soft-reset pulse of the \op{RESET} opcode, so the
host can recover the engine over SPI without a physical reset (the RAM stays intact).
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=7mm]
\node[fnblockA,minimum width=22mm](ss){\code{spi\_slave}};
\node[fnblockA,right=8mm of ss,minimum width=22mm](se){\code{spi\_engine}};
\node[fnblockT,below=8mm of se,minimum width=26mm](sq){\code{layer\_sequencer}};
\node[fnblockD,right=10mm of se,minimum width=24mm](mux){ctrl MUX\\{\scriptsize on \code{seq\_busy}}};
\node[fnblock,below=8mm of mux,minimum width=26mm](nm){\code{neuron\_memory}};
\node[fnblockD,right=10mm of mux,minimum width=22mm](arb){\code{mem\_arbiter}};
\node[fnblockA,right=8mm of arb,minimum width=26mm](mem){PSRAM chain};
\draw[fnarrow] (ss)--(se);
\draw[fnarrow] (se)--(mux);
\draw[fnarrow] (sq)--(mux);
\draw[fnarrow] (mux)--(nm);
\draw[fnarrow] (se.south) to[bend right=10] (arb.north west);
\draw[fnarrow] (nm)--(arb);
\draw[fnarrow] (sq.east) to[bend right=20] (arb.south west);
\draw[fnbus] (arb)--(mem);
\end{tikzpicture}
\end{center}
The multiplexer switches the control lines of \code{neuron\_memory} between the sequencer
(while \code{seq\_busy} is high) and the direct path of \code{spi\_engine} (legacy
single-layer mode), returning the engine to the direct path at the end of the sequence.
\begin{fnnote}[End-to-end verification]
\code{spi\_neuron\_top} is verified in simulation with real PSRAM
(\code{psram\_model.v}, no mock): RESET/READ\_CONFIG/WRITE\_RAM/READ\_RAM/SET\_BASE/
START/STATUS/READ\_OUTPUT and \op{RUN\_NETWORK} are exercised purely over simulated SPI
(ch.~\ref{ch:impl}).
\end{fnnote}
@@ -1,169 +0,0 @@
\chapter[ECP5 implementation]{ECP5 implementation and characterization}
\label{ch:impl}
\section{Flow and verification}
The project is verified on two complementary planes: functional \textbf{simulation} with
Icarus Verilog (signed algebra, products, accumulation, groups, bias, ReLU, saturation,
busy/done signals) and real \textbf{implementation} with Yosys (synthesis) $+$
nextpnr-ecp5 (place\&route, timing) $+$ Project~Trellis (\code{ecppack}).
\begin{tabularx}{\textwidth}{L{5.0cm} C{3.0cm} Y}
\toprule
\rowh \thd{Verification stage} & \thd{Outcome} & \thd{Covers} \\
\midrule
Functional RTL & \PASS & datapath correctness \\
\rowa Parametric simulation & \PASS & configuration sweep \\
ECP5 synthesis & \PASS & synthesizability, mapping \\
\rowa Placement / Routing & \PASS & LUT/FF/DSP, timing \\
Bitstream (\code{ecppack}) & \PASS & full flow, 0 errors (P2 and P8) \\
\bottomrule
\end{tabularx}
\begin{fnnote}[End-to-end toolchain through the bitstream]
The full flow RTL $\to$ Yosys $\to$ nextpnr-ecp5 $\to$ \code{ecppack} produces a valid
bitstream for P2 and P8, \textbf{0 errors at every stage}. Header verified byte-by-byte:
\code{Part: LFE5U-45F-8CABGA381}, the target's real part number, not a placeholder. Only
\emph{generation} is verified: no physical-hardware test in this session.
\end{fnnote}
\section{Datapath benchmark (256$\times$4)}
Configuration: INT8/INT32, \code{N\_INPUTS}=256, \code{N\_NEURONS}=4, variable
\code{PARALLEL}, 80~MHz target, device \code{LFE5U-45F-8BG381C} ($-8$). The test buses
are generated \emph{inside} the benchmark wrapper so as not to expose thousands of I/Os;
the top-level exposes only \code{clk/rst/start/y\_bus/busy/done}.
\begin{tabularx}{\textwidth}{C{1.4cm} C{1.8cm} C{1.4cm} C{1.6cm} C{1.6cm} C{1.5cm} C{1.4cm}}
\toprule
\rowh \thd{PAR} & \thd{tot MAC} & \thd{DSP} & \thd{Fmax} & \thd{Tcrit} & \thd{80\,MHz} & \thd{LUT4} \\
\midrule
16 & 64 & 64/72 & 52.13 & 19.18 & \FAIL & $\approx$2531 \\
\rowa 8 & 32 & 32/72 & 61.71 & 16.20 & \FAIL & --- \\
4 & 16 & 16/72 & 75.01 & 13.33 & \FAIL & 804 \\
\rowa 2 & 8 & 8/72 & 87.88 & 11.38 & \PASS & 481 \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
Fmax and Tcrit in MHz and ns. Total MACs $=$ PARALLEL$\times$4 neurons.\end{center}
\subsection{Fmax and throughput versus parallelism}
\begin{center}
\begin{tikzpicture}
\begin{axis}[
width=0.62\textwidth,height=6.0cm,
axis y line*=left, axis x line=bottom,
xlabel={\footnotesize PARALLEL}, ylabel={\footnotesize Fmax [MHz]},
xtick={2,4,8,16}, xmode=log, log basis x=2,
ymin=40,ymax=95, ytick={40,55,70,85},
tick label style={font=\scriptsize}, label style={font=\footnotesize},
grid=major, grid style={fnRule!40},
legend style={font=\scriptsize,at={(0.5,-0.28)},anchor=north,legend columns=2}]
\addplot[fnTeal,mark=*,thick,mark options={fill=fnTeal}]
coordinates {(2,87.88)(4,75.01)(8,61.71)(16,52.13)};
\addlegendentry{Fmax}
\draw[fnAmber,dashed,thick] (axis cs:2,80)--(axis cs:16,80);
\node[font=\scriptsize,text=fnAmber] at (axis cs:11,82.5){80 MHz target};
\end{axis}
\begin{axis}[
width=0.62\textwidth,height=6.0cm,
axis y line*=right, axis x line=none,
xmode=log, log basis x=2, xmin=2,xmax=16,
ylabel={\footnotesize throughput [G\,MAC/s]},
ymin=0,ymax=3.6, ytick={0,1,2,3},
tick label style={font=\scriptsize}, label style={font=\footnotesize}]
\addplot[fnBlue,mark=square*,thick,mark options={fill=fnBlue}]
coordinates {(2,0.703)(4,1.20)(8,1.97)(16,3.34)};
\label{plt:tp}
\end{axis}
\end{tikzpicture}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
Fundamental trade-off: as PARALLEL grows, Fmax drops (deeper routing/tree) but the
theoretical throughput rises. The blue line (squares) is the throughput
$\approx$MAC/cycle$\times$Fmax.\end{center}
\subsection{Interpretation}
Reducing \code{PARALLEL} lowers simultaneous MACs, DSPs, adder-tree depth and routing
congestion, so Fmax rises; but the number of groups increases and hence the latency.
Frequency alone is not enough to choose: what matters is the overall throughput
$\approx$MAC/cycle$\times$frequency.
\begin{fnnote}[Architectural choices]
\code{PARALLEL=8} is the candidate for the throughput-oriented V1: exactly 32~simultaneous
MACs with 4 neurons, DSP at $\approx$44\%, leaving resources for controller, buffers, SPI
and future pipelines. \code{PARALLEL=2} is the frequency-oriented reference: 87.88~MHz,
the only one to exceed the 80~MHz target, but it requires 128 groups for a 256-input
neuron.
\end{fnnote}
\subsection{Critical path and the 100~MHz limit}
The 100~MHz target is not met (best result 87.88~MHz with P2). The limit is
\emph{temporal}, not one of occupancy: with P2 the FPGA is barely used (DSP $\approx$11\%,
LUT $\approx$1\%). The critical path runs through weight FF $\to$ \code{MULT18X18D} $\to$
products $\to$ adder/carry $\to$ \code{acc\_next} $\to$ ReLU/saturation $\to$ output FF.
Exceeding 100~MHz will require one or more internal pipelines, not yet necessary to
proceed.
\section{Full integrated system}
Real synthesis of \code{spi\_neuron\_top} (SPI + arbiter + \code{neuron\_memory} +
\code{graph\_engine} + PSRAM chain), speed grade $-8$. Before timing closure the integrated
system missed the 80~MHz target (P2 $\approx$55~MHz, P8 $\approx$45~MHz), with a critical
path entirely inside \code{neuron\_parallel}.
\subsection{Cause: the saturation/ReLU carry chain}
Resource usage is not the cause (device below 10\% everywhere). The integrated system's
critical path is the \textbf{\code{CCU2C} carry chain of the saturation/ReLU comparator} in
\code{neuron\_parallel.v} --- \emph{not} SPI, arbiter, PSRAM, nor the Type~\#2 modules. The
saturation was written as an arithmetic comparison (\code{acc > 127}, \code{acc < -128}),
mapped by the synthesizer onto a 32-bit subtractor with a long carry chain.
\subsection{Timing closure (2026-09-03)}
An explicit waiver of the ``datapath untouchable'' rule for a separate timing-closure task,
with the single constraint of \textbf{bit-exact equivalence} across the whole regression.
Two steps:
\begin{itemize}
\item \textbf{Step 1 --- saturation/ReLU as bit-test.} A signed 32-bit value fits INT8 iff
\code{acc[31:7]} are all equal: an AND/OR reduction over a bit slice instead of a 32-bit
carry chain. A correct, bit-exact-verified simplification; real logic gain, but on its own
submerged by placement noise.
\item \textbf{Step 2 --- pipeline register} between accumulate and activation (\code{+1}
cycle of latency per neuron, absorbed by the \code{start}/\code{done} handshake, transparent
to callers). This is the decisive step.
\end{itemize}
\begin{tabularx}{\textwidth}{L{4.6cm} C{2.6cm} C{2.4cm} Y}
\toprule
\rowh \thd{Config} & \thd{Before} & \thd{After} & \thd{$\Delta$} \\
\midrule
P2, real \code{.lpf} & 54.58 & \textbf{75.30} & $+38\%$ \\
\rowa P2, 5-seed sweep & 55.59 & 73.38--75.55 & robust \\
P8, unconstrained & 45.47 & \textbf{60.26} & $+33\%$ \\
\rowa P8, 5-seed sweep & 43.15--50.48 & 60.26--68.87 & non-overlapping \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
Fmax in MHz, real place\&route (\code{nextpnr-ecp5}). Robust across 5 seeds, not attributable
to placement luck.\end{center}
\begin{fnnote}[Stop criterion and real margin]
80~MHz is not reached (75.30~MHz at P2, 94\% of target) but the gain is large and real
($+38\%$/$+33\%$). The next step (the \code{MULT18X18D} output register, which would touch
\code{mac\_unit.v}) was left out: 80~MHz is \emph{headroom} toward the real \code{.lpf}, not
an operating requirement. With the planned 16~MHz oscillator, even the worst measured number
($\approx$45~MHz at P8) has $2.8\times$ of margin.
\textbf{Superseded 2026-09-04}: after adding the flash subsystem (ch.~\ref{ch:spi}
§\ref{sec:flashspi}, ch.~\ref{ch:roadmap}), Fmax for the full system (P2, same real pinout +
3 new flash signals) was 66.68~MHz, critical path still on the same
\code{neuron\_parallel} accumulator chain identified above --- not a new bottleneck, the
difference from 75.30~MHz was placement/routing noise from the added pins/logic.
\textbf{Updated again the same day (Phase F7)}: made the flash SPI bus genuinely
independent (dropped the \code{CCLK}/\code{USRMCLK} reuse, added a 4th ordinary
\code{flash\_sclk} pin), Fmax re-measured \textbf{67.91~MHz} (slight improvement, critical
path confirmed still identical). Margin on the 16~MHz oscillator: $4.2\times$.
\end{fnnote}
\begin{fnnote}[Separate future optimization]
Independent of timing closure: the \code{x\_mem}/\code{w\_mem} arrays of \code{neuron\_memory}
are still inferred as distributed RAM on LUTs instead of \code{DP16KD}. Moving them to block
RAM would free LUTs and is a Phase~7 candidate --- but it was not on the critical path
resolved here.
\end{fnnote}
@@ -1,326 +0,0 @@
\chapter[Hardware design and pinout]{Hardware design and signal map}
\label{ch:hw}
\begin{fnnote}[Pinout status --- assigned and verified]
A real \code{.lpf} now exists (\code{synth/ecp5/spi\_neuron\_top.lpf}) with the top-level's
\textbf{57 signals} assigned to concrete CABGA381 balls, \textbf{verified by a full
0-error \code{nextpnr-ecp5} place\&route} (no longer \code{-{}-lpf-allow-unconstrained}). The
balls come from Project~Trellis's device database (\code{iodb.json}, the same nextpnr uses)
and were independently validated against §4.3.2 of the official Lattice datasheet (per-bank
GPIO counts: exact match on 6 of 7 banks, off by 1 ball on bank~3, immaterial since no
assigned signal uses it). \code{TRELLIS\_IO}: 57/245 (23\%). Current-build Fmax (full system
incl. flash subsystem with an independent SPI bus, Phase F7, 2026-09-04) \textbf{67.91~MHz},
critical path confirmed still on \code{neuron\_parallel}'s accumulator chain, unchanged from
earlier builds (ch.~\ref{ch:impl}). Pin-by-pin summary at the front of the document
(pp.~2--3). The boot config-SPI and JTAG balls do not appear here: they are dedicated
fixed-function pins with no corresponding RTL port, nextpnr never requires them (0 errors),
they matter only for the PCB schematic.
\end{fnnote}
\section{Target device}
\begin{tabularx}{\textwidth}{L{4.2cm}Y}
\toprule
\rowh \thd{Parameter} & \thd{Value} \\
\midrule
Device & Lattice ECP5 \code{LFE5U-45F-8BG381C} \\
\rowa Package & CABGA381 (381 balls) \\
Speed grade & $-8$ (the fastest of the ECP5 family) \\
\rowa Resources & $\approx$44k LUT/FF, 72$\times$\code{MULT18X18D}, \code{DP16KD} block RAM \\
Usable I/O & $\approx$232 balls out of 381 (rest: power/ground/NC) \\
\bottomrule
\end{tabularx}
\section{Pin budget}
The project requires about 60 signals out of $\approx$232 usable I/Os: ample margin
($>$170 free pins), so the board is not pin-constrained.
\begin{tabularx}{\textwidth}{Y C{2.2cm}}
\toprule
\rowh \thd{Function} & \thd{Pins} \\
\midrule
PSRAM (22 address, 16 data, 6 control) & up to 44 \\
\rowa Application SPI (\code{sclk/mosi/miso/cs\_n}) & 4 \\
Clock, reset & 2 \\
\rowa Host attention pins (\code{irq\_n}, \code{data\_ready\_n}) & 2 \\
Flash runtime SPI bus (\code{flash\_sclk/flash\_mosi/flash\_miso/flash\_cs\_n}, ordinary GPIO, fully independent bus --- Phase F7) & 4 \\
\rowa JTAG (bring-up / debug, recommended) & 4 \\
\midrule
\rowh \thd{Total} & \thd{$\approx$60} \\
\bottomrule
\end{tabularx}
\section{Signal map (top-level \texttt{spi\_neuron\_top}) --- real balls}
Real assignment of the top-level's 57 signals, verified by place\&route, \textbf{an
individual ball for every bit} (never a bus range). I/O standard: LVCMOS33 (3.3~V I/O
supply). The balls come from the real place\&route-verified \code{.lpf}. A compact summary
of the same table also appears at the front of the document (pp.~2--3).
\renewcommand{\arraystretch}{1.1}
\begin{tabularx}{\textwidth}{L{3.0cm} C{1.0cm} C{1.9cm} C{1.0cm} Y}
\toprule
\rowh \thd{Signal} & \thd{Dir} & \thd{Ball} & \thd{Bank} & \thd{Function} \\
\midrule
\multicolumn{5}{l}{\textit{\color{fnDark}Clock and reset (bank 7, left edge)}}\\
\code{clk} & IN & H5 & 7 & System clock on pad \code{GR\_PCLK7\_0} (dedicated global clock). \\
\rowa \code{rst} & IN & B4 & 7 & Global synchronous reset, active high. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Application SPI (bank 7, opposite the PSRAM bus)}}\\
\code{sclk} & IN & B5 & 7 & SPI clock (CPOL=0, CPHA=0). \\
\rowa \code{mosi} & IN & C5 & 7 & Master-Out Slave-In. \\
\code{miso} & OUT & A3 & 7 & Master-In Slave-Out (driven on the falling edge). \\
\rowa \code{cs\_n} & IN & B3 & 7 & Active-low chip-select. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Host attention pins (bank 7, active-low, level)}}\\
\code{data\_ready\_n} & OUT & C3 & 7 & Low while a result awaits reading (mirrors \code{STATUS.done}, clear on STATUS read). \\
\rowa \code{irq\_n} & OUT & C4 & 7 & Low if the graph load-time guard has tripped (mirrors \code{STATUS.err}); clears only on \code{RESET} or a fresh \code{run\_start}, \emph{not} on a STATUS read. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Flash subsystem --- independent SPI bus toward the onboard W25Q128JV (bank 7, Phases F1-F7)}}\\
\code{flash\_sclk} & OUT & E3 & 7 & SPI clock toward the flash --- ordinary GPIO, no config primitive involved (Phase F7). \\
\rowa \code{flash\_mosi} & OUT & D3 & 7 & Master-Out Slave-In toward the flash. \\
\code{flash\_miso} & IN & D5 & 7 & Master-In Slave-Out from the flash. \\
\rowa \code{flash\_cs\_n} & OUT & E4 & 7 & Flash chip-select, active low. \\
\multicolumn{5}{l}{\textit{\color{fnDark}PSRAM address bus \code{psram\_a[21:0]} --- 22 individual balls (bank 2)}}\\
\code{psram\_a[0]} & OUT & E16 & 2 & PSRAM A0 \\
\rowa \code{psram\_a[1]} & OUT & F16 & 2 & PSRAM A1 \\
\code{psram\_a[2]} & OUT & D18 & 2 & PSRAM A2 \\
\rowa \code{psram\_a[3]} & OUT & E17 & 2 & PSRAM A3 \\
\code{psram\_a[4]} & OUT & E18 & 2 & PSRAM A4 \\
\rowa \code{psram\_a[5]} & OUT & F18 & 2 & PSRAM A5 \\
\code{psram\_a[6]} & OUT & F17 & 2 & PSRAM A6 \\
\rowa \code{psram\_a[7]} & OUT & G16 & 2 & PSRAM A7 \\
\code{psram\_a[8]} & OUT & G18 & 2 & PSRAM A8 \\
\rowa \code{psram\_a[9]} & OUT & H16 & 2 & PSRAM A9 \\
\code{psram\_a[10]} & OUT & H17 & 2 & PSRAM A10 \\
\rowa \code{psram\_a[11]} & OUT & H18 & 2 & PSRAM A11 \\
\code{psram\_a[12]} & OUT & J16 & 2 & PSRAM A12 \\
\rowa \code{psram\_a[13]} & OUT & J17 & 2 & PSRAM A13 \\
\code{psram\_a[14]} & OUT & C20 & 2 & PSRAM A14 \\
\rowa \code{psram\_a[15]} & OUT & D19 & 2 & PSRAM A15 \\
\code{psram\_a[16]} & OUT & E19 & 2 & PSRAM A16 \\
\rowa \code{psram\_a[17]} & OUT & E20 & 2 & PSRAM A17 \\
\code{psram\_a[18]} & OUT & F19 & 2 & PSRAM A18 \\
\rowa \code{psram\_a[19]} & OUT & F20 & 2 & PSRAM A19 \\
\code{psram\_a[20]} & OUT & G20 & 2 & PSRAM A20 \\
\rowa \code{psram\_a[21]} & OUT & H20 & 2 & PSRAM A21 \\
\code{psram\_a[22]} & OUT & P18 & 3 & Always 0 (byte$\to$word shift): NC on the board. \\
\multicolumn{5}{l}{\textit{\color{fnDark}PSRAM data bus \code{psram\_dq[15:0]} --- 16 individual balls (banks 2 and 3)}}\\
\rowa \code{psram\_dq[0]} & IO & K18 & 2 & PSRAM DQ0 \\
\code{psram\_dq[1]} & IO & C18 & 2 & PSRAM DQ1 (dual-function ball, used as ordinary GPIO). \\
\rowa \code{psram\_dq[2]} & IO & D17 & 2 & PSRAM DQ2 \\
\code{psram\_dq[3]} & IO & D20 & 2 & PSRAM DQ3 \\
\rowa \code{psram\_dq[4]} & IO & G19 & 2 & PSRAM DQ4 \\
\code{psram\_dq[5]} & IO & J18 & 2 & PSRAM DQ5 \\
\rowa \code{psram\_dq[6]} & IO & J19 & 2 & PSRAM DQ6 \\
\code{psram\_dq[7]} & IO & J20 & 2 & PSRAM DQ7 \\
\rowa \code{psram\_dq[8]} & IO & K19 & 2 & PSRAM DQ8 \\
\code{psram\_dq[9]} & IO & K20 & 2 & PSRAM DQ9 \\
\rowa \code{psram\_dq[10]} & IO & L17 & 3 & PSRAM DQ10 \\
\code{psram\_dq[11]} & IO & M18 & 3 & PSRAM DQ11 \\
\rowa \code{psram\_dq[12]} & IO & M17 & 3 & PSRAM DQ12 \\
\code{psram\_dq[13]} & IO & N16 & 3 & PSRAM DQ13 \\
\rowa \code{psram\_dq[14]} & IO & N18 & 3 & PSRAM DQ14 \\
\code{psram\_dq[15]} & IO & P17 & 3 & PSRAM DQ15 (bidirectional tri-state data bus, \code{dq\_oe} = direction). \\
\multicolumn{5}{l}{\textit{\color{fnDark}PSRAM control (bank 3)}}\\
\rowa \code{psram\_ce\_n} & OUT & N17 & 3 & Chip enable, active low. \\
\code{psram\_oe\_n} & OUT & R16 & 3 & Output enable (read). \\
\rowa \code{psram\_we\_n} & OUT & R17 & 3 & Write enable (write). \\
\code{psram\_lb\_n} & OUT & T16 & 3 & Lower-byte enable (DQ[7:0]). \\
\rowa \code{psram\_ub\_n} & OUT & N19 & 3 & Upper-byte enable (DQ[15:8]). \\
\code{psram\_zz\_n} & OUT & N20 & 3 & Sleep/snooze (inactive=high in operation). \\
\bottomrule
\end{tabularx}
\renewcommand{\arraystretch}{1.25}
\begin{fnnote}[Board signals not exposed as RTL ports]
Not ports of \code{spi\_neuron\_top} but required at board level: the \textbf{configuration
SPI} lines to the onboard NOR flash (\code{PROGRAMN}/\code{INITN}/\code{DONE}/\code{CCLK}\ldots,
the datasheet's ``Miscellaneous Dedicated Pins'') and the 4 \textbf{JTAG} lines
(\code{TCK}/\code{TMS}/\code{TDI}/\code{TDO}), the \textbf{oscillator} on the \code{PCLK}
pad, the \textbf{power supplies}. Their ball numbers are not in the Lattice datasheet
(separate file) but are not needed here: dedicated pins with no RTL port, nextpnr never
requires them (0 errors), they matter only for the PCB schematic.
\end{fnnote}
\begin{fnwarn}[Application SPI separate from configuration SPI]
The application SPI (\code{sclk/mosi/miso/cs\_n}) must land on ordinary I/Os,
\textbf{never} on the configuration-SPI pins: the config-SPI clock pin is not reusable as
a general-purpose input after configuration without a board-level workaround. Keeping them
physically separate avoids that problem.
\end{fnwarn}
\section{Per-bank allocation (real die geometry)}
The placement follows the die-edge geometry (from Trellis's \code{globals.json},
ball~$\to$~(col,row)~$\to$~bank): banks \textbf{2 and 3} sit contiguously along the chip's
\textbf{right} edge and together hold the entire PSRAM bus (44+1 signals) --- exactly the
``one or two adjacent banks'' recommended. Bank \textbf{7} (\textbf{left} edge, physically
opposite the PSRAM bus) holds the application SPI and clock/reset, deliberately on the far
side so the two buses do not cross. \code{clk} is on the dedicated pad \code{H5}
(\code{GR\_PCLK7\_0}). Where a bank ran out of plain balls (part of \code{psram\_dq}), the
next dual-function ball was used as ordinary GPIO, confirmed usable by the real place\&route.
\begin{tabularx}{\textwidth}{Y C{1.6cm} L{4.4cm}}
\toprule
\rowh \thd{Signal group} & \thd{\# pins} & \thd{Bank (real)} \\
\midrule
PSRAM addresses \code{psram\_a[21:0]} & 22 & bank 2 (right edge) \\
\rowa PSRAM data \code{psram\_dq[15:0]} & 16 & banks 2 + 3 (adjacent) \\
PSRAM control (ce/oe/we/lb/ub/zz) & 6 & bank 3 \\
\rowa Application SPI & 4 & bank 7 (left edge) \\
Host attention pins (\code{irq\_n}, \code{data\_ready\_n}) & 2 & bank 7 \\
\rowa Independent flash SPI bus (\code{flash\_sclk/flash\_mosi/flash\_miso/flash\_cs\_n}) & 4 & bank 7 \\
Clock / reset & 2 & bank 7, \code{clk} on \code{GR\_PCLK7\_0} \\
\rowa Boot config SPI / JTAG & --- & dedicated pins (outside RTL, PCB only) \\
\bottomrule
\end{tabularx}
\section{PSRAM subsystem}
The \code{psram\_controller.v} controller implements an \textbf{asynchronous parallel}
interface (address bus, 16-bit data, \code{ce\_n/oe\_n/we\_n} and byte-lanes
\code{lb\_n/ub\_n}, plus \code{zz\_n}) with an access latency of \textbf{70~ns} wired as
$\lceil 70\,\text{ns}\times f_{clk}\rceil$. It is an asynchronous-SRAM-style bus, not QSPI.
\begin{tabularx}{\textwidth}{L{3.0cm}Y}
\toprule
\rowh \thd{Role} & \thd{Component} \\
\midrule
Working memory & ISSI \code{IS66WVE4M16EBLL-70BLI} --- 64\,Mbit parallel PSRAM (4M$\times$16, 8~MB), async, 70~ns, an exact match to the controller timing. \\
\rowa Fallback & ISSI \code{IS61WV6416DBLL} / \code{IS61WV102416BLL} (true async SRAM, drop-in on the same signals, \code{zz\_n} inactive, $\sim$10~ns, lower density). \\
Persistent storage & Winbond \code{W25Q128JV} --- 16~MB SPI NOR flash for bitstream, weights, bias, network metadata. \\
\bottomrule
\end{tabularx}
\subsection{PSRAM connection (FPGA-exclusive)}
The PSRAM is driven \textbf{exclusively by the FPGA} through \code{psram\_controller.v}: no
external master touches the bus. The external host (RPi/ESP32/MCU) only speaks SPI to the
FPGA and never touches these lines. Pin-by-pin connection FPGA~$\leftrightarrow$~ISSI
\code{IS66WVE4M16EBLL-70BLI}:
\begin{tabularx}{\textwidth}{L{3.6cm} L{3.0cm} Y}
\toprule
\rowh \thd{FPGA signal} & \thd{PSRAM pin} & \thd{Function} \\
\midrule
\code{psram\_a[21:0]} & A0--A21 & Address bus (22 lines, 8~MB word address). \\
\rowa \code{psram\_dq[15:0]} & DQ0--DQ15 & Bidirectional data bus (tri-state, \code{dq\_oe}=direction). \\
\code{psram\_ce\_n} & CE\# & Chip enable (active low). \\
\rowa \code{psram\_oe\_n} & OE\# & Output enable (read). \\
\code{psram\_we\_n} & WE\# & Write enable (write). \\
\rowa \code{psram\_lb\_n} & LB\# & Lower-byte enable (DQ[7:0]). \\
\code{psram\_ub\_n} & UB\# & Upper-byte enable (DQ[15:8]). \\
\rowa \code{psram\_zz\_n} & ZZ\# & Sleep/snooze (held high in operation). \\
\bottomrule
\end{tabularx}
PSRAM supply: \textbf{3.3~V} (BLL variant), on the same I/O rail as banks 2/3 to which it is
wired (ch.~\ref{ch:hw}, real balls). Decoupling per supply pin per the ISSI datasheet.
\section{Clock}
\label{sec:clock}
There is no PLL in the RTL yet: \code{CLK\_FREQ\_MHZ} is a \emph{timing parameter} (it
feeds the PSRAM access formulas), not a clock generator. The mounted oscillator drives
\code{clk} directly. Recommendation: a 16~MHz MEMS oscillator (SiTime SiT2001B family),
well below the 67.91~MHz Fmax of the full integrated system (incl. flash subsystem,
ch.~\ref{ch:impl}). \code{CLK\_FREQ\_MHZ} must
be set to the real value of the mounted oscillator, otherwise the PSRAM timing comes out
wrong.
\section{Power}
A \textbf{three-rail} tree (the Lattice eval board's SERDES section is not needed and is
omitted: no 1.2~V \code{VCCA}/\code{VCCHTX}):
\begin{tabularx}{\textwidth}{L{3.4cm} C{2.0cm} Y}
\toprule
\rowh \thd{Rail} & \thd{Voltage} & \thd{Feeds / regulator} \\
\midrule
\code{VCC} (core) & 1.1~V & FPGA core logic. Buck \code{TLV62568}, $\geq$600~mA. \\
\rowa \code{VCCIO0/2/3/6/7} & 3.3~V & I/O of all used banks + PSRAM. Buck \code{TLV62568}, 1~A. \\
\code{VCCAUX} & 2.5~V & FPGA auxiliary. LDO \code{TLV73325}, 10~mA. \\
\bottomrule
\end{tabularx}
Decoupling: at least one capacitor per supply pin + bulk per rail, per the Lattice ECP5
hardware checklist. Input: external 12~V (or match the bucks to the source).
\section{Configuration and programming}
\label{sec:config}
Writing the FPGA ``map'' (bitstream) happens through dedicated silicon pins, \textbf{not}
RTL top-level ports. Default mode: \textbf{MSPI} --- automatic boot from the NOR flash at
power-on (standalone product); JTAG available for development.
\subsection{JTAG (development / debug)}
\begin{tabularx}{\textwidth}{L{3.0cm} C{2.2cm} Y}
\toprule
\rowh \thd{Signal} & \thd{Ball\textsuperscript{$\dagger$}} & \thd{Function} \\
\midrule
\code{TCK} & T5 & Test clock. \\
\rowa \code{TDI} & R5 & Test data in. \\
\code{TDO} & V4 & Test data out. \\
\rowa \code{TMS} & U5 & Test mode select. \\
\bottomrule
\end{tabularx}
\subsection{Config-SPI to boot flash}
The FPGA loads the bitstream from the \textbf{Winbond \code{W25Q128JV}} (128~Mbit SPI NOR,
Quad read) at power-on. The flash subsystem (\code{rtl/flash\_slot\_manager.v}, Phases
F1-F7, ch.~\ref{ch:impl}) uses the \textbf{same physical flash} for network weights/bias/
metadata at runtime, FPGA-exclusive access: after configuration, the FPGA regains control
of the chip through a fully independent 4-wire SPI bus, \code{flash\_sclk/flash\_mosi/
flash\_miso/flash\_cs\_n} (all ordinary GPIO, pp.~2--3 and §``Signal map'' --- no ECP5
config primitive involved, Phase F7) --- this still implies a board-level dual connection
(the flash's DI/DO/CS/CLK pins wired both to the dedicated boot pins below and to these 4
ordinary balls, since it is the same physical chip serving both roles), not yet captured in
a schematic (none exists yet, see the checklist below).
\begin{tabularx}{\textwidth}{L{3.4cm} C{2.2cm} Y}
\toprule
\rowh \thd{Signal} & \thd{Ball\textsuperscript{$\dagger$}} & \thd{Function} \\
\midrule
\code{CCLK/MCLK/SCK} & U3 & Configuration clock. \\
\rowa \code{DQ0\_MOSI} & W2 & Config data (MOSI). \\
\code{DQ1\_MISO} & V2 & Config data (MISO). \\
\rowa \code{BUSY\_CSSPIN} & R2 & Flash chip-select. \\
\code{DQ2 / DQ3} & Y2 / W1 & Quad-read lines. \\
\rowa \code{PROGRAMN} & W3 & Start reconfiguration (button, active low). \\
\code{INITN} & V3 & Init / configuration error (LED). \\
\rowa \code{DONE} & Y3 & Configuration complete (LED). \\
\code{CFGMDN[2:0]} & R4/T4/U4 & Mode select (see below). \\
\bottomrule
\end{tabularx}
\subsection{Configuration modes (\texttt{CFGMDN})}
\begin{tabularx}{\textwidth}{L{4.0cm} C{4.0cm} Y}
\toprule
\rowh \thd{Mode} & \thd{CFGMDN[2:0]} & \thd{Use} \\
\midrule
MSPI (boot from flash) & \code{010} & \textbf{Default} --- standalone. \\
\rowa SSPI (slave SPI) & \code{001} & Config from external host. \\
SCM (slave serial) & \code{101} & Serial config. \\
\rowa SPCM (slave parallel) & \code{111} & 8-bit parallel config. \\
\bottomrule
\end{tabularx}
\begin{fnwarn}[Configuration balls to verify on the 45F]
\textsuperscript{$\dagger$}The JTAG and config-SPI balls listed here are the \emph{reference}
from the Lattice eval board (85F device). JTAG and config-SPI are dedicated, largely fixed
pins in the ECP5 family, but the exact positions on the \code{LFE5U-45F-8BG381C} target must
be confirmed against the Lattice 45F pinout file (Diamond/Radiant or the Trellis database)
before committing them to the schematic, as already done for the application signals
(ch.~\ref{ch:hw}).
\textbf{Distinct from this open item} (do not conflate the two): the flash subsystem's own
runtime SPI pins (\code{flash\_sclk}, \code{flash\_mosi}, \code{flash\_miso},
\code{flash\_cs\_n} --- Phases F1-F6, made fully independent in Phase F7) \textbf{are} real,
pinned, place\&route-verified ordinary GPIO on bank 7 --- \textbf{no pin shared with any
ECP5 config primitive}: an earlier version reused the boot \code{CCLK} pad for SCLK via
\code{USRMCLK}, dropped in Phase F7 (\code{USRMCLK} utilisation in the current full-system
synthesis is 0/1, confirming it is no longer used at all).
\end{fnwarn}
\section{Open tasks before schematic capture}
\begin{itemize}
\item[\OK] \code{ADDR\_WIDTH}=23 (full 8~MB) across all modules and testbenches.
\item[\OK] Real \code{.lpf} with the CABGA381 ball assignment, place\&route-verified with
0 errors (\code{synth/ecp5/spi\_neuron\_top.lpf}, 57 signals incl. flash subsystem).
\item[\OK] Boot/persistence flash subsystem (Phases F1-F7): SPI master, copy engine,
CRC32 slot catalog, fully independent 4-wire SPI bus, real synthesis at 0 errors, Fmax
67.91~MHz (\code{WORKLOG.md}).
\item[$\square$] Confirm PSRAM/SPI signal integrity at the actually mounted clock.
\item[$\square$] Board-level dual-wiring diagram for the flash's DI/DO/CS/CLK pins (dedicated
boot pins + the flash subsystem's 4 ordinary balls) --- not yet captured in a schematic.
\item[$\square$] Choice of the JTAG connector footprint.
\item[$\square$] Schematic capture (KiCad or other): no schematic exists yet for this
device/package combination.
\end{itemize}
@@ -1,80 +0,0 @@
\chapter{Quick reference}
\label{ch:ref}
\section{SPI opcodes}
\begin{tabularx}{\textwidth}{C{1.4cm} L{3.2cm} C{2.4cm} Y}
\toprule
\rowh \thd{Value} & \thd{Name} & \thd{Response} & \thd{Summary} \\
\midrule
\op{0x00} & NOP & --- & idle \\
\rowa \op{0x01} & WRITE\_RAM & --- & PSRAM block write \\
\op{0x02} & READ\_RAM & \code{len} B & PSRAM block read \\
\rowa \op{0x0F} & RESET & --- & engine reset + STATUS latch \\
\op{0x10} & SET\_BASE & --- & set base/register (sel 0..10) \\
\rowa \op{0x20} & START & --- & single-layer start \\
\op{0x21} & STATUS & 1 B & busy(live)/done(sticky) \\
\rowa \op{0x22} & READ\_OUTPUT & N\_NEURONS B & \code{y\_bus} \\
\op{0x23} & RUN\_NETWORK & --- & multi-layer start \\
\rowa \op{0x30} & READ\_CONFIG & 11 B & configuration record \\
\bottomrule
\end{tabularx}
\section{STATUS byte}
\begin{center}
\begin{tikzpicture}[font=\scriptsize]
\foreach \i/\lbl [count=\x from 0] in {7/0,6/0,5/0,4/0,3/0,2/0,1/{done},0/{busy}}{
\node[fnreg,minimum width=13mm,minimum height=9mm] (b\x) at (\x*13mm,0) {\lbl};
\node[font=\tiny,text=fnGrey,above=0.5mm of b\x] {bit \i};
}
\node[fill=fnAmber,text=white,rounded corners=1pt,inner sep=1.5pt,font=\tiny]
at (b7.center){reserved = 0};
\node[fill=fnTeal,text=white,rounded corners=1pt,inner sep=1.5pt,font=\tiny]
at (b6.center){};
\end{tikzpicture}
\end{center}
\code{done} is sticky, clear-on-read; \code{busy} is live.
\section{SET\_BASE selectors}
\begin{multicols}{2}\footnotesize
\begin{itemize}
\item 0 --- \code{x\_base}
\item 1 --- \code{w\_base}
\item 2 --- \code{bias\_addr}
\item 3 --- \code{table\_base}
\item 4 --- \code{buf\_a\_base}
\columnbreak
\item 5 --- \code{buf\_b\_base}
\item 6 --- \code{activation} (single-layer)
\item 7 --- \code{n\_inputs\_real} (single-layer)
\item 8 --- \code{n\_neurons\_real} (single-layer)
\item 9 --- \code{num\_neurons\_graph} (Type \#2)
\item 10 --- \code{n\_out} (Type \#2)
\end{itemize}
\end{multicols}
\section{Descriptor table (11 bytes/layer, MSB-first)}
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=0mm]
\node[fnreg,minimum width=20mm,minimum height=8mm](a){\code{w\_base}\\3B};
\node[fnreg,minimum width=20mm,minimum height=8mm,right=0mm of a](b){\code{bias\_addr}\\3B};
\node[fnreg,minimum width=14mm,minimum height=8mm,right=0mm of b](c){\code{act}\\1B};
\node[fnreg,minimum width=22mm,minimum height=8mm,right=0mm of c](d){\code{n\_inputs\_real}\\2B};
\node[fnreg,minimum width=22mm,minimum height=8mm,right=0mm of d](e){\code{n\_neurons\_real}\\2B};
\end{tikzpicture}
\end{center}
\section{Build parameters}
\begin{multicols}{2}\footnotesize
\begin{itemize}
\item \code{DATA\_WIDTH} --- 8 (INT8)
\item \code{ACC\_WIDTH} --- 32 (INT32)
\item \code{N\_INPUTS} --- max inputs
\item \code{N\_NEURONS} --- max neurons
\item \code{PARALLEL} --- simultaneous MACs
\columnbreak
\item \code{N\_LAYERS} --- max layers
\item \code{ADDR\_WIDTH} --- 23 (8 MB)
\item \code{MEM\_DATA\_WIDTH} --- 16
\item \code{CLK\_FREQ\_MHZ} --- PSRAM timing
\end{itemize}
\end{multicols}
@@ -1,64 +0,0 @@
\chapter{Roadmap and development status}
\label{ch:roadmap}
\section{Development phases}
\begin{tabularx}{\textwidth}{C{1.2cm} L{4.6cm} C{1.8cm} Y}
\toprule
\rowh \thd{Phase} & \thd{Title} & \thd{Status} & \thd{Content} \\
\midrule
1 & Parametric layer & \OK & inputs/neurons/parallelism, accumulation, bias, ReLU; test 32$\times$4/P=8. \\
\rowa 2 & Parameter sweep & \OK & multiple configurations incl. non-multiple and degenerate; elaboration guard added. \\
3 & Memory architecture & \OK & \code{neuron\_memory} single/multi-neuron, real PSRAM tested; multi-layer buffers $\to$ Phase~5. \\
\rowa 4 & SPI interface & \OK & \code{spi\_slave}+\code{spi\_engine}, 17 opcodes incl. flash subsystem, Fmax checked at full-system level. \\
5 & Multi-layer network & \OK$^\dagger$ & \code{layer\_sequencer}, configurable activations, runtime width; real toolchain checked. \\
\rowa 6 & Host software & planned & Linux and ESP32 drivers on the same protocol. \\
7 & Optimization & in progress & timing closure done (55$\to$75~MHz); PSRAM page-mode done (gather bandwidth +42\%); $x$/$w$ block RAM remains. \\
\rowa 8 & Hardware training (opt.) & future & backprop, gradients, weight update. \\
9 & Flash subsystem (F1-F7) & \OK & dedicated SPI master, flash$\leftrightarrow$PSRAM copy engine, 16-slot catalog with CRC32, fully independent 4-wire SPI bus (F7), 8 opcodes (\op{0x40}--\op{0x47}, ch.~\ref{ch:spi} §\ref{sec:flashspi}); real synthesis 0 errors. \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
$\dagger$ RTL, unit tests and end-to-end over simulated SPI complete; timing closure done:
75.30~MHz (P2) / 60.26~MHz (P8) at the time of Phase~5, bit-exact across the whole
regression; Fmax of the full system after Phase~9 (incl. independent flash subsystem):
\textbf{67.91~MHz} (ch.~\ref{ch:impl}).\end{center}
\section{Component status}
\begin{tabularx}{\textwidth}{Y C{4.2cm}}
\toprule
\rowh \thd{Component} & \thd{Status} \\
\midrule
Parametric neural layer & \OK{} working \\
\rowa Parametric inputs/neurons/parallelism & \OK \\
Accumulation, bias, ReLU & \OK \\
\rowa 32$\times$4 / P=8 validation & \OK \\
Dedicated RAM (interface + controller + INT8 access) & \OK{} tested on real PSRAM \\
\rowa SPI interface (17 opcodes incl. RUN\_NETWORK + flash) & \OK{} Fmax at full-system level \\
Dual SPI & future \\
\rowa Multi-layer engine & \OK{} timing closure 75.30~MHz (P2) at the time of Phase~5 \\
Configurable activations (ACT\_NONE/ACT\_RELU) & \OK \\
\rowa Runtime network width (one bitstream, any topology) & \OK{} measured savings \\
Type \#2 graph network (act\_buffer, graph\_engine, netasm) & \OK{} RTL + tests + synthesis \\
\rowa CABGA381 pinout (real \code{.lpf}, 57 signals incl. flash) & \OK{} place\&route-verified, 0 errors \\
PSRAM page-mode (G7) & \OK{} done (37.53 cycles/edge, bandwidth +42\%) \\
\rowa Flash subsystem (SPI master, copy engine, CRC32 catalog, independent bus F7) & \OK{} real synthesis 0 errors, Fmax 67.91~MHz \\
Real bitstream (\code{ecppack}, P2/P8) & \OK{} 0 errors, part LFE5U-45F-8CABGA381 \\
\rowa Linux / ESP32 host driver & planned \\
Hardware training & future \\
\bottomrule
\end{tabularx}
\section{Architectural principle (summary)}
\begin{fnspec}[Foundation of the project]
The FPGA implements the neural machine and owns its own RAM; the host configures and uses
the machine. A build fixes the \emph{ceiling} (max layers, max width, PARALLEL); the host
configures the \emph{actual} network --- number of layers, per-layer width, per-layer
activation, trained parameters --- entirely at runtime, over SPI, into the FPGA's local
memory. A single bitstream serves any topology up to that ceiling.
\end{fnspec}
\section{Long-term vision}
The final goal is a reusable hardware block integrable into different future projects:
the host platform can change (Linux, ESP32, MCU, PC) without changing the fundamental
architecture of the engine. The FPGA becomes a dedicated neural computation peripheral,
optimized for the topology required by each application.
@@ -1,96 +0,0 @@
\chapter[Modules and toolchain]{Modules, ports and toolchain}
\label{ch:appmod}
\section{List of RTL modules}
\begin{tabularx}{\textwidth}{L{3.4cm} C{2.0cm} Y}
\toprule
\rowh \thd{File} & \thd{Type} & \thd{Role} \\
\midrule
\code{rtl/mac\_unit.v} & combinational & single multiply-accumulator \\
\rowa \code{rtl/mac8.v} & combinational & parallel MAC + balanced adder tree \\
\code{rtl/neuron\_parallel.v} & FSM & neuron: groups, bias, activation, saturation \\
\rowa \code{rtl/layer.v} & structural & N\_NEURONS neurons in parallel \\
\code{rtl/neuron\_memory.v} & FSM & memory/neuron bridge, neuron loop \\
\rowa \code{rtl/layer\_sequencer.v} & FSM & multi-layer sequencing, ping-pong \\
\code{rtl/int8\_memory\_access.v} & FSM & byte $\leftrightarrow$ word conversion \\
\rowa \code{rtl/memory\_interface.v} & FSM & req/ready handshake \\
\code{rtl/psram\_controller.v} & FSM & async 70~ns physical PSRAM bus \\
\rowa \code{rtl/mem\_arbiter.v} & arbiter & 3 ports, priority B$>$C$>$A \\
\code{rtl/spi\_slave.v} & FSM & SPI Mode 0 physical layer + CDC \\
\rowa \code{rtl/spi\_engine.v} & FSM & opcode + register bank \\
\code{rtl/act\_buffer.v} & block RAM & DP16KD activation buffer (Type \#2) \\
\rowa \code{rtl/graph\_engine.v} & FSM & graph-network engine (Type \#2) \\
\code{rtl/spi\_neuron\_top.v} & top & full integration \\
\rowa \code{rtl/memory\_model.v} & model & behavioral RAM (sim) \\
\bottomrule
\end{tabularx}
\section{Ports of the top-level \texttt{spi\_neuron\_top}}
See the complete signal-by-signal table in ch.~\ref{ch:hw}. In summary: clock and reset
(\code{clk}, \code{rst}); application SPI (\code{sclk}, \code{mosi}, \code{miso},
\code{cs\_n}); PSRAM bus (\code{psram\_a[22:0]}, \code{psram\_dq[15:0]},
\code{psram\_ce\_n/oe\_n/we\_n/lb\_n/ub\_n/zz\_n}).
\section{Toolchain}
\begin{tabularx}{\textwidth}{L{3.6cm} L{3.4cm} Y}
\toprule
\rowh \thd{Tool} & \thd{Version} & \thd{Use} \\
\midrule
Yosys & 0.68+post & RTL synthesis $\to$ JSON netlist, ECP5 mapping \\
\rowa nextpnr-ecp5 & 0.11.1-19-g8dbcee5 & placement, routing, timing \\
Project Trellis & install & \code{ecppack}/\code{ecppll}/\code{ecpbram} \\
\rowa Icarus Verilog & \code{-g2012} & functional simulation \\
\bottomrule
\end{tabularx}
\subsection{Main nextpnr parameters}
\begin{lstlisting}[language=,basicstyle=\ttfamily\scriptsize]
--45k selects LFE5U-45F
--package CABGA381 package
--speed 8 speed grade -8
--json <netlist> netlist from Yosys
--lpf <constraints> pin constraints (currently empty)
--lpf-allow-unconstrained allows unconstrained I/Os (benchmark)
--freq 80 80 MHz timing target
\end{lstlisting}
\subsection{Simulation example}
\begin{lstlisting}[language=,basicstyle=\ttfamily\scriptsize]
iverilog -g2012 -Ptb.PARALLEL=16 -o sim/parametric_256x4_p16 \
sim/parametric_tb.v rtl/mac_unit.v rtl/mac8.v \
rtl/neuron_parallel.v rtl/layer.v
vvp sim/parametric_256x4_p16
\end{lstlisting}
\section{Main testbenches}
\begin{tabularx}{\textwidth}{L{5.4cm} Y}
\toprule
\rowh \thd{Testbench} & \thd{Coverage} \\
\midrule
\code{parametric\_tb.v} & 256$\times$4 datapath, accumulate/bias/ReLU/saturation cases \\
\rowa \code{parameter\_sweep\_tb.v} & sweep of valid configurations \\
\code{neuron\_parallel\_tb.v} & activations, runtime width (T7) \\
\rowa \code{neuron\_memory\_tb.v} / \code{\_multi\_tb.v} & single/multi-neuron memory integration, real PSRAM (T5) \\
\code{psram\_controller\_tb.v} & PSRAM controller \\
\rowa \code{psram\_page\_mode\_tb.v} & page bursts, page crossing, close on WRITE/$t_{CEM}$ timeout, byte-enable changes (§~5.5) \\
\code{spi\_slave\_tb.v} & SPI physical layer (4 tests) \\
\rowa \code{spi\_engine\_tb.v} & opcodes, registers (10+ tests) \\
\code{spi\_neuron\_top\_tb.v} & end-to-end, real PSRAM over simulated SPI \\
\rowa \code{spi\_neuron\_top\_runnetwork\_tb.v} & RUN\_NETWORK 2-layer end-to-end \\
\code{layer\_sequencer\_tb.v} & 2-layer sequence, ping-pong, byte-exact copy \\
\bottomrule
\end{tabularx}
\vfill
\begin{center}
\begin{tikzpicture}
\node[draw=fnRule,rounded corners=3pt,inner sep=8pt,fill=fnLight,text width=15.5cm]{
\footnotesize\color{fnGrey}
This datasheet is generated from the RTL code, the documentation and the benchmarks
present in the repository \texttt{github.com/manvalan/FPGA-Neural} as of \datasheetdate.
The Fmax, resource usage and throughput values are those reported in the repository
measurements (real \texttt{.lpf} already assigned and place\&route-verified,
ch.~\ref{ch:hw}) and must be re-verified on any substantial RTL change or as the
Phase~7 timing closure, still in progress, continues (ch.~\ref{ch:roadmap}).};
\end{tikzpicture}
\end{center}
@@ -1,184 +0,0 @@
% ======================================================================
% FPGA-Neural Datasheet -- preamble / stile
% ======================================================================
\usepackage[T1]{fontenc}
\usepackage[utf8]{inputenc}
\usepackage[italian,provide=*]{babel}
\usepackage{helvet}
\renewcommand{\familydefault}{\sfdefault}
\usepackage{courier}
\usepackage{microtype}
\usepackage[a4paper,top=2.4cm,bottom=2.3cm,left=2.2cm,right=2.2cm,headheight=15pt]{geometry}
\usepackage[table]{xcolor}
\usepackage{graphicx}
\usepackage{booktabs}
\usepackage{tabularx}
\usepackage{longtable}
\usepackage{array}
\usepackage{ltablex}
\keepXColumns
\usepackage{multirow}
\usepackage{multicol}
\usepackage{enumitem}
\usepackage{amsmath}
\usepackage{amssymb}
\usepackage{ragged2e}
% ---------- Palette ----------------------------------------------------
\definecolor{fnDark}{HTML}{0B2E4F} % blu profondo (primario)
\definecolor{fnBlue}{HTML}{15629B} % blu medio
\definecolor{fnTeal}{HTML}{0E8F8A} % accento teal
\definecolor{fnAmber}{HTML}{C9761B} % accento ambra
\definecolor{fnRed}{HTML}{B22C34} % fail / warning
\definecolor{fnGreen}{HTML}{2E7D32} % pass / ok
\definecolor{fnGrey}{HTML}{5B6B78}
\definecolor{fnLight}{HTML}{EEF3F7} % sfondo chiaro
\definecolor{fnLight2}{HTML}{E2ECF3}
\definecolor{fnRule}{HTML}{9FB4C4}
\definecolor{codebg}{HTML}{F5F7F9}
\definecolor{codekw}{HTML}{15629B}
\definecolor{codecom}{HTML}{5B6B78}
\definecolor{codestr}{HTML}{0E8F8A}
% ---------- Titoli -----------------------------------------------------
\usepackage{titlesec}
\titleformat{\chapter}[display]
{\normalfont\bfseries\color{fnDark}}
{\filright\Large\color{fnTeal}CAPITOLO \thechapter}
{6pt}
{\Huge\filright}
[\vspace{2pt}{\color{fnRule}\titlerule[1.3pt]}]
\titlespacing*{\chapter}{0pt}{6pt}{18pt}
\titleformat{\section}
{\normalfont\large\bfseries\color{fnDark}}{\thesection}{0.6em}{}
\titleformat{\subsection}
{\normalfont\bfseries\color{fnBlue}}{\thesubsection}{0.6em}{}
\titleformat{\subsubsection}
{\normalfont\bfseries\color{fnGrey}}{\thesubsubsection}{0.6em}{}
\titlespacing*{\section}{0pt}{12pt}{4pt}
% ---------- Header / footer -------------------------------------------
\usepackage{fancyhdr}
\pagestyle{fancy}
\fancyhf{}
\renewcommand{\headrulewidth}{0.6pt}
\renewcommand{\footrulewidth}{0.4pt}
\renewcommand{\headrule}{\hbox to\headwidth{\color{fnRule}\leaders\hrule height \headrulewidth\hfill}}
\renewcommand{\footrule}{\hbox to\headwidth{\color{fnRule}\leaders\hrule height \footrulewidth\hfill}}
\renewcommand{\chaptermark}[1]{\markboth{#1}{}}
\fancyhead[L]{\small\color{fnDark}\textbf{FPGA-Neural}}
\fancyhead[R]{\footnotesize\color{fnGrey}\nouppercase{\leftmark}}
\fancyfoot[L]{\small\color{fnGrey}Datasheet~\textbf{Rev.\,\datasheetrev}}
\fancyfoot[C]{\small\color{fnGrey}manvalan/FPGA-Neural}
\fancyfoot[R]{\small\color{fnGrey}\thepage}
\fancypagestyle{plain}{\fancyhf{}%
\fancyfoot[L]{\small\color{fnGrey}Datasheet~\textbf{Rev.\,\datasheetrev}}%
\fancyfoot[C]{\small\color{fnGrey}manvalan/FPGA-Neural}%
\fancyfoot[R]{\small\color{fnGrey}\thepage}%
\renewcommand{\headrulewidth}{0pt}}
% ---------- tcolorbox --------------------------------------------------
\usepackage[most]{tcolorbox}
\tcbuselibrary{skins,breakable}
% Box "nota"
\newtcolorbox{fnnote}[1][Nota]{
enhanced, breakable, colback=fnLight, colframe=fnTeal,
boxrule=0.4pt, left=8pt, right=8pt, top=5pt, bottom=5pt,
fonttitle=\bfseries\color{white}, coltitle=white,
attach boxed title to top left={xshift=6pt,yshift=-3pt},
boxed title style={colback=fnTeal,boxrule=0pt,arc=1pt}, title={#1}}
% Box "attenzione"
\newtcolorbox{fnwarn}[1][Attenzione]{
enhanced, breakable, colback=fnLight, colframe=fnAmber,
boxrule=0.4pt, left=8pt, right=8pt, top=5pt, bottom=5pt,
fonttitle=\bfseries\color{white}, coltitle=white,
attach boxed title to top left={xshift=6pt,yshift=-3pt},
boxed title style={colback=fnAmber,boxrule=0pt,arc=1pt}, title={#1}}
% Box "registro/parametro"
\newtcolorbox{fnspec}[1][Specifica]{
enhanced, breakable, colback=white, colframe=fnBlue,
boxrule=0.7pt, left=8pt, right=8pt, top=5pt, bottom=5pt, arc=1.5pt,
fonttitle=\bfseries\color{white}, coltitle=white,
attach boxed title to top left={xshift=6pt,yshift=-3pt},
boxed title style={colback=fnBlue,boxrule=0pt,arc=1pt}, title={#1}}
% ---------- listings (Verilog) ----------------------------------------
\usepackage{listings}
\lstdefinestyle{verilog}{
language=Verilog,
backgroundcolor=\color{codebg},
basicstyle=\ttfamily\scriptsize,
keywordstyle=\color{codekw}\bfseries,
commentstyle=\color{codecom}\itshape,
stringstyle=\color{codestr},
numbers=left, numberstyle=\tiny\color{fnGrey}, numbersep=7pt,
showstringspaces=false, breaklines=true, frame=leftline,
framerule=1.2pt, rulecolor=\color{fnTeal},
xleftmargin=12pt, framexleftmargin=10pt, tabsize=2,
morekeywords={logic,always_ff,always_comb,localparam,signed,genvar,generate,endgenerate}
}
\lstset{style=verilog}
% ---------- Tabelle ----------------------------------------------------
\newcolumntype{L}[1]{>{\raggedright\arraybackslash}p{#1}}
\newcolumntype{C}[1]{>{\centering\arraybackslash}p{#1}}
\newcolumntype{R}[1]{>{\raggedleft\arraybackslash}p{#1}}
\newcolumntype{Y}{>{\raggedright\arraybackslash}X}
\renewcommand{\arraystretch}{1.25}
\arrayrulecolor{fnRule}
% intestazione tabella colorata
\newcommand{\thd}[1]{\textbf{\color{white}#1}}
\newcommand{\rowh}{\rowcolor{fnDark}}
\newcommand{\rowa}{\rowcolor{fnLight}}
% ---------- Caption ----------------------------------------------------
\usepackage{caption}
\captionsetup{font=small,labelfont={bf,color=fnTeal},labelsep=period}
% ---------- TikZ / pgfplots -------------------------------------------
\usepackage{tikz}
\usetikzlibrary{arrows.meta,positioning,calc,shapes.geometric,shapes.misc,
fit,backgrounds,chains,decorations.pathreplacing,decorations.markings,
matrix,shadows.blur}
\usepackage{pgfplots}
\pgfplotsset{compat=1.17}
\usepackage{tikz-timing}
% stili di blocco riusabili
\tikzset{
fnblock/.style={draw=fnBlue,fill=fnLight,rounded corners=2pt,
minimum height=9mm,minimum width=24mm,align=center,font=\small,
inner sep=4pt,line width=0.7pt},
fnblockT/.style={fnblock,draw=fnTeal,fill=fnLight2},
fnblockD/.style={fnblock,draw=fnDark,fill=fnDark,text=white},
fnblockA/.style={fnblock,draw=fnAmber,fill=white},
fnreg/.style={draw=fnGrey,fill=white,minimum height=8mm,align=center,
font=\footnotesize,inner sep=3pt},
fnstate/.style={draw=fnBlue,fill=fnLight,circle,minimum size=13mm,
align=center,font=\scriptsize,line width=0.7pt},
fnarrow/.style={-{Stealth[length=2.6mm]},line width=0.8pt,draw=fnDark},
fnarrowT/.style={-{Stealth[length=2.6mm]},line width=0.8pt,draw=fnTeal},
fnbus/.style={-{Stealth[length=3mm]},line width=1.6pt,draw=fnBlue},
fnlbl/.style={font=\scriptsize\itshape,fill=white,inner sep=1pt,text=fnGrey}
}
% ---------- varie ------------------------------------------------------
\newcommand{\reg}[1]{\texttt{\textbf{#1}}}
\newcommand{\sig}[1]{\texttt{#1}}
\newcommand{\op}[1]{\texttt{\color{fnBlue}#1}}
\newcommand{\PASS}{\textcolor{fnGreen}{\textbf{PASS}}}
\newcommand{\FAIL}{\textcolor{fnRed}{\textbf{FAIL}}}
\newcommand{\OK}{\textcolor{fnGreen}{\textbf{OK}}}
\newcommand{\code}[1]{\texttt{#1}}
\usepackage{enumitem}
\setlist{noitemsep,topsep=2pt,leftmargin=1.4em}
\usepackage[hidelinks,colorlinks=true,linkcolor=fnBlue,urlcolor=fnTeal,
citecolor=fnBlue]{hyperref}
@@ -1,54 +0,0 @@
\thispagestyle{plain}
\noindent
\begin{tikzpicture}
\node[fill=fnDark,text=white,rounded corners=2pt,inner sep=6pt,
minimum width=\textwidth,anchor=west]
{\large\bfseries Pinout summary --- scope and honesty note};
\end{tikzpicture}
\vspace{6pt}
\noindent
{\footnotesize
V2's top-level module, \code{neural\_multiprocessor.v}, has been
synthesized and placed\&routed \textbf{unconstrained}
(\code{nextpnr-ecp5 --lpf-allow-unconstrained}) throughout this project's
own real-toolchain characterization: every Fmax/resource number in this
datasheet is real and measured, but \textbf{no ball-by-ball pin
assignment (\code{.lpf}) has been generated for V2's top level in this
revision}. Unlike V1's own pinout chapter (which reports a real,
\code{iodb.json}-verified ball map from a constrained place\&route run),
this chapter reports what is \textbf{honestly known} and nothing
invented.
}
\vspace{6pt}
\begin{fnnote}[What is real and reusable]
V2's PSRAM-facing pins (\code{psram\_a}, \code{psram\_dq},
\code{psram\_ce\_n/oe\_n/we\_n/lb\_n/ub\_n/zz\_n}) drive the exact same,
real, unmodified V1 backend chain (\code{memory\_interface.v} $\to$
\code{psram\_controller.v}) as V1's own \code{spi\_neuron\_top}. If V2 is
deployed on the same board, \textbf{V1's own real, verified ball
assignment for these signals (ch.~10 of the V1 datasheet) applies
unchanged} --- the controller was never touched, so its pin requirements
did not change either.
\end{fnnote}
\begin{fnwarn}[What is NOT yet real]
The node-registration bus (\code{reg\_valid}, \code{reg\_node\_id},
\code{reg\_required}, \code{reg\_producer\_ids}, \code{reg\_x\_base},
\code{reg\_w\_base}, \code{reg\_n\_tiles}, \code{reg\_result\_addr},
\code{reg\_ready}) has no assigned physical pins in this revision: every
V2 measurement to date drove this bus directly from a Verilator
testbench or an unconstrained synthesis top-level, never through a real
host-facing SPI (or other) interface with its own placed pinout. Framing
this bus as a real, deployable host interface (analogous to V1's SPI
Mode~0 slave) is explicitly \textbf{future work} --- see
ch.~\ref{ch:roadmap}.
\end{fnwarn}
\vspace{4pt}
\noindent
{\footnotesize\color{fnGrey}
Logical (not physical) port list and field widths: ch.~\ref{ch:regs}
(``Register-level interface''). Real PSRAM signal reuse and board
wiring: ch.~\ref{ch:hw}.\par}
@@ -1,55 +0,0 @@
\chapter{Top-level module}
\label{ch:toplevel}
\section{\texttt{neural\_multiprocessor.v}}
The real, hardware-facing top level: \code{dataflow\_core.v} (Dependency
Manager $+$ Neural Director $+$ \code{N\_SLOTS}$\times$(Memory Manager $+$
Neural Processor) $+$ Activation Cache) with its \code{N\_SLOTS}$+$1
Memory Backend Interface ports funneled through \code{slot\_mem\_arbiter.v}
down to the real, unmodified V1 PSRAM chain
(\code{memory\_interface.v} $\to$ \code{psram\_controller.v}).
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.2cm} C{1.6cm} Y}
\toprule
\rowh \thd{Port} & \thd{Dir} & \thd{Width} & \thd{Function} \\
\midrule
\code{clk}, \code{rst} & IN & 1 & System clock, synchronous reset. \\
\rowa \code{reg\_valid} & IN & 1 & Node registration request (ch.~\ref{ch:host}). \\
\code{reg\_ready} & OUT & 1 & This node id's table slot is \code{EMPTY}. \\
\rowa \code{reg\_node\_id} & IN & $\lceil\log_2\text{N\_NODES}\rceil$ & Node id. \\
\code{reg\_required} & IN & $\lceil\log_2(\text{MAX\_DEPS}{+}1)\rceil$ & Producer count. \\
\rowa \code{reg\_producer\_ids} & IN & \code{MAX\_DEPS}$\times\lceil\log_2\text{N\_NODES}\rceil$ & Packed producer id list. \\
\code{reg\_x\_base}, \code{reg\_w\_base}, \code{reg\_result\_addr} & IN & \code{ADDR\_WIDTH} each & Job descriptor addresses. \\
\rowa \code{reg\_n\_tiles} & IN & 16 & Tile count. \\
\code{psram\_a} & OUT & \code{ADDR\_WIDTH} & PSRAM address bus (real V1 controller, unmodified). \\
\rowa \code{psram\_dq} & INOUT & \code{PSRAM\_DATA\_WIDTH} & PSRAM bidirectional data bus. \\
\code{psram\_ce\_n}, \code{psram\_oe\_n}, \code{psram\_we\_n}, \code{psram\_lb\_n}, \code{psram\_ub\_n}, \code{psram\_zz\_n} & OUT & 1 each & PSRAM control, identical to V1's own real, verified signal set. \\
\bottomrule
\end{tabularx}
\begin{fnnote}[No \texttt{int8\_memory\_access.v} in this datapath]
Earlier milestones instantiated V1's \code{int8\_memory\_access.v}
between the arbiter and \code{memory\_interface.v}. Post word-burst
rewrite (ch.~\ref{ch:mem}), it is no longer instantiated here --- the
file itself is untouched (still frozen V1); V2 simply reuses one layer
lower in the same frozen stack.
\end{fnnote}
\section{Internal hierarchy}
\noindent\code{neural\_multiprocessor.v}
\begin{itemize}[leftmargin=2.4em]
\footnotesize
\item \code{u\_dataflow\_core} : \code{dataflow\_core.v}
\begin{itemize}
\item \code{u\_dep\_mgr} : \code{dependency\_manager.v}
\item \code{u\_director} : \code{neural\_director.v}
\item \code{GEN\_SLOT[0..N\_SLOTS-1]}: \code{memory\_manager.v} $+$ \code{neural\_processor.v}
\begin{itemize}
\item[--] \code{u\_prefetch} : \code{prefetch\_engine.v} (weights only, word-level)
\end{itemize}
\item \code{u\_activation\_cache} : \code{activation\_cache.v}
\end{itemize}
\item \code{u\_arbiter} : \code{slot\_mem\_arbiter.v} (\code{N\_SLOTS}$+$1 ports)
\item \code{u\_memif} : \code{memory\_interface.v} (frozen V1)
\item \code{u\_psram\_ctrl} : \code{psram\_controller.v} (frozen V1)
\end{itemize}
@@ -1,184 +0,0 @@
% ======================================================================
% FPGA-Neural Datasheet -- preamble / stile
% ======================================================================
\usepackage[T1]{fontenc}
\usepackage[utf8]{inputenc}
\usepackage[english]{babel}
\usepackage{helvet}
\renewcommand{\familydefault}{\sfdefault}
\usepackage{courier}
\usepackage{microtype}
\usepackage[a4paper,top=2.4cm,bottom=2.3cm,left=2.2cm,right=2.2cm,headheight=15pt]{geometry}
\usepackage[table]{xcolor}
\usepackage{graphicx}
\usepackage{booktabs}
\usepackage{tabularx}
\usepackage{longtable}
\usepackage{array}
\usepackage{ltablex}
\keepXColumns
\usepackage{multirow}
\usepackage{multicol}
\usepackage{enumitem}
\usepackage{amsmath}
\usepackage{amssymb}
\usepackage{ragged2e}
% ---------- Palette ----------------------------------------------------
\definecolor{fnDark}{HTML}{0B2E4F} % blu profondo (primario)
\definecolor{fnBlue}{HTML}{15629B} % blu medio
\definecolor{fnTeal}{HTML}{0E8F8A} % accento teal
\definecolor{fnAmber}{HTML}{C9761B} % accento ambra
\definecolor{fnRed}{HTML}{B22C34} % fail / warning
\definecolor{fnGreen}{HTML}{2E7D32} % pass / ok
\definecolor{fnGrey}{HTML}{5B6B78}
\definecolor{fnLight}{HTML}{EEF3F7} % sfondo chiaro
\definecolor{fnLight2}{HTML}{E2ECF3}
\definecolor{fnRule}{HTML}{9FB4C4}
\definecolor{codebg}{HTML}{F5F7F9}
\definecolor{codekw}{HTML}{15629B}
\definecolor{codecom}{HTML}{5B6B78}
\definecolor{codestr}{HTML}{0E8F8A}
% ---------- Titoli -----------------------------------------------------
\usepackage{titlesec}
\titleformat{\chapter}[display]
{\normalfont\bfseries\color{fnDark}}
{\filright\Large\color{fnTeal}CHAPTER \thechapter}
{6pt}
{\Huge\filright}
[\vspace{2pt}{\color{fnRule}\titlerule[1.3pt]}]
\titlespacing*{\chapter}{0pt}{6pt}{18pt}
\titleformat{\section}
{\normalfont\large\bfseries\color{fnDark}}{\thesection}{0.6em}{}
\titleformat{\subsection}
{\normalfont\bfseries\color{fnBlue}}{\thesubsection}{0.6em}{}
\titleformat{\subsubsection}
{\normalfont\bfseries\color{fnGrey}}{\thesubsubsection}{0.6em}{}
\titlespacing*{\section}{0pt}{12pt}{4pt}
% ---------- Header / footer -------------------------------------------
\usepackage{fancyhdr}
\pagestyle{fancy}
\fancyhf{}
\renewcommand{\headrulewidth}{0.6pt}
\renewcommand{\footrulewidth}{0.4pt}
\renewcommand{\headrule}{\hbox to\headwidth{\color{fnRule}\leaders\hrule height \headrulewidth\hfill}}
\renewcommand{\footrule}{\hbox to\headwidth{\color{fnRule}\leaders\hrule height \footrulewidth\hfill}}
\renewcommand{\chaptermark}[1]{\markboth{#1}{}}
\fancyhead[L]{\small\color{fnDark}\textbf{FPGA-Neural}}
\fancyhead[R]{\footnotesize\color{fnGrey}\nouppercase{\leftmark}}
\fancyfoot[L]{\small\color{fnGrey}Datasheet~\textbf{Rev.\,\datasheetrev}}
\fancyfoot[C]{\small\color{fnGrey}manvalan/FPGA-Neural}
\fancyfoot[R]{\small\color{fnGrey}\thepage}
\fancypagestyle{plain}{\fancyhf{}%
\fancyfoot[L]{\small\color{fnGrey}Datasheet~\textbf{Rev.\,\datasheetrev}}%
\fancyfoot[C]{\small\color{fnGrey}manvalan/FPGA-Neural}%
\fancyfoot[R]{\small\color{fnGrey}\thepage}%
\renewcommand{\headrulewidth}{0pt}}
% ---------- tcolorbox --------------------------------------------------
\usepackage[most]{tcolorbox}
\tcbuselibrary{skins,breakable}
% Box "nota"
\newtcolorbox{fnnote}[1][Note]{
enhanced, breakable, colback=fnLight, colframe=fnTeal,
boxrule=0.4pt, left=8pt, right=8pt, top=5pt, bottom=5pt,
fonttitle=\bfseries\color{white}, coltitle=white,
attach boxed title to top left={xshift=6pt,yshift=-3pt},
boxed title style={colback=fnTeal,boxrule=0pt,arc=1pt}, title={#1}}
% Box "attenzione"
\newtcolorbox{fnwarn}[1][Warning]{
enhanced, breakable, colback=fnLight, colframe=fnAmber,
boxrule=0.4pt, left=8pt, right=8pt, top=5pt, bottom=5pt,
fonttitle=\bfseries\color{white}, coltitle=white,
attach boxed title to top left={xshift=6pt,yshift=-3pt},
boxed title style={colback=fnAmber,boxrule=0pt,arc=1pt}, title={#1}}
% Box "registro/parametro"
\newtcolorbox{fnspec}[1][Specification]{
enhanced, breakable, colback=white, colframe=fnBlue,
boxrule=0.7pt, left=8pt, right=8pt, top=5pt, bottom=5pt, arc=1.5pt,
fonttitle=\bfseries\color{white}, coltitle=white,
attach boxed title to top left={xshift=6pt,yshift=-3pt},
boxed title style={colback=fnBlue,boxrule=0pt,arc=1pt}, title={#1}}
% ---------- listings (Verilog) ----------------------------------------
\usepackage{listings}
\lstdefinestyle{verilog}{
language=Verilog,
backgroundcolor=\color{codebg},
basicstyle=\ttfamily\scriptsize,
keywordstyle=\color{codekw}\bfseries,
commentstyle=\color{codecom}\itshape,
stringstyle=\color{codestr},
numbers=left, numberstyle=\tiny\color{fnGrey}, numbersep=7pt,
showstringspaces=false, breaklines=true, frame=leftline,
framerule=1.2pt, rulecolor=\color{fnTeal},
xleftmargin=12pt, framexleftmargin=10pt, tabsize=2,
morekeywords={logic,always_ff,always_comb,localparam,signed,genvar,generate,endgenerate}
}
\lstset{style=verilog}
% ---------- Tabelle ----------------------------------------------------
\newcolumntype{L}[1]{>{\raggedright\arraybackslash}p{#1}}
\newcolumntype{C}[1]{>{\centering\arraybackslash}p{#1}}
\newcolumntype{R}[1]{>{\raggedleft\arraybackslash}p{#1}}
\newcolumntype{Y}{>{\raggedright\arraybackslash}X}
\renewcommand{\arraystretch}{1.25}
\arrayrulecolor{fnRule}
% intestazione tabella colorata
\newcommand{\thd}[1]{\textbf{\color{white}#1}}
\newcommand{\rowh}{\rowcolor{fnDark}}
\newcommand{\rowa}{\rowcolor{fnLight}}
% ---------- Caption ----------------------------------------------------
\usepackage{caption}
\captionsetup{font=small,labelfont={bf,color=fnTeal},labelsep=period}
% ---------- TikZ / pgfplots -------------------------------------------
\usepackage{tikz}
\usetikzlibrary{arrows.meta,positioning,calc,shapes.geometric,shapes.misc,
fit,backgrounds,chains,decorations.pathreplacing,decorations.markings,
matrix,shadows.blur}
\usepackage{pgfplots}
\pgfplotsset{compat=1.17}
\usepackage{tikz-timing}
% stili di blocco riusabili
\tikzset{
fnblock/.style={draw=fnBlue,fill=fnLight,rounded corners=2pt,
minimum height=9mm,minimum width=24mm,align=center,font=\small,
inner sep=4pt,line width=0.7pt},
fnblockT/.style={fnblock,draw=fnTeal,fill=fnLight2},
fnblockD/.style={fnblock,draw=fnDark,fill=fnDark,text=white},
fnblockA/.style={fnblock,draw=fnAmber,fill=white},
fnreg/.style={draw=fnGrey,fill=white,minimum height=8mm,align=center,
font=\footnotesize,inner sep=3pt},
fnstate/.style={draw=fnBlue,fill=fnLight,circle,minimum size=13mm,
align=center,font=\scriptsize,line width=0.7pt},
fnarrow/.style={-{Stealth[length=2.6mm]},line width=0.8pt,draw=fnDark},
fnarrowT/.style={-{Stealth[length=2.6mm]},line width=0.8pt,draw=fnTeal},
fnbus/.style={-{Stealth[length=3mm]},line width=1.6pt,draw=fnBlue},
fnlbl/.style={font=\scriptsize\itshape,fill=white,inner sep=1pt,text=fnGrey}
}
% ---------- varie ------------------------------------------------------
\newcommand{\reg}[1]{\texttt{\textbf{#1}}}
\newcommand{\sig}[1]{\texttt{#1}}
\newcommand{\op}[1]{\texttt{\color{fnBlue}#1}}
\newcommand{\PASS}{\textcolor{fnGreen}{\textbf{PASS}}}
\newcommand{\FAIL}{\textcolor{fnRed}{\textbf{FAIL}}}
\newcommand{\OK}{\textcolor{fnGreen}{\textbf{OK}}}
\newcommand{\code}[1]{\texttt{#1}}
\usepackage{enumitem}
\setlist{noitemsep,topsep=2pt,leftmargin=1.4em}
\usepackage[hidelinks,colorlinks=true,linkcolor=fnBlue,urlcolor=fnTeal,
citecolor=fnBlue]{hyperref}
-132
View File
@@ -1,132 +0,0 @@
Perfetto, il CSV è quello giusto: FPGA-SC-02034-3-0-ECP5U-45-Pinout.csv, rev. 3.0. L'ho letto e possiamo finalmente fare il pinout esatto per il LFE5U-45F caBGA381, senza deduzioni.
Ho anche verificato la documentazione Lattice aggiornata: per LFE5U VCC = 1,1 V ±5%, VCCAUX = 2,5 V ±5%, VCCIO da 1,2 a 3,3 V e VCCIO8 è il banco associato alla configurazione.
Pinout essenziale BG381
Ball Funzione Alimentazione / collegamento
H13 VCC 1.1 V
J13 VCC 1.1 V
K13 VCC 1.1 V
L13 VCC 1.1 V
M13 VCC 1.1 V
N13 VCC 1.1 V
H12 VCC 1.1 V
N12 VCC 1.1 V
H11 VCC 1.1 V
N11 VCC 1.1 V
H10 VCC 1.1 V
N10 VCC 1.1 V
H9 VCC 1.1 V
N9 VCC 1.1 V
H8 VCC 1.1 V
J8 VCC 1.1 V
K8 VCC 1.1 V
L8 VCC 1.1 V
M8 VCC 1.1 V
N8 VCC 1.1 V
F6 VCCAUX 2.5 V
P6 VCCAUX 2.5 V
F15 VCCAUX 2.5 V
P15 VCCAUX 2.5 V
F9, F10 VCCIO0 Bank 0
F11, F12 VCCIO1 Bank 1
H14, H15, J15 VCCIO2 Bank 2
L14, L15, M15 VCCIO3 Bank 3
L6, L7, M6 VCCIO6 Bank 6
H6, H7, J6 VCCIO7 Bank 7
P9, P10 VCCIO8 Bank 8 / configuration
Nota: il CSV elenca anche i pin VSS/VSSIO; vanno tutti collegati correttamente al piano GND, senza lasciarne nessuno scollegato.
JTAG / programmazione
Questi sono particolarmente importanti perché sono ball esatti del BG381:
Ball Pin Bank Collegamento
R5 TDI 40 JTAG TDI
T5 TCK 40 JTAG TCK
U5 TMS 40 JTAG TMS
V4 TDO 40 JTAG TDO
W3 PROGRAMN 8 controllo configurazione
V3 INITN 8 stato configurazione
Y3 DONE 8 configurazione completata
Il bank 8 è quello associato all'interfaccia di configurazione; Lattice indica esplicitamente VCCIO8 come alimentazione da dimensionare in funzione dell'interfaccia di configurazione utilizzata.
CCLK
Dal CSV:
Ball Pin Bank
U3 CCLK 8
Attenzione però: CCLK non è automaticamente il clock di sistema della nostra rete neurale. È il clock associato alla configurazione; il clock operativo della FPGA va identificato separatamente nel percorso dell'oscillatore/PLL.
CFG[2:0] (selezione modalità di boot)
Dal CSV, tutti banco 8:
Ball Pin Note
U4 CFG_0 CFGMDN0
T4 CFG_1 CFGMDN1
R4 CFG_2 CFGMDN2
Per boot automatico da flash #2 (MSPI): CFG[2:0]=[0,1,0] (letto CFG2,CFG1,CFG0) → CFG_2 a GND, CFG_1 a pull-up 110kΩ verso VCCIO8, CFG_0 a GND (dato reale, Lattice FPGA-TN-02039-2.3 §6.1.1, Tabella 6.3). Pin resi modificabili via jumper/resistori 0Ω, non hardwired fissi.
Pin dual-function MSPI verso flash #2 (boot)
Dal CSV, ball dual-function del banco 8, NON pin JTAG/dedicati separati — sono ball PIO ordinari con funzione secondaria sysCONFIG:
Ball Pin CSV Funzione MSPI
R2 PB15A: HOLDN/DI/BUSY/CSSPIN/CEN CSSPIN (chip select verso flash #2), + 4.7kΩ pull-up a VCCIO8
W2 PB11B: D0/MOSI/IO0 D0/MOSI verso flash #2
V2 PB11A: D1/MISO/IO1 D1/MISO verso flash #2
U3 CCLK (vedi sopra) MCLK verso flash #2, pull-up debole interna
Questi 4 ball (insieme a PROGRAMN/INITN/DONE sopra) collegano l'FPGA esclusivamente alla flash di boot — unico chip flash presente nel design attuale.
**Aggiornamento 2026-09-07 — Flash #1 (dati rete neurale) rimossa**: era stata realmente integrata (RTL V1 `flash_copy_engine.v`/`flash_slot_manager.v` istanziato, adapter nuovo, opcode SPI dedicato, testbench dedicato, verificata bit-exact) sui ball B2/E2/F2/F3 (banco 7). **Rimossa di nuovo** su scelta esplicita dell'utente: degradava il timing reale di N_SLOTS=4 (8/8→3/8 PASS a 64MHz) e la frequenza di clock è stata giudicata più importante della persistenza locale dei pesi — l'ESP32 può ricaricarli ad ogni sessione via il protocollo SPI applicativo. Revert pulito (`git revert`, commit `59901a4`, recuperabile in futuro). I ball B2/E2/F2/F3 sono di nuovo liberi (nessuna funzione riservata). Vedi `decisions.log` DEC-0041 (design originale) e DEC-0042 (rimozione + recupero del timing).
Decoupling
Qui preferisco essere molto preciso: non voglio inventare una quantità di condensatori per “ogni VCC”.
La Hardware Checklist ufficiale Lattice richiede esplicitamente:
VCC 1,1 V con PCB plane e adeguato decoupling;
VCCAUX 2,5 V;
VCCIO 1,23,3 V;
VCCIO8 coerente con l'interfaccia di configurazione.
Per la nostra scheda farei quindi una rete di decoupling distribuita, non un condensatore unico per ogni ball:
LFE5U-45F
┌──────────────┐
1V1_CORE ─────┤ VCC │
│ │
2V5_AUX ──────┤ VCCAUX │
│ │
3V3/2V5/etc ──┤ VCCIOx │
│ │
GND ──────────┤ VSS │
└──────────────┘
con i condensatori fisicamente distribuiti attorno al BGA, con quelli ad alta frequenza il più vicino possibile alle rispettive vie/pad di alimentazione.
Una cosa importante
Adesso che abbiamo il CSV ufficiale, posso fare un lavoro ancora più utile: estrarre tutti i 381 ball del BG381 e generarti una tabella Markdown completa divisa in:
VCC
VCCAUX
VCCIO08
VSS/VSSIO
JTAG
PROGRAMN/INITN/DONE
clock
GPIO utilizzabili
pin da NON utilizzare
BOM decoupling consigliata
header JTAG completo
Quella diventerebbe praticamente la nostra scheda di riferimento hardware per il PCB V2.
@@ -850,9 +850,9 @@ module tb #(
wait (u_nmp.u_sdram_backend.u_sdram_ctrl.state == u_nmp.u_sdram_backend.u_sdram_ctrl.S_IDLE); wait (u_nmp.u_sdram_backend.u_sdram_ctrl.state == u_nmp.u_sdram_backend.u_sdram_ctrl.S_IDLE);
@(posedge clk); @(posedge clk);
// STEP19 official memory map (hardware/v2/docs/MEMORY_ARCHITECTURE.md): // Official V2 memory map (datasheet ch.5): weights @ 0x010000,
// weights @ 0x010000, activations @ 0x200000, results @ 0x300000 -- // activations @ 0x200000, results @ 0x300000 -- non-overlapping
// non-overlapping 1MB-aligned regions in the single 8MB SDRAM. // 1MB-aligned regions in the single SDRAM.
run_dense_layer("D-Stress", 256, 16, 16'd400, 26'h200000, 26'h010000, 26'h300000, 1'b0); run_dense_layer("D-Stress", 256, 16, 16'd400, 26'h200000, 26'h010000, 26'h300000, 1'b0);
// FPGA_DATA_READY check: the whole graph (256 nodes) just // FPGA_DATA_READY check: the whole graph (256 nodes) just