chore: remove root-level V1 duplicates, superseded by hardware/v1/ freeze

hardware/v1/ was created (dc0b331) as a frozen snapshot of the V1
project that then lived at the repo root (rtl/, sim/, synth/, tools/,
docs/). Root received zero further commits to those files after the
freeze -- confirmed byte-identical to the hardware/v1/ copy for every
file removed here. Root was the "before", hardware/v1/ is the
curated, canonical "after".

Removed (all verified exact-hash duplicates of hardware/v1/ content):
  - rtl/ (20 files, 100% covered by hardware/v1/rtl/)
  - tools/{netasm,pinout,run_regression.py,flash_catalog,validation,
    fpga_benchmark.py} (19 files, 100% covered by hardware/v1/tools/;
    tools/neural_sim/ kept -- unique, post-freeze, no counterpart)
  - sim/*.v (47 testbenches, 100% covered by hardware/v1/sim/; the
    ~38 remaining sim/ entries are compiled binaries and .vcd
    waveform dumps, left as a separate cleanup decision)
  - synth/ecp5/{p2,p4,p8,post_fix_verify} (25 files, exact duplicates
    of hardware/v1/synthesis/; the other ~84 synth/ecp5/* experiment
    build directories are historical artifacts never carried into the
    freeze, left as a separate decision)
  - WORKLOG.md (duplicate of hardware/v1/docs/WORKLOG.md)
  - docs/{FPGA-Neural-Datapatch-Benchmark,FPGA-Neural-Hardware-Design,
    FPGA-NeuralNetwork-Engine}.md, docs/validation/*.md (18 files),
    docs/FPGA-Neural-Datasheet-{EN,IT}.pdf -- all exact duplicates of
    hardware/v1/docs/ content
  - hardware/v1/docs/DatasheetLatex/ (24 files) -- exact duplicate of
    hardware/v2/docs/datasheet/files/docs/datasheet/en/ (discovered
    during this audit; not the same DatasheetLatex already removed
    from hardware/v2/docs/ in an earlier commit)

Moved (genuine, unique, post-freeze V2 content -- not duplicated
anywhere, just living in the wrong/legacy root docs/ location):
  - docs/architecture/*.md -> hardware/v2/docs/architecture/
  - docs/pinouts.md, docs/FPGA_NEURAL_V2_DATASHEET.md,
    docs/FPGA_NEURAL_V2_SCHEMATIC.md,
    docs/FPGA-Neural-V2-Datasheet-EN.pdf -> hardware/v2/docs/

Left untouched (separate decisions, not part of this cleanup):
  - docs/FPGA-Neural-Flash-Subsystem-Verification.md, docs/
    v2-description.md -- orphaned root-only content, no duplicate
    found anywhere, but also not part of the reviewed plan
  - synth/ecp5/* experiment dirs and sim/*_sim + sim/*.vcd build
    artifacts -- not literal duplicates, flagged as candidates for a
    future, separate cleanup pass

Verified no functional breakage: grepped all remaining scripts/docs
for references to every removed path -- only prose/comment mentions
found, no executable imports or build-script paths broken.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
This commit is contained in:
2026-09-07 05:30:17 +02:00
co-authored by Claude Sonnet 5
parent e1cefd13db
commit 81a9619214
184 changed files with 0 additions and 1023781 deletions
Binary file not shown.
@@ -0,0 +1,186 @@
# FPGA-Neural V2 — Datasheet
**Status: DRAFT / PRE-RELEASE.** This datasheet documents the INTENDED
V2 board architecture as of STEP20. It does **not** certify a finished,
release-ready design — see §11 Limitations and
`hardware/v2/docs/OPEN_ITEMS.md` for the current, real blocker list.
Do not read any statement here as "physically validated" unless it
says so explicitly.
## 1. General
FPGA-Neural V2 is an embedded neural-network accelerator built around a
Lattice ECP5 FPGA and a single external SDRAM. It executes small,
dependency-graph-structured INT8 neural networks (dense layers, DAGs)
using a Neural Multiprocessor of parallel MAC engines, streaming
weight/activation tiles from one external SDRAM chip that also holds
results.
Architecture stack (top to bottom): SPI host interface → job
registration → Dependency Manager / Neural Director → N parallel
Neural Processors → Memory Manager / streaming tile delivery → Unified
SDRAM Backend → one physical SDRAM.
## 2. FPGA
| Item | Value | Basis |
|---|---|---|
| Part | Lattice LFE5U-45F | DESIGN DECISION |
| Package | CABGA381 | DESIGN DECISION |
| Speed grade | -8 | DESIGN DECISION |
| Ordering part number | LFE5U-45F-8BG381C | DESIGN DECISION (standard Lattice ordering suffix for this grade/package; not independently cross-checked against a live distributor listing this session) |
| Logic (post-synthesis, N=4, frozen STEP19 compute core) | TRELLIS_FF=6425, TRELLIS_COMB=6023, MULT18X18D=32, DP16KD=0 | VERIFIED (real Yosys synthesis, STEP19) |
| I/O used (frozen STEP19 top, no physical host bus) | 149/245 TRELLIS_IO | VERIFIED (real nextpnr-ecp5 P&R, STEP19) |
| I/O used (this step's new board-level top, SPI + osc + reset + SDRAM) | not yet synthesized this round | OPEN — see §11 |
Operating assumption: single clock domain, no CDC beyond the SPI
bridge's own double-flop synchronizers and the reset synchronizer
(§5).
## 3. Neural accelerator
| Parameter | Value |
|---|---|
| N_PROCESSORS | 4 (frozen reference; N=2 also validated; N=8 is a future evolution) |
| P_IN (MAC width) | 8 |
| Data representation | INT8 operands |
| Accumulator | INT32, ReLU + INT8 saturate on output |
| MAC architecture | 8-wide parallel MAC, balanced adder tree (`neural_processor.v`, unchanged since before this freeze) |
| Processor parallelism | N independent Neural Processors, one dependency-graph node in flight per processor |
| Supported memory traffic | weights (read-only, 64-bit packed fetch, cached), activations (read, byte-maskable), results (write, byte-maskable) — all through the SAME single SDRAM |
**RTL capability vs. software/API capability:** the RTL executes one
pre-compiled dependency graph (nodes with producer/consumer edges,
fixed tile counts) registered via 108 bits of per-job configuration
(node id, dependency list, activation/weight/result base addresses,
tile count). There is no on-chip graph compiler, no floating point, no
training — job graphs and addresses are computed off-chip and loaded
via the host interface (§9).
## 4. Unified memory
```
┌─────────────────────┐
│ FPGA ECP5 │
│ │
│ 4x Neural Engines │
│ │ │
│ v │
│ Unified SDRAM │
│ Backend / Arbiter │
└─────────┬───────────┘
│ 16-bit SDRAM bus
v
┌─────────────────────┐
│ AS4C4M16SA-6TIN │
│ Weights │
│ Activations │
│ Results │
└─────────────────────┘
```
| Item | Value | Basis |
|---|---|---|
| Device | Alliance Memory AS4C4M16SA-6TIN | DESIGN DECISION (STEP16-19) |
| Capacity | 4M x 16 (8MB) | DATASHEET VALUE |
| Data width | 16-bit (DQ[15:0]) + DQM[1:0] byte mask | DATASHEET VALUE |
| Addressing | BA[1:0] (4 banks) + A[11:0] (row/col, multiplexed) | DATASHEET VALUE |
| Clock | shared with FPGA system clock (§5) | DESIGN DECISION |
| Initialization/refresh | real, RTL-implemented power-up wait + mode-register-set + periodic AUTO REFRESH (`sdram_controller.v`) | VERIFIED (real refresh events observed in simulation, STEP16-19) |
| Arbitration | single physical port, 2-way logical split: W (weight, read-only, cached) / AR (activation+result, read/write, byte-maskable), each internally arbitrated across N processors by a generic, reused `slot_mem_arbiter` | VERIFIED (STEP19 bit-exact regression, reconfirmed via Verilator this step — see errors.log ERR-0024) |
| Official V2 memory map | weights @0x010000, activations @0x200000, results @0x300000, all within the single 8MB space, 1MB-aligned | DESIGN DECISION |
PSRAM is **not** part of V2. The V1 PSRAM controller (`hardware/v1/rtl/psram_controller.v`) is not instantiated anywhere in the V2 physical path.
## 5. Clock / PLL
```
16 MHz OSCILLATOR
|
v
ECP5 PLL (EHXPLLL)
CLKI_DIV=1 CLKFB_DIV=4 CLKOP_DIV=9
FEEDBK_PATH=CLKOP VCO=576MHz
|
v
FPGA SYSTEM CLOCK
64 MHz
(real, tool-generated ratio: 16 * 4 / 1, CLKOP_DIV=9 -> 576/9=64)
```
| Item | Value | Basis |
|---|---|---|
| Oscillator | 16 MHz (board-level, prior project record) | DESIGN DECISION (part number: TBD — not selected this session) |
| PLL primitive | EHXPLLL (`ecp5_pll_sys_clk.v`) | VERIFIED design-time via Project Trellis `ecppll` v1.4 (real tool, real parameters) |
| Generated system clock | 64 MHz | DESIGN DECISION, chosen over 80MHz because STEP19's own multi-seed P&R data showed only 1/8 seeds closing timing at >=80MHz on the compute-only core, and the new board-level top adds more logic still; 64MHz is not yet itself confirmed by P&R on the NEW top (see §11) |
| PLL lock | `locked` output, feeds `reset_sync.v` | DESIGN DECISION; NOT simulatable (Lattice EHXPLLL has no open sim model) — real lock behavior is a real-hardware-only characterization, see §10 |
| Timing constraints | none yet written for the new board-level top | OPEN — see §11 |
## 6. Interfaces
### SPI host interface (`spi_host_bridge.v`)
Mode 0 (CPOL=0/CPHA=0), MSB-first, one opcode per CS-low period.
Opcodes: `0x10` WRITE_JOB (job registration, 15-byte payload), `0x01`
WRITE_MEM / `0x02` READ_MEM (raw, word-addressed SDRAM access via a
second arbitrated port), `0x20` STATUS, `0x0F` RESET. Verified in
isolation (18/18, `tb_spi_host_bridge.v`). **Not yet verified
end-to-end under realistic multi-job pacing** — see §11/ERR-0025.
The 110-pin `reg_*` bus used by V2's own internal simulation
testbenches is a testbench-only convenience and is **not** the
physical interface.
### JTAG
Standard ECP5 JTAG (TDI/TDO/TCK/TMS), always available regardless of
configuration boot mode, per Lattice's own standard requirement.
### Configuration
Standard ECP5 PROGRAMN/INITN/DONE/CCLK. Boot-mode/flash-part decision:
OPEN (see §11).
## 7. Electrical
Rail voltage requirements are DATASHEET VALUEs (from real device
datasheets); no regulator part numbers, current budget, or decoupling
values are finalized this round. Full detail:
`hardware/v2/docs/POWER_ARCHITECTURE.md`.
## 8. Pinout
Full table: `hardware/v2/docs/PINOUT.md`. Summary: 37 real SDRAM
signals + clk/rst are ball-assigned and P&R-verified (STEP19, against
the STEP19 compute-only top). The board-level top added this step
(SPI + oscillator + reset pins) has **not** had its own ball
assignment or P&R run yet.
## 9. Mechanical / board assumptions
None assumed beyond the package footprint implied by CABGA381. No PCB
dimensions, connector placement, or stack-up are specified — that is
schematic/PCB-capture work, not yet started (see
`hardware/v2/docs/SCHEMATIC_READINESS.md`).
## 10. Programming / first power-on
JTAG programming is standard. A first-power-on procedure exists at
`hardware/v2/docs/FIRST_POWER_ON.md` (procedure only — not executed
against real hardware, since no board has been fabricated).
## 11. Limitations (real, current, as of this datasheet's own writing)
- **The physical SPI host interface is NOT proven end-to-end
correct.** A real, disclosed defect (errors.log ERR-0025 Part B)
produces wrong results when two jobs are dispatched with realistic
SPI pacing, even though registration itself is confirmed correct.
This is the single largest open item.
- The board-level top (`fpga_neural_v2_top.v`) has not been through
synthesis or P&R this round — deliberately, since running the real
toolchain against RTL known to compute wrong answers would not be a
meaningful result.
- No PCB, schematic capture, or fabricated hardware exists. Nothing in
this document should be read as "physically validated."
- Regulator, configuration-flash, and connector part numbers are not
selected.
- The STEP19 compute+memory core (raw `reg_*` interface, no SPI
bridge) IS bit-exact verified (N=2 and N=4, 256/256, reconfirmed via
Verilator this session) and remains the actual, working reference
design underneath this datasheet's own described board architecture.
@@ -0,0 +1,160 @@
# FPGA-Neural V2 — Reference Schematic (textual)
**No KiCad schematic was generated this session.** No RTL-to-schematic
or netlist-to-KiCad automation tool is available in this environment,
and the project's own separate, pre-existing KiCad PCB directory
(`FPGA-Neural/FPGA-Neural/FPGA-Neural/`) is an unrelated, independently
tracked project (its own nested `.git`, near-empty as of last check) —
it was not touched, and this document does not assume its contents.
This is a textual/ASCII reference schematic: a real starting point for
PCB capture, not a substitute for one. All ball assignments below are
the real, P&R-verified ones from `hardware/v2/constraints/
v2_unified.lpf` (STEP19) unless marked otherwise.
## 1. Top-level block diagram
```
+---------------------------+
| HOST MCU |
| SPI |
+------------+----------------+
|
v
+----------------------------------------------------------+
| ECP5 FPGA (LFE5U-45F-8BG381) |
| |
| +--------------+ +---------------------------+ |
| | SPI Host |---->| Register / Control | |
| | Bridge | | (job registration) | |
| +--------------+ +-------------+-------------+ |
| | |
| +-------------v-------------+ |
| | Neural Accelerator (N=4) | |
| | Processor 0..3 | |
| +-------------+-------------+ |
| | |
| +-------------v-------------+ |
| | Unified SDRAM Backend | |
| +-------------+-------------+ |
+----------------------------------------------------------+
|
16-bit SDRAM bus
v
+----------------------------+
| AS4C4M16SA-6TIN |
| Weights / Activations / |
| Results |
+----------------------------+
16 MHz osc --> ECP5 PLL (EHXPLLL) --> 64 MHz system clock
Power rails --> POR/supervisor --> FPGA reset, SDRAM init
Configuration flash + JTAG connector (see 5/6)
```
## 2. SDRAM connection table (real, P&R-verified balls)
| Signal | Ball | Bank | I/O std (assumed) | Direction |
|---|---|---|---|---|
| CLK (shared w/ system clk) | H5 | — | LVCMOS33 | FPGA -> SDRAM |
| CKE | B5 | 7 | LVCMOS33 | FPGA -> SDRAM |
| CS_N | C5 | 7 | LVCMOS33 | FPGA -> SDRAM |
| RAS_N | C4 | 7 | LVCMOS33 | FPGA -> SDRAM |
| CAS_N | A3 | 7 | LVCMOS33 | FPGA -> SDRAM |
| WE_N | B3 | 7 | LVCMOS33 | FPGA -> SDRAM |
| BA[0] | E4 | 7 | LVCMOS33 | FPGA -> SDRAM |
| BA[1] | C3 | 7 | LVCMOS33 | FPGA -> SDRAM |
| A[0..11] | D5,D3,F4,E5,E3,F5,A2,B1,C2,C1,D2,D1 | 7 | LVCMOS33 | FPGA -> SDRAM |
| DQ[0..15] | E1,G5,H3,J5,K3,K2,H1,J1,K1,K4,L4,L5,M5,M4,N4,N5 | 7/6 | LVCMOS33 | bidirectional |
| DQM[0..1] | P5,N3 | 6 | LVCMOS33 | FPGA -> SDRAM |
Full source: `hardware/v2/constraints/v2_unified.lpf`. LVCMOS33 is
assumed to match the SDRAM's own real 3.3V requirement and matches
banks 6/7's real VCCIO range per `docs/pinouts.md` — not yet
independently cross-checked at the schematic/PCB level (WARNING, not
BLOCKER).
## 3. Clock schematic
```
16MHz OSC ---> CLKI (H5, reused from V1's own real LPF)
|
+-----v------+
| EHXPLLL | CLKI_DIV=1, CLKFB_DIV=4, CLKOP_DIV=9
| (hard IP) | FEEDBK_PATH=CLKOP, VCO=576MHz
+-----+------+
| CLKOP = 64MHz
v
FPGA system clock (feeds compute, SDRAM ctrl, SPI bridge)
|
+-----v------+
| reset_sync | <-- ext POR (active-low) + PLL LOCK
+-----+------+
v
rst (sync-deassert, feeds every synchronous block)
```
Oscillator part number: **TBD** (not selected this session — a real
16MHz, 3.3V HCMOS clock oscillator in a standard SMD package is the
intended class of part; no specific manufacturer/part number is
claimed without a real datasheet lookup performed this session).
## 4. Power schematic (rails only — no regulator parts selected)
```
3.3V/1.1V/2.5V rails (regulators: TBD)
| | |
v v v
VCCIO VCC(core) VCCAUX
(banks (real ball (real ball
6/7=SDRAM cluster, cluster,
I/O, etc) see see
POWER_ARCH POWER_ARCH
.md) .md)
|
v
SDRAM VDD/VDDQ (3.3V, DATASHEET VALUE per AS4C4M16SA-6TIN)
```
Full rail table, decoupling guidance, and current-budget status:
`hardware/v2/docs/POWER_ARCHITECTURE.md` (unchanged this step — no new
power work performed).
## 5. Configuration / JTAG schematic
```
FPGA
|-- TDI/TDO/TCK/TMS --> JTAG connector (standard pinout, always
| available regardless of boot mode)
|-- PROGRAMN/INITN/DONE/CCLK --> configuration flash (part: TBD) or
JTAG-only bring-up (decision: OPEN)
```
No configuration-flash part has been selected; JTAG-only bring-up
remains a valid fallback and is documented as such in
`hardware/v2/docs/CONFIGURATION.md`-equivalent content inside
`OPEN_ITEMS.md` (a dedicated `CONFIGURATION.md` was not created this
round — tracked as an open item, not silently dropped).
## 6. Host interface schematic
```
Host MCU --SPI--> FPGA: spi_sclk, spi_mosi, spi_miso, spi_cs_n
```
No ball assignment exists yet for these 4 signals (the board-level
top was not run through P&R this session — see the datasheet's own
§11 Limitations). Pull-up on `spi_cs_n` (idle-high) is the standard,
expected design decision for a single-master SPI bus; not yet placed
in any real LPF.
## 7. What this schematic deliberately does NOT claim
- No KiCad artifact. No PCB. No fabricated board.
- No ball assignment for the new SPI/oscillator/reset pins (P&R not
run against the new board-level top this session, since the design
has a known, unresolved functional defect — see errors.log
ERR-0025 Part B).
- No regulator, flash, or connector part numbers.
This document is a real, honest starting point for PCB capture, not a
finished schematic.
@@ -0,0 +1,92 @@
# NMS Activation Fill Controller Timing (STEP14 Part B)
Status: fixed, real post-P&R verified, bit-exact, adopted. Full data:
`hardware/v2/reports/step14_activation_timing.csv`. Full narrative:
`hardware/v2/logs/experiments.log` (EXP-0029, 0030, 0031),
`decisions.log` (DEC-0026, DEC-0027).
## B1 — Exact critical path (not assumed)
Mined directly from the real nextpnr-ecp5 P&R report for
`nms_neural_multiprocessor_stream.v` at N_SLOTS=4
(Fmax=55.22 MHz, FAIL @ 80 MHz). Full path, 18.11 ns total (6.25 ns
logic + 11.85 ns routing):
```
SOURCE: u_act_fill.resident_tag[11] (register Q)
-> COMBINATIONAL, chained, NO register in between:
(1) max_n_tiles computation, nms_activation_fill_ctrl.v:92
(N_SLOTS-wide running-max fold, each iteration gated by a
23-bit tag-equality check) -- long CCU2C carry chain
(2) resident_count < max_n_tiles comparison, line 165
(the ST_IDLE refill/continue decision) -- ANOTHER 16-bit
magnitude-comparison carry chain, feeding directly off (1)
in the SAME cycle
(3) into pf_start's own next-state logic
DESTINATION: u_act_fill.pf_addr's clock-enable (CE) pin
```
Two full 16-bit magnitude comparisons sit in **one** combinational
cone across **one** clock edge. This confirms, at the exact RTL-line
level, the failure class DEC-0016/EXP-0022 predicted analytically
("O(N_SLOTS) unpipelined combinational scan feeding directly into a
control decision") — but precisely localizes it to the comparison
logic (lines 92 and 165), *not* the priority-encoder
(`desired_valid`/`desired_x_base`, lines 77-86), which does not appear
in this critical path at all.
## B2 — Scaling behavior
The bottleneck is the `max_n_tiles` running-max fold: an imperative
`for` loop creates a data dependency between iterations (`max_n_tiles`
after iteration *i* depends on iteration *i-1*), which Yosys
synthesizes as a sequentially-chained carry structure — inherently
O(N_SLOTS) deep, not O(log N_SLOTS). At N_SLOTS=4 the chain reached
6.25 ns logic + 11.85 ns routing; at N_SLOTS=8 it doubles again (see
below).
## B3 — Minimum fix (two iterations, evidence-driven)
**v2** (one pipeline stage: register `max_n_tiles` before its use in
the `resident_count` comparison): Fmax 55.22 → 72.78 MHz (+31.8%) —
real improvement, still fails 80 MHz. Re-tracing showed the *remaining*
critical path was entirely inside `max_n_tiles`'s own computation
(now feeding its own register), confirming the fix needed to go one
level deeper.
**v3** (second stage: register each slot's tag-equality/masking result
first — independent per-slot work, no N_SLOTS-dependent chain — *then*
fold the already-registered, already-masked values): Fmax 55.22 →
**106.81 MHz** (+93.4%). **PASSES** 80 MHz with real margin. Resource
cost: LUT4 -5.5%, FF +1.4% (2 added pipeline registers), CCU2C
unchanged.
## B4 — No serialization reintroduced
Verified directly: N_SLOTS=2 bit-exact regression test (D-Stress, real
V1 PSRAM chain) gives **numerically identical** cycle count and
sustained MAC/cycle before and after the fix (185270/185270 cycles,
0.1769/0.1769 MAC/cycle). The 3 total cycles of added latency apply
only to the rare, tile-refill-boundary-only decision — never to the
real-time per-tile consumption path (already fully decoupled by
STEP13's own streaming manager). Higher Fmax, zero throughput cost —
satisfying B4's explicit requirement.
## N=8 (exploratory)
`nms_activation_fill_ctrl_v3.v` at N_SLOTS=8: DSP=64/72 (89%, FEASIBLE),
LUT4=4653, FF=10855 (both comfortably FEASIBLE). **Fmax=52.25 MHz,
FAILS 80 MHz** — the v3 fix's second stage (the max-fold itself) is
still O(N_SLOTS)-deep; at N=8 it is twice as deep as at N=4 and becomes
dominant again. This is expected: v3 shifted the crossover point, it
did not eliminate the underlying dependency. A genuine balanced-tree
reduction (or a pipeline scaling with log₂(N_SLOTS) rather than a flat
2-stage split) would be required for N=8 — not undertaken this round
(N=8 is explicitly exploratory; the limiting resource (Fmax, not
DSP/LUT/FF/BRAM) is precisely identified and quantified, per spec).
## Adoption
`nms_activation_fill_ctrl_v3.v` is adopted as the reference activation
fill controller for N_SLOTS≥4 configurations (DEC-0027). The original
and the insufficient v2 are preserved for reference.
@@ -0,0 +1,122 @@
# NMS Continuous Tile Stream — Memory Manager Redesign (STEP13)
Status: implemented, bit-exact verified, synthesized. **Adopted** as
the new reference NMS memory-manager configuration (DEC-0025). Full
data: `hardware/v2/nms/reports/batch_processor_{sweep.csv,summary.md}`.
Full narrative: `hardware/v2/logs/experiments.log` (EXP-0025 through
EXP-0028), `decisions.log` (DEC-0024, DEC-0025).
## Why this file is not `neural_processor_batch.v`
The governing brief for this STEP asked for a "batch/continuous
neuron execution model" — multiple neurons processed per dispatch, or
a continuous neuron stream — to amortize the ~68.5-cycles/neuron
non-memory floor found in EXP-0024. Before writing any RTL, Step 1
required tracing the actual RTL to find exactly where those cycles
go, rather than assuming.
That trace (EXP-0025, an isolated testbench with `neural_processor.v`
+ `nms_memory_manager_pf.v` driven with zero real memory latency
anywhere) found: **93.4% of the floor is explained by a 4-cycles/tile
serialization bug inside the memory manager's own `ST_RUN` state**,
not by per-job dispatch overhead (only 6.6%). `ST_RUN` implements
operand delivery as a strictly sequential chain —
`read_issued → read_ready → present → consumed` — with zero overlap
between consecutive tiles, even though:
- the local activation/weight SRAMs (`nms_activation_replicated.v`,
`nms_weight_packed.v`) have only a 1-cycle `rd_en`-to-data latency;
- `neural_processor.v`'s own `operand_ready` is held continuously high
through the whole tile-loading phase — its datapath is explicitly
designed (per its own header comment) to accept a new tile every
cycle while previous tiles drain through the adder tree/accumulator.
Neither side of this interface requires 4 cycles/tile. It is purely
an artifact of the memory manager's own un-pipelined FSM. **The fix is
therefore a continuous per-tile streaming redesign of the memory
manager, not a neuron-batching scheme — hence
`nms_memory_manager_stream.v`, not `neural_processor_batch.v`.**
`neural_processor.v` itself required no modification.
## Design: `nms_memory_manager_stream.v`
Drop-in replacement for `nms_memory_manager_pf.v` (identical external
interface, same `weight_prefetch_engine.v` instance, same outer job
FSM `ST_IDLE`/`ST_WAIT_RESULT`/`ST_WRITE_RES`/`ST_DONE`). Only
`ST_RUN`'s internal operand-delivery logic differs:
- `rd_ptr` — the read-**issue** pointer (which tile's SRAM read has
been, or is about to be, issued), independent of and normally one
tile ahead of `tile_idx` (the **consumption** pointer, i.e. how many
tiles `neural_processor.v` has actually accepted).
- A 1-deep skid buffer (`buf_valid`/`buf_input`/`buf_weight`/
`buf_last`) holds one tile's fully-read SRAM data, presented to NP
as `operand_valid`/`input_data`/`weight_data`/`tile_last`.
- Every cycle: if a read issued last cycle is landing now (1-cycle
SRAM latency), it's captured into the skid buffer; independently, a
new read is issued for `rd_ptr` whenever legal (in bounds, weight +
activation ready) **and** the buffer will not overflow (empty, or
being drained this same cycle).
Since `operand_ready` stays high throughout the tile-loading phase,
the skid buffer drains every cycle it's full, so a new read can be
issued every cycle too — sustained ~1 cycle/tile, down from 4.
`tile_idx` (the consumption pointer) is still what feeds
`weight_prefetch_engine.v`'s own `consumed_count` port — its external
contract is unchanged; only the local SRAM read-issue pointer
(`rd_ptr`) is new, and it can run up to one tile ahead of `tile_idx`
(the skid buffer's own depth).
## Verification chain (all real, none assumed)
1. **EXP-0025**: isolated zero-latency trace of the *old* design —
established the 4-cycles/tile floor and its 93.4% share of
EXP-0024's real measured floor.
2. **EXP-0026**: same isolated trace against the *new* design — the
fix works exactly as designed (confirmed cycle-by-cycle), but
total cycles barely move (81→80), because it immediately hits a
*second*, previously-masked bottleneck: `weight_prefetch_engine.v`'s
own word-fetch rate is *also* exactly 4 cycles/tile (P_IN=8 bytes ÷
16-bit bus = 4 word-transactions, 1 cycle/word minimum even at
zero real latency) — a bus-**width** ceiling, structurally
different from an FSM-serialization ceiling, that happens to
coincide numerically today.
3. **EXP-0027**: a direct control experiment — a scratch variant with
weight-fetch bypassed (always-ready) isolates the new design's
*own* ceiling: a clean 1 cycle/tile (100% of `neural_processor.v`'s
theoretical per-tile rate), vs. the old design's hard 4-cycles/tile
cap under the identical bypass. This is the direct proof that the
fix removes a real, 4× architectural ceiling — it was just masked
by a coincidentally-equal second bottleneck.
4. **EXP-0028**: full real-system integration
(`nms_dataflow_core_stream.v``nms_neural_multiprocessor_stream.v`,
real V1 PSRAM chain) — bit-exact PASS, 256/256 neurons, D-Stress
workload identical to EXP-0022/0024. Real cycle count: 185270 vs.
185398 (`_pf` baseline), -0.07% — confirms the "masked, zero net
benefit today" prediction exactly. Real synthesis + P&R: N=1
Fmax=142.92 MHz (+3.7% vs. baseline), N=2 Fmax=92.57 MHz (-2.8%,
still comfortably above 80 MHz), resource cost within ±6%. N=4:
55.22 MHz, FAILS 80 MHz — but for the *pre-existing*,
already-documented `nms_activation_fill_ctrl.v` priority-scan
regression (EXP-0022), unrelated to and unaffected by this fix.
## Outcome and adoption
**Outcome B** (helps, but another bottleneck appears — see
DEC-0025 and `batch_processor_summary.md` for the full nine-question
final decision). `nms_memory_manager_stream.v` is adopted as the new
reference configuration: it is a strict improvement (bit-exact,
resource-neutral, no measured downside) and is **required groundwork**
for any future PSRAM bandwidth increase to actually translate into a
throughput gain — without it, a wider/faster memory would immediately
hit the old FSM's 4-cycles/tile ceiling and realize only 25% of its
potential benefit. The original `nms_memory_manager.v` and
`nms_memory_manager_pf.v` remain preserved, unmodified, for A/B/C
reference. Neuron-batching (the brief's original Model B/C) was not
pursued — evidence showed it addresses only 6.6% of the real floor and
would deliver no measurable benefit today for the identical reason
(weight-fetch-rate-bound). N=4/N=8 viability remains blocked by two
independent issues neither addressed by this STEP: external PSRAM
bandwidth, and the activation fill controller's own Fmax regression —
both flagged as future work.
@@ -0,0 +1,145 @@
# NMS Real Weight Prefetch Engine (STEP11)
Status: implemented, bit-exact verified, benchmarked against the real
V1 PSRAM chain, synthesized. **Not adopted as the default NMS
configuration** — see Outcome/Recommendation below. Full data:
`hardware/v2/nms/reports/nms_prefetch_sweep.csv`,
`nms_prefetch_summary.md`; full narrative:
`hardware/v2/logs/experiments.log` (EXP-0023, EXP-0024),
`decisions.log` (DEC-0023), `errors.log` (ERR-0015).
## Problem
The "Current NMS" baseline (`nms_memory_manager.v`, backed by
`prefetch_engine.v`) measured `prefetch_effectiveness≈0%` and
`weight_stall≈92.5%` at N_SLOTS=2 (EXP-0022). Tracing the actual RTL
(not assuming from filenames) showed the real gap: `prefetch_engine.v`
is a single-shot FSM (`ST_IDLE`/`ST_READ_W`/`ST_DONE`) that can only
have **one fetch in flight at a time**, and `nms_memory_manager.v`'s
own restart logic only re-triggers the next tile's fetch once the
*previous* tile's fetch has fully completed and the FSM has returned
to idle — paying a real per-tile control-plane restart cost on every
tile boundary. The gap was never insufficient lookahead *distance*
(the old design already tried to fetch as far ahead as `n_tiles`
allowed); it was zero *outstanding-request depth*.
## Real backend constraint
`memory_interface.v``psram_controller.v` (V1, reused verbatim,
never modified) is a fire-and-forget, **one-transaction-in-flight**
protocol: a single `mem_req` pulse, wait for `mem_ready`, and that IS
the whole transaction. No wire-level pipelining is physically possible
against a real single PSRAM port. So "multiple outstanding requests"
cannot mean multiple simultaneous word transactions — it means
eliminating the *control-plane* overhead paid at every tile boundary
and letting the fetch stream run continuously across tiles, queueing
up to `PREFETCH_DISTANCE` tiles of lookahead ahead of consumption.
## Design: `weight_prefetch_engine.v`
Two monotonic counters fully describe the engine (tiles are always
fetched in strict sequential order, never reordered or re-fetched, so
no per-tile state array is needed):
- `fetch_tile`/`fetch_word` — the next word to request (or the word
currently in flight).
- `ready_count` — tiles 0..`ready_count`-1 are fully resident in the
weight SRAM.
`consumed_count` (the consumer's own tile index, `nms_memory_manager_pf.v`'s
`tile_idx`) bounds a configurable lookahead window:
`window_limit = consumed_count + PREFETCH_DISTANCE`; the engine may
fetch tile K only if `K < n_tiles` **and** `K < window_limit`.
The core mechanism: on `mem_ready && req_outstanding`, the just-completed
word is committed **and**, in the same cycle, the very next request is
issued — either the same tile's next word, or (at a tile boundary) the
next tile's first word — giving zero-gap streaming across tile
boundaries against a backend that only ever has one word in flight.
(An earlier draft used mutually-exclusive `if/else-if` branches for
"commit" vs. "issue next", which reintroduced a 1-cycle gap between
*every* word, not just tile boundaries; fixed by merging both into one
branch — see `weight_prefetch_engine.v`'s own header comment.)
## Integration: the "_pf" A/B variants
Per the explicit "preserve the current NMS baseline" constraint, the
new engine was integrated into parallel `_pf`-suffixed files, leaving
the originals untouched:
- `nms_memory_manager_pf.v` — drop-in replacement for
`nms_memory_manager.v`'s external interface; internally swaps the
private `prefetch_engine.v` instance for `weight_prefetch_engine.v`,
and changes `can_present`'s weight-ready check from
`tile_idx < wgt_fetched` to `tile_idx < wgt_ready_count`.
- `nms_dataflow_core_pf.v` — mirrors `nms_dataflow_core.v`, adds a
`PREFETCH_DISTANCE` parameter, instantiates `nms_memory_manager_pf`.
- `nms_neural_multiprocessor_pf.v` — mirrors
`nms_neural_multiprocessor.v`, instantiates `nms_dataflow_core_pf`.
Both the baseline (`nms_neural_multiprocessor.v`) and the prefetch
variant (`nms_neural_multiprocessor_pf.v`) remain in the repository
side by side; neither supersedes the other.
## Verification
`hardware/v2/nms/sim/tb_weight_prefetch.v` — isolated correctness
testbench: real `sim_word_mem` (configurable extra latency), real
`nms_weight_packed.v` production SRAM, bit-exact fill-pattern checking.
Covers `n_tiles ∈ {0,1,2,PFD,PFD+1,MAX_TILES-1,MAX_TILES}`, back-to-back
jobs with no explicit reset, a dedicated windowing-cap test (frozen
consumer, confirms `ready_count` stops exactly at
`min(PFD,MAX_TILES)`), and (post-ERR-0015) a large-PFD regression case.
10/10 (9/9 at PFD≥MAX_TILES) tests pass bit-exact across
PFD∈{1,2,4,8,32} and under injected extra memory latency.
`hardware/v2/nms/sim/tb_nms_dstress_pf.v` — full real-integration
benchmark: identical D-Stress workload/golden-model/correctness
criteria as `tb_nms_dstress.v` (EXP-0022), instantiating
`nms_neural_multiprocessor_pf` with a `PFD_CFG` parameter, plus new
testbench-only instrumentation for `weight_stall_cycles` and
`prefetch_effectiveness` (tiles consumed with zero weight-blocking
cycles beforehand / total tiles consumed — the exact STEP11
definition). All runs pass 256/256 neurons bit-exact vs. the golden
model.
## ERR-0015: a real bug found and fixed
The initial `window_limit` computation truncated the
`PREFETCH_DISTANCE` *parameter itself* to `CNTW` bits
(`PREFETCH_DISTANCE[CNTW-1:0]`) before adding it to `consumed_count`.
At `MAX_TILES=16` (`CNTW=5` bits), `PFD=32` truncates to 0, making
`window_limit == consumed_count` forever and deadlocking the engine
completely (0/256 neurons ever completed, 0% PSRAM utilization).
Fixed by computing `window_limit` and its comparisons in a fixed
32-bit width, using the untruncated parameter value. Regression-tested
in `tb_weight_prefetch.v`. Full writeup: `errors.log` ERR-0015.
## Results and outcome
See `nms_prefetch_summary.md` for the full comparison table and the
nine explicitly-answered final-report questions. In short:
- **N_SLOTS=1** (no port contention): a real, reproducible **-10.3%**
cycle-count improvement (PFD=1 → PFD≥2), then a complete plateau —
deeper buffering gives zero further benefit. Sustained MAC/cycle
reaches only 2.8% of the theoretical target.
- **N_SLOTS=2** (this project's own primary reference configuration,
real shared-port contention via `slot_mem_arbiter`): **zero
measurable benefit** at any PREFETCH_DISTANCE from 1 to 16 — all
runs are statistically indistinguishable from each other and from
the pre-STEP11 baseline. The single physical PSRAM port is already
saturated (90.5% busy, unchanged from baseline) by natural two-slot
contention before any lookahead scheme can act.
**Final decision: Outcome B (N_SLOTS=1, partial) / Outcome C
(N_SLOTS=2, failure against the 90% criterion).** The mechanism is
correct and does measurably hide latency when the port has spare
capacity; it cannot manufacture bandwidth out of an already-saturated
single physical port. Reaching the STEP11 target would require ~36×
(N=1) to ~82× (N=2) more real PSRAM bandwidth — a hardware-level
constraint, not an RTL-scheduling one. Per DEC-0023, the new engine is
**not** recommended as the default NMS configuration; both variants
are preserved for reference. The evidence-backed next step (real PSRAM
bandwidth — wider bus, multiple independent banks, or a faster backing
technology) is flagged as future work, not undertaken this round.
@@ -0,0 +1,91 @@
# NMS Weight Datapath Scaling (STEP14 Part A)
Status: architectural requirement established and proven (simulation),
**not realizable on real hardware today** (fixed 16-bit physical
PSRAM). Full data: `hardware/v2/reports/step14_weight_scaling.csv`.
Full narrative: `hardware/v2/logs/experiments.log` (EXP-0032, EXP-0033),
`decisions.log` (DEC-0028).
## Question answered
*At what weight-path width does the processor stop being fundamentally
starved by weight delivery?* **64 bits** — exactly `P_IN × DATA_WIDTH`
(8 × 8). Proven by direct cycle-exact simulation, not assumed.
## What was built
`weight_prefetch_engine_wide.v` — a parameterized (`MEM_DATA_WIDTH`)
generalization of the real `weight_prefetch_engine.v`'s continuous
cross-tile-boundary streaming design, simulation-only/exploratory
(same status as `ideal_memory_model.v`). `WORDS_PER_TILE =
ceil(P_IN*DATA_WIDTH / MEM_DATA_WIDTH)`, clamped to a minimum of 1.
`nms_memory_manager_stream_wide.v` pairs it with STEP13's own streaming
memory manager unchanged (A2's requirement), on a *separate* logical
wide port from the real 16-bit result-write-back port.
A real bug was found and fixed during development: address stepping
initially used `WORDS_PER_TILE × BYTES_PER_WORD` as the inter-tile
byte stride, which over-counts whenever the bus is wider than one full
tile (the 128-bit case, `WORDS_PER_TILE=1` but `BYTES_PER_WORD=16`
while the tile itself is only 8 bytes) — this skips over the next
tile's actual data in the packed backing store. Fixed by defining
`TILE_BYTES = TILE_BITS/8` as the canonical, width-independent stride.
## Results (bit-exact + ideal-memory cycle count)
| Width | Words/tile | Steady-state cycles/tile | 16-tile job total |
|---|---|---|---|
| 16-bit | 4 | 4 | 80 |
| 32-bit | 2 | 2 | 48 |
| **64-bit** | **1** | **1** | **32** |
| 128-bit | 1 | 1 | 32 |
All four widths pass bit-exact correctness (9/9 tests each, including
under injected extra memory latency). 64-bit achieves a clean,
cycle-exact **1 cycle/tile** — 100% of `neural_processor.v`'s own
theoretical per-tile acceptance rate, exactly matching the streaming
memory manager's own ceiling (EXP-0027, STEP13). 128-bit gives **zero**
further benefit: a bus wider than one full tile still delivers exactly
one tile per transaction in this single-tile-per-request design (no
multi-tile bursting was attempted).
## The critical distinction: logical vs. physical bandwidth (A5)
STEP14 explicitly warned against assuming a wider logical interface
means the real memory can deliver it. It cannot, here: **the real V1
PSRAM chain is fixed at 16 bits** — a real chip
(ISSI IS66WVE4M16EBLL-70BLI, x16), not an RTL parameter. The
already-existing, already-verified `weight_prefetch_engine.v` (real,
used throughout STEP11-13) *is* exactly what a "64-bit logical / 16-bit
physical" packing adapter would produce: it assembles one 64-bit
logical tile from 4 real sequential 16-bit word transactions. Its real,
repeatedly-measured result is 4 cycles/tile — identical to the ideal
16-bit row above, because the real transaction count is unchanged
regardless of what the logical interface upstream claims. **A logical
wide interface backed by a physically-narrow bus delivers exactly the
narrow bus's own throughput.** No new "packing adapter" module was
built for this reason — the real engine already demonstrates the
answer, conclusively, without further RTL.
## Answer to the primary research questions
- **Is 16-bit weight delivery fundamentally insufficient for P_IN=8?**
Yes — it costs 4 cycles/tile, 4× the achievable minimum.
- **Is 32-bit enough?** No — still 2× the achievable minimum (2
cycles/tile).
- **Is 64-bit the natural architectural point?** Yes, exactly — proven
cycle-exact, not approximate.
- **Does wider logical delivery actually improve real throughput?**
**Not on this hardware.** Realizing the 64-bit ideal requires a
matching *physical* bandwidth increase (a real 64-bit-wide external
bus, or multiple parallel 16-bit PSRAM chips banked together) — a
board/silicon-level change, outside this project's own RTL scope.
## Recommendation
The 64-bit requirement is now precisely quantified and should inform
any future hardware revision (wider PSRAM, multiple banks). No RTL
change is warranted on the current board: `weight_prefetch_engine.v`
(real, 16-bit) remains the correct, already-optimal implementation
given the fixed physical bus width — STEP13's streaming-manager fix
already extracts everything available from the real interface.
+107
View File
@@ -0,0 +1,107 @@
Perfetto, il CSV è quello giusto: FPGA-SC-02034-3-0-ECP5U-45-Pinout.csv, rev. 3.0. L'ho letto e possiamo finalmente fare il pinout esatto per il LFE5U-45F caBGA381, senza deduzioni.
Ho anche verificato la documentazione Lattice aggiornata: per LFE5U VCC = 1,1 V ±5%, VCCAUX = 2,5 V ±5%, VCCIO da 1,2 a 3,3 V e VCCIO8 è il banco associato alla configurazione.
Pinout essenziale BG381
Ball Funzione Alimentazione / collegamento
H13 VCC 1.1 V
J13 VCC 1.1 V
K13 VCC 1.1 V
L13 VCC 1.1 V
M13 VCC 1.1 V
N13 VCC 1.1 V
H12 VCC 1.1 V
N12 VCC 1.1 V
H11 VCC 1.1 V
N11 VCC 1.1 V
H10 VCC 1.1 V
N10 VCC 1.1 V
H9 VCC 1.1 V
N9 VCC 1.1 V
H8 VCC 1.1 V
J8 VCC 1.1 V
K8 VCC 1.1 V
L8 VCC 1.1 V
M8 VCC 1.1 V
N8 VCC 1.1 V
F6 VCCAUX 2.5 V
P6 VCCAUX 2.5 V
F15 VCCAUX 2.5 V
P15 VCCAUX 2.5 V
F9, F10 VCCIO0 Bank 0
F11, F12 VCCIO1 Bank 1
H14, H15, J15 VCCIO2 Bank 2
L14, L15, M15 VCCIO3 Bank 3
L6, L7, M6 VCCIO6 Bank 6
H6, H7, J6 VCCIO7 Bank 7
P9, P10 VCCIO8 Bank 8 / configuration
Nota: il CSV elenca anche i pin VSS/VSSIO; vanno tutti collegati correttamente al piano GND, senza lasciarne nessuno scollegato.
JTAG / programmazione
Questi sono particolarmente importanti perché sono ball esatti del BG381:
Ball Pin Bank Collegamento
R5 TDI 40 JTAG TDI
T5 TCK 40 JTAG TCK
U5 TMS 40 JTAG TMS
V4 TDO 40 JTAG TDO
W3 PROGRAMN 8 controllo configurazione
V3 INITN 8 stato configurazione
Y3 DONE 8 configurazione completata
Il bank 8 è quello associato all'interfaccia di configurazione; Lattice indica esplicitamente VCCIO8 come alimentazione da dimensionare in funzione dell'interfaccia di configurazione utilizzata.
CCLK
Dal CSV:
Ball Pin Bank
U3 CCLK 8
Attenzione però: CCLK non è automaticamente il clock di sistema della nostra rete neurale. È il clock associato alla configurazione; il clock operativo della FPGA va identificato separatamente nel percorso dell'oscillatore/PLL.
Decoupling
Qui preferisco essere molto preciso: non voglio inventare una quantità di condensatori per “ogni VCC”.
La Hardware Checklist ufficiale Lattice richiede esplicitamente:
VCC 1,1 V con PCB plane e adeguato decoupling;
VCCAUX 2,5 V;
VCCIO 1,23,3 V;
VCCIO8 coerente con l'interfaccia di configurazione.
Per la nostra scheda farei quindi una rete di decoupling distribuita, non un condensatore unico per ogni ball:
LFE5U-45F
┌──────────────┐
1V1_CORE ─────┤ VCC │
│ │
2V5_AUX ──────┤ VCCAUX │
│ │
3V3/2V5/etc ──┤ VCCIOx │
│ │
GND ──────────┤ VSS │
└──────────────┘
con i condensatori fisicamente distribuiti attorno al BGA, con quelli ad alta frequenza il più vicino possibile alle rispettive vie/pad di alimentazione.
Una cosa importante
Adesso che abbiamo il CSV ufficiale, posso fare un lavoro ancora più utile: estrarre tutti i 381 ball del BG381 e generarti una tabella Markdown completa divisa in:
VCC
VCCAUX
VCCIO08
VSS/VSSIO
JTAG
PROGRAMN/INITN/DONE
clock
GPIO utilizzabili
pin da NON utilizzare
BOM decoupling consigliata
header JTAG completo
Quella diventerebbe praticamente la nostra scheda di riferimento hardware per il PCB V2.