Files
FPGA-Neural/CLAUDE.md
T
micheleandClaude Sonnet 5 43a12379a5 feat: MILESTONE - real activation-fetch engine, full N=2 system verified on real DDR3 (EXP-0079)
Closes the last major disclosed functional gap: packed_slot.v's
activation data was read through a combinational stand-in since
EXP-0062. New act_tile_fetch.v reads activation tiles directly from
DDR3 (no on-chip buffering needed, unlike weights -- activation data
has no reuse), sharing each slot's existing ctrl port with its own
weight-prefetch engine. Real memory layout: one full BURST_LEN=8-word
burst per tile, deliberately avoiding any runtime-indexed part-select
given this project's thin P&R timing margin (EXP-0078).

Verified at three levels: act_tile_fetch.v alone (6/6), packed_slot.v
with real preloaded activation data (9/9), and the full N=2 system
against real DDR3 via xsim (8/8, 0 errors) -- the first time this
project's compute path has been verified end-to-end with real DDR3
for both weights and activations.

Retired hardware/v3/rtl/n2_system_top.v and its testbench (pre-DDR3
SDR-placeholder era, fully superseded by n2_system_ddr3_top.v).

Also: docs/PHYSICAL_REALIZATION.md (real pinout/parts/timing/protocol
reference for the physical board) and CLAUDE.md (persistent project
instructions for future Claude Code sessions).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MUG92aM9m68TRc4rG55BcC
2026-09-20 08:13:15 +02:00

102 lines
6.1 KiB
Markdown

# FPGA-Neural — project instructions for Claude Code
Hardware neural accelerator for a custom PCB: bare **Xilinx XC7A100T-CSG324-2**
(Artix-7) chip + real DDR3, designed and assembled by the user themselves —
never a Digilent/dev-board purchase. An ESP32 is the host/central processor,
talking to the FPGA over a dedicated SPI bus (FPGA is slave there) and, via
the FPGA, through to a separate config flash used only for FPGA bootstrapping
(FPGA is master on that second, physically distinct SPI bus).
Active branch: **`v3-artix7`**. `hardware/v3/` is the current, real target.
`hardware/v2/` is the archived ECP5 baseline (frozen, DSP-count-limited,
superseded — do not build on it, some of its RTL is still *reused*
unmodified by v3, e.g. `layer_prefetch_ctrl.v`/`layer_weight_buffer.v`).
`hardware/v1/` is older still, reference only.
## Read first
- `docs/PHYSICAL_REALIZATION.md` — every real pin assignment, part number,
timing number, and memory-layout convention needed for the physical board
and for host (ESP32) firmware. Keep it in sync with reality — if a pin
assignment or timing number changes, update this file in the same commit.
- `hardware/v2/logs/experiments.log` — the real project history, one
`EXP-NNNN` entry per real experiment/change (context/method/result/
decision/next_action). Read the tail before starting new work; append a
new entry for anything non-trivial you do. This log — not memory, not
chat history — is the authoritative record of what's been tried and why.
## Toolchains (real paths, already working — don't re-diagnose from scratch)
- **Vivado 2026.1**: `source /home/michele/tools_cache/Xilinx/2026.1/Vivado/settings64.sh`,
then `export LD_LIBRARY_PATH="/home/michele/tools_cache/Xilinx/2026.1/Vivado/lib/lnx64.o/Ubuntu/24:$LD_LIBRARY_PATH"`
(this machine's Ubuntu is too new for Vivado's own OS detection; the
LD_LIBRARY_PATH points at Vivado's own bundled compat libs — not a real
distro package, must be set every session).
- **OSS CAD Suite** (iverilog/vvp for fast plain-Verilog sims, no Xilinx
primitives): `source /home/michele/tools_cache/oss-cad-suite/environment`.
- Real Vivado project: `Vivado/NeuralProcessor/NeuralProcessor.xpr` — the
MIG DDR3 IP lives there, real, already generated for the exact part.
## Hard-won lessons (do not re-derive these the slow way)
- **Vivado's imported source copies go stale silently.** If a project file
under `NeuralProcessor.srcs/sources_1/imports/...` was ever edited on disk
*after* being added to the project, diff it against the live
`hardware/v3/...` source before trusting any P&R result — `add_files`/
`update_compile_order` do NOT auto-refresh it, and a stale copy produces
no error, just silently wrong (old) synthesis results (EXP-0078). Prefer
adding new files so they stay a direct reference (check `IS_GLOBAL_INCLUDE`/
the file's own path isn't under `imports/`) rather than get copied.
- **Testbench stimulus must use nonblocking assignment (`<=`), not blocking
(`=`), when driving a DUT's inputs from a separate `always`/`initial`
block.** Blocking assignment races the DUT's own `posedge`-triggered
always block under Icarus and can silently corrupt data OR miss a one-shot
pulse entirely (causing a real hang) — hit and fixed repeatedly (EXP-0073,
0075, 0077) before this became standing practice. If a new Icarus
testbench shows shuffled/duplicated fields or an inexplicable hang,
suspect this class of bug before assuming the RTL is wrong.
- **Never use a runtime-indexed part-select** (`data[idx*W +: W]` where `idx`
is a signal, not a constant) on a wide bus in anything synthesizable — a
known real Fmax killer (`weight_tile_gather.v`'s own header, EXP-0061).
Use fixed shift-concat, or (as `act_tile_fetch.v` does, EXP-0079) design
the memory layout so only a fixed slice is ever needed. The current real
P&R timing margin is thin (WNS +0.013ns, EXP-0078) — there is no slack to
absorb a new critical path.
- **A one-shot-pulse requester on a shared/arbitrated bus must see its own
grant the SAME cycle its own `active` signal first asserts** — a
registered/one-cycle-late grant silently loses the request forever
(EXP-0066's own real bug, now a standing design rule for every arbiter/
requester pair in this project).
- **`xvlog`/`iverilog` need `-sv`/`-g2012`** respectively to accept
SystemVerilog-only syntax (e.g. `'0`) even in a plain `.v` file — prefer
just not using SV-only syntax in synthesizable RTL (Vivado's `synth_design`
has no such escape hatch at all).
- **Verify real component availability (LCSC) before committing to a part**
— the user has asked for this explicitly more than once. Don't guess
availability or specs from training data; search when it matters.
- **Real, measured numbers only — never estimate/guess a timing or
performance figure and present it as fact.** Out-of-context synthesis is
not a real signoff; only a real in-context `place_design`/`route_design`
run on the actual top-level module counts. If a number is a projection
(not measured), say so explicitly and show the real numbers it's built
from.
## Working discipline
- Fork before promote: don't edit an already-verified, in-use RTL file in
place for a new experiment — copy/fork it, verify the fork, then decide
whether to promote it. (Established V2-era convention, still followed in
V3.)
- One variable at a time: verify a new module in isolation before wiring it
into a larger system; verify the larger system before trusting a P&R
number built on top of it.
- Root-cause every anomaly via hierarchical signal tracing — never guess or
paper over an unexplained result. Several real bugs in this project were
found exactly this way, not by inspection.
- After ANY RTL change to logic that's part of the real synthesis target
(`hardware/v3/rtl/n2_system_ddr3_top.v` and its dependents), re-run a real
P&R before claiming it's still timing-clean — the margin is thin enough
that this is not optional caution, it's load-bearing.
- Commit messages end with the attribution lines already configured for this
session (Co-Authored-By + Claude-Session) — keep using them.