First real RTL piece of the SPI interface (docs §8.1 protocol
draft): the physical layer only -- Mode 0 (CPOL=0, CPHA=0),
MSB-first, byte-level shift register with a 3-stage CDC synchronizer
for SCLK/MOSI/CS_N (the SPI master clock is asynchronous to the
FPGA system clock). Exposes rx_byte/rx_valid, tx_byte/tx_byte_req,
and cs_active/cs_start/cs_end to the (not yet written) protocol
engine.
Documented an important consumer contract on tx_byte_req: it is a
prefetch hint (fires once extra after the last byte of every
transaction, since the slave cannot know in advance whether the
master will keep clocking), not a "byte consumed" event -- a
consumer must advance any stateful pointer (e.g. a RAM read address)
on rx_valid instead, which fires exactly once per real byte
transferred.
sim/spi_slave_tb.v: bit-banged SPI master BFM (4 tests: single byte,
multi-byte in one CS period, back-to-back transactions, slower
SCLK). Two testbench-only bugs found and fixed during bring-up (RTL
itself needed no functional change beyond the tx_byte_req contract
comment): the BFM was advancing its tx queue on tx_byte_req instead
of rx_valid (see contract above), and inter-test reset pulses raced
against posedge clk (blocking `rst=1` landing on the same simulation
time as a clock edge) -- fixed by asserting/deasserting reset on
negedge clk instead.
Verified two ways: Icarus Verilog (4/4 tests pass) and the real
ECP5 toolchain used for prior benchmarks (Yosys 0.68 synth: 0
problems, 41 FF / 55 LUT4, no latches; nextpnr-ecp5 --45k --package
CABGA381 --speed 8 --freq 80: PASS, Fmax 403.23 MHz; ecppack:
bitstream generated with no errors).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQV3vS9TXaGDJ5cRfnfidt
Phase 4 (SPI Interface) only had a high-level conceptual sequence
(RESET/CONFIGURE/LOAD.../START/WAIT/READ) with no concrete opcodes,
framing, or register map -- not enough to start RTL from. Added
docs/FPGA-NeuralNetwork-Engine.md §8.1 with a concrete v1 draft:
- SPI Mode 0, MSB-first, one opcode byte per CS-low transaction.
- Explicit length field on WRITE_RAM/READ_RAM (chosen over
CS-edge-delimited streaming: simpler controller, just a byte
counter).
- READ_CONFIG opcode exposing N_INPUTS/N_NEURONS/PARALLEL/
ADDR_WIDTH/DATA_WIDTH at runtime, so one host firmware build can
target different bitstreams.
- RESET kept as its own opcode (0x0F), distinct from NOP.
- STATUS.done documented as required to be a STICKY, clear-on-read
bit in the SPI register bank: neuron_memory.done is a one-cycle
pulse that a slow SPI poll would almost certainly miss otherwise.
Opcode values themselves are marked explicitly as draft/example,
not frozen -- only the framing rules and the two decisions above are
meant to stick going into Phase 4 RTL work.
No RTL or testbench changes in this commit; design-only.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQV3vS9TXaGDJ5cRfnfidt
neuron_memory.v only handled a single neuron. Added an N_NEURONS
parameter (default 1, fully backward compatible) and a memory-bound
neuron loop: X is read once (shared layer input), and for each
neuron in turn W and bias are re-read from PSRAM and fed to a
single, reused neuron_parallel instance -- no change to the
validated compute datapath (neuron_parallel/mac8/mac_unit).
Addressing follows layer.v's neuron-major convention: neuron n's
weights live at w_base + n*N_INPUTS bytes, its bias at
bias_addr + n. Output changed from a single `y` port to a packed
`y_bus` (DATA_WIDTH*N_NEURONS bits, neuron-major), matching
layer.v's y_bus.
- rtl/neuron_memory.v: N_NEURONS parameter, neuron_index/
w_group_base/bias_group_addr tracking, y_reg[] array assembled
into y_bus, STATE_WAIT_N now loops back to STATE_READ_W for the
next neuron instead of finishing after one.
- sim/neuron_memory_tb.v: updated to the new y_bus port
(N_NEURONS=1 explicit); all 5 existing tests still pass unchanged,
confirming backward compatibility.
- sim/neuron_memory_multi_tb.v: new end-to-end test (full
memory_interface + psram_controller + psram_model stack) with
N_NEURONS=3, validating per-neuron addressing and a single done
pulse at the end of the sequence (scale, larger value, ReLU).
- Full regression re-run: all existing testbenches still pass.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQV3vS9TXaGDJ5cRfnfidt
Both Phase 2 findings (docs/FPGA-NeuralNetwork-Engine.md) shared one
root cause: GROUPS = N_INPUTS / PARALLEL is integer division. When
N_INPUTS is not an exact multiple of PARALLEL, the remainder inputs
were silently dropped from the accumulation (wrong result, no
error); when PARALLEL > N_INPUTS, GROUPS = 0 and the controller's
terminal condition was never met, hanging the neuron forever.
Added a single elaboration-time guard to rtl/neuron_parallel.v: a
`generate` block instantiates a deliberately undefined module when
N_INPUTS % PARALLEL != 0, forcing a hard failure in both simulation
and synthesis instead of a silent wrong answer or a deadlock. Valid
configurations are unaffected (the branch is never elaborated). The
validated datapath (mac8/mac_unit/accumulation/ReLU/saturation) is
untouched -- this is authorized as a scoped exception to the
"core is fixed, do not touch" project policy, for this guard only.
- sim/neuron_parallel_guard_negative_nonmultiple_tb.v and
sim/neuron_parallel_guard_negative_degenerate_tb.v: negative tests
that must fail to elaborate; verified both fail with the expected
"Unknown module type" error.
- sim/parameter_sweep_tb.v: rewritten to valid-configs-only (the
three configs that used to demonstrate truncation/hang no longer
compile, by design); added PARALLEL=2 and PARALLEL=4 configs,
the two best-performing values from
docs/FPGA-Neural-Datapatch-Benchmark.md.
- Full regression re-run after the RTL change: all existing
testbenches still pass unchanged.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQV3vS9TXaGDJ5cRfnfidt
Roadmap Phase 2 asks to validate N_INPUTS/N_NEURONS/PARALLEL
combinations, including non-exact-multiple configurations. Added
sim/parameter_sweep_tb.v with 5 configs (two exact-multiple sanity
checks, two non-exact-multiple, one degenerate PARALLEL>N_INPUTS),
using a cycle-count watchdog instead of a blocking wait so a hanging
config is reported rather than hanging the simulation.
Findings (RTL unchanged, core datapath left untouched):
- GROUPS = N_INPUTS / PARALLEL truncates: when N_INPUTS is not an
exact multiple of PARALLEL, the remainder inputs are silently
never summed (confirmed 30/8 -> 6 dropped, 20/16 -> 4 dropped).
- PARALLEL > N_INPUTS gives GROUPS=0, and the controller's
group_index == GROUPS-1 terminal condition is never met: the
neuron hangs forever (confirmed via watchdog timeout).
Documented both as findings under Phase 2 in
docs/FPGA-NeuralNetwork-Engine.md for follow-up in Phase 3/7.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQV3vS9TXaGDJ5cRfnfidt
Both testbenches instantiated their DUTs with a FRAC_BITS parameter and
Q8.8 fixed-point 16-bit values, which no longer exist in rtl/neuron_parallel.v
(now plain INT8, DATA_WIDTH=8, hardcoded +127 saturation, ReLU-only clamp).
This made both tests fail elaboration ("parameter FRAC_BITS not found").
Rewrote both benches with integer INT8 stimuli and expectations matching
the current core (no RTL changes): neuron_parallel_tb covers a mixed
vector, ReLU, positive saturation, and mixed values with a boundary
negative bias; layer_tb covers 8 neurons exercising scale, ReLU,
saturation, bias-only, and a sparse weight pattern across groups.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQV3vS9TXaGDJ5cRfnfidt
preload_vector and preload_weights wrote using base as a word address
instead of a byte address (base + (k>>1) instead of (base>>1) + (k>>1)),
misaligning X/W data in PSRAM relative to what int8_memory_access expects.
Also adds a PATTERN test (X=1..32) to exercise mixed even/odd byte reads
and catch this class of bug going forward.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQV3vS9TXaGDJ5cRfnfidt