Files
FPGA-Neural/rtl/act_buffer.v
T
micheleandClaude Sonnet 5 55c827bedf feat: PSRAM page-mode reads + graph engine (Type #2) + real pinout/IRQ pins
PSRAM page-mode read burst support in psram_controller.v: enables the
ISSI IS66WVE4M16EBLL-70BLI's page mode via its configuration-register
software-access sequence at boot (disabled by default on the real
chip), then keeps CE#/OE# asserted after a read so a same-page
continuation only pays tAPA (20ns) instead of a full tAA (70ns)
random access, with automatic tCEM-safe session closing. Only a WRITE
closes the page -- byte-enable changes do not, since
int8_memory_access.v alternates them on nearly every access and an
early implementation attempt that treated them as a close condition
measured a real regression (53.25->61.25 cycles/edge) before being
corrected (53.25->37.53 cycles/edge, +42% gather bandwidth).
sim/psram_model.v gained independent tAPA/tAA and tCEM enforcement
(with a real Verilog same-timestep event-ordering race found and
fixed via a #0 sync) so the regression proves real timing compliance,
not just data correctness. New sim/psram_page_mode_tb.v; full 26-file
regression suite re-run clean. Real nextpnr-ecp5 Fmax re-measured on
the full spi_neuron_top system: 75.73MHz (P2, up from 55.59MHz) and
65.13MHz (P8) -- still under the 80MHz target but not regressed, with
the critical path confirmed (not assumed) to remain entirely inside
neuron_parallel's accumulate chain, never psram_controller.

Also includes this session's other already-validated work: the graph
engine (Type #2 sparse-graph network: act_buffer, graph_engine,
netasm host assembler), real CABGA381 pinout (.lpf, place&route
verified) and physical IRQ_N/DATA_READY_N pins, and Phase 7 timing
closure logs -- all previously uncommitted, documented in WORKLOG.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LH3jPeJ3eFMfF2v8SQhpkk
2026-09-03 17:12:05 +02:00

65 lines
2.5 KiB
Verilog

`timescale 1ns/1ps
// ================================================================
// ACT_BUFFER (Phase G1 -- Graph engine, Type #2 network)
//
// Global activation buffer for the sparse-graph datapath: one INT8
// slot per signal id, id 0..N_in-1 are the network's external
// inputs, every other id is the output of exactly one neuron (its
// out_id). graph_engine's gather stage reads this buffer by src_id;
// each neuron's result is written back here at its own out_id.
//
// Dual-port, byte-addressed:
// Port A - synchronous write (neuron output / input copy-in).
// Port B - synchronous READ, REGISTERED: rd_data reflects rd_addr
// from the PREVIOUS clock edge, not the current one (one
// cycle of latency). Callers must account for this --
// see graph_engine.v's gather stage.
//
// This is the plain always-block-per-port inference idiom Yosys'
// ECP5 memory_bram pass maps onto a Lattice DP16KD block RAM (single
// clock, independent read/write address, synchronous read with no
// output reset): no `rst` port is provided on purpose, matching what
// a real DP16KD offers and keeping this off the LUT-RAM path that a
// register-array-with-reset idiom would force it down.
//
// N_TOTAL is the V1 ceiling on distinct signal ids (§2 of the spec:
// 4096, 16-bit id space would allow up to 65536 without a format
// change). ADDR_WIDTH is derived, not passed in, so every caller
// stays consistent with N_TOTAL automatically.
// ================================================================
module act_buffer #(
parameter N_TOTAL = 4096,
parameter DATA_WIDTH = 8,
localparam ADDR_WIDTH = $clog2(N_TOTAL)
)(
input wire clk,
// ------------------------------------------------------------
// Port A: write
// ------------------------------------------------------------
input wire wr_en,
input wire [ADDR_WIDTH-1:0] wr_addr,
input wire signed [DATA_WIDTH-1:0] wr_data,
// ------------------------------------------------------------
// Port B: read (registered, 1-cycle latency)
// ------------------------------------------------------------
input wire [ADDR_WIDTH-1:0] rd_addr,
output reg signed [DATA_WIDTH-1:0] rd_data
);
reg signed [DATA_WIDTH-1:0] mem [0:N_TOTAL-1];
always @(posedge clk) begin
if (wr_en)
mem[wr_addr] <= wr_data;
end
always @(posedge clk) begin
rd_data <= mem[rd_addr];
end
endmodule