V2.0.0 hardware freeze - single SDRAM

FASE #1 hardware freeze for FPGA-Neural V2, N4/P8, single external
SDRAM (Alliance Memory AS4C4M16SA-6TIN) serving weights, activations,
and results through one physical sdram_controller.v instance. Removes
the PSRAM dependency (hardware/v1/rtl/psram_controller.v +
memory_interface.v) from the V2 physical path entirely -- V1 itself
remains fully unmodified, the golden reference.

New RTL: sdram_unified_backend.v (2-way W/AR arbitration over one
SDRAM controller, real per-byte DQM write masking added to
sdram_controller.v for correct single-byte result writes with no
read-modify-write), nms_neural_multiprocessor_sdram_unified.v (the
frozen top-level). Two real bugs found and fixed via full-system
testing before being accepted (ERR-0023): a deadlock and an off-by-one
data-shift bug in the new arbitration logic.

Real results: N=4 and N=2 D-Stress bit-exact (256/256 neurons), 40
real AUTO REFRESH events interleaved with zero corruption, real
Yosys+nextpnr-ecp5 synthesis/P&R for LFE5U-45F-8CABGA381 (149/245
TRELLIS_IO, a real 45-pin reduction from the prior dual-memory
design). Timing is MARGINAL (1/8 P&R seeds >=80MHz), reported honestly
rather than masked by the best seed.

Real, sourced ball-level pinout for the SDRAM bus + clk/rst (39/149
signals, P&R-verified) using the official Lattice ECP5U-45 pinout CSV
found on disk during this step's own pre-commit review -- corrects an
earlier draft that wrongly assumed no real pinout data was available.

Chip readiness: NO. Real, disclosed blockers remain (no physical host
interface exists yet -- the RTL's own reg_* ports are a 110-pin raw
test-harness bus; clock source/PLL decision; power/configuration
component selection) -- see hardware/v2/docs/{HARDWARE_FREEZE,
CHIP_READINESS,OPEN_ITEMS}.md for the complete, itemized status.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
This commit is contained in:
2026-09-06 13:39:55 +02:00
co-authored by Claude Sonnet 5
parent 5c9ec618d3
commit 8e014d8d49
208 changed files with 3000390 additions and 0 deletions
@@ -0,0 +1,54 @@
// ============================================================
// Neural Memory System (NMS) -- STEP4/5 candidate A: REPLICATED
// activation memory.
//
// One full copy of the shared activation vector's tile storage per
// slot (N_SLOTS independent single-write/single-read BRAMs). A shared
// fill engine broadcasts each filled tile to EVERY copy on the same
// cycle (one PSRAM-side write, N_SLOTS on-chip writes) -- after fill,
// every slot's own read port is completely private: zero contention,
// ever, by construction (no arbitration logic at all on the read
// side). Real cost is N_SLOTS x the single-copy storage; this file
// exists to MEASURE that real DP16KD/LUT/Fmax cost against Candidate
// B (nms_activation_banked.v) rather than assume replication is too
// expensive a priori (EXP-0019/DEC-0019).
// ============================================================
module nms_activation_replicated #(
parameter DATA_WIDTH = 8,
parameter P_IN = 8,
parameter N_SLOTS = 4,
parameter MAX_TILES = 16,
parameter TIW = (MAX_TILES <= 1) ? 1 : $clog2(MAX_TILES)
)(
input clk,
input rst,
// ---- fill port: one write, broadcast to every copy ----
input fill_we,
input [TIW-1:0] fill_addr,
input [DATA_WIDTH*P_IN-1:0] fill_data,
// ---- per-slot private read port ----
input [N_SLOTS-1:0] rd_en,
input [N_SLOTS*TIW-1:0] rd_addr_flat,
output [N_SLOTS*DATA_WIDTH*P_IN-1:0] rd_data_flat
);
genvar g;
generate
for (g = 0; g < N_SLOTS; g = g + 1) begin : GEN_COPY
reg [DATA_WIDTH*P_IN-1:0] mem [0:MAX_TILES-1];
reg [DATA_WIDTH*P_IN-1:0] rd_data_reg;
always @(posedge clk) begin
if (fill_we)
mem[fill_addr] <= fill_data;
if (rd_en[g])
rd_data_reg <= mem[rd_addr_flat[g*TIW +: TIW]];
end
assign rd_data_flat[g*DATA_WIDTH*P_IN +: DATA_WIDTH*P_IN] = rd_data_reg;
end
endgenerate
endmodule