Begins the V2 Neural Multiprocessor / Dataflow architecture per docs/v2-description.md, per explicit user request to freeze V1 and start V2 development, copying from V1 what's needed. Scaffold: - hardware/v1/: byte-exact, read-only copy of the current V1 codebase (rtl, testbenches, tools, constraints, a representative subset of synthesis results, and reference docs) -- verified identical via diff/cmp against the live top-level tree before being made filesystem-read-only. The live top-level tree is untouched and remains the project's "production" V1 (see hardware/v1/README.md and hardware/v2/logs/decisions.log DEC-0001 for why copy-not-move). - hardware/v2/: mandatory structure (rtl/sim/constraints/synthesis/ reports/scripts/logs/docs) plus the full logging system required by the spec (development/architecture/simulation/synthesis/timing/ benchmark/decisions/experiments/errors.log). M1 -- Neural Processor (hardware/v2/rtl/neural_processor.v): - 8-stage pipelined perceptron unit (P_IN=8): input align, 8 multipliers, 3-level adder tree, accumulator, bias+activation, INT8 saturation. Genuine 1-tile/cycle throughput, not just a wider combinational datapath. - 7-state FSM (NP_IDLE..NP_ERROR per docs/v2-description.md §6, with 4 baseline states merged into NP_WAIT_OPERANDS -- see decisions.log DEC-0002); valid/ready/data/last stream interfaces per §7. - Bit-exact vs the frozen hardware/v1/rtl/neuron_parallel.v + mac8.v + mac_unit.v: 7/7 tests pass (hardware/v2/sim/tb_neural_processor.v), covering regular/mixed-sign/extreme-INT8 vectors, both activations, a zero-idle-gap back-to-back-tiles throughput check, and an 8-tile job -- verified with Verilator (see below for why). - Real synthesis + place&route (Yosys + nextpnr-ecp5): 0 CHECK problems, Fmax 183.12 MHz at ACC_WIDTH=32 (PASS at 80MHz, ~3x V1's isolated PARALLEL=8 Fmax of 61.71 MHz) and 176.21 MHz at ACC_WIDTH=24 (a user-requested comparison experiment, also bit-exact-verified; see experiments.log EXP-0001/EXP-0002 and benchmark.log). Three real bugs found and resolved during M1 development (full diagnostic record in errors.log): - Two independent, reproducible Icarus Verilog v13.0 scheduling defects (ERR-0001, ERR-0002) that silently produced wrong simulation results for standard sequential Verilog -- confirmed via Verilator 5.050 giving correct results on the same minimal repros. Verilator is now the trusted simulator for hardware/v2/ (decisions.log DEC-0004); Icarus's affected protocol-violation check was removed from the RTL and deferred architecturally to the Neural Director (DEC-0003) rather than chased further. - One real RTL bug (ERR-0003): last0 wasn't gated like valid0, letting a "last tile" tag leak into the pipeline ahead of its actual valid tile on back-to-back jobs. Fixed and verified. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
327 lines
15 KiB
Verilog
327 lines
15 KiB
Verilog
`timescale 1ns/1ps
|
|
|
|
// ================================================================
|
|
// FLASH_MODEL - behavioral model of the Winbond W25Q128JV(-DTR)
|
|
// SPI NOR flash, SIMULATION ONLY (not synthesizable, not intended
|
|
// to be instantiated anywhere but a testbench).
|
|
//
|
|
// This is an independent-oracle support model for F1-F6 of the
|
|
// flash subsystem (see WORKLOG.md): it exists so RTL testbenches
|
|
// have something to talk to that enforces the datasheet's write
|
|
// rules in hardware-like fashion (erase-to-FF, program-can-only-
|
|
// clear-bits, WIP timing, 256B page-program wraparound), rather
|
|
// than a trivial RAM that would silently accept violations the
|
|
// real chip would corrupt on. It is NOT the oracle for "is the
|
|
// design's behavior correct" (that oracle is the Python reference
|
|
// model + datasheet numbers, per WORKLOG.md §A.1) -- it is the
|
|
// thing-under-test's counterparty, playing the role of the flash.
|
|
//
|
|
// Unlike rtl/spi_slave.v (which crosses an async SPI clock into
|
|
// the FPGA's own `clk` domain via double-flop sync), this model has
|
|
// no domain of its own to defend -- it IS the SPI-clocked device,
|
|
// so sampling/driving directly on posedge/negedge of `sclk` (as the
|
|
// real chip's internal logic does) is the correct and simplest
|
|
// idiom here, not a shortcut.
|
|
//
|
|
// Datasheet: local copy at
|
|
// /Users/michelebigi/Development/HubAudio/datasheets/W25Q128JVS.pdf
|
|
// title page identifies it as "W25Q128JV-DTR" (Double Transfer Rate
|
|
// variant). All rules modeled here come from that document's
|
|
// STANDARD SPI instruction set (§8.1.2 Table 1, p.26) and AC
|
|
// Electrical Characteristics (§9.6, p.90), which are shared across
|
|
// the whole W25Q128JV family (DTR adds extra instructions/modes on
|
|
// top; it does not change the base standard-SPI behavior modeled
|
|
// here) -- see the JEDEC ID note below for the one confirmed
|
|
// discrepancy between this PDF and the actual populated part.
|
|
//
|
|
// Modeled (datasheet citations inline at point of use below):
|
|
// - RDID (9Fh) §8.2.30 p.70 -- manufacturer/type/capacity, 3 bytes
|
|
// - READ (03h) Table 1 p.26 -- sequential byte read, address wraps
|
|
// at the modeled DEPTH boundary
|
|
// - WREN (06h) §8.2.1 p.30 -- sets WEL
|
|
// - PP (02h) Table 1 p.26, -- program (AND-only, never sets bits),
|
|
// note 3 p.29 wraps to page start past byte 256
|
|
// - SE (20h) Table 1 p.26 -- 4KB sector erase -> all-FF
|
|
// - RDSR-1(05h) §7.1.1/7.1.2 -- bit0=BUSY(WIP), bit1=WEL (TOC order
|
|
// p.15 on p.2 lists BUSY before WEL; this
|
|
// bit assignment is also universal
|
|
// across the whole Winbond SPI-NOR
|
|
// family -- flagged as the one bit-
|
|
// layout fact taken from convention
|
|
// rather than a single explicit
|
|
// consolidated bit table in this
|
|
// particular local PDF, see WORKLOG.md
|
|
// F1 entry)
|
|
// - tPP/tSE timing: §9.6 p.90, MAX (worst-case) values used, scaled
|
|
// by TIME_SCALE (see parameter) so simulation doesn't burn real
|
|
// milliseconds of event-queue time. Using MAX rather than TYP is
|
|
// deliberate: a design that only polls WIP a couple of times
|
|
// (tuned to the typical-case latency) must still pass against the
|
|
// worst case, or it is not actually correct.
|
|
//
|
|
// NOT modeled (explicit limitation, see WORKLOG.md §A.6):
|
|
// - Fast Read / Dual / Quad / QPI / DTR instructions (out of scope
|
|
// for this design's F1, which only uses standard single-SPI).
|
|
// - Status Register-2/3, block/sector protect bits, individual
|
|
// block locks, /WP, /HOLD, power-down (B9h/ABh) -- the FPGA has
|
|
// exclusive, trusted control of this flash, so write protection
|
|
// is out of scope; this model always accepts writes once WEL is
|
|
// set, like a factory-default (WPS=0, all BP bits 0) part.
|
|
// - Analog/electrical timing (rise/fall times, setup/hold margins,
|
|
// signal integrity) -- not reachable in a digital behavioral model.
|
|
// - True mid-operation power-loss: real NOR flash can leave a page
|
|
// PARTIALLY programmed if power is cut mid-instruction. This
|
|
// model's commit is a single non-blocking assignment burst after
|
|
// the modeled tPP/tSE delay, so the power-loss hook (below) can
|
|
// only represent "operation aborted before it committed at all",
|
|
// not "committed halfway" -- still sufficient to exercise the
|
|
// catalog's CRC+valid_flag rejection path (F4), but coarser than
|
|
// real silicon.
|
|
// ================================================================
|
|
|
|
module flash_model #(
|
|
parameter DEPTH = 32'h0002_0000, // 128KB modeled (32x 4KB sectors).
|
|
// §A.6: NOT the real 16MB (2^24) --
|
|
// mirrors sim/psram_model.v's own
|
|
// precedent of a reduced DEPTH for
|
|
// simulation speed. Covers the
|
|
// catalog sector + several data
|
|
// sectors for F2-F5 without the
|
|
// multi-second iverilog elaboration
|
|
// cost a full 16MB reg array would add.
|
|
parameter integer TIME_SCALE = 100000 // divides real ms-scale datasheet
|
|
// timing (see tPP/tSE below) into
|
|
// simulation-friendly ns
|
|
)(
|
|
input wire sclk,
|
|
input wire mosi,
|
|
output reg miso,
|
|
input wire cs_n
|
|
);
|
|
|
|
localparam [31:0] SECTOR_BYTES = 32'd4096; // Sector Erase (4KB), Table 1 p.26 "20h"
|
|
|
|
// ------------------------------------------------------------
|
|
// JEDEC ID: EF4018h (plain, non-DTR W25Q128JV), NOT the 7018h
|
|
// this DTR-titled PDF's own §8.1.1 table (p.24) shows.
|
|
//
|
|
// ASSUMPTION FLAGGED (§A.1/§A.6): docs/FPGA-Neural-Hardware-Design.md
|
|
// and docs/FPGA-Neural-Datapatch-Benchmark.md both specify the
|
|
// populated part as "W25Q128JV" / "W25Q128JVS" -- WITHOUT the
|
|
// -DTR suffix -- and those hardware docs are the confirmed BOM
|
|
// source (per project memory). The plain (non-DTR) part's
|
|
// publicly-documented JEDEC ID is EF4018h (memory type 40h);
|
|
// this local PDF is for the DTR-capable die (memory type 70h),
|
|
// a different part number that happens to share the same base
|
|
// filename. Rather than silently trusting whichever byte the
|
|
// one local PDF prints, this model uses the value that matches
|
|
// the part actually specified in the project's own hardware
|
|
// docs, and flags the mismatch explicitly (see WORKLOG.md F1
|
|
// entry) -- real hardware bring-up MUST confirm the populated
|
|
// chip's actual RDID response against this constant.
|
|
// ------------------------------------------------------------
|
|
localparam [7:0] JEDEC_MFR = 8'hEF;
|
|
localparam [7:0] JEDEC_MEMTYPE = 8'h40;
|
|
localparam [7:0] JEDEC_CAPACITY = 8'h18; // 128Mbit -- §8.1.1 p.24, agrees between DTR/non-DTR
|
|
|
|
localparam [7:0] OP_WREN = 8'h06;
|
|
localparam [7:0] OP_READ = 8'h03;
|
|
localparam [7:0] OP_PP = 8'h02;
|
|
localparam [7:0] OP_SE = 8'h20;
|
|
localparam [7:0] OP_RDSR1 = 8'h05;
|
|
localparam [7:0] OP_RDID = 8'h9F;
|
|
|
|
reg [7:0] mem [0:DEPTH-1];
|
|
|
|
integer i;
|
|
initial begin
|
|
// Erase state is 0xFF everywhere -- §8.2.18 p.56: Sector
|
|
// Erase sets every bit in the addressed region to 1. A
|
|
// freshly-elaborated model matches a freshly-blanked part.
|
|
for (i = 0; i < DEPTH; i = i + 1)
|
|
mem[i] = 8'hFF;
|
|
end
|
|
|
|
reg wel;
|
|
reg busy;
|
|
reg [7:0] opcode;
|
|
reg [23:0] addr;
|
|
reg [31:0] bit_count; // total bits received since CS fell
|
|
reg [7:0] shift_out;
|
|
|
|
reg pending_pp, pending_se;
|
|
reg [23:0] pending_addr;
|
|
reg [8:0] pp_nbytes;
|
|
reg [7:0] pp_bytes [0:255];
|
|
|
|
localparam realtime TPP_MAX_NS = 3_000_000.0 / TIME_SCALE; // 3ms MAX, §9.6 p.90
|
|
localparam realtime TSE_MAX_NS = 400_000_000.0 / TIME_SCALE; // 400ms MAX, §9.6 p.90
|
|
|
|
initial begin
|
|
wel = 1'b0; busy = 1'b0; miso = 1'b0;
|
|
pending_pp = 1'b0; pending_se = 1'b0; bit_count = 0;
|
|
end
|
|
|
|
// ============================================================
|
|
// Transaction start: reset framing state. Every instruction is
|
|
// exactly one CS-low period (§8 intro, p.24: "Instructions are
|
|
// initiated with the falling edge of Chip Select").
|
|
// ============================================================
|
|
always @(negedge cs_n) begin
|
|
bit_count <= 0;
|
|
opcode <= 8'h00;
|
|
end
|
|
|
|
// ============================================================
|
|
// MOSI sampled on the rising edge of CLK (mode 0), matching
|
|
// every timing diagram in the datasheet (Fig. 7/28/30/43a).
|
|
// ============================================================
|
|
always @(posedge sclk) begin
|
|
if (!cs_n) begin
|
|
|
|
if (bit_count < 8) begin
|
|
opcode <= {opcode[6:0], mosi};
|
|
end else if (bit_count < 32) begin
|
|
addr <= {addr[22:0], mosi};
|
|
end else begin
|
|
// Data phase: PP captures incoming bytes; READ's
|
|
// outgoing bytes are handled entirely in the
|
|
// falling-edge block below (this model does not
|
|
// echo MOSI during READ, matching "DI High
|
|
// Impedance" shown for read-type instructions).
|
|
if (opcode == OP_PP && wel) begin
|
|
pp_bytes[(bit_count - 32) >> 3][7 - ((bit_count - 32) & 3'h7)] <= mosi;
|
|
end
|
|
end
|
|
|
|
bit_count <= bit_count + 1;
|
|
end
|
|
end
|
|
|
|
// ============================================================
|
|
// MISO driven on the falling edge of CLK, one bit ahead of the
|
|
// next rising-edge sample (mode 0), exactly as every read-type
|
|
// timing diagram in the datasheet shows (e.g. Fig. 43a: DO
|
|
// changes shortly after each falling edge).
|
|
// ============================================================
|
|
always @(negedge sclk) begin
|
|
if (!cs_n) begin
|
|
case (opcode)
|
|
|
|
OP_RDID: begin
|
|
if (bit_count < 16) miso <= JEDEC_MFR[15 - bit_count];
|
|
else if (bit_count < 24) miso <= JEDEC_MEMTYPE[23 - bit_count];
|
|
else if (bit_count < 32) miso <= JEDEC_CAPACITY[31 - bit_count];
|
|
else miso <= 1'b0;
|
|
end
|
|
|
|
OP_RDSR1: begin
|
|
// bit0=BUSY, bit1=WEL -- see header note.
|
|
if (bit_count < 8) miso <= 1'b0;
|
|
else case ((bit_count - 8) & 3'h7)
|
|
3'd7: miso <= busy; // MSB-first shift-out of {6'b0,WEL,BUSY}
|
|
3'd6: miso <= wel;
|
|
default: miso <= 1'b0;
|
|
endcase
|
|
end
|
|
|
|
OP_READ: begin
|
|
if (bit_count < 32) begin
|
|
miso <= 1'b0;
|
|
end else if (bit_count == 32) begin
|
|
if (addr >= DEPTH) begin
|
|
$display("[flash_model] FATAL: READ out of modeled range 0x%06h (DEPTH=0x%06h) at t=%0t", addr, DEPTH, $time);
|
|
$fatal(1);
|
|
end
|
|
shift_out <= mem[addr];
|
|
miso <= mem[addr][7];
|
|
end else if (((bit_count - 32) & 3'h7) == 0) begin
|
|
shift_out <= mem[(addr + ((bit_count - 32) >> 3)) % DEPTH];
|
|
miso <= mem[(addr + ((bit_count - 32) >> 3)) % DEPTH][7];
|
|
end else begin
|
|
shift_out <= {shift_out[6:0], 1'b0};
|
|
miso <= shift_out[6];
|
|
end
|
|
end
|
|
|
|
default: miso <= 1'b0;
|
|
|
|
endcase
|
|
end
|
|
end
|
|
|
|
// ============================================================
|
|
// Transaction end: byte-boundary-gated side effects (§8 intro,
|
|
// p.24: writes/erases that don't end on a byte boundary are
|
|
// ignored entirely -- modeled by requiring bit_count to be an
|
|
// exact multiple of 8 with the right minimum length below).
|
|
// ============================================================
|
|
always @(posedge cs_n) begin
|
|
|
|
if (!busy && opcode == OP_WREN && bit_count == 8) begin
|
|
wel <= 1'b1;
|
|
end
|
|
|
|
if (!busy && wel && opcode == OP_PP
|
|
&& bit_count >= 40 && ((bit_count - 32) & 3'h7) == 0) begin
|
|
pp_nbytes <= (bit_count - 32) >> 3;
|
|
pending_addr <= addr;
|
|
pending_pp <= 1'b1;
|
|
busy <= 1'b1;
|
|
wel <= 1'b0;
|
|
end
|
|
|
|
if (!busy && wel && opcode == OP_SE && bit_count == 32) begin
|
|
if (addr >= DEPTH || addr[11:0] != 12'h000) begin
|
|
// §A.3 negative test: Sector Erase to an address not
|
|
// aligned to a 4KB sector boundary. Real Winbond
|
|
// parts silently erase the containing sector (low
|
|
// address bits don't-care); this model instead
|
|
// FATALs so a design bug that relies on that
|
|
// silent behavior (rather than aligning addresses
|
|
// itself, per WORKLOG.md's erase-before-write design
|
|
// decision) is caught, not masked.
|
|
$display("[flash_model] FATAL: Sector Erase to unaligned/out-of-range address 0x%06h at t=%0t", addr, $time);
|
|
$fatal(1);
|
|
end
|
|
pending_addr <= addr;
|
|
pending_se <= 1'b1;
|
|
busy <= 1'b1;
|
|
wel <= 1'b0;
|
|
end
|
|
|
|
end
|
|
|
|
// ============================================================
|
|
// WIP timers: commit the write/erase after the modeled MAX
|
|
// (worst-case) datasheet duration. Committing all at once here
|
|
// (rather than incrementally) is what makes the power-loss hook
|
|
// below meaningful: forcing `pending_pp`/`pending_se` and `busy`
|
|
// low from a testbench (via hierarchical reference) before this
|
|
// block's delay expires means the commit loop never runs at all
|
|
// -- see the NOT-modeled note in the header for the fidelity
|
|
// limit of that hook.
|
|
// ============================================================
|
|
|
|
always @(posedge pending_pp) begin
|
|
# (TPP_MAX_NS);
|
|
if (pending_pp) begin // still armed: not aborted by a power-loss hook
|
|
for (i = 0; i < pp_nbytes; i = i + 1)
|
|
mem[pending_addr + i] <= mem[pending_addr + i] & pp_bytes[i]; // AND-only, Table1/§8.2.16 p.53
|
|
busy <= 1'b0;
|
|
end
|
|
pending_pp <= 1'b0;
|
|
end
|
|
|
|
always @(posedge pending_se) begin
|
|
# (TSE_MAX_NS);
|
|
if (pending_se) begin
|
|
for (i = 0; i < SECTOR_BYTES; i = i + 1)
|
|
mem[pending_addr + i] <= 8'hFF; // erase -> all-ones, §8.2.18 p.56
|
|
busy <= 1'b0;
|
|
end
|
|
pending_se <= 1'b0;
|
|
end
|
|
|
|
endmodule
|