feat(v2): M8 PSRAM integration - real V1 backend shared across concurrent slots
neural_multiprocessor.v wraps dataflow_core.v (M7, unmodified) around the real, unmodified V1 PSRAM backend chain (int8_memory_access -> memory_interface -> psram_controller), funneling N_SLOTS independent Memory Backend Interface ports through a new generic N-port arbiter (slot_mem_arbiter.v) inspired by (not copied from) V1's own mem_arbiter.v. Real concurrent-slot simulation immediately surfaced a genuine bug (ERR-0008): memory_manager/prefetch_engine's byte-level backend protocol is fire-and-forget (a single-cycle mem_req pulse with no accept handshake) - correct for M4's direct 1:1 connection, but a naive arbiter silently drops a pulse arriving while the shared bus is owned by another slot, hanging that slot forever. Fixed with a per-port pending-request latch, the same "queue, don't drop" idiom already used by memory_manager's own pf_pending register (ERR-0006). Verified (Verilator): 4/4 PASS with 2 slots genuinely contending for one real PSRAM port (444 cycles). No regression on M4's own testbench. Real synthesis + nextpnr-ecp5 P&R (no harness needed - real PSRAM pins keep the top-level at 157 pins): 0 problems, Fmax 142.45 MHz, PASS at 80MHz. Arbitration policy is fixed lowest-index priority, not fairness- balanced (DEC-0010) - consistent with every other "simplest correct policy first" scheduling choice in this roadmap, revisited only if M9's real measurement shows starvation matters. Logged: simulation/synthesis/timing/benchmark/decisions (DEC-0010)/ experiments (EXP-0009)/errors (ERR-0008)/development.log, ROADMAP.md updated. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
This commit is contained in:
@@ -103,3 +103,17 @@ uses 32/72 (44%), consistent with DEC-0005's finding that DSP, not
|
||||
LUT/FF, is the first resource to saturate as concurrency grows (M2's
|
||||
own N_PROCESSORS=8 measurement: 88%). BRAM=0 on both is expected --
|
||||
M3's buffers are not wired into dataflow_core yet (DEC-0009).
|
||||
|
||||
[2026-09-05] M8 Neural Multiprocessor top (real standalone synthesis +
|
||||
P&R, no harness needed -- real PSRAM pins keep the bare top-level
|
||||
pin count at 157, under the TRELLIS_IO budget)
|
||||
|
||||
| Module (config) | Fmax (POST-P&R) | LUT4 | CCU2C | FF | DSP | BRAM |
|
||||
|--------------------------------------|------------------|------|-------|------|-----|------|
|
||||
| neural_multiprocessor (N_SLOTS=2) | 142.45 MHz | 3145 | 388 | 3659 | 16 | 0 |
|
||||
|
||||
Compare to M7's dataflow_core alone (N_SLOTS=2): 165.15 MHz / LUT4=2127
|
||||
/ CCU2C=248 / FF=2505 / DSP=16. Adding the real V1 PSRAM chain +
|
||||
slot_mem_arbiter costs ~1000 LUT4/140 CCU2C/1150 FF and drops Fmax by
|
||||
~23 MHz (165.15 -> 142.45) -- both real, measured costs of real PSRAM
|
||||
integration, not assumed.
|
||||
|
||||
Reference in New Issue
Block a user