feat(v2): M7 Dataflow Core - full M1-M6 integration, wake-up loop closed end-to-end
dataflow_core.v integrates dependency_manager (M6) -> neural_director (M5) -> N_SLOTS x (memory_manager (M4) + neural_processor (M1)) for the first time. A slot's completion (via neural_director's new slot_node_id tracking, an additive port) feeds back as a producer_done event to dependency_manager, waking up any node that depended on it - closing the dataflow loop without external glue. Verified end-to-end (Verilator) on a 3-node DAG: two independent nodes plus a third depending on both, confirmed to dispatch only after both genuinely complete via real neural_processor computation. 4/4 PASS. Real synthesis + nextpnr-ecp5 P&R via a synthesis-only timing harness (bare per-slot backend ports exceed the LFE5U-45F's TRELLIS_IO budget, same pattern as ERR-0005): N_SLOTS=2 -> 165.15 MHz, N_SLOTS=4 -> 133.19 MHz, both PASS at 80MHz, 0 synthesis problems. Scope explicitly deferred to M8 (DEC-0009): M3's BRAM buffers not wired in yet, per-slot Memory Backend Interface ports not arbitrated to one shared PSRAM master yet - both need real measured data before committing to a design, not guessed at here. Logged: simulation/synthesis/timing/benchmark/decisions (DEC-0009)/ experiments (EXP-0008)/errors (ERR-0007, a Yosys chparam-ordering build quirk, not an RTL bug)/development.log, ROADMAP.md updated. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
This commit is contained in:
@@ -54,3 +54,14 @@ Fmax: 250.50 MHz -- PASS at 80MHz (real place&route measurement)
|
||||
no harness needed), real nextpnr-ecp5 --45k --package CABGA381
|
||||
--speed 8 --freq 80 --lpf-allow-unconstrained
|
||||
Fmax: 155.30 MHz -- PASS at 80MHz (real place&route measurement)
|
||||
|
||||
[2026-09-05] EXP-0008 -- dataflow_core (via harness_dataflow_core.v,
|
||||
see errors.log ERR-0005 for why a harness was needed), real
|
||||
nextpnr-ecp5 --45k --package CABGA381 --speed 8 --freq 80
|
||||
--lpf-allow-unconstrained
|
||||
N_SLOTS=2: Fmax = 165.15 MHz -- PASS at 80MHz (real place&route)
|
||||
N_SLOTS=4: Fmax = 133.19 MHz -- PASS at 80MHz (real place&route)
|
||||
Fmax drops as N_SLOTS grows (more concurrent memory_manager+
|
||||
neural_processor instances competing for the same routing fabric
|
||||
around the shared neural_director/dependency_manager hub) -- both
|
||||
configs still clear the 80MHz target with real margin.
|
||||
|
||||
Reference in New Issue
Block a user