# V2 errors log -- solo append, mai troncato/sovrascritto (vedi README.md) # Nessuna entry ancora -- popolato incrementalmente man mano che avanza lo sviluppo V2. ERR-0001 (Icarus Verilog v13.0 toolchain bug, TASK/SCOPE-ENTRY DESYNC) DATE: 2026-09-05 MODULE: hardware/v2/sim/tb_neural_processor.v (M1 testbench development) SYMPTOM: a task (or any named `begin:label` block) whose FIRST executable statement is a blocking assignment to a signal read by another module's `always @(posedge clk)`, when the task/block is entered immediately after a time-consuming statement in the caller with no intervening `@(posedge clk)`, can make that FIRST assignment invisible to the DUT at the very next clock edge (the DUT's own always block behaves as if the signal never changed). Confirmed at V1's own frozen `neuron_parallel.v` (unmodified, already certified) via a minimal 3-statement task (`v1_start=1; @(posedge clk); v1_start=0;`) -- `busy` never asserted. REPRODUCTION: /tmp/mt6.v-style repro (not committed, transient scratch file) -- see conversation record for the exact minimal case. DIAGNOSIS METHOD: bisected from the full dual-DUT testbench down to a standalone ~15-line repro, ruling out RTL, port connections, and operator precedence one at a time. WORKAROUND: always begin such a task with an explicit `@(posedge clk);` before its first assignment (matches the pre-existing convention in hardware/v1/sim's own tasks, e.g. neuron_parallel_tb.v's run_neuron, which is presumably why V1's own test suite was never affected). STATUS: WORKAROUND APPLIED in hardware/v2/sim/tb_neural_processor.v's run_case. NOT reported upstream (out of scope for this session). See ERR-0004 for the broader consequence of this finding. ERR-0002 (Icarus Verilog v13.0 toolchain bug, SPURIOUS CONDITION EVALUATION) DATE: 2026-09-05 MODULE: hardware/v2/rtl/neural_processor.v (protocol-violation guard, removed -- see decisions.log DEC-0003) SYMPTOM: `if (operand_valid && !operand_ready && (state-is-one-of-four))` inside `always @(posedge clk)` evaluated TRUE at an edge where `operand_valid` was independently confirmed (via $display in the same timestep, and via the testbench's own port-connected signal) to be 0. Bisected term-by-term: even `if (operand_valid && !operand_ready)` alone, and even `if (operand_valid)` alone with explicit `== 1'b1` comparisons, still fired spuriously. Confirmed NOT a precedence issue (parens are unambiguous) and NOT specific to this exact expression shape (multiple simplified variants all reproduced it). CROSS-CHECK: root-caused further via a minimal 2-state FSM (`if (go) st<=B;`) with NO relation to the removed guard -- Icarus failed to transition on an ODD-numbered testbench clock edge (`repeat(3)` before the pulse) but succeeded on an EVEN-numbered one (`repeat(4)`), reproduced identically with both `always #5 clk=~clk` and `initial ... forever #5 clk=~clk` clock generators. VERILATOR 5.050 gives the CORRECT result for the same repro in both cases. This suggests ERR-0001 and ERR-0002 are two symptoms of the same underlying VVP scheduling defect (edge-count/thread-parity dependent), not two unrelated bugs. STATUS: the offending RTL block (protocol-violation detection) was REMOVED rather than chased further -- see DEC-0003. Root cause not fully isolated (documented honestly, not overclaimed). ERR-0003 (real RTL bug in hardware/v2/rtl/neural_processor.v, FOUND AND FIXED) DATE: 2026-09-05 MODULE: hardware/v2/rtl/neural_processor.v, stage 0 (input alignment/register) SYMPTOM: back-to-back single-tile jobs (and some multi-tile jobs) produced result_data=0 instead of the correct value, while the internal `y7` register (one stage upstream of the FSM's capture) showed the CORRECT value one cycle later than `valid7` first asserted. ROOT CAUSE: `last0 <= tile_last;` was unconditional, while `valid0 <= operand_valid && operand_ready;` was correctly gated. A master asserting `tile_last` before `operand_ready` rises (legal valid-before-ready behavior) let a "last" tag propagate through the pipeline (last1, last_tree[], last5, last6) with NO corresponding valid tile behind it, arriving at stage 7 one cycle ahead of the real valid/data pair and causing the FSM to capture a stale/wrong `y7`. EVIDENCE: isolated to a single-DUT, no-task, no-V1 repro (hardware/v2/sim/tb_neural_processor.v run under Verilator, with a cycle-by-cycle dump of valid5/last5/valid6/last6/valid7/y7) -- `last5=1` while `valid5=0` on the same cycle, confirmed the desync's exact origin at stage 0. FIX: `last0 <= (operand_valid && operand_ready) ? tile_last : 1'b0;` -- last0 is now gated identically to valid0. VERIFICATION: full 7-test bit-exact-vs-V1 regression (hardware/v2/sim/tb_neural_processor.v under Verilator) -- 7/7 PASS after the fix, including the back-to-back and single-tile-after- multi-tile cases that exposed it. STATUS: FIXED, verified. ERR-0004 (methodology consequence of ERR-0001/ERR-0002) DATE: 2026-09-05 NOTE: this session's V1 certification campaign (docs/validation/, hardware/v1/docs/validation/) was verified exclusively with Icarus Verilog v13.0, the ONLY simulator available on this machine at the time. ERR-0001/ERR-0002 show that v13.0 has at least one real, reproducible scheduling defect around clock-edge/task-entry timing. V1's own testbenches were NOT observed to trigger it in this session (V1's `run_neuron`-style tasks already begin with `@(posedge clk)`, which incidentally avoids ERR-0001's trigger condition), and V1 remains frozen/untouched regardless. This is flagged here for honesty, not to imply V1's certification is wrong -- re-verifying the full V1 suite under Verilator was explicitly OUT OF SCOPE for this V2-kickoff session (V1 is frozen, not to be touched) and was not performed. See decisions.log DEC-0004. STATUS: OPEN CAVEAT, not actioned in this session by design. ERR-0005 (synthesis measurement artifact, WORKED AROUND, not an RTL bug) DATE: 2026-09-05 MODULE: hardware/v2/rtl/neural_processor_array.v SYMPTOM: synthesizing neural_processor_array as a bare top-level module (every per-processor job/operand/result field exposed as a real TRELLIS_IO pin) works at N_PROCESSORS=1 but fails place&route at N_PROCESSORS=2 with "Unable to place cell ...$tr_io, no BELs remaining to implement cell type 'TRELLIS_IO'". ROOT CAUSE: not a logic/timing limit -- the LFE5U-45F-8BG381 package has 245 TRELLIS_IO pins total; the array's wide per-processor buses (input_data/weight_data alone are DATA_WIDTH*P_IN*N_PROCESSORS bits) exceed that budget once N_PROCESSORS>=2, purely because these ports have no on-chip consumer yet (the Memory Manager/M4 and Neural Director/M5 that will drive them in the real system don't exist yet). WORKAROUND: hardware/v2/synthesis/harness_neural_processor_array.v -- a synthesis-only wrapper (NOT part of rtl/, not a functional deliverable) that drives all wide buses from an internal free- running LFSR and reduces outputs to a small checksum, keeping only clk/rst/seed/checksum as real top-level pins. See its own header comment and experiments.log EXP-0003 for the resulting real resource/Fmax numbers. STATUS: WORKED AROUND. Will become moot once M4/M5 exist and the array is synthesized as part of a larger design with on-chip ports instead of a bare top-level module. ERR-0006 (real RTL bugs in hardware/v2/rtl/memory_manager.v, FOUND AND FIXED) DATE: 2026-09-05 MODULE: hardware/v2/rtl/memory_manager.v SYMPTOM: end-to-end M4 testbench (real V1 PSRAM backend + real M1 neural_processor) hung permanently partway through the first multi-tile job -- bank_ready for the "current" bank never became 1, even though the byte-level backend clearly kept completing read transactions (observed via cycle-by-cycle hierarchical tracing of memory_manager/prefetch_engine internal state). ROOT CAUSES (three, found together during the same debugging session): 1. The tile-N+2 prefetch request, queued on tile-N's handoff, could be issued (pf_start asserted) while prefetch_engine was STILL mid-fetch for tile-N+1 -- there was no single-in-flight-request discipline at all in the first draft. Fixed by adding a single-entry pf_pending register: requests are queued, not issued directly, and a dedicated rule launches the queued request only once the engine reports free. 2. Even with that queue, pf_busy does not read 1 until the cycle AFTER pf_start was first observed by prefetch_engine (its own fetch_busy<=1 lags its own fetch_start sampling by one clock) -- checking only `!pf_busy` left a genuine one-cycle window where a second queued request would fire on top of the one just launched, silently overwriting pf_target_bank (and pf_x_addr/ pf_w_addr) for the fetch already in flight. The address corruption was harmless (prefetch_engine had already latched the correct address into its own state that same edge), but pf_target_bank corruption meant the eventually-completed fetch's real data got filed into the WRONG bank's bank_ready/bank_x/bank_w, permanently starving the bank actually needed next. Fixed by gating the issue rule on `!pf_busy && !pf_start` (the extra term closes exactly this one-cycle window). 3. A combinational mux selecting between prefetch_engine's own backend wires and the result-write-back FSM's wires was gated on `state == MM_WRITE_RESULT`, but wr_mem_req (asserted while state==MM_WRITE_RESULT) only becomes valid the FOLLOWING cycle, i.e. while state==MM_DONE -- the mux therefore selected the wrong source for the one cycle the write request pulse was actually high, silently dropping the PSRAM write entirely. Fixed by widening the mux's select condition to cover both states. DIAGNOSIS METHOD: cycle-by-cycle hierarchical signal dumps (mm.state, tile_idx, bank_ready, pf_busy, pf_pending, pf_target_bank, prefetch_engine's own state) under Verilator, printed only on signal-change to keep the trace readable, comparing against the hand-derived expected sequence of events for a 3-tile job. VERIFICATION: hardware/v2/sim/tb_memory_manager.v -- 3/3 tests PASS after all three fixes, including a 5-tile job (steady-state double-buffer swap across more than 2 tiles) and independent PSRAM read-back of the written-back result byte (not just internal signal inspection). STATUS: FIXED, verified end-to-end with the real (unmodified) V1 PSRAM backend chain and a real M1 neural_processor. ERR-0007 (Yosys usage quirk, WORKED AROUND, not an RTL bug) DATE: 2026-09-05 MODULE: hardware/v2/synthesis/harness_dataflow_core.v (build script) SYMPTOM: `chparam -set N_SLOTS 2 dataflow_core` (setting the parameter directly on the NON-top child module, before running `synth_ecp5 -top harness_dataflow_core`) synthesizes with no visible error from the chparam/hierarchy commands themselves, but `synth_ecp5` then fails with "Module `\dataflow_core' referenced in module `\harness_dataflow_core' in cell `\dut' is not part of the design" -- even though a standalone `hierarchy -top harness_dataflow_core` run (no synth_ecp5) with the exact same chparam succeeds. ROOT CAUSE: harness_dataflow_core.v's own instantiation of dataflow_core explicitly overrides N_SLOTS via its own local parameter (`.N_SLOTS(N_SLOTS)`) -- chparam on the child module's DEFAULT is therefore always shadowed at that instantiation site regardless of its value, and synth_ecp5's own internal re-hierarchy pass (distinct from a standalone `hierarchy` call) does not reconcile a chparam'd-but-never-actually-used child default the same way, dropping the generic module reference instead. WORKAROUND: set the parameter on the TOP module being synthesized instead (`chparam -set N_SLOTS 2 harness_dataflow_core`), letting its own instantiation forward the value down to dataflow_core as designed. Confirmed working for both N_SLOTS=2 and N_SLOTS=4. STATUS: WORKED AROUND. A build-script ordering detail, not a defect in dataflow_core.v or harness_dataflow_core.v themselves -- noted here so a future N_SLOTS sweep (M9/M10) does not re-trip over it. ERR-0008 (real RTL bug in the first draft of hardware/v2/rtl/slot_mem_arbiter.v, FOUND AND FIXED) DATE: 2026-09-05 MODULE: hardware/v2/rtl/slot_mem_arbiter.v (M8, new) SYMPTOM: hardware/v2/sim/tb_neural_multiprocessor.v -- node0 (slot 0) completes correctly (result=48), but node1 (slot 1, dispatched concurrently to node0) never completes; its byte-level backend request appears to simply vanish, and the slot hangs forever (watchdog timeout at 20000 cycles, result stays 0). ROOT CAUSE: memory_manager.v/prefetch_engine.v's own byte-level backend protocol (mem_req/mem_wr/mem_addr/mem_wdata/mem_rdata/ mem_ready) is FIRE-AND-FORGET: mem_req is asserted for exactly ONE clock cycle per byte transaction, with no separate "request accepted" acknowledgment -- only `mem_ready` (transaction COMPLETION) exists. M4's own testbench (tb_memory_manager.v) never exposed this because it connects exactly ONE memory_manager directly to int8_memory_access, which is always idle and therefore always able to accept that single pulse the instant it fires. The first arbiter draft only granted a port while its s_req was LIVE that same cycle -- if slot 1's one-cycle pulse arrived on a cycle where the arbiter was already owned by slot 0, the pulse was gone the very next cycle with no record of it ever having happened, and slot 1's prefetch_engine sat in ST_READ_X/ST_READ_W waiting forever for a mem_ready that could never arrive (its request never reached int8_memory_access at all). DIAGNOSIS METHOD: ran the M8 testbench, observed node0 (whichever slot the Director happened to grant the shared bus to first) complete while node1 (the other, concurrently-dispatched slot) hung; traced the fire-and-forget nature of mem_req directly in prefetch_engine.v's own state machine (`mem_req <= 1'b1;` appearing only inside single-cycle state-transition branches, unconditionally cleared to 0 every other cycle) -- confirmed the arbiter's naive "grant only while req is live" logic could not possibly catch a pulse arriving during contention. FIX: every incoming s_req pulse is now LATCHED into a per-port `pending` register (capturing wr/addr/wdata the same cycle), regardless of arbiter state -- the same single-entry "queue, don't drop the request" idiom already used by memory_manager's own pf_pending register (ERR-0006 fix #1). Grants are drawn from `pending`, never from a live s_req directly. This adds a uniform minimum 1-cycle latency to every byte transaction (a real, honestly measured cost of sharing one PSRAM port across N_SLOTS -- see timing.log/benchmark.log EXP-0009), but never drops a request regardless of contention. VERIFICATION: hardware/v2/sim/tb_neural_multiprocessor.v -- 4/4 PASS after the fix (444 cycles end-to-end, vs the buggy draft's 20000- cycle watchdog timeout). hardware/v2/sim/tb_memory_manager.v (M4, untouched) re-run unchanged -- still 3/3 PASS, confirming the fix is entirely contained inside the new arbiter module. STATUS: FIXED, verified end-to-end with the real (unmodified) V1 PSRAM backend chain and real concurrent multi-slot contention.