Consolidates real, already-measured data from hardware/v1/ (frozen, pre-certified) and V2's own M1-M8 logs into the §32-mandated comparison table, on an apples-to-apples basis: both full systems (V1's spi_neuron_top post_fix_verify vs V2's neural_multiprocessor N_SLOTS=2), both PARALLEL=8/P_IN=8, both using the real unmodified V1 PSRAM backend. Headline, all real measurements: V2 full-system Fmax 142.45 MHz POST-P&R (PASS at 80MHz) vs V1's 68.65 MHz (FAIL at 80MHz); 166 vs 209 real simulated cycles for one neuron's 8-input dot product through the same real PSRAM chain (2.6x wall-clock speedup); peak MAC/cycle 16 (N_SLOTS=2 concurrent slots, real contention already demonstrated in EXP-0009) vs V1's 8 (single sequential core); lower LUT/FF despite V2 already including full dependency-graph scheduling that V1 has none of. 9 of the table's 12 rows carry real sourced numbers; stall %/memory utilization/processor utilization are reported as NOT MEASURED rather than approximated (DEC-0011) - a real number needs dedicated cycle-accounting instrumentation neither system has had built for it yet, and approximating from partial data would violate §30's "no invented results" rule. Deferred to M10, which needs exactly this data to decide what to optimize. No new RTL this milestone - pure data consolidation, logged as EXP-0010. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
466 lines
25 KiB
Plaintext
466 lines
25 KiB
Plaintext
# V2 experiments log -- REGISTRO PRINCIPALE (docs/v2-description.md §25).
|
|
# Un EXP-XXXX per ogni esperimento end-to-end (config -> sim/synth/timing ->
|
|
# risultato). ID mai riutilizzati, anche per i FAIL. Ogni EXP rimanda a
|
|
# reports/experiments/EXP-XXXX/ (experiment.yaml, simulation.log,
|
|
# synthesis.log, timing.log, results.txt, notes.md).
|
|
#
|
|
# Nessun esperimento ancora eseguito.
|
|
|
|
EXP-0001
|
|
timestamp: 2026-09-05T12:03:12Z
|
|
git_commit: 07a48e401f56d395f9d1a6e22f1b5ebc6b598e64 (+ uncommitted M1 work)
|
|
session: v2-M1-neural-processor
|
|
module: hardware/v2/rtl/neural_processor.v
|
|
configuration: DATA_WIDTH=8, P_IN=8, ACC_WIDTH=32
|
|
action: first M1 implementation -- 8-stage pipelined perceptron unit,
|
|
bit-exact vs hardware/v1/rtl/neuron_parallel.v + mac8.v + mac_unit.v
|
|
command (sim, Verilator): verilator --binary --timing -j 0 -Wno-fatal
|
|
--top-module tb -o /tmp/vtb_np hardware/v1/rtl/neuron_parallel.v
|
|
hardware/v1/rtl/mac8.v hardware/v1/rtl/mac_unit.v
|
|
hardware/v2/rtl/neural_processor.v hardware/v2/sim/tb_neural_processor.v
|
|
&& /tmp/vtb_np
|
|
command (synth): yosys -p "synth_ecp5 -json
|
|
hardware/v2/synthesis/neural_processor_p8/top.json -top neural_processor"
|
|
hardware/v2/rtl/neural_processor.v
|
|
command (timing): nextpnr-ecp5 --45k --package CABGA381 --speed 8
|
|
--freq 80 --json hardware/v2/synthesis/neural_processor_p8/top.json
|
|
--lpf-allow-unconstrained --textcfg .../top.config
|
|
result:
|
|
SIMULATED: 7/7 tests PASS, bit-exact vs V1 (functional +
|
|
extreme-INT8 + back-to-back-tiles + 64-input/8-tile coverage).
|
|
SYNTHESIZED: 0 CHECK problems, 8 MULT18X18D, 533 TRELLIS_FF,
|
|
96 CCU2C, 55 LUT4 (36 benign integer-loop-variable warnings, see
|
|
hardware/v2/logs/synthesis.log).
|
|
POST-P&R: Fmax = 183.12 MHz, PASS at 80MHz target (real
|
|
nextpnr-ecp5 measurement, hardware/v2/synthesis/neural_processor_p8/
|
|
nextpnr.log). Compare V1 isolated PARALLEL=8: 61.71 MHz
|
|
(hardware/v1/synthesis/p8/) -- ~3x higher Fmax for the pipelined
|
|
single-processor datapath alone (not yet a system-level
|
|
comparison -- no Memory Manager/Director/multi-processor overhead
|
|
included at this milestone).
|
|
errors: ERR-0001, ERR-0002, ERR-0003 (see errors.log) encountered and
|
|
resolved/worked around during development.
|
|
decision: see decisions.log DEC-0002, DEC-0003, DEC-0004.
|
|
next_action: EXP-0002 (ACC_WIDTH=24 comparison), then M2 (Processor
|
|
Array, N_PROCESSORS sweep).
|
|
|
|
EXP-0002
|
|
timestamp: 2026-09-05T12:03:12Z
|
|
git_commit: 07a48e401f56d395f9d1a6e22f1b5ebc6b598e64 (+ uncommitted M1 work)
|
|
session: v2-M1-neural-processor
|
|
module: hardware/v2/rtl/neural_processor.v
|
|
configuration: DATA_WIDTH=8, P_IN=8, ACC_WIDTH=24 (vs EXP-0001's
|
|
ACC_WIDTH=32 baseline) -- requested explicitly by the user as a test
|
|
variant ("come test nel perceptrone crea una versione con
|
|
accumulatori a 24 bit anziche' a 32 bit").
|
|
reason: ACC_WIDTH is already fully parametric throughout
|
|
neural_processor.v (no code duplication needed); 24 bits is a real,
|
|
non-degenerate choice -- worst-case realistic accumulation (e.g. 256
|
|
INT8 inputs all at +-127/+-128) needs at most 23 signed bits
|
|
(256 * 128 = 32768, |value| <= 2^22 for the largest single-tile-
|
|
chain sums this project's configs reach), so 24 bits carries a
|
|
legitimate 1-bit margin, not an arbitrary/unsafe truncation.
|
|
command (sim, Verilator): same as EXP-0001 with a copy of
|
|
tb_neural_processor.v with ACC_WIDTH localparam changed to 24 (V1's
|
|
own neuron_parallel instantiated at ACC_WIDTH=24 too, for a true
|
|
bit-exact comparison at the SAME width -- V1's ACC_WIDTH is also a
|
|
module parameter, no V1 modification needed).
|
|
command (synth): yosys -p "read_verilog
|
|
hardware/v2/rtl/neural_processor.v; chparam -set ACC_WIDTH 24
|
|
neural_processor; synth_ecp5 -json
|
|
hardware/v2/synthesis/neural_processor_p8_acc24/top.json -top
|
|
neural_processor"
|
|
command (timing): nextpnr-ecp5, same flags as EXP-0001, json from
|
|
the ACC_WIDTH=24 build.
|
|
result:
|
|
SIMULATED: 7/7 tests PASS, bit-exact vs V1 at ACC_WIDTH=24 (no
|
|
overflow in any test vector, as expected from the margin analysis
|
|
above).
|
|
SYNTHESIZED: 0 CHECK problems, 8 MULT18X18D (unchanged, multiplier
|
|
width is DATA_WIDTH-driven not ACC_WIDTH-driven), 509 TRELLIS_FF
|
|
(was 533 at ACC_WIDTH=32, -24 FF as expected for 8 fewer
|
|
accumulator bits carried through ~3 pipeline-stage copies),
|
|
88 CCU2C (was 96), 49 LUT4 (was 55).
|
|
POST-P&R: Fmax = 176.21 MHz (was 183.12 MHz at ACC_WIDTH=32) --
|
|
REAL measurement, both from nextpnr-ecp5. The 24-bit build is
|
|
~4% SLOWER despite fewer resources -- almost certainly placement
|
|
noise (consistent with this project's established finding,
|
|
hardware/v1/docs/WORKLOG.md "Timing closure", that seed-to-seed
|
|
placement variance on this device dominates small logic-width
|
|
differences), NOT attributed to a real architectural effect
|
|
without a seed sweep to confirm. Reported as-measured, not
|
|
overinterpreted -- see hardware/v2/logs/benchmark.log.
|
|
errors: none.
|
|
decision: ACC_WIDTH=32 remains the M1 default (matches V1 exactly for
|
|
ongoing bit-exact comparisons); ACC_WIDTH=24 confirmed as a viable,
|
|
correctly-functioning, slightly-smaller alternative, not adopted as
|
|
default without a proper seed sweep (out of scope for this single
|
|
comparison run).
|
|
next_action: revisit ACC_WIDTH choice during M10 (Optimization) with a
|
|
real seed sweep, not before.
|
|
|
|
EXP-0003
|
|
timestamp: 2026-09-05T14:30:00Z
|
|
git_commit: dc0b331 (+ uncommitted M2 work)
|
|
session: v2-M2-processor-array
|
|
module: hardware/v2/rtl/neural_processor_array.v
|
|
configuration: N_PROCESSORS in {1,2,4,8}, P_IN=8, ACC_WIDTH=32 each
|
|
action: M2 -- functional array + real N_PROCESSORS resource/timing sweep
|
|
command (sim, Verilator): verilator --binary --timing -j 0 -Wno-fatal
|
|
--top-module tb -o /tmp/vtb_array hardware/v2/rtl/neural_processor.v
|
|
hardware/v2/rtl/neural_processor_array.v
|
|
hardware/v2/sim/tb_neural_processor_array.v && /tmp/vtb_array
|
|
command (synth/timing): see hardware/v2/synthesis/harness_n{1,2,4,8}/
|
|
(yosys.log, nextpnr.log) -- synthesized via
|
|
hardware/v2/synthesis/harness_neural_processor_array.v, a
|
|
synthesis-only timing harness (see its own header comment and
|
|
errors.log ERR-0005 for why the array cannot be synthesized as a
|
|
bare top-level module beyond N_PROCESSORS=1 without it).
|
|
result:
|
|
SIMULATED (N_PROCESSORS=4, tb_neural_processor_array.v): 7/7 tests
|
|
PASS -- single-processor sanity, 4 processors launched the SAME
|
|
cycle with different tile counts (finish at different times,
|
|
proving true concurrency), and a staggered-start test (processor 1
|
|
launched mid-way through processor 0's 6-tile job, both complete
|
|
correctly and independently, confirming §18/§34's "un processor
|
|
bloccato non deve bloccare gli altri").
|
|
SYNTHESIZED (resource scaling, harness): perfectly linear in
|
|
N_PROCESSORS, confirming no unintended resource sharing:
|
|
N=1: LUT=59 FF=409 MULT18X18D=8 CCU2C=96
|
|
N=2: LUT=106 FF=786 MULT18X18D=16 CCU2C=192
|
|
N=4: LUT=207 FF=1540 MULT18X18D=32 CCU2C=384
|
|
N=8: LUT=374 FF=3048 MULT18X18D=64 CCU2C=768
|
|
0 CHECK problems in every configuration.
|
|
POST-P&R (real nextpnr-ecp5, --45k --package CABGA381 --speed 8
|
|
--freq 80): Fmax PASS at 80MHz in every configuration:
|
|
N=1: 159.11 MHz N=2: 149.59 MHz
|
|
N=4: 151.01 MHz N=8: 134.70 MHz
|
|
Fmax decreases gently with N (routing congestion), never close to
|
|
failing the 80MHz target up to N=8.
|
|
REAL RESOURCE CEILING FOUND (not assumed, measured via nextpnr's own
|
|
device utilisation report): MULT18X18D usage is 22%/44%/88% of the
|
|
LFE5U-45F's 72 total DSPs at N=2/4/8 respectively, while LUT4/FF
|
|
usage stays under 6% even at N=8. **DSP (not LUT/FF/routing) is
|
|
the first hard ceiling as N_PROCESSORS grows at P_IN=8** --
|
|
N_PROCESSORS=9 would already exceed the device's 72 MULT18X18D
|
|
budget at P_IN=8, before accounting for any multipliers the rest
|
|
of a real system (Memory Manager, PSRAM path, etc.) might need.
|
|
See decisions.log DEC-0005 and benchmark.log.
|
|
errors: ERR-0005 (toplevel-pin-count synthesis artifact, worked around
|
|
with the timing harness -- see errors.log); a first harness attempt
|
|
fed every processor and every MAC lane identical LFSR-derived data,
|
|
which Yosys correctly (from pure logic-equivalence) collapsed via
|
|
CSE down to 1 processor's worth of multipliers regardless of N --
|
|
fixed by giving each processor AND each of its P_IN MAC lanes a
|
|
distinct bit-rotated data source, confirmed by the corrected,
|
|
properly-linear MULT18X18D counts above.
|
|
decision: see decisions.log DEC-0005 (DSP is the binding constraint,
|
|
not LUT/FF -- informs how the N_PROCESSORS x P_IN trade-off should
|
|
be explored going forward).
|
|
next_action: M3 -- activation_buffer.v / weight_buffer.v /
|
|
result_buffer.v.
|
|
|
|
EXP-0004
|
|
timestamp: 2026-09-05T15:15:00Z
|
|
git_commit: 3026dcd (+ uncommitted M3 work)
|
|
session: v2-M3-buffers
|
|
module: hardware/v2/rtl/activation_buffer.v, weight_buffer.v,
|
|
result_buffer.v
|
|
configuration: activation_buffer/result_buffer DEPTH in {4096, 256};
|
|
weight_buffer DEPTH (in tiles) in {512, 64}, P_IN=8, DATA_WIDTH=8
|
|
(TILE_WIDTH=64 bits)
|
|
action: M3 -- three parametric dual-port BRAM-inferring buffers
|
|
(Input/Weight/Result of the §12 data-plane diagram), reusing the
|
|
proven inference idiom from hardware/v1/rtl/act_buffer.v (sync
|
|
write, sync REGISTERED read, no reset on the read register).
|
|
command (sim, Verilator): verilator --binary --timing -j 0 -Wno-fatal
|
|
--top-module tb -o /tmp/vtb_buffers hardware/v2/rtl/
|
|
activation_buffer.v hardware/v2/rtl/weight_buffer.v
|
|
hardware/v2/rtl/result_buffer.v hardware/v2/sim/tb_buffers.v &&
|
|
/tmp/vtb_buffers
|
|
command (synth): yosys -p "read_verilog <file>; chparam -set DEPTH
|
|
<N> <module>; synth_ecp5 -json .../top.json -top <module>" for each
|
|
of the 6 (module, depth) combinations.
|
|
command (timing): nextpnr-ecp5 --45k --package CABGA381 --speed 8
|
|
--freq 80 --lpf-allow-unconstrained, default-depth configs only.
|
|
result:
|
|
SIMULATED: 10/10 tests PASS -- write-then-read correctness incl.
|
|
extreme INT8 (-128/127) round-tripping exactly, back-to-back
|
|
writes to different addresses not disturbing earlier entries, and
|
|
weight_buffer's full 64-bit TILE_WIDTH round-tripping (not just
|
|
one lane).
|
|
SYNTHESIZED (real BRAM mapping, not assumed):
|
|
activation_buffer DEPTH=4096: 2 DP16KD, 37 LUT4, 30 FF
|
|
activation_buffer DEPTH=256: 1 DP16KD, 21 LUT4, 26 FF
|
|
weight_buffer DEPTH=512: 2 DP16KD, 88 LUT4, 139 FF
|
|
weight_buffer DEPTH=64: 2 DP16KD, 73 LUT4, 136 FF
|
|
result_buffer DEPTH=4096: 2 DP16KD, 37 LUT4, 30 FF
|
|
result_buffer DEPTH=256: 1 DP16KD, 21 LUT4, 26 FF
|
|
0 CHECK problems in all 6 configurations -- every one correctly
|
|
inferred DP16KD block RAM, none fell back to LUT-RAM.
|
|
REAL, NON-OBVIOUS FINDING (§14: "non assumere che buffer piu' grandi
|
|
siano automaticamente migliori"): shrinking weight_buffer's DEPTH
|
|
8x (512->64) did NOT reduce DP16KD usage at all (2->2) -- its
|
|
64-bit TILE_WIDTH (P_IN*DATA_WIDTH), not its depth, is what forces
|
|
2 physical block RAMs (a single DP16KD's usable port width in the
|
|
density this needs is narrower than 64 bits). Depth-only buffer
|
|
sizing decisions are the wrong lever for THIS buffer; width
|
|
(P_IN) is. activation_buffer/result_buffer, being byte-wide, DO
|
|
show the expected depth-proportional DP16KD count (2 -> 1).
|
|
POST-P&R (default-depth configs): all PASS at 80MHz with large
|
|
margin (287-367 MHz range, real place&route) -- these buffers are
|
|
not a timing concern in isolation.
|
|
errors: none.
|
|
decision: keep DEPTH parametric as specified, but document (this
|
|
entry + benchmark.log) that weight_buffer's real BRAM cost is driven
|
|
by P_IN (its width), not DEPTH -- relevant when M4/M9 size these
|
|
buffers against a real workload.
|
|
next_action: M4 -- memory_manager.v + prefetch_engine.v (wires these
|
|
three buffers + the array together, PSRAM backend reused unmodified
|
|
from V1 per §15).
|
|
|
|
EXP-0005
|
|
timestamp: 2026-09-05T16:00:00Z
|
|
git_commit: 5f0d7f1 (+ uncommitted M4 work)
|
|
session: v2-M4-memory-manager
|
|
module: hardware/v2/rtl/memory_manager.v, prefetch_engine.v
|
|
configuration: P_IN=8, ADDR_WIDTH=23, real V1 backend chain
|
|
(int8_memory_access -> memory_interface -> psram_controller ->
|
|
psram_model, ALL unmodified), real M1 neural_processor
|
|
action: M4 -- end-to-end integration: memory_manager double-buffers
|
|
prefetch of X/W tiles from PSRAM, feeds a real neural_processor,
|
|
writes the computed result back to PSRAM.
|
|
command (sim, Verilator): verilator --binary --timing -j 0 -Wno-fatal
|
|
--top-module tb -o /tmp/vtb_mm hardware/v1/rtl/int8_memory_access.v
|
|
hardware/v1/rtl/memory_interface.v hardware/v1/rtl/psram_controller.v
|
|
hardware/v1/sim/psram_model.v hardware/v2/rtl/neural_processor.v
|
|
hardware/v2/rtl/prefetch_engine.v hardware/v2/rtl/memory_manager.v
|
|
hardware/v2/sim/tb_memory_manager.v && /tmp/vtb_mm
|
|
command (synth): yosys -p "synth_ecp5 -json .../top.json -top
|
|
memory_manager" hardware/v2/rtl/memory_manager.v
|
|
hardware/v2/rtl/prefetch_engine.v (standalone, for real resource
|
|
counts); harness_memory_manager.v (see errors.log ERR-0005 pattern)
|
|
for real P&R Fmax, since the bare module exceeds the device's
|
|
TRELLIS_IO budget as a top-level (same class of artifact as the
|
|
Processor Array, not a logic limit).
|
|
result:
|
|
SIMULATED: 3/3 jobs PASS end-to-end -- 3-tile (bias=0, ACT_RELU,
|
|
saturates to 127), 1-tile (no saturation, y=32), and 5-tile
|
|
(steady-state double-buffer swap across more than 2 tiles, y=40)
|
|
-- each verified by an INDEPENDENT PSRAM read-back of the
|
|
written result byte (not just internal signal inspection), with
|
|
"poison" bytes surrounding the real operand regions to catch any
|
|
off-by-one addressing (none found).
|
|
Cycle counts (real, PSRAM power-up already excluded): 3-tile job =
|
|
446 cycles, 1-tile job = 166 cycles, 5-tile job = 728 cycles.
|
|
Roughly ~140-150 cycles/tile at this PSRAM timing (dominated by
|
|
the real ~70ns TAA access latency modeled in psram_model.v, not
|
|
by memory_manager's own control overhead) -- a real, measured
|
|
number, not estimated.
|
|
SYNTHESIZED (standalone, real resource count): 0 CHECK problems,
|
|
851 LUT4, 789 TRELLIS_FF, 108 CCU2C, 0 DSP (expected -- no
|
|
multiplication in this module).
|
|
POST-P&R (via harness, real Fmax): 165.86 MHz, PASS at 80MHz with
|
|
large margin.
|
|
errors: ERR-0006 (three real RTL bugs found and fixed during bring-up
|
|
-- see errors.log for full detail: a missing single-in-flight-
|
|
request discipline, a one-cycle pf_busy blind spot, and an
|
|
off-by-one state mux for the write-back path). ERR-0005's pin-count
|
|
artifact recurred for this module too (worked around the same way).
|
|
decision: see decisions.log DEC-0006 (single prefetch engine + pending
|
|
register is sufficient for this milestone's scope; a real backend
|
|
arbiter is deferred until multiple processors/jobs actually need to
|
|
share one memory_manager).
|
|
next_action: M5 -- neural_director.v (first-free scheduling), wiring
|
|
job dispatch to potentially multiple (memory_manager, neural_
|
|
processor) pairs instead of the single hardcoded pair tested here.
|
|
|
|
EXP-0006
|
|
timestamp: 2026-09-05T17:00:00Z
|
|
git_commit: 175f697 (+ uncommitted M5 work)
|
|
session: v2-M5-neural-director
|
|
module: hardware/v2/rtl/neural_director.v
|
|
configuration: N_SLOTS=2 (sim), N_SLOTS=4 (synth default),
|
|
QUEUE_DEPTH=4 (sim), ADDR_WIDTH=23
|
|
action: M5 -- first-free job scheduler dispatching to N_SLOTS
|
|
(memory_manager, neural_processor) pairs, with a parametric-depth
|
|
ready queue.
|
|
command (sim, Verilator): verilator --binary --timing -j 0 -Wno-fatal
|
|
--top-module tb -o /tmp/vtb_dir hardware/v2/rtl/neural_processor.v
|
|
hardware/v2/rtl/prefetch_engine.v hardware/v2/rtl/memory_manager.v
|
|
hardware/v2/rtl/neural_director.v hardware/v2/sim/tb_neural_director.v
|
|
&& /tmp/vtb_dir
|
|
command (synth): yosys -p "synth_ecp5 -json .../top.json -top
|
|
neural_director" hardware/v2/rtl/neural_director.v (standalone, real
|
|
resource counts); harness_neural_director.v (see errors.log
|
|
ERR-0005 pattern) for real P&R Fmax.
|
|
result:
|
|
SIMULATED (N_SLOTS=2, tb_neural_director.v): 4/4 tests PASS --
|
|
3 jobs submitted to 2 slots (first two dispatch immediately,
|
|
first-free; third correctly WAITS in the ready queue until a slot
|
|
frees, then auto-dispatches), each result independently verified;
|
|
a deliberate 2-long-job burst forces the ready queue to genuinely
|
|
fill (job_in_ready correctly deasserts after exactly QUEUE_DEPTH
|
|
queued jobs while both slots are kept busy) and recover once
|
|
slots/queue drain.
|
|
SYNTHESIZED (standalone, N_SLOTS=4): 0 CHECK problems, 382 LUT4,
|
|
366 TRELLIS_FF, 4 CCU2C, 0 DSP.
|
|
POST-P&R (via harness, real Fmax): 250.50 MHz, PASS at 80MHz with
|
|
large margin.
|
|
errors: two real testbench bugs found and fixed during bring-up (not
|
|
RTL bugs): (1) sim_byte_mem's DEPTH parameter (1024) was smaller
|
|
than the test's own address range (up to 0x703=1795), an
|
|
out-of-bounds array access silently returning garbage; (2) the
|
|
initial completion-wait loop exited as soon as ANY ONE of three
|
|
jobs' result bytes changed, not all three -- fixed by counting
|
|
job_out_done pulses instead of polling result memory directly.
|
|
decision: see decisions.log DEC-0007 (reduced 4-state FSM; dependency
|
|
handling deferred to M6; per-slot independent behavioral memory
|
|
instead of shared real PSRAM, deferred to when backend arbitration
|
|
is actually needed).
|
|
next_action: M6 -- dependency_manager.v (ready/waiting queue,
|
|
dependency counters, wake-up, producer tracking) -- the first
|
|
milestone where job READINESS itself, not just free-slot dispatch,
|
|
becomes the Director's actual gating condition.
|
|
|
|
EXP-0007
|
|
timestamp: 2026-09-05T18:00:00Z
|
|
git_commit: 2e4cedc (+ uncommitted M6 work)
|
|
session: v2-M6-dependency-manager
|
|
module: hardware/v2/rtl/dependency_manager.v
|
|
configuration: N_NODES=8 (sim), N_NODES=16 (synth default),
|
|
MAX_DEPS=4, ADDR_WIDTH=23
|
|
action: M6 -- dependency-count tracking table (node_id/state/
|
|
required_dependencies/resolved_dependencies/producer_ids, §10
|
|
exact field list), first-found-ready dispatch to the Director.
|
|
command (sim, Verilator): verilator --binary --timing -j 0 -Wno-fatal
|
|
--top-module tb -o /tmp/vtb_dep hardware/v2/rtl/dependency_manager.v
|
|
hardware/v2/sim/tb_dependency_manager.v && /tmp/vtb_dep
|
|
command (synth/timing): yosys -p "synth_ecp5 -json .../top.json -top
|
|
dependency_manager" hardware/v2/rtl/dependency_manager.v;
|
|
nextpnr-ecp5 --45k --package CABGA381 --speed 8 --freq 80
|
|
--lpf-allow-unconstrained (no timing harness needed this time --
|
|
module's ports fit within the TRELLIS_IO budget as a bare top-level).
|
|
result:
|
|
SIMULATED: 4/4 tests PASS on a small hand-built DAG (node0, node1:
|
|
no dependencies; node2: depends on BOTH node0 and node1 --
|
|
"dipendenze multiple"; node3: depends on node0 ALONE -- "risultati
|
|
condivisi... piu' consumer"): node0/node1 dispatch immediately;
|
|
node3 becomes READY the cycle node0's producer_done arrives (before
|
|
node1 completes); node2 stays WAITING until BOTH node0 AND node1
|
|
have completed, confirmed by an explicit negative check (still
|
|
WAITING after only one of its two dependencies resolved).
|
|
SYNTHESIZED: 0 CHECK problems, 763 LUT4, 474 TRELLIS_FF, 0 DSP,
|
|
0 CCU2C.
|
|
POST-P&R (real, no harness needed): Fmax = 155.30 MHz, PASS at
|
|
80MHz.
|
|
errors: one testbench syntax error (nested nonblocking nested-
|
|
replication `{(N){M{1'b0}}}` needs an extra brace level, Verilator
|
|
correctly rejected it) -- fixed by building reg_producer_ids via
|
|
explicit bit-slice assignment instead of one big concatenation
|
|
expression. Not an RTL bug.
|
|
decision: see decisions.log DEC-0008 (no value forwarding yet, no
|
|
slot reclamation yet -- both explicitly deferred, not missing by
|
|
oversight).
|
|
next_action: M7 -- dataflow_core.v, integrating Director + Dependency
|
|
Manager + Memory Manager + Processor Array + Buffers into one top-
|
|
level module for the first time.
|
|
|
|
[2026-09-05] EXP-0008 -- hardware/v2/rtl/dataflow_core.v (M7, full
|
|
M1-M6 integration)
|
|
test: hardware/v2/sim/tb_dataflow_core.v -- a 3-node DAG (node0/node1
|
|
independent, node2 depends on BOTH) run through the REAL
|
|
dependency_manager -> neural_director -> N_SLOTS x (memory_manager +
|
|
neural_processor) chain end-to-end for the first time, each slot
|
|
backed by its own independent behavioral byte memory (shared real
|
|
PSRAM arbitration explicitly deferred to M8, decisions.log DEC-0009)
|
|
simulator: Verilator 5.050 (--binary --timing)
|
|
PASS/FAIL:
|
|
SIMULATED: 4/4 PASS -- node0=48, node1=8 (correct real
|
|
neural_processor computations via the full stack), node2=40
|
|
dispatched only after BOTH node0 and node1 genuinely completed
|
|
(continuously polled every cycle, not just checked at the end).
|
|
SYNTHESIZED (via harness_dataflow_core.v -- see errors.log ERR-0005):
|
|
N_SLOTS=2: 0 CHECK problems, LUT4=2127, CCU2C=248, TRELLIS_FF=2505,
|
|
MULT18X18D=16, DP16KD=0.
|
|
N_SLOTS=4: 0 CHECK problems, LUT4=3953, CCU2C=500, TRELLIS_FF=4688,
|
|
MULT18X18D=32, DP16KD=0.
|
|
POST-P&R (real, harness-based): N_SLOTS=2 Fmax=165.15 MHz,
|
|
N_SLOTS=4 Fmax=133.19 MHz -- both PASS at 80MHz.
|
|
errors: one Yosys build-script usage quirk (chparam ordering against
|
|
a non-top module vs synth_ecp5's own internal re-hierarchy pass) --
|
|
see errors.log ERR-0007. Not an RTL bug; no dataflow_core.v or
|
|
harness_dataflow_core.v source change needed, only the build command
|
|
itself.
|
|
decision: see decisions.log DEC-0009 (no M3 buffers wired in yet, no
|
|
shared-PSRAM arbitration across slots yet -- both explicitly
|
|
deferred to M8/a future measurement-driven decision, not missing by
|
|
oversight).
|
|
next_action: M8 -- PSRAM integration. Wire the real (unmodified) V1
|
|
PSRAM backend chain (int8_memory_access -> memory_interface ->
|
|
psram_controller) end-to-end through dataflow_core, and design/
|
|
measure whatever N_SLOTS>1 arbitration across ONE physical PSRAM
|
|
port actually requires.
|
|
|
|
[2026-09-05] EXP-0009 -- hardware/v2/rtl/neural_multiprocessor.v (M8,
|
|
real PSRAM integration)
|
|
test: hardware/v2/sim/tb_neural_multiprocessor.v -- same 3-node DAG as
|
|
EXP-0008, now routed through the REAL, UNMODIFIED V1 PSRAM backend
|
|
chain shared across N_SLOTS=2 genuinely concurrent memory_manager
|
|
instances via the new slot_mem_arbiter.v
|
|
simulator: Verilator 5.050 (--binary --timing)
|
|
PASS/FAIL:
|
|
SIMULATED: 4/4 PASS after fixing a real dropped-request bug in the
|
|
arbiter's first draft (errors.log ERR-0008) -- 444 cycles
|
|
end-to-end. hardware/v2/sim/tb_memory_manager.v (M4, untouched)
|
|
re-confirmed 3/3 PASS, no regression.
|
|
SYNTHESIZED (real standalone top-level, no harness needed -- 157
|
|
port bits, real PSRAM pins keep it under the TRELLIS_IO budget):
|
|
0 CHECK problems, LUT4=3145, CCU2C=388, TRELLIS_FF=3659,
|
|
MULT18X18D=16, DP16KD=0.
|
|
POST-P&R (real): Fmax=142.45 MHz -- PASS at 80MHz.
|
|
errors: one real RTL bug (errors.log ERR-0008) -- the first arbiter
|
|
draft silently dropped a request pulse arriving during contention;
|
|
fixed with a per-port pending-request latch, the same "queue, don't
|
|
drop" idiom already used by memory_manager's own pf_pending register
|
|
(ERR-0006).
|
|
decision: see decisions.log DEC-0010 (fixed lowest-index priority
|
|
arbitration, not fairness-balanced -- consistent with every other
|
|
scheduling policy chosen so far in this roadmap; revisit only if
|
|
M9's real measurement shows starvation actually matters).
|
|
next_action: M9 -- Full benchmark. Produce the V1-vs-V2 comparison
|
|
table mandated by §32 (Fmax, LUT, FF, DSP, BRAM, MAC/cycle,
|
|
cycles/neuron, neurons/s, stall %, memory/processor utilization,
|
|
effective MAC/s), each number labeled THEORETICAL/SIMULATED/
|
|
SYNTHESIZED/POST-P&R per §30.
|
|
|
|
[2026-09-05] EXP-0010 -- M9 Full benchmark: V1 vs V2 comparison table
|
|
(docs/v2-description.md §32)
|
|
test: not a new simulation -- a consolidation of real, already-
|
|
measured data from hardware/v1/ (frozen, pre-certified) and V2's own
|
|
M1-M8 logs (EXP-0001..EXP-0009), placed side-by-side on an
|
|
apples-to-apples basis (both full-system, both PARALLEL=8/P_IN=8,
|
|
both using the real unmodified V1 PSRAM backend).
|
|
result: see benchmark.log's own M9 entry for the full 12-row table.
|
|
Headline: V2's full system (neural_multiprocessor, N_SLOTS=2) is
|
|
142.45 MHz POST-P&R (PASS @80MHz) vs V1's 68.65 MHz (FAIL @80MHz);
|
|
166 vs 209 real simulated cycles for one neuron's 8-input dot
|
|
product through the real PSRAM chain -- 2.6x wall-clock speedup,
|
|
measured, not assumed.
|
|
errors: none this milestone (pure data consolidation, no new RTL).
|
|
decision: see decisions.log DEC-0011 -- 3 of the 12 §32 rows (stall %,
|
|
memory utilization, processor utilization) are reported as NOT
|
|
MEASURED rather than approximated, since a real number would require
|
|
dedicated cycle-accounting instrumentation neither system has had
|
|
built for it yet; approximating from partial data would violate
|
|
§30's "no invented results" rule.
|
|
next_action: M10 -- Optimization, using the REAL data gathered in M1-
|
|
M9 (not blind guessing): revisit memory_manager's +1-cycle/tile
|
|
overhead (DEC-0006), the ACC_WIDTH 24-vs-32 Fmax question (EXP-0002,
|
|
inconclusive without a seed sweep), the N_SLOTS x P_IN trade-off
|
|
given the DSP ceiling (DEC-0005), dependency_manager's node-slot
|
|
reclamation (DEC-0008), and slot_mem_arbiter's fixed-priority
|
|
fairness question (DEC-0010) -- plus building the stall %/
|
|
utilization instrumentation DEC-0011 deferred, since M10 is exactly
|
|
where that data becomes actionable.
|