Implements M3: three parametric dual-port buffers for the §12 data-plane (Input/Weight/Result), reusing the proven BRAM-inference idiom from the frozen hardware/v1/rtl/act_buffer.v (synchronous write, synchronous REGISTERED read, no reset on the read register -- keeps Yosys off the LUT-RAM path). Verified with Verilator: 10/10 tests pass (write-then-read correctness, extreme INT8 round-tripping, weight_buffer's full 64-bit tile width round-tripping, undisturbed re-reads). Real synthesis at two depths per module (6 configs total): 0 CHECK problems, every configuration correctly infers DP16KD (never LUT-RAM). Non-obvious real finding: weight_buffer's BRAM cost is driven by its P_IN*DATA_WIDTH tile width, not its DEPTH -- an 8x depth reduction (512->64) left DP16KD usage unchanged at 2, while activation_buffer/result_buffer (byte-wide) scale as naively expected (2->1). All default-depth configs PASS at 80MHz with large margin (287-367 MHz) via real nextpnr-ecp5 place&route. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
39 lines
2.1 KiB
Plaintext
39 lines
2.1 KiB
Plaintext
# V2 synthesis log -- solo append, mai troncato/sovrascritto (vedi README.md)
|
|
# Nessuna entry ancora -- popolato incrementalmente man mano che avanza lo sviluppo V2.
|
|
|
|
[2026-09-05] EXP-0001 -- neural_processor (P_IN=8, ACC_WIDTH=32)
|
|
LUT: 55 FF: 533 DSP(MULT18X18D): 8 BRAM: 0 CCU2C: 96
|
|
CHECK: 0 problems. 36 warnings, all "multiple conflicting drivers for
|
|
neural_processor.\gi" -- benign Yosys quirk for an `integer` used as
|
|
a synthesizable for-loop index in stage 0's unrolled always block,
|
|
not a real multi-driver conflict (cross-verified functionally
|
|
correct on 2 independent simulators). Log: hardware/v2/synthesis/
|
|
neural_processor_p8/yosys.log
|
|
|
|
[2026-09-05] EXP-0002 -- neural_processor (P_IN=8, ACC_WIDTH=24)
|
|
LUT: 49 FF: 509 DSP(MULT18X18D): 8 BRAM: 0 CCU2C: 88
|
|
CHECK: 0 problems, same 36 benign warnings as EXP-0001.
|
|
Log: hardware/v2/synthesis/neural_processor_p8_acc24/yosys.log
|
|
|
|
[2026-09-05] EXP-0003 -- neural_processor_array via
|
|
harness_neural_processor_array.v (synthesis-only wrapper, see
|
|
errors.log ERR-0005), N_PROCESSORS in {1,2,4,8}, P_IN=8
|
|
N=1: LUT=59 FF=409 MULT18X18D=8 CCU2C=96
|
|
N=2: LUT=106 FF=786 MULT18X18D=16 CCU2C=192
|
|
N=4: LUT=207 FF=1540 MULT18X18D=32 CCU2C=384
|
|
N=8: LUT=374 FF=3048 MULT18X18D=64 CCU2C=768
|
|
CHECK: 0 problems in all 4 configurations. Perfectly linear scaling in
|
|
N confirms no unintended cross-processor resource sharing (a first,
|
|
flawed harness attempt fed identical data to every processor/lane
|
|
and Yosys silently deduplicated down to 1x regardless of N -- caught
|
|
by checking for exactly this linearity before trusting the numbers).
|
|
|
|
[2026-09-05] EXP-0004 -- activation_buffer/weight_buffer/result_buffer,
|
|
2 depths each
|
|
activation_buffer D=4096: LUT=37 FF=30 DP16KD=2 (D=256: LUT=21 FF=26 DP16KD=1)
|
|
weight_buffer D=512: LUT=88 FF=139 DP16KD=2 (D=64: LUT=73 FF=136 DP16KD=2)
|
|
result_buffer D=4096: LUT=37 FF=30 DP16KD=2 (D=256: LUT=21 FF=26 DP16KD=1)
|
|
CHECK: 0 problems, all 6 configs correctly infer DP16KD (no LUT-RAM
|
|
fallback). weight_buffer's DP16KD count does NOT drop with depth
|
|
(width-bound, not depth-bound -- see decisions.log / benchmark.log).
|