Files
FPGA-Neural/hardware/v2/reports/step17_cycle_decomposition.csv
micheleandClaude Sonnet 5 8e014d8d49 V2.0.0 hardware freeze - single SDRAM
FASE #1 hardware freeze for FPGA-Neural V2, N4/P8, single external
SDRAM (Alliance Memory AS4C4M16SA-6TIN) serving weights, activations,
and results through one physical sdram_controller.v instance. Removes
the PSRAM dependency (hardware/v1/rtl/psram_controller.v +
memory_interface.v) from the V2 physical path entirely -- V1 itself
remains fully unmodified, the golden reference.

New RTL: sdram_unified_backend.v (2-way W/AR arbitration over one
SDRAM controller, real per-byte DQM write masking added to
sdram_controller.v for correct single-byte result writes with no
read-modify-write), nms_neural_multiprocessor_sdram_unified.v (the
frozen top-level). Two real bugs found and fixed via full-system
testing before being accepted (ERR-0023): a deadlock and an off-by-one
data-shift bug in the new arbitration logic.

Real results: N=4 and N=2 D-Stress bit-exact (256/256 neurons), 40
real AUTO REFRESH events interleaved with zero corruption, real
Yosys+nextpnr-ecp5 synthesis/P&R for LFE5U-45F-8CABGA381 (149/245
TRELLIS_IO, a real 45-pin reduction from the prior dual-memory
design). Timing is MARGINAL (1/8 P&R seeds >=80MHz), reported honestly
rather than masked by the best seed.

Real, sourced ball-level pinout for the SDRAM bus + clk/rst (39/149
signals, P&R-verified) using the official Lattice ECP5U-45 pinout CSV
found on disk during this step's own pre-commit review -- corrects an
earlier draft that wrongly assumed no real pinout data was available.

Chip readiness: NO. Real, disclosed blockers remain (no physical host
interface exists yet -- the RTL's own reg_* ports are a 110-pin raw
test-harness bus; clock source/PLL decision; power/configuration
component selection) -- see hardware/v2/docs/{HARDWARE_FREEZE,
CHIP_READINESS,OPEN_ITEMS}.md for the complete, itemized status.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
2026-09-06 13:39:55 +02:00

1.8 KiB

1metricn_slots2n_slots4unitclassification
2total_cycles5216149430cyclesINTEGRATED BENCHMARK
3tiles_delivered40964096tilesINTEGRATED BENCHMARK
4slot_cycle_budget104322197720slot-cycles (N_SLOTS*total_cycles)DERIVED
5useful_mac_cycles40804078slot-cycles (tile-delivery events summed across slots)INTEGRATED BENCHMARK
6useful_mac_cycles_pct3.912.06percent of slot_cycle_budgetDERIVED
7weight_stall_cycles86464179756slot-cyclesINTEGRATED BENCHMARK
8weight_stall_pct82.8890.91percent of slot_cycle_budgetDERIVED
9per_slot_idle_cycles_sum10721772slot-cycles (sum of total_cycles-busy_cycles per slot)DERIVED
10per_slot_idle_pct1.030.90percent of slot_cycle_budgetDERIVED
11unaccounted_residual_slot_cycles1270612114slot-cyclesDERIVED
12unaccounted_residual_pct12.186.13percent of slot_cycle_budget (plausibly activation-wait + pipeline/tile-boundary bubbles + dispatch overhead -- NOT separately isolated this round; also includes a small ~16-tile/slot-cycle discrepancy between the two independent tile-delivery counters used, itself unexplained and disclosed rather than papered over)DERIVED
13startup_cycles5656cycles (before first tile delivered anywhere)INTEGRATED BENCHMARK
14drain_cycles3232cycles (after last tile delivered until completion)INTEGRATED BENCHMARK
15sustained_mac_per_cycle0.62820.6629MAC/cycleINTEGRATED BENCHMARK
16theoretical_peak_mac_per_cycle1632MAC/cycleTHEORETICAL
17processor_utilization_pct3.932.07percent (sustained/theoretical_peak)DERIVED
18active_slots_0_pct0.030.03percent of cyclesINTEGRATED BENCHMARK
19active_slots_1_pct2.000.03percent of cyclesINTEGRATED BENCHMARK
20active_slots_2_pct97.970.66percent of cyclesINTEGRATED BENCHMARK
21active_slots_3_pctNA2.07percent of cyclesINTEGRATED BENCHMARK
22active_slots_4_pctNA97.22percent of cyclesINTEGRATED BENCHMARK