chore: remove root-level V1 duplicates, superseded by hardware/v1/ freeze
hardware/v1/ was created (dc0b331) as a frozen snapshot of the V1
project that then lived at the repo root (rtl/, sim/, synth/, tools/,
docs/). Root received zero further commits to those files after the
freeze -- confirmed byte-identical to the hardware/v1/ copy for every
file removed here. Root was the "before", hardware/v1/ is the
curated, canonical "after".
Removed (all verified exact-hash duplicates of hardware/v1/ content):
- rtl/ (20 files, 100% covered by hardware/v1/rtl/)
- tools/{netasm,pinout,run_regression.py,flash_catalog,validation,
fpga_benchmark.py} (19 files, 100% covered by hardware/v1/tools/;
tools/neural_sim/ kept -- unique, post-freeze, no counterpart)
- sim/*.v (47 testbenches, 100% covered by hardware/v1/sim/; the
~38 remaining sim/ entries are compiled binaries and .vcd
waveform dumps, left as a separate cleanup decision)
- synth/ecp5/{p2,p4,p8,post_fix_verify} (25 files, exact duplicates
of hardware/v1/synthesis/; the other ~84 synth/ecp5/* experiment
build directories are historical artifacts never carried into the
freeze, left as a separate decision)
- WORKLOG.md (duplicate of hardware/v1/docs/WORKLOG.md)
- docs/{FPGA-Neural-Datapatch-Benchmark,FPGA-Neural-Hardware-Design,
FPGA-NeuralNetwork-Engine}.md, docs/validation/*.md (18 files),
docs/FPGA-Neural-Datasheet-{EN,IT}.pdf -- all exact duplicates of
hardware/v1/docs/ content
- hardware/v1/docs/DatasheetLatex/ (24 files) -- exact duplicate of
hardware/v2/docs/datasheet/files/docs/datasheet/en/ (discovered
during this audit; not the same DatasheetLatex already removed
from hardware/v2/docs/ in an earlier commit)
Moved (genuine, unique, post-freeze V2 content -- not duplicated
anywhere, just living in the wrong/legacy root docs/ location):
- docs/architecture/*.md -> hardware/v2/docs/architecture/
- docs/pinouts.md, docs/FPGA_NEURAL_V2_DATASHEET.md,
docs/FPGA_NEURAL_V2_SCHEMATIC.md,
docs/FPGA-Neural-V2-Datasheet-EN.pdf -> hardware/v2/docs/
Left untouched (separate decisions, not part of this cleanup):
- docs/FPGA-Neural-Flash-Subsystem-Verification.md, docs/
v2-description.md -- orphaned root-only content, no duplicate
found anywhere, but also not part of the reviewed plan
- synth/ecp5/* experiment dirs and sim/*_sim + sim/*.vcd build
artifacts -- not literal duplicates, flagged as candidates for a
future, separate cleanup pass
Verified no functional breakage: grepped all remaining scripts/docs
for references to every removed path -- only prose/comment mentions
found, no executable imports or build-script paths broken.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
This commit is contained in:
@@ -0,0 +1,145 @@
|
||||
# NMS Real Weight Prefetch Engine (STEP11)
|
||||
|
||||
Status: implemented, bit-exact verified, benchmarked against the real
|
||||
V1 PSRAM chain, synthesized. **Not adopted as the default NMS
|
||||
configuration** — see Outcome/Recommendation below. Full data:
|
||||
`hardware/v2/nms/reports/nms_prefetch_sweep.csv`,
|
||||
`nms_prefetch_summary.md`; full narrative:
|
||||
`hardware/v2/logs/experiments.log` (EXP-0023, EXP-0024),
|
||||
`decisions.log` (DEC-0023), `errors.log` (ERR-0015).
|
||||
|
||||
## Problem
|
||||
|
||||
The "Current NMS" baseline (`nms_memory_manager.v`, backed by
|
||||
`prefetch_engine.v`) measured `prefetch_effectiveness≈0%` and
|
||||
`weight_stall≈92.5%` at N_SLOTS=2 (EXP-0022). Tracing the actual RTL
|
||||
(not assuming from filenames) showed the real gap: `prefetch_engine.v`
|
||||
is a single-shot FSM (`ST_IDLE`/`ST_READ_W`/`ST_DONE`) that can only
|
||||
have **one fetch in flight at a time**, and `nms_memory_manager.v`'s
|
||||
own restart logic only re-triggers the next tile's fetch once the
|
||||
*previous* tile's fetch has fully completed and the FSM has returned
|
||||
to idle — paying a real per-tile control-plane restart cost on every
|
||||
tile boundary. The gap was never insufficient lookahead *distance*
|
||||
(the old design already tried to fetch as far ahead as `n_tiles`
|
||||
allowed); it was zero *outstanding-request depth*.
|
||||
|
||||
## Real backend constraint
|
||||
|
||||
`memory_interface.v` → `psram_controller.v` (V1, reused verbatim,
|
||||
never modified) is a fire-and-forget, **one-transaction-in-flight**
|
||||
protocol: a single `mem_req` pulse, wait for `mem_ready`, and that IS
|
||||
the whole transaction. No wire-level pipelining is physically possible
|
||||
against a real single PSRAM port. So "multiple outstanding requests"
|
||||
cannot mean multiple simultaneous word transactions — it means
|
||||
eliminating the *control-plane* overhead paid at every tile boundary
|
||||
and letting the fetch stream run continuously across tiles, queueing
|
||||
up to `PREFETCH_DISTANCE` tiles of lookahead ahead of consumption.
|
||||
|
||||
## Design: `weight_prefetch_engine.v`
|
||||
|
||||
Two monotonic counters fully describe the engine (tiles are always
|
||||
fetched in strict sequential order, never reordered or re-fetched, so
|
||||
no per-tile state array is needed):
|
||||
|
||||
- `fetch_tile`/`fetch_word` — the next word to request (or the word
|
||||
currently in flight).
|
||||
- `ready_count` — tiles 0..`ready_count`-1 are fully resident in the
|
||||
weight SRAM.
|
||||
|
||||
`consumed_count` (the consumer's own tile index, `nms_memory_manager_pf.v`'s
|
||||
`tile_idx`) bounds a configurable lookahead window:
|
||||
`window_limit = consumed_count + PREFETCH_DISTANCE`; the engine may
|
||||
fetch tile K only if `K < n_tiles` **and** `K < window_limit`.
|
||||
|
||||
The core mechanism: on `mem_ready && req_outstanding`, the just-completed
|
||||
word is committed **and**, in the same cycle, the very next request is
|
||||
issued — either the same tile's next word, or (at a tile boundary) the
|
||||
next tile's first word — giving zero-gap streaming across tile
|
||||
boundaries against a backend that only ever has one word in flight.
|
||||
(An earlier draft used mutually-exclusive `if/else-if` branches for
|
||||
"commit" vs. "issue next", which reintroduced a 1-cycle gap between
|
||||
*every* word, not just tile boundaries; fixed by merging both into one
|
||||
branch — see `weight_prefetch_engine.v`'s own header comment.)
|
||||
|
||||
## Integration: the "_pf" A/B variants
|
||||
|
||||
Per the explicit "preserve the current NMS baseline" constraint, the
|
||||
new engine was integrated into parallel `_pf`-suffixed files, leaving
|
||||
the originals untouched:
|
||||
|
||||
- `nms_memory_manager_pf.v` — drop-in replacement for
|
||||
`nms_memory_manager.v`'s external interface; internally swaps the
|
||||
private `prefetch_engine.v` instance for `weight_prefetch_engine.v`,
|
||||
and changes `can_present`'s weight-ready check from
|
||||
`tile_idx < wgt_fetched` to `tile_idx < wgt_ready_count`.
|
||||
- `nms_dataflow_core_pf.v` — mirrors `nms_dataflow_core.v`, adds a
|
||||
`PREFETCH_DISTANCE` parameter, instantiates `nms_memory_manager_pf`.
|
||||
- `nms_neural_multiprocessor_pf.v` — mirrors
|
||||
`nms_neural_multiprocessor.v`, instantiates `nms_dataflow_core_pf`.
|
||||
|
||||
Both the baseline (`nms_neural_multiprocessor.v`) and the prefetch
|
||||
variant (`nms_neural_multiprocessor_pf.v`) remain in the repository
|
||||
side by side; neither supersedes the other.
|
||||
|
||||
## Verification
|
||||
|
||||
`hardware/v2/nms/sim/tb_weight_prefetch.v` — isolated correctness
|
||||
testbench: real `sim_word_mem` (configurable extra latency), real
|
||||
`nms_weight_packed.v` production SRAM, bit-exact fill-pattern checking.
|
||||
Covers `n_tiles ∈ {0,1,2,PFD,PFD+1,MAX_TILES-1,MAX_TILES}`, back-to-back
|
||||
jobs with no explicit reset, a dedicated windowing-cap test (frozen
|
||||
consumer, confirms `ready_count` stops exactly at
|
||||
`min(PFD,MAX_TILES)`), and (post-ERR-0015) a large-PFD regression case.
|
||||
10/10 (9/9 at PFD≥MAX_TILES) tests pass bit-exact across
|
||||
PFD∈{1,2,4,8,32} and under injected extra memory latency.
|
||||
|
||||
`hardware/v2/nms/sim/tb_nms_dstress_pf.v` — full real-integration
|
||||
benchmark: identical D-Stress workload/golden-model/correctness
|
||||
criteria as `tb_nms_dstress.v` (EXP-0022), instantiating
|
||||
`nms_neural_multiprocessor_pf` with a `PFD_CFG` parameter, plus new
|
||||
testbench-only instrumentation for `weight_stall_cycles` and
|
||||
`prefetch_effectiveness` (tiles consumed with zero weight-blocking
|
||||
cycles beforehand / total tiles consumed — the exact STEP11
|
||||
definition). All runs pass 256/256 neurons bit-exact vs. the golden
|
||||
model.
|
||||
|
||||
## ERR-0015: a real bug found and fixed
|
||||
|
||||
The initial `window_limit` computation truncated the
|
||||
`PREFETCH_DISTANCE` *parameter itself* to `CNTW` bits
|
||||
(`PREFETCH_DISTANCE[CNTW-1:0]`) before adding it to `consumed_count`.
|
||||
At `MAX_TILES=16` (`CNTW=5` bits), `PFD=32` truncates to 0, making
|
||||
`window_limit == consumed_count` forever and deadlocking the engine
|
||||
completely (0/256 neurons ever completed, 0% PSRAM utilization).
|
||||
Fixed by computing `window_limit` and its comparisons in a fixed
|
||||
32-bit width, using the untruncated parameter value. Regression-tested
|
||||
in `tb_weight_prefetch.v`. Full writeup: `errors.log` ERR-0015.
|
||||
|
||||
## Results and outcome
|
||||
|
||||
See `nms_prefetch_summary.md` for the full comparison table and the
|
||||
nine explicitly-answered final-report questions. In short:
|
||||
|
||||
- **N_SLOTS=1** (no port contention): a real, reproducible **-10.3%**
|
||||
cycle-count improvement (PFD=1 → PFD≥2), then a complete plateau —
|
||||
deeper buffering gives zero further benefit. Sustained MAC/cycle
|
||||
reaches only 2.8% of the theoretical target.
|
||||
- **N_SLOTS=2** (this project's own primary reference configuration,
|
||||
real shared-port contention via `slot_mem_arbiter`): **zero
|
||||
measurable benefit** at any PREFETCH_DISTANCE from 1 to 16 — all
|
||||
runs are statistically indistinguishable from each other and from
|
||||
the pre-STEP11 baseline. The single physical PSRAM port is already
|
||||
saturated (90.5% busy, unchanged from baseline) by natural two-slot
|
||||
contention before any lookahead scheme can act.
|
||||
|
||||
**Final decision: Outcome B (N_SLOTS=1, partial) / Outcome C
|
||||
(N_SLOTS=2, failure against the 90% criterion).** The mechanism is
|
||||
correct and does measurably hide latency when the port has spare
|
||||
capacity; it cannot manufacture bandwidth out of an already-saturated
|
||||
single physical port. Reaching the STEP11 target would require ~36×
|
||||
(N=1) to ~82× (N=2) more real PSRAM bandwidth — a hardware-level
|
||||
constraint, not an RTL-scheduling one. Per DEC-0023, the new engine is
|
||||
**not** recommended as the default NMS configuration; both variants
|
||||
are preserved for reference. The evidence-backed next step (real PSRAM
|
||||
bandwidth — wider bus, multiple independent banks, or a faster backing
|
||||
technology) is flagged as future work, not undertaken this round.
|
||||
Reference in New Issue
Block a user