- PHYSICAL_REALIZATION.md: replace stale "1 tile = 1 burst" layout description with the real EXP-0081/0082 "2 tiles = 1 burst" convention; add EXP-0082 signoff row and history table. - ARCHITECTURE_ANALYSIS.md: mark §5.1 (denser activation packing) DONE with real re-measured numbers (bandwidth ceiling fraction 25%->50%, WNS +0.030->+0.068ns); add §5.4, the real device-data-backed comparison of 32-bit single-channel widening vs a second independent DDR3 channel (decided: 32-bit widening, per real DQS/bank pin-conflict analysis); update scaling-path recommendation to reflect the user's final directive (widen channel -> build DDRManager -> N=2/4/8/16 tests, N=8 target, N=16 documentary). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MUG92aM9m68TRc4rG55BcC
461 lines
27 KiB
Markdown
461 lines
27 KiB
Markdown
# FPGA-Neural V3 — Architecture Analysis: Timing, Bottlenecks, and Recommended Interventions
|
||
|
||
Scope: the current, real, P&R-verified V3 design (`hardware/v3/`, branch
|
||
`v3-artix7`), updated through EXP-0082 (denser activation packing, real
|
||
P&R: WNS +0.068ns). Every number in this document is either directly
|
||
measured (real simulation trace, real P&R report) or a calculation built
|
||
from directly-measured building blocks — the two are labeled explicitly
|
||
throughout. Nothing here is guessed.
|
||
|
||
**Status note (post EXP-0082)**: §5.1 (denser activation packing) described
|
||
below as a *recommendation* is now **DONE and real-P&R-verified** — see the
|
||
"DONE" marker in that section and the updated bandwidth numbers in §3. The
|
||
document originally analyzed the pre-fix state; it's kept below (marked
|
||
historical) because the comparison is itself informative, then updated with
|
||
the real post-fix numbers throughout.
|
||
|
||
---
|
||
|
||
## 1. Executive summary
|
||
|
||
The single most important finding of this analysis: **the system is DDR3
|
||
memory-bandwidth-bound, not DSP-bound, already at N=1 core** — and this is
|
||
true *before* considering any core-count scaling. Adding more compute cores
|
||
(N=4/8/16) without first addressing memory bandwidth would not increase real
|
||
throughput; it would only add more cores contending for the same saturated
|
||
DDR3 channel.
|
||
|
||
| Metric | Value | Source |
|
||
|---|---|---|
|
||
| Real DDR3 back-to-back burst bandwidth (physical channel) | **1.24 GB/s** (9.92 Gbps) | measured, real JEDEC trace (§3.1) — unchanged by packing, this is a physical-channel limit |
|
||
| Real DDR3 bandwidth needed for ONE core at peak DSP throughput | **4.96 GB/s** | calculated from measured DSP rate + memory layout (§3.2) |
|
||
| → DDR3 can sustain, pre-EXP-0081 packing (1 tile/burst) | **~25%** of one core's peak compute throughput | §3.2, historical |
|
||
| → DDR3 can sustain, post-EXP-0081/0082 packing (2 tiles/burst, DONE) | **~50%** of one core's peak compute throughput | §3.2, current, real |
|
||
| Real P&R timing margin (WNS) | **+0.068 ns** | measured, EXP-0082 real P&R (improved from +0.030ns pre-packing) |
|
||
| DSP48E1 headroom for scaling | 224/240 free (93%) | measured, real P&R utilization |
|
||
|
||
The DSP headroom is real and large. The memory-bandwidth ceiling is real,
|
||
and **denser activation packing (§5.1) has already doubled the real
|
||
achievable fraction of it** — from ~25% to ~50% of one core's peak DSP
|
||
throughput, with the real P&R margin *improving*, not degrading, as a side
|
||
effect. This was free leverage and it's now banked.
|
||
|
||
Even at ~50%, DDR3 is still the limiting resource, not DSP count — the next
|
||
real interventions, per the user's own explicit direction, are: (a) widening
|
||
the physical DDR3 channel from 16-bit to 32-bit (§5.5 — doubles the physical
|
||
1.24 GB/s ceiling itself, unlike §5.1 which only reduced waste against a
|
||
fixed ceiling), and (b) an intelligent DDRManager (§5.2) to hide latency via
|
||
orchestrator-driven prefetch. Both are required together — a wider channel
|
||
without a smarter prefetcher still stalls on latency; a smarter prefetcher
|
||
against a 16-bit channel still hits the same physical bandwidth wall.
|
||
|
||
---
|
||
|
||
## 2. Real signoff history (all real P&R runs to date)
|
||
|
||
| EXP | What changed | WNS (ns) | LUTs | DSP48E1 | Notes |
|
||
|---|---|---|---|---|---|
|
||
| 0059 | isolated packed core, out-of-context | n/a (isolated) | 507 | 8 | first real P&R, out-of-context only |
|
||
| 0074 | first real in-context P&R: DDR3 + pins + register file not yet added | +0.040 | 5140 | 16 | first trustworthy board-accurate number |
|
||
| 0076 | + register file, + pin constraints, + SPI physical-layer fix | +0.056 | 5173 | 16 | margin improved slightly (P&R is not perfectly monotonic run to run) |
|
||
| 0078 | + config-flash bridge (real STARTUPE2 placement) | +0.013 | 5213 | 16 | margin dropped — real added logic |
|
||
| 0079 | + real activation-fetch engine (`act_tile_fetch.v`) | +0.030 | 5379 | 16 | pre-packing baseline |
|
||
| 0082 | + denser activation packing (2 tiles/burst, EXP-0081) | **+0.068** | 5437 | 16 | current, final, trustworthy number — margin IMPROVED despite added mux logic |
|
||
|
||
**Observation**: WNS does not move monotonically with LUT count (0.056 →
|
||
0.013 → 0.030 → 0.068 while LUTs only ever grow) — this is normal P&R behavior
|
||
(placer/router heuristics find different solutions each run, small logic
|
||
changes can shift which path is critical). **Do not extrapolate a trend
|
||
line from 3-4 data points** — the only safe practice is a fresh real P&R
|
||
after every real change, which this project already does.
|
||
|
||
**DSP48E1 has stayed at 16 across every real change since EXP-0059's core
|
||
design was fixed** — confirms the packed-DSP MAC design (§4.1) is the
|
||
efficient, stable part of this architecture; all the margin pressure has
|
||
come from *control/glue logic* (arbitration, the flash bridge, the
|
||
activation engine's FSM), not from the compute datapath itself.
|
||
|
||
---
|
||
|
||
## 3. The real memory-bandwidth bottleneck (the analysis's central finding)
|
||
|
||
### 3.1 Real measured DDR3 throughput
|
||
|
||
From the actual `ddr3_model.sv` JEDEC command trace captured during the
|
||
EXP-0079 real simulation run (`tb_n2_system_ddr3.v` via `xsim`), two
|
||
back-to-back `Read` commands to the same open row:
|
||
|
||
```
|
||
Read bank 0 col 000 @ 7090615 ps
|
||
Read bank 0 col 000 @ 7103515 ps (delta = 12900 ps = 12.9 ns)
|
||
```
|
||
|
||
One `BURST_LEN=8` transaction moves 8 × 16 bits = 128 bits. Real measured
|
||
throughput for same-row back-to-back bursts:
|
||
|
||
```
|
||
128 bits / 12.9 ns = 9.92 Gbps = 1.24 GB/s
|
||
```
|
||
|
||
This exactly matches the *theoretical* peak for a 16-bit DDR3 interface at
|
||
310.078 MHz (16 bits × 2 (DDR) × 310.078 MHz = 9.92 Gbps) — confirming the
|
||
real controller achieves its theoretical ceiling for the best case
|
||
(same-row, no row switches). This is the **best-case** number; anything
|
||
requiring a row change (Activate/Precharge) is real-measured to cost more
|
||
(§3.3).
|
||
|
||
### 3.2 Real compute-side bandwidth requirement
|
||
|
||
Each packed core (EXP-0059's design, unchanged since) uses 8 DSP48E1, each
|
||
computing 2 packed INT8 MACs per cycle (lane A + lane B sharing one resident
|
||
weight) = 16 MACs/cycle/core. At the real measured 155.039 MHz compute clock:
|
||
|
||
```
|
||
16 MACs/cycle × 155.039 × 10^6 cycles/s = 2.48 GMAC/s per core (real, from measured Fmax)
|
||
```
|
||
|
||
Each compute cycle consumes 1 activation byte per MAC lane (16 bytes total:
|
||
8 for lane A, 8 for lane B).
|
||
|
||
**Historical (EXP-0079, "1 tile = 1 burst")**: each tile fetch moved a full
|
||
16-byte burst for only 8 useful bytes:
|
||
|
||
```
|
||
2 tiles (A+B) × 16 bytes/burst = 32 bytes moved DDR3 traffic → 16 MACs
|
||
= 2 real DDR3 bytes moved per MAC operation
|
||
2.48 GMAC/s × 2 bytes/MAC = 4.96 GB/s needed per core
|
||
```
|
||
|
||
**Current (EXP-0081/0082, "2 tiles = 1 burst", DONE and real-P&R-verified)**:
|
||
two tiles now share one 16-byte burst, halving the moved-bytes-per-useful-byte
|
||
ratio:
|
||
|
||
```
|
||
2 tiles (A+B) × 16 bytes/burst ÷ 2 tiles-per-burst = 16 bytes moved DDR3 traffic → 16 MACs
|
||
= 1 real DDR3 byte moved per MAC operation (halved)
|
||
2.48 GMAC/s × 1 byte/MAC = 2.48 GB/s needed per core (halved)
|
||
```
|
||
|
||
### 3.3 The real gap
|
||
|
||
```
|
||
Historical: 1.24 GB/s available vs 4.96 GB/s needed per core → ~25% sustainable
|
||
Current: 1.24 GB/s available vs 2.48 GB/s needed per core → ~50% sustainable
|
||
```
|
||
|
||
The physical channel ceiling (1.24 GB/s, §3.1) did **not** change — §5.1's
|
||
fix reduced waste against a fixed ceiling, it did not raise the ceiling
|
||
itself. Even at ~50%, DDR3 remains the binding constraint, not DSP count,
|
||
even in the BEST case (zero row-switch overhead, one core, nothing else
|
||
sharing the bus). Raising the *physical* ceiling requires a wider channel
|
||
(§5.5) or a second channel — both analyzed below.
|
||
|
||
This ceiling gets **worse**, not better, with:
|
||
- **Row switches**: real measured Activate→Read latency is 16.125 ns
|
||
(5 DDR3 clock cycles at 3.225 ns = real CAS-latency-5 timing, matches the
|
||
MIG's own configured CL=5). A tile fetch that requires a fresh row
|
||
activation costs ~16-30 ns instead of the 12.9 ns same-row case — a
|
||
real, measured 25-130% penalty per row switch.
|
||
- **More cores sharing the one DDR3 channel** (N=2 today, N=4/8/16
|
||
proposed): the 1.24 GB/s ceiling is shared across ALL active requesters,
|
||
not per-core. Adding cores divides an already-insufficient budget further.
|
||
- **Weight fetching** (amortized, but not free): each job pair's weight
|
||
prefetch (8 bursts for a 128-byte layer) adds real DDR3 traffic on top of
|
||
activation fetching, though this cost is shared across the M reuse
|
||
positions and becomes negligible for large M.
|
||
|
||
**Conclusion**: this was a genuine architectural ceiling, not a tuning
|
||
problem, and half of it has now been recovered for free. The 2× byte-
|
||
overhead from the original "1 tile = 1 full burst" convention (chosen in
|
||
EXP-0079 specifically to avoid a runtime-indexed part-select, given the
|
||
then-already-thin timing margin) was confirmed, with real numbers, to be the
|
||
single most expensive design decision in the memory path — and has since
|
||
been fixed (§5.1, DONE, EXP-0081/0082) by registering the select bit at
|
||
request time instead of avoiding the select entirely. The remaining gap
|
||
(~50% sustainable, not 100%) is now a *physical channel width* problem, not
|
||
a packing-waste problem — see §5.5.
|
||
|
||
---
|
||
|
||
## 4. Module-by-module review
|
||
|
||
### 4.1 Compute core (`mac2_dsp_packed.v`, `neural_processor_packed.v`)
|
||
|
||
Real, stable, efficient. 8 DSP48E1/core unchanged since EXP-0059. The
|
||
2-INT8-MAC-per-DSP48E1 packing technique is the correct choice for this INT8
|
||
workload — real utilization (16/240 DSP = 6.67% at N=2) confirms there is no
|
||
DSP-side pressure at all; **all scaling headroom is here, and all scaling
|
||
risk is elsewhere** (§3, §4.5).
|
||
|
||
No real change recommended here. This is the part of the design that is
|
||
*not* the bottleneck.
|
||
|
||
### 4.2 Weight-reuse path (`layer_prefetch_ctrl.v`, `layer_weight_buffer.v`,
|
||
`weight_tile_gather.v`)
|
||
|
||
Reused unmodified from V2 (ECP5 era), real and well-verified across many
|
||
EXPs (0057/0058/0061/0062 and every integration test since). Correctly
|
||
amortizes DDR3 traffic across M reuse positions — this part of the design
|
||
already does the "fetch once, use many times" optimization the activation
|
||
path currently lacks (§5.1's recommendation follows the SAME philosophy).
|
||
|
||
No real change recommended; this module is a good template for how the
|
||
activation path should evolve.
|
||
|
||
### 4.3 Activation-fetch path (`act_tile_fetch.v`, EXP-0079, updated EXP-0081/0082)
|
||
|
||
Real, correct (verified 3 levels deep, §2 of EXP-0079's own log entry; the
|
||
EXP-0081 packing change re-verified at all 3 levels again — `tb_act_tile_fetch.v`
|
||
8/8, `tb_packed_slot.v` 9/9, `tb_n2_system_ddr3.v` 8/8, plus real P&R). Two
|
||
real design choices identified in this analysis:
|
||
|
||
1. **1 tile = 1 full burst (2× byte overhead)** — **FIXED (EXP-0081/0082,
|
||
DONE)**: now 2 tiles share 1 burst via a request-time-registered select
|
||
bit (`sel_lat`), avoiding the runtime-indexed-part-select Fmax risk while
|
||
still halving DDR3 waste. Real P&R confirms margin improved, not
|
||
degraded (+0.030ns → +0.068ns). See §5.1.
|
||
2. **Lane A then lane B, sequential, per tile**: still doubles the real
|
||
number of DDR3 transactions (and row-switch risk) versus a design that
|
||
could fetch both lanes in a single wider transaction when they happen to
|
||
be adjacent in memory. Not changed — flagged for future work; the
|
||
DDRManager (§5.2) and 32-bit channel widening (§5.5) are higher-leverage
|
||
and were prioritized first per the user's explicit direction.
|
||
|
||
### 4.4 Scheduling (`neural_director_packed.v`) and arbitration
|
||
(`sdram_arbiter_n.v`)
|
||
|
||
Real, correct, and — importantly for §5.2 — **`neural_director_packed.v`
|
||
already maintains a real queue of pending jobs** (`QUEUE_DEPTH=8`,
|
||
`q_x_base`/`q_w_base`/etc arrays). This is directly relevant to the
|
||
DDRManager/prefetch proposal (§5.2): the information needed to "know what's
|
||
coming next" already exists in this module, it just isn't currently used for
|
||
anything beyond pairing/dispatch decisions.
|
||
|
||
`sdram_arbiter_n.v`'s combinational-first-grant design (EXP-0066) is real,
|
||
proven, and already generalized to N-way (verified at N=3, EXP-0069) — no
|
||
real change needed to scale its own requester count for a DDRManager
|
||
addition or for N=4/8/16 scaling, though the arbiter's own **fairness
|
||
policy** (lowest-index-wins, a deliberate simplicity choice, EXP-0066) would
|
||
need re-examination if a DDRManager starts issuing speculative/anticipatory
|
||
requests that could starve a real, urgent request — see §5.2's own caveat.
|
||
|
||
### 4.5 Host interface (`spi_host_bridge_v3.v`, `flash_spi_master.v`,
|
||
`host_mem_bridge.v`)
|
||
|
||
Real, verified, low resource cost, not on any critical performance path
|
||
(host commands are inherently much slower than the internal compute/memory
|
||
loop). No bottleneck here. Not a scaling concern.
|
||
|
||
### 4.6 Missing: result-writeback engine
|
||
|
||
Still genuinely absent (disclosed since `packed_slot.v`'s own original
|
||
header, unchanged through EXP-0079). Currently `result_data_a/b` are literal
|
||
top-level pins — functional at N=2 (32 pins), but this is the **exact same
|
||
class of mistake already caught once** for activation data (EXP-0074: ~360
|
||
pins nearly exceeded the whole package's I/O budget). At N=16 this port
|
||
alone would need 8 bits × 2 lanes × 16 cores = 256 pins — **a real, hard
|
||
blocker for any scaling beyond a handful of cores**, independent of the
|
||
memory-bandwidth ceiling in §3. Recommended fix in §5.3.
|
||
|
||
---
|
||
|
||
## 5. Recommended interventions, ranked by real leverage
|
||
|
||
### 5.1 [Highest leverage] Denser activation packing — attack the real bandwidth ceiling directly — **DONE (EXP-0081/0082)**
|
||
|
||
**What was built**: 2 tiles (even/odd tile index) now share one
|
||
`BURST_LEN=8` burst instead of one tile per burst, halving real DDR3
|
||
bytes-per-MAC from 2 to 1. Real, measured effect: the achievable fraction
|
||
of one core's peak DSP throughput rose from ~25% to ~50% (§3.2/§3.3,
|
||
current numbers).
|
||
|
||
**Why this was avoided in EXP-0079**: doing so naively requires a
|
||
runtime-indexed part-select (which half of the burst response to use,
|
||
selected by a runtime tile-index bit) — the same anti-pattern
|
||
`weight_tile_gather.v` (EXP-0061) already flagged as a real Fmax risk, and
|
||
the P&R margin was already thin (+0.013ns) when that original design
|
||
decision was made.
|
||
|
||
**The safe path that was actually built**: the tile-index LSB is registered
|
||
into `sel_lat` at *request* time (`S_IDLE`, the same cycle `tcnt` is
|
||
latched) — many `ui_clk` cycles before the real DDR3 round-trip completes
|
||
and `ctrl_rdata` becomes valid. The eventual data-select mux therefore
|
||
selects on an already-long-stable registered bit, never one racing the read
|
||
data. **Real P&R confirms this is genuinely timing-safe, not just
|
||
functionally correct**: margin *improved* from +0.030ns to +0.068ns
|
||
(EXP-0082), despite the added mux logic (LUTs 5379→5437). Full detail:
|
||
`hardware/v3/rtl/act_tile_fetch.v` header, `docs/PHYSICAL_REALIZATION.md` §4,
|
||
`hardware/v2/logs/experiments.log` EXP-0081/EXP-0082.
|
||
|
||
**Verification**: `tb_act_tile_fetch.v` (8/8 PASS, covers even/odd-in-same-
|
||
burst, new-burst crossing, back-to-back alternation), `tb_packed_slot.v`
|
||
(9/9 PASS, bit-identical numeric results to the pre-change run),
|
||
`tb_n2_system_ddr3.v` (8/8 PASS, real xsim against real `ddr3_model.sv`,
|
||
JEDEC trace confirmed to show no more half-burst zero-padding).
|
||
|
||
### 5.2 [Complementary, addresses latency not bandwidth] DDRManager with orchestrator-driven prefetch (user's proposal)
|
||
|
||
**The real problem this solves**: even within whatever bandwidth ceiling
|
||
§5.1 establishes, the CURRENT design only ever requests a tile the moment
|
||
`packed_slot.v`'s own FSM reaches `S_TILEREQ` for it — meaning the DSPs
|
||
stall waiting for that fetch's real latency (§3.3: 12.9-30+ ns) every
|
||
single tile, with no overlap between "fetching tile N+1" and "computing on
|
||
tile N". A DDRManager that issues tile N+1's fetch WHILE tile N is still
|
||
computing would hide that latency almost entirely (compute time per tile,
|
||
1/155.039MHz ≈ 6.4ns per cycle, vs a real fetch latency of 12.9-30+ns —
|
||
today's design is very likely stalling the DSPs for the majority of real
|
||
time, an real, additional cost on top of §3's raw bandwidth ceiling).
|
||
|
||
**What it does NOT solve**: §3.2's bandwidth ceiling is a hard physical
|
||
limit (bytes/second the DDR3 channel can physically move) — prefetching
|
||
earlier doesn't move more bytes per second, it only avoids IDLE gaps where
|
||
the channel is free but nothing is queued to use it. **§5.1 and §5.2 are
|
||
complementary, not alternatives** — §5.1 reduces bytes needed per MAC, §5.2
|
||
ensures the channel is never idle when there's real bandwidth budget
|
||
available and useful work queued. Do both, in this order (§5.1 first, since
|
||
it raises the ceiling §5.2 will then use more fully).
|
||
|
||
**Concrete design sketch** (informed by what already exists in this
|
||
codebase, not a from-scratch proposal):
|
||
|
||
- `neural_director_packed.v` already queues up to `QUEUE_DEPTH=8` pending
|
||
jobs, each with a known `x_base`/`w_base`/`n_tiles` — this is exactly the
|
||
"reservation" information a DDRManager needs. No new bookkeeping is
|
||
required at the Director level; a DDRManager would READ this existing
|
||
queue, not need the Director to change its own job-acceptance logic.
|
||
- A new module (name suggestion: `ddr_prefetch_mgr.v`) would sit between
|
||
`sdram_arbiter_n.v` and the per-slot `act_tile_fetch.v`/
|
||
`layer_prefetch_ctrl.v` instances, with a small staging buffer per slot
|
||
(double-buffered, matching `layer_weight_buffer.v`'s own already-proven
|
||
double-buffer pattern) — while `packed_slot.v` computes on the CURRENT
|
||
tile, the manager issues the request for the NEXT tile into the "other"
|
||
buffer, swapping on completion.
|
||
- **Real caveat, not glossed over**: this adds real arbitration complexity
|
||
— a prefetched-but-not-yet-consumed request competing with another slot's
|
||
genuinely urgent request needs a real priority policy, not just
|
||
`sdram_arbiter_n.v`'s current lowest-index-wins simplicity (§4.4). A
|
||
speculative prefetch that turns out to be wrong (e.g., the Director
|
||
reorders/never dispatches that queued job) also wastes real bandwidth —
|
||
needs a real cancellation/staleness mechanism, not assumed away.
|
||
- **Recommended validation before committing engineering time**: build a
|
||
minimal version scoped to ONE slot's OWN activation-tile look-ahead
|
||
(prefetch tile N+1 while computing tile N, using the double-buffer
|
||
pattern above) before attempting the full "reserve across the whole
|
||
Director queue" version — matches this project's own "one variable at a
|
||
time" discipline, and would give a real, measured stall-reduction number
|
||
to justify (or not) the added complexity of the full design.
|
||
|
||
### 5.3 [Blocking for any real scaling] Result-writeback engine
|
||
|
||
Must exist before N>2 is even attemptable (§4.6) — result data needs to go
|
||
into DDR3 (or through the SPI status/register path for small result sets),
|
||
never as N-scaled literal top-level pins again. Same architectural shape as
|
||
the weight-fetch path, in reverse (write instead of read) — a reasonable,
|
||
bounded scope, and a real prerequisite, not optional polish.
|
||
|
||
### 5.4 [Decided] Widening the physical DDR3 channel: 32-bit single channel vs. a second independent 16-bit channel
|
||
|
||
**The real question**: §5.1 halved *waste* against a fixed 1.24 GB/s
|
||
physical ceiling; it did not raise the ceiling itself. Getting past ~50%
|
||
sustained DSP utilization requires more physical bytes/second, which means
|
||
either (a) widening the existing channel from 16-bit to 32-bit data width,
|
||
or (b) adding a second, independent 16-bit DDR3 channel. Both roughly
|
||
double the real 1.24 GB/s ceiling to ~2.48 GB/s. The user asked for an
|
||
honest comparison, not a diplomatically-balanced non-answer — here it is,
|
||
based on **real device data**, not guessed.
|
||
|
||
**Real device data** (queried directly from the actual Vivado part database
|
||
for this exact part/package, XC7A100T-**CSG324**): this package has only
|
||
**5 total I/O banks** — 14 (56 pins), 15 (56 pins), 16 (11 pins), 34 (56
|
||
pins), 35 (56 pins). All report `BANK_TYPE=BT_HIGH_RANGE` (Artix-7 has no
|
||
separate "HP" bank class the way some other families do). Banks **14 and
|
||
15** each expose **8 DQS-capable pin pairs** — the same memory-PHY
|
||
signature already used by the real, placed DDR3 controller on banks 34/35.
|
||
This means a second, independent DDR3 channel is *physically plausible* on
|
||
this package (the DQS-capable pins exist), but:
|
||
|
||
| Factor | 32-bit single channel | Second independent 16-bit channel |
|
||
|---|---|---|
|
||
| Real ceiling gain | ~2× (1.24 → ~2.48 GB/s) | ~2× (aggregate, same total) |
|
||
| Pin cost | Reuses/extends the existing MIG's own bank(s); no new bank claimed | Would claim banks 14 **and/or** 15 (the only banks with free DQS-capable pins) |
|
||
| **Conflict with already-placed I/O** | None | **Real, direct**: the management-SPI bus (bank 15: A15/B16/B17/A16) and the config-flash bus (bank 14: K17/K18/L13) are already placed in exactly the banks that would need to host a second channel. Only bank 16 (11 pins) would remain free — not enough margin for either SPI bus, let alone both. |
|
||
| Controller logic cost | One MIG instance, wider data path (mostly automatic — the MIG wizard regenerates CAS/CWL/MMCM ratios for the new width) | A full second MIG instance: second calibration sequence, second `ui_clk` domain, and critically the **DDRManager/arbiter would need to become channel-aware**, not just requester-aware — real added complexity on top of §5.2's own design, not a simplification |
|
||
| Real risk given current thin margin (+0.068ns) | Lower — one controller, one clock domain, incremental change to an already-proven design | Higher — two independent PHYs, two calibration state machines, cross-channel coordination logic, all new |
|
||
| PCB impact | None (same pins, same DDR3 part, different bus width usage — **the physical board the user is designing does not need to change** for this) | Would require re-routing/relocating whichever board-level bus (SPI mgmt or flash) currently occupies bank 14/15 pins — real PCB-level rework, not just RTL |
|
||
|
||
**Recommendation (honest, not deferential, as requested)**: **32-bit single-
|
||
channel widening**, not a second independent channel. The bandwidth gain is
|
||
identical, but the 32-bit path has zero pin conflicts with already-placed,
|
||
already-verified I/O (SPI management bus, config-flash bus), a much smaller
|
||
real risk profile against the current thin timing margin, doesn't require
|
||
a second full MIG/calibration instance, and — most importantly for the
|
||
user's own board — needs **no PCB changes**, since it reuses the DDR3 part's
|
||
own existing data pins at a wider access width rather than claiming new
|
||
banks. A second channel's only real advantage (aggregate bandwidth could in
|
||
principle scale further with a 3rd/4th channel later) doesn't apply here —
|
||
this package genuinely has no more free DQS-capable banks to grow into
|
||
after banks 14/15/34/35, so there's no future-proofing benefit being given
|
||
up. **User confirmed this recommendation and it is the decided path
|
||
forward.**
|
||
|
||
**What this requires (not yet done, real, disclosed)**: the real Xilinx MIG
|
||
"Customize IP" wizard must be re-run interactively (Data Width 16→32 AND
|
||
Input Clock Period both changed in the *same* wizard session, since both
|
||
require the wizard's own JEDEC/PLL calculator to recompute CAS Latency/CWL/
|
||
MMCM ratios correctly — this is **not** safe to hand-edit in `mig_a.prj` the
|
||
way the earlier `TargetFPGA` speed-grade field was, per this project's own
|
||
established discipline). This is a real, outstanding, user-gated
|
||
prerequisite before §5.2's DDRManager and any N>2 scaling test can use the
|
||
wider channel.
|
||
|
||
### 5.5 Scaling path recommendation (real numbers, not a guess) — updated per user's final directive
|
||
|
||
Given §3's real bandwidth ceiling: **scaling core count alone, without
|
||
§5.2/§5.4, provides no real additional throughput past whatever N already
|
||
saturates the (post-§5.1) ~2.48 GB/s-equivalent demand ceiling** —
|
||
back-of-envelope, using §3.2's current numbers (2.48 GB/s needed per core
|
||
at peak, 1.24 GB/s physically available today pre-widening), that's already
|
||
close to N≈1 in the worst case and at most N≈2 in the best (zero-row-switch)
|
||
case on the current 16-bit channel. **Building N=4/8/16 before §5.2/§5.4 are
|
||
in place would very likely show near-IDENTICAL real throughput to N=2** — a
|
||
real, wasted engineering cycle the analysis recommends avoiding.
|
||
|
||
**Decided real order of work** (§5.1 already done; this reflects the user's
|
||
own explicit final direction — 32-bit widening, then a complete DDRManager,
|
||
then N=2/4/8/16 tests, real target N=8, N=16 built specifically to document
|
||
where/how it breaks rather than to succeed):
|
||
1. ~~§5.3 (result-writeback)~~ / ~~§5.1 (denser activation packing)~~ — §5.1
|
||
**DONE** (EXP-0081/0082). §5.3 remains a genuine blocker for N>2 and must
|
||
land before any scaling test that needs real result data out of more than
|
||
2 cores' worth of pins.
|
||
2. §5.4 (32-bit channel widening) — **decided**, user-gated on a real
|
||
interactive MIG wizard session (Data Width + Input Clock Period changed
|
||
together). Doubles the physical ceiling itself, which §5.1 alone could
|
||
not do.
|
||
3. §5.2 (DDRManager, "evoluto e completo" per the user's own spec) — build
|
||
against the widened channel so its benefit is measured against the real
|
||
final bandwidth budget, not the pre-widening one.
|
||
4. Real N=2/4/8/16 tests, each with its own real P&R signoff (margin is
|
||
thin, §2 — do not assume a prior N's timing closure predicts the next).
|
||
N=8 is the real target configuration; N=16 is expected to expose real
|
||
bus/arbitration/timing limits and is built specifically to document that
|
||
breakdown, not to be a viable production configuration.
|
||
|
||
---
|
||
|
||
## 6. Summary table: what's real vs. what's a calculation
|
||
|
||
| Claim | Status |
|
||
|---|---|
|
||
| WNS/WHS/LUT/DSP numbers throughout | **Measured** (real Vivado P&R reports, current: EXP-0082) |
|
||
| DDR3 back-to-back burst throughput (1.24 GB/s) | **Measured** (real `ddr3_model.sv` JEDEC trace) — physical channel limit, unchanged by §5.1's packing fix |
|
||
| Real Activate→Read latency (16.125 ns) | **Measured** (same trace) |
|
||
| Per-core compute throughput (2.48 GMAC/s) | **Calculated** from measured Fmax (155.039MHz) + known, fixed DSP-packing factor |
|
||
| Bandwidth needed per core, pre-packing (4.96 GB/s) | **Calculated**, historical (EXP-0079 layout) |
|
||
| Bandwidth needed per core, post-packing (2.48 GB/s) | **Calculated** from the real, as-built EXP-0081/0082 memory layout — current |
|
||
| "~25% of peak sustainable" (pre-packing) / "~50%" (post-packing, current) | **Calculated** ratios; post-packing figure re-verified against real P&R (EXP-0082) and real simulation (§5.1) |
|
||
| I/O bank/DQS pin counts for XC7A100T-CSG324 (banks 14/15/16/34/35) | **Measured** — queried directly from the real Vivado part database for this exact part/package, used in §5.4's dual-channel-vs-widening analysis |
|
||
| Row-switch penalty as a fraction of real workloads | **Not measured** — depends on host-chosen memory layout, flagged as an open question, not asserted |
|
||
| DDRManager's real stall-reduction benefit | **Not measured** — no prototype exists yet; §5.2 recommends building a minimal version specifically to get this real number |
|
||
| 32-bit widening's real post-change bandwidth/timing numbers | **Not measured** — requires the user's own interactive MIG wizard session (§5.4); this document's ~2.48 GB/s figure is a doubling projection, not yet re-verified by real P&R/simulation |
|