docs: consolidate all V2 datasheets into one current, complete document

The repository had accumulated multiple, contradictory "current state"
documents for V2 hardware: an old V1 IT/EN datasheet copy nested inside
hardware/v2/docs/datasheet/, a stray untracked duplicate at repo root
(docs/DatasheetLatex/), and a second, much older documentation track
(hardware/v2/docs/*.md: PRE_PCB_VERIFICATION.md, PRE_PCB_CLOSURE_4POINT.md,
MEMORY_UPGRADE_64MB_N8.md, and 10 more) describing an earlier PSRAM/
N_SLOTS<=2 milestone alongside the real, current SDRAM/N_SLOTS=4 board.
The LaTeX datasheet's own front matter (features/pinout cover pages) and
chapter 9 (benchmarks) were themselves still describing that obsolete
architecture, contradicting the real, current chapters 5/7/10 elsewhere
in the same document.

This commit:
- Flattens hardware/v2/docs/datasheet/files/docs/datasheet/v2-en/* up to
  hardware/v2/docs/datasheet/ (was 4 levels of redundant nesting).
- Removes the old V1 IT/EN LaTeX copies and the stray root-level
  duplicate entirely (recoverable from git history, not from disk).
- Preserves the real component reference PDFs (ECP5 eval board, ISSI
  PSRAM, programming cables) under datasheet/references/.
- Removes 13 superseded hardware/v2/docs/*.md status documents after
  folding every real, unique fact they contained into the datasheet:
  SPI max verified clock (12MHz, exact 12.8MHz CDC edge), SDRAM directed
  boundary test (21/21 PASS), 16MHz oscillator MPN (ECS-3225MV-160-BN-TR),
  and the real FPGA<->SDRAM ball mapping cross-check.
- Rewrites the datasheet's own front matter, ch.4 (parameters), ch.8
  (top-level module -- was documenting the wrong, non-physical top
  entirely), and ch.9 (benchmarks) to describe the current, real SDRAM/
  N_SLOTS=4 production board, while keeping the real PSRAM-era chapters
  as clearly-labeled history rather than deleting correctly-measured
  work.
- Fixes a title-page tikzpicture that was clipped off the page edge
  (pre-existing, unrelated to this change) by scaling it to fit.

Net: 85 files changed, -8814/+498 lines. hardware/v2/docs/ now contains
exactly one current datasheet plus FIRST_POWER_ON.md (a bring-up
runbook, not a duplicate spec).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
This commit is contained in:
2026-09-09 00:41:06 +02:00
co-authored by Claude Sonnet 5
parent b7b4ada99f
commit 07f8ec0460
69 changed files with 494 additions and 5896 deletions
Binary file not shown.
@@ -41,19 +41,21 @@
% --- compact block diagram on the title page --- % --- compact block diagram on the title page ---
\begin{center} \begin{center}
\resizebox{\textwidth}{!}{%
\begin{tikzpicture}[node distance=7mm and 10mm] \begin{tikzpicture}[node distance=7mm and 10mm]
\node[fnblockD,minimum width=26mm] (host) {HOST\\{\scriptsize graph loader}}; \node[fnblockD,minimum width=26mm] (host) {HOST\\{\scriptsize graph loader}};
\node[fnblockT,right=14mm of host,minimum width=30mm] (dm) {Dependency\\Manager}; \node[fnblockT,right=14mm of host,minimum width=30mm] (dm) {Dependency\\Manager};
\node[fnblockT,right=14mm of dm,minimum width=28mm] (dir) {Neural\\Director}; \node[fnblockT,right=14mm of dm,minimum width=28mm] (dir) {Neural\\Director};
\node[fnblock,right=14mm of dir,minimum width=34mm] (slots) {N\_SLOTS $\times$ (Memory\\Manager $+$ Neural Proc.)}; \node[fnblock,right=14mm of dir,minimum width=34mm] (slots) {N\_SLOTS $\times$ (Memory\\Manager $+$ Neural Proc.)};
\node[fnblock,right=14mm of slots,minimum width=24mm] (ram) {PSRAM\\{\scriptsize 8\,MB, real V1 chain}}; \node[fnblock,right=10mm of slots,minimum width=20mm] (ram) {SDRAM\\{\scriptsize 64\,MB}};
\draw[fnbus] (host) -- (dm); \draw[fnbus] (host) -- (dm);
\draw[fnbus] (dm) -- (dir); \draw[fnbus] (dm) -- (dir);
\draw[fnbus] (dir) -- (slots); \draw[fnbus] (dir) -- (slots);
\draw[fnbus] (slots) -- node[fnlbl,above]{16-bit word} (ram); \draw[fnbus] (slots) -- node[fnlbl,above]{16-bit word} (ram);
\node[below=1mm of slots,font=\scriptsize\itshape,text=fnGrey] \node[below=1mm of slots,font=\scriptsize\itshape,text=fnGrey]
{computation entirely on-chip, dependency graph resolved autonomously}; {computation entirely on-chip, dependency graph resolved autonomously};
\end{tikzpicture} \end{tikzpicture}%
}
\end{center} \end{center}
\vfill \vfill
@@ -63,14 +65,18 @@
\footnotesize \footnotesize
\textbf{\color{fnDark}Reference target device:} Lattice ECP5 \code{LFE5U-45F-8BG381C} \textbf{\color{fnDark}Reference target device:} Lattice ECP5 \code{LFE5U-45F-8BG381C}
(speed grade $-8$, CABGA381) --- identical device and board as V1.\\[2pt] (speed grade $-8$, CABGA381) --- identical device and board as V1.\\[2pt]
\textbf{\color{fnDark}Recommended configuration:} INT8/INT32, \code{P\_IN}=8, \textbf{\color{fnDark}Production configuration:} INT8/INT32, \code{P\_IN}=8,
\code{N\_SLOTS}=2 (real, measured net win --- see ch.~\ref{ch:impl2}), same real, \code{N\_SLOTS}=4 (real 8/8-seed timing closure at 64\,MHz --- see
unmodified V1 PSRAM backend, ISSI \code{IS66WVE4M16EBLL-70BLI}.\\[2pt] ch.~\ref{ch:hw}), single unified SDR SDRAM (Alliance Memory
\code{AS4C32M16SB-7BIN}, 64\,MB), real board-level pinout and KiCad
schematic/BOM.\\[2pt]
\textbf{\color{fnDark}Status:} RTL verified in real Verilator simulation and real \textbf{\color{fnDark}Status:} RTL verified in real Verilator simulation and real
synthesis + place\&route (Yosys + nextpnr-ecp5). Full benchmark campaign, two synthesis + place\&route (Yosys + nextpnr-ecp5). Full benchmark campaign, two
post-campaign memory optimizations, and a full alternative memory-subsystem post-campaign memory optimizations, an alternative memory-subsystem
redesign (the Neural Memory System, ch.~\ref{ch:nms}) complete and measured. redesign that became the current architecture (the Neural Memory System,
Document describing the project as of \datasheetdate. ch.~\ref{ch:nms}), and a real, board-level schematic/BOM verification pass
(ch.~\ref{ch:hw}) all complete and measured. Document describing the
project as of \datasheetdate.
}; };
\end{tikzpicture} \end{tikzpicture}
\end{center} \end{center}
@@ -79,8 +85,10 @@ Document describing the project as of \datasheetdate.
Project author: Michele Bigi \textbullet{} MIKILAB / manvalan.\\ Project author: Michele Bigi \textbullet{} MIKILAB / manvalan.\\
This datasheet documents V2 of the RTL code, documentation and benchmarks This datasheet documents V2 of the RTL code, documentation and benchmarks
present in the repository \texttt{github.com/manvalan/FPGA-Neural}. V1 remains present in the repository \texttt{github.com/manvalan/FPGA-Neural}. V1 remains
frozen and unmodified as the project's golden functional/performance reference; frozen and unmodified as the project's golden functional/performance
it is documented in a separate datasheet.\par} reference; its own datasheet previously lived alongside this one in this
repository and was consolidated out of the working tree as part of a
2026-09-09 documentation cleanup (recoverable from git history).\par}
\end{titlepage} \end{titlepage}
% ====================================================================== % ======================================================================
@@ -89,7 +97,7 @@ it is documented in a separate datasheet.\par}
\input{chapters/00-features} \input{chapters/00-features}
% ====================================================================== % ======================================================================
% PINOUT SUMMARY (honesty note -- no real ball assignment for V2 yet) % PINOUT SUMMARY (real, board-verified ball assignment)
% ====================================================================== % ======================================================================
\newpage \newpage
\input{chapters/00b-pinout} \input{chapters/00b-pinout}
@@ -14,10 +14,11 @@ and unmodified as the project's golden reference). Where V1 executes one
neuron at a time under host-driven SPI control, V2 registers a neuron at a time under host-driven SPI control, V2 registers a
\textbf{dependency graph of neurons} and keeps \code{N\_SLOTS} independent \textbf{dependency graph of neurons} and keeps \code{N\_SLOTS} independent
Neural Processor $+$ Memory Manager pairs busy concurrently, resolving data Neural Processor $+$ Memory Manager pairs busy concurrently, resolving data
dependencies and hiding PSRAM latency in hardware, without host dependencies and hiding memory latency in hardware, without host
intervention once a graph is loaded. Computation (INT8 MAC, ReLU, intervention once a graph is loaded. Computation (INT8 MAC, ReLU,
saturation) is bit-exact identical to V1's own datapath; what changed is saturation) is bit-exact identical to V1's own datapath; what changed is
everything \emph{around} it.} everything \emph{around} it, including, mid-project, the external memory
device itself (\S\ref{sec:sdram-mem-addendum}).}
\vspace{8pt} \vspace{8pt}
\begin{multicols}{2} \begin{multicols}{2}
@@ -29,22 +30,21 @@ everything \emph{around} it.}
once every producer it depends on has genuinely completed --- verified once every producer it depends on has genuinely completed --- verified
for 1-hop shared-producer/multi-consumer graphs and 2-hop transitive for 1-hop shared-producer/multi-consumer graphs and 2-hop transitive
(diamond) graphs. (diamond) graphs.
\item \code{N\_SLOTS} independent \textbf{Neural Processor + Memory Manager} \item \code{N\_SLOTS}=4 independent \textbf{Neural Processor + Memory
pairs (default recommended: \textbf{2}), each running the identical Manager} pairs (production baseline), each running the identical
8-stage INT8 pipeline inherited from V1. 8-stage INT8 pipeline inherited from V1.
\item \textbf{Word-level burst memory backend}: fetches move a full 16-bit \item \textbf{Single unified SDRAM}: one external SDR SDRAM device serves
PSRAM word per transaction instead of one byte, reusing weights, activations, AND results through one arbitrated backend
\code{memory\_interface.v}/\code{psram\_controller.v} directly and its (\code{sdram\_unified\_backend.v}) --- no PSRAM, no second physical
already-implemented page-mode support --- \textbf{2.24--2.37$\times$} memory device, in the current, frozen hardware path.
real wall-clock speedup, measured. \item \textbf{Real physical host transport}: a placed, ball-assigned SPI
\item \textbf{Shared on-chip activation cache}: a vector of activations Mode~0 slave (\code{spi\_host\_bridge.v}) plus a real
shared by many neurons of the same layer is fetched from PSRAM \code{FPGA\_DATA\_READY} completion pin --- both verified on real
\emph{once}, not once per neuron --- a further real \code{nextpnr-ecp5} place\&route, not just in simulation.
\textbf{1.66--2.00$\times$} cycle reduction on shared-input workloads. \item \textbf{Real, board-level verification}: a real KiCad schematic
\item Same \textbf{real, unmodified V1 PSRAM backend} throughout capture, a real exported BOM, and real component selections
(\code{memory\_interface.v}, \code{psram\_controller.v}) --- V1 remains (regulators, oscillator, configuration flash) all cross-checked
the frozen golden reference and was never altered to make V2 look against this datasheet --- not merely a simulated design.
faster.
\item \textbf{Real, measured} characterization at every step: Verilator \item \textbf{Real, measured} characterization at every step: Verilator
RTL simulation, Yosys synthesis, real \code{nextpnr-ecp5} RTL simulation, Yosys synthesis, real \code{nextpnr-ecp5}
place\&route --- no theoretical number reported without a matching place\&route --- no theoretical number reported without a matching
@@ -56,15 +56,17 @@ everything \emph{around} it.}
{\color{fnDark}\large\bfseries Honest, measured limitations}\\[2pt] {\color{fnDark}\large\bfseries Honest, measured limitations}\\[2pt]
{\footnotesize {\footnotesize
\begin{itemize}[leftmargin=1.1em] \begin{itemize}[leftmargin=1.1em]
\item The system is \textbf{memory-bound}, not compute-bound: real compute- \item \code{N\_SLOTS}=8 is \textbf{functionally correct but not
to-memory-wait ratio on the order of 1:170--1:220. A single shared timing-closed}: only 3/8 tested placement seeds pass 64\,MHz ---
PSRAM port saturates at $\approx$90\% utilization regardless of deferred, not production-frozen (\S\ref{sec:clock-closure-current}).
\code{N\_SLOTS}$\ge$2 --- real parallel scaling beyond 2 slots is \item \textbf{Hold-time closure is a genuine, disclosed tool-chain
essentially flat for large workloads. limitation}: no \code{pytrellis}/vendor static-timing-analysis path
\item \code{N\_SLOTS=4} is \textbf{not recommended}: it delivers no is available in this environment to check min-delay/hold, only
additional real throughput once the shared PSRAM port saturates, setup (\S\ref{sec:clock-closure-current}).
and with the activation cache active it \textbf{fails the 80\,MHz \item \textbf{FPGA dynamic power/current draw is not measured}: no ECP5
timing target outright} (65.01\,MHz measured). power estimator is available in this toolchain; regulator sizing
uses datasheet-based engineering margin, not a computed budget
(\S\ref{sec:power-addendum}).
\item Fixed, lowest-index-priority arbitration (Director and memory \item Fixed, lowest-index-priority arbitration (Director and memory
arbiter alike) is not fairness-balanced --- a real, measured arbiter alike) is not fairness-balanced --- a real, measured
per-slot workload imbalance exists under sustained contention. per-slot workload imbalance exists under sustained contention.
@@ -74,22 +76,23 @@ everything \emph{around} it.}
{\color{fnDark}\large\bfseries Target \& toolchain}\\[2pt] {\color{fnDark}\large\bfseries Target \& toolchain}\\[2pt]
{\footnotesize {\footnotesize
\begin{itemize}[leftmargin=1.1em] \begin{itemize}[leftmargin=1.1em]
\item FPGA: Lattice ECP5 \code{LFE5U-45F-8BG381C} ($-8$, CABGA381) --- same \item FPGA: Lattice ECP5 \code{LFE5U-45F-8BG381C} ($-8$, commercial grade,
target device as V1. 381-ball caBGA, 0.8\,mm pitch) --- same target device as V1.
\item Synthesis: Yosys; place\&route: real \code{nextpnr-ecp5}. \item SDRAM: Alliance Memory \code{AS4C32M16SB-7BIN} (512\,Mbit/64\,MB,
4M$\times$16, 54-ball FBGA).
\item Synthesis: Yosys; place\&route: real \code{nextpnr-ecp5} 0.11.1.
\item Simulation: Verilator 5.050 (\code{--binary --timing}) --- adopted \item Simulation: Verilator 5.050 (\code{--binary --timing}) --- adopted
for V2 after two independent Icarus Verilog v13.0 scheduling for V2 after two independent Icarus Verilog v13.0 scheduling
defects were found and reproduced on minimal repros (V1's own defects were found and reproduced on minimal repros (V1's own
certification, performed separately, was unaffected). certification, performed separately, was unaffected).
\item PSRAM: ISSI \code{IS66WVE4M16EBLL-70BLI} (64\,Mb, 4M$\times$16),
real chain reused byte-for-byte from V1.
\end{itemize}} \end{itemize}}
\end{multicols} \end{multicols}
\vspace{2pt} \vspace{2pt}
% --- key parameter table --- % --- key parameter table ---
\noindent \noindent
{\small\color{fnDark}\bfseries Key parameters (recommended configuration, real measured data)} {\small\color{fnDark}\bfseries Key parameters (production configuration,
real measured data)}
\vspace{2pt} \vspace{2pt}
\noindent \noindent
@@ -100,11 +103,12 @@ everything \emph{around} it.}
Data precision & INT8 (signed) & \code{DATA\_WIDTH}=8, identical to V1 \\ Data precision & INT8 (signed) & \code{DATA\_WIDTH}=8, identical to V1 \\
\rowa Accumulator & INT32 (signed) & \code{ACC\_WIDTH}=32 \\ \rowa Accumulator & INT32 (signed) & \code{ACC\_WIDTH}=32 \\
Dot-product width & 8 & \code{P\_IN}=8 parallel MAC lanes per neuron \\ Dot-product width & 8 & \code{P\_IN}=8 parallel MAC lanes per neuron \\
\rowa Recommended concurrency & \code{N\_SLOTS}=2 & real, measured net win; see ch.~\ref{ch:impl2} \\ \rowa Production concurrency & \code{N\_SLOTS}=4 & real, 8/8-seed timing closure; see \S\ref{sec:clock-closure-current} \\
Fmax, full system (\code{N\_SLOTS}=2) & 87.72~MHz & real place\&route, word-burst + activation cache active \\ System clock & 64\,MHz & 16\,MHz oscillator $\to$ \code{EHXPLLL} PLL; 80\,MHz confirmed NO-GO (genuine regenerated PLL, 0/8 seeds) \\
\rowa cycles/neuron (1 neuron, 8 inputs, real PSRAM) & 166 (V1: 209) & \textbf{2.6$\times$} real wall-clock speedup vs V1 \\ \rowa Fmax, \code{N\_SLOTS}=4 (real P\&R, 8 seeds) & worst 64.55\,MHz / best 72.37\,MHz & production baseline, 8/8 PASS \\
Combined real speedup vs baseline (\code{N\_SLOTS}=2) & \textbf{2.45$\times$} & word-burst $+$ activation cache, D-Stress workload \\ D-Stress regression (256 neurons) & 49,927 cycles, 256/256 bit-exact & 780\,\textmu s wall-clock @ 64\,MHz \\
\rowa Address space & 23~bit (byte) & \code{ADDR\_WIDTH}=23, unchanged from V1 \\ \rowa SPI host clock, verified & 12\,MHz recommended (12.8\,MHz hard CDC edge) & simulation-verified, real margin below the deterministic edge \\
Address space & 26~bit (byte), single SDRAM & \code{ADDR\_WIDTH}=26 \\
\bottomrule \bottomrule
\end{tabularx} \end{tabularx}
@@ -112,28 +116,29 @@ Combined real speedup vs baseline (\code{N\_SLOTS}=2) & \textbf{2.45$\times$} &
\noindent \noindent
{\small\color{fnDark}\bfseries System block diagram} {\small\color{fnDark}\bfseries System block diagram}
\begin{center} \begin{center}
\resizebox{\textwidth}{!}{%
\begin{tikzpicture}[node distance=6mm and 9mm,font=\footnotesize] \begin{tikzpicture}[node distance=6mm and 9mm,font=\footnotesize]
\node[fnblockD,minimum width=24mm,minimum height=15mm] (host){HOST\\{\scriptsize registers a node graph}}; \node[fnblockD,minimum width=24mm,minimum height=15mm] (host){HOST (SPI)\\{\scriptsize registers a node graph}};
\node[fnblockT,right=14mm of host,minimum width=30mm,minimum height=13mm] (dm){Dependency\\Manager}; \node[fnblockT,right=14mm of host,minimum width=30mm,minimum height=13mm] (dm){Dependency\\Manager};
\node[fnblockT,right=14mm of dm,minimum width=28mm,minimum height=13mm] (dir){Neural\\Director}; \node[fnblockT,right=14mm of dm,minimum width=28mm,minimum height=13mm] (dir){Neural\\Director};
\node[fnreg,fill=white,right=14mm of dir,minimum width=30mm,minimum height=20mm] (slots){ \node[fnreg,fill=white,right=14mm of dir,minimum width=30mm,minimum height=20mm] (slots){
\begin{tabular}{c} \begin{tabular}{c}
N\_SLOTS $\times$ \\ N\_SLOTS=4 $\times$ \\
Memory Manager \\ Memory Manager \\
$+$ Neural Processor $+$ Neural Processor
\end{tabular}}; \end{tabular}};
\node[fnblockA,below=9mm of dir,minimum width=28mm,minimum height=11mm] (cache){Activation\\Cache}; \node[fnblock,right=14mm of slots,minimum width=26mm,minimum height=15mm] (ram){SDRAM 64\,MB\\{\scriptsize unified backend}};
\node[fnblock,right=14mm of slots,minimum width=22mm,minimum height=15mm] (ram){PSRAM 8\,MB\\{\scriptsize real V1 backend}};
\draw[fnbus] (host) -- (dm); \draw[fnbus] (host) -- (dm);
\draw[fnbus] (dm) -- node[fnlbl,above]{ready node} (dir); \draw[fnbus] (dm) -- node[fnlbl,above]{ready node} (dir);
\draw[fnbus] (dir) -- (slots); \draw[fnbus] (dir) -- (slots);
\draw[fnarrowT] (slots.south) |- (cache.east); \draw[fnbus] (slots) -- node[fnlbl,above]{W / AR ports} (ram);
\draw[fnarrowT] (cache.north) |- node[fnlbl,above]{producer done} (dm.south); \draw[fnarrowT] (slots.south) |- ++(0,-4mm) -| node[fnlbl,below]{producer done} (dm.south);
\draw[fnbus] (slots) -- node[fnlbl,above]{16-bit word} (ram); \end{tikzpicture}%
\draw[fnbus] (cache.east) -- ++(6mm,0) |- ([yshift=-2mm]ram.south); }
\end{tikzpicture}
\end{center} \end{center}
\begin{center}\footnotesize\itshape\color{fnGrey} \begin{center}\footnotesize\itshape\color{fnGrey}
A slot's completion feeds back to the Director (frees the slot) and to the A slot's completion feeds back to the Director (frees the slot) and to the
Dependency Manager (wakes up any node waiting on it) --- closing the Dependency Manager (wakes up any node waiting on it) --- closing the
dataflow loop entirely on-chip.\end{center} dataflow loop entirely on-chip. \code{FPGA\_DATA\_READY} (ball G3) goes high
once every registered node has both resolved and dispatched
(\S\ref{sec:host-addendum}).\end{center}
+67
View File
@@ -0,0 +1,67 @@
\thispagestyle{plain}
\noindent
\begin{tikzpicture}
\node[fill=fnDark,text=white,rounded corners=2pt,inner sep=6pt,
minimum width=\textwidth,anchor=west]
{\large\bfseries Pinout summary --- real, board-verified};
\end{tikzpicture}
\vspace{6pt}
\noindent
{\footnotesize
V2's top-level module, \code{fpga\_neural\_v2\_top.v}, has a complete,
real ball assignment: every signal --- SDRAM bus, SPI host transport,
clock/reset, \code{FPGA\_DATA\_READY}, JTAG, configuration mode straps,
and the boot flash's dedicated MSPI pins --- carries a real CABGA381 ball
site, sourced from the official Lattice pinout CSV (rev 3.0) and
cross-checked against Project Trellis's own \code{iodb.json}. This
supersedes an earlier V2 milestone in which the board-level top had been
placed only \textbf{unconstrained}; a full, constrained \code{.lpf} now
exists (\code{hardware/v2/constraints/v2\_board\_top.lpf}) and every
Fmax number in this datasheet (\S\ref{sec:clock-closure-current}) is
measured against it.
}
\vspace{6pt}
\begin{fnnote}[What is real]
Every ball in the summary table below is placed, P\&R-confirmed, and
cross-checked against a real, exported KiCad schematic and BOM
(\S\ref{sec:schematic-capture}--\ref{sec:bom}) --- not a simulation-only
placeholder. No PSRAM signals exist anywhere in this revision: the
single external memory is SDR SDRAM (\S\ref{sec:sdram-mem-addendum}).
\end{fnnote}
\begin{fnwarn}[What remains open]
FPGA dynamic power/current draw has not been measured post-implementation
(no ECP5 power estimator is available in this toolchain), so exact
decoupling/regulator sizing uses datasheet-based engineering margin, not
a computed budget. Hold-time closure is a genuine tool-chain limitation
(no min-delay analysis path available) --- setup timing is fully
verified. See ch.~\ref{ch:hw} for the complete, disclosed list.
\end{fnwarn}
\vspace{6pt}
\noindent
{\small\color{fnDark}\bfseries Ball summary (see ch.~\ref{ch:hw} for the
complete, per-signal table)}
\vspace{2pt}
\noindent
\begin{tabularx}{\textwidth}{L{3.4cm}L{2.4cm}Y}
\toprule
\rowh \thd{Interface} & \thd{Ball count} & \thd{Notes} \\
\midrule
SDRAM bus (A[0:12], BA[0:1], DQ[0:15], DQM[0:1], CKE/CS\#/RAS\#/CAS\#/WE\#) & 35 & Bank 6/7, real, P\&R-confirmed \\
\rowa SPI host transport (\code{sclk}/\code{mosi}/\code{miso}/\code{cs\_n}) & 4 & Bank 6/7, plain GPIO \\
\code{FPGA\_DATA\_READY}, \code{osc\_clk}, \code{ext\_rst\_n}, \code{sdram\_clk}, \code{pll\_locked} & 5 & Bank 6/7 \\
\rowa JTAG (TCK/TMS/TDI/TDO) & 4 & Bank 40, to ESP32 \\
Config control (PROGRAMN/INITN/DONE) + CFG[2:0] straps & 6 & Bank 8 \\
\rowa Boot-flash dedicated MSPI (CSSPIN/MCLK/D0/D1) & 4 & Bank 8, dual-function \\
\bottomrule
\end{tabularx}
\vspace{4pt}
\noindent
{\footnotesize\color{fnGrey}
Complete per-signal ball tables and the real KiCad schematic/BOM: ch.~\ref{ch:hw}.
Logical (not physical) register-level port list: ch.~\ref{ch:regs}.\par}
@@ -26,11 +26,15 @@ runs autonomously --- no per-neuron host intervention.
ReLU/linear activation with saturation --- \code{neural\_processor.v} ReLU/linear activation with saturation --- \code{neural\_processor.v}
is a direct, bit-exact-verified port of V1's own is a direct, bit-exact-verified port of V1's own
\code{neuron\_parallel.v}/\code{mac8.v}/\code{mac\_unit.v}. \code{neuron\_parallel.v}/\code{mac8.v}/\code{mac\_unit.v}.
\item The real PSRAM backend: \code{memory\_interface.v} and \item V1's own PSRAM backend files (\code{memory\_interface.v},
\code{psram\_controller.v} are reused \textbf{byte-for-byte, \code{psram\_controller.v}) remain byte-for-byte, unmodified
unmodified} from V1 throughout every V2 milestone --- including copies throughout the repository --- V1 itself, as a tree
the two post-campaign optimizations (ch.~\ref{ch:mem}). V1 itself, (\code{hardware/v1/}), is frozen and was never touched.
as a tree (\code{hardware/v1/}), is frozen and was never touched. \textbf{Not currently part of V2's physical board}, however: the
project has since replaced external memory with a single SDR
SDRAM device (\S\ref{sec:sdram-mem-addendum}); the PSRAM-era
chapters that follow document real, correctly-measured work for
the architecture it was measured on, not the current board.
\item The target device (Lattice ECP5 \code{LFE5U-45F-8BG381C}) and the \item The target device (Lattice ECP5 \code{LFE5U-45F-8BG381C}) and the
real-toolchain-only measurement discipline: every number in this real-toolchain-only measurement discipline: every number in this
datasheet is labelled \textsc{Theoretical}, \textsc{Simulated}, datasheet is labelled \textsc{Theoretical}, \textsc{Simulated},
@@ -70,8 +74,21 @@ also accounted for, \code{N\_SLOTS}=4 measures as \emph{slower} in real
wall-clock time than \code{N\_SLOTS}=1 for the largest workload tested wall-clock time than \code{N\_SLOTS}=1 for the largest workload tested
--- more hardware parallelism made that specific configuration worse, --- more hardware parallelism made that specific configuration worse,
not better, because the bottleneck was never compute. This finding not better, because the bottleneck was never compute. This finding
directly shaped both post-campaign optimizations in ch.~\ref{ch:mem} and directly shaped both post-campaign optimizations in ch.~\ref{ch:mem}.
the \code{N\_SLOTS}=2 recommendation carried throughout this datasheet.
\begin{fnwarn}[Architecture changed since this finding: SDRAM, not PSRAM]
This memory-bound finding was measured on the PSRAM-era architecture
described above. The project has since replaced PSRAM with a single
SDR SDRAM device (\S\ref{sec:sdram-mem-addendum}) and closed on
\textbf{\code{N\_SLOTS}=4 as the production configuration} --- chosen
primarily because it is the largest slot count that reliably closes
real timing (8/8 seeds @ 64\,MHz, ch.~\ref{ch:hw}
\S\ref{sec:clock-closure-current}), not from a re-run of this specific
utilization/scaling study. Whether the SDRAM backend's own
utilization/saturation ratio matches the PSRAM-era $\approx$90\% figure
above has \textbf{not been independently re-measured} --- disclosed as
an open item, not assumed to carry over.
\end{fnwarn}
\begin{fnnote}[Reproducibility] \begin{fnnote}[Reproducibility]
Every real number in this datasheet traces to a specific, append-only Every real number in this datasheet traces to a specific, append-only
@@ -1,6 +1,18 @@
\chapter{Architecture} \chapter{Architecture}
\label{ch:arch} \label{ch:arch}
\begin{fnnote}[Scheduling core unchanged; memory backend and slot count
have]
\code{dependency\_manager.v} and \code{neural\_director.v} (this
chapter's own subject) are identical between the PSRAM-era milestone
described below and the current, real SDRAM board --- the scheduling
logic itself did not change. What changed since is the memory backend
(single SDR SDRAM, not PSRAM, \S\ref{sec:sdram-mem-addendum}), the
absence of the shared \textbf{Activation Cache} module from the current
physical top (ch.~\ref{ch:toplevel}), and the production slot count
(\code{N\_SLOTS}=4, not 2).
\end{fnnote}
\section{Module map} \section{Module map}
\begin{center} \begin{center}
\begin{tikzpicture}[node distance=7mm and 11mm,font=\footnotesize] \begin{tikzpicture}[node distance=7mm and 11mm,font=\footnotesize]
@@ -26,9 +38,11 @@
\end{tikzpicture} \end{tikzpicture}
\end{center} \end{center}
\begin{center}\footnotesize\itshape\color{fnGrey} \begin{center}\footnotesize\itshape\color{fnGrey}
N\_SLOTS=2 shown (the recommended configuration); the architecture is PSRAM-era diagram, N\_SLOTS=2 shown; the architecture is parametric in
parametric in N\_SLOTS. Every arrow is a real signal path verified in N\_SLOTS. Every arrow is a real signal path verified in Verilator
Verilator simulation and real Yosys/nextpnr-ecp5 synthesis.\end{center} simulation and real Yosys/nextpnr-ecp5 synthesis. The current, real
board (N\_SLOTS=4, single SDRAM, no Activation Cache module) is shown
in ch.~\ref{ch:toplevel}'s own hierarchy listing.\end{center}
\section{Dependency Manager} \section{Dependency Manager}
Holds a table of \code{N\_NODES} job descriptors, each tracking: node Holds a table of \code{N\_NODES} job descriptors, each tracking: node
@@ -2,33 +2,38 @@
\label{ch:param} \label{ch:param}
\section{Build parameters (synthesis-time)} \section{Build parameters (synthesis-time)}
\begin{fnwarn}[Current, real board parameters (\code{fpga\_neural\_v2\_top.v})]
The table below reflects the real, current SDRAM-architecture top
level. The PSRAM-era \S\S\ref{ch:mem} chapters below this one describe
an earlier, real, correctly-measured milestone with different defaults
(notably \code{ADDR\_WIDTH}=23 and a PSRAM data-bus parameter) ---
superseded, not deleted, since that data remains accurate for the
architecture it was measured on.
\end{fnwarn}
\begin{tabularx}{\textwidth}{L{3.0cm} C{1.8cm} Y} \begin{tabularx}{\textwidth}{L{3.0cm} C{1.8cm} Y}
\toprule \toprule
\rowh \thd{Parameter} & \thd{Default} & \thd{Meaning} \\ \rowh \thd{Parameter} & \thd{Default} & \thd{Meaning} \\
\midrule \midrule
\code{DATA\_WIDTH} & 8 & Data width (INT8), unchanged from V1. \\ \code{DATA\_WIDTH} & 8 & Data width (INT8), unchanged from V1. \\
\rowa \code{ACC\_WIDTH} & 32 & Accumulator width; \textbf{24 recommended} for new P\_IN=8 configurations (ch.~\ref{ch:datapath}). \\ \rowa \code{ACC\_WIDTH} & 32 & Accumulator width. \\
\code{P\_IN} & 8 & Parallel MAC lanes per neuron per tile; must be even (word-level burst constraint, ch.~\ref{ch:mem}). \\ \code{P\_IN} & 8 & Parallel MAC lanes per neuron per tile; must be even (word-level burst constraint, ch.~\ref{ch:mem}). \\
\rowa \code{ADDR\_WIDTH} & 23 & Byte-address width (8~MB), unchanged from V1. \\ \rowa \code{ADDR\_WIDTH} & 26 & Byte-address width (widened 23$\to$26 for the 64\,MB SDRAM device, DEC-0039). \\
\code{N\_SLOTS} & 4 (RTL default) & Concurrent Memory Manager$+$Neural Processor pairs. \textbf{2 recommended} --- see the honesty note below. \\ \code{N\_SLOTS} & 4 & Concurrent Memory Manager$+$Neural Processor pairs. \textbf{Production configuration} --- real 8/8-seed timing closure at 64\,MHz (ch.~\ref{ch:hw} \S\ref{sec:clock-closure-current}). \\
\rowa \code{N\_NODES} & 16 & Dependency Manager node-table depth. Sized to the largest node-id range a graph will ever use; never reclaimed (ch.~\ref{ch:arch}). \\ \rowa \code{N\_NODES} & 16 & Dependency Manager node-table depth. Sized to the largest node-id range a graph will ever use; never reclaimed (ch.~\ref{ch:arch}). \\
\code{MAX\_DEPS} & 4 & Maximum producers a single node can list. \\ \code{MAX\_DEPS} & 4 & Maximum producers a single node can list. \\
\rowa \code{QUEUE\_DEPTH} & 8 & Neural Director's own ready-job FIFO depth. \\ \rowa \code{QUEUE\_DEPTH} & 8 & Neural Director's own ready-job FIFO depth. \\
\code{MAX\_TILES} & 16 (internal, activation\_cache.v) & Longest activation vector the shared cache can hold; not yet exposed as a top-level parameter. \\ \code{MAX\_TILES} & 16 & Longest activation/weight tile run a job can request. \\
\rowa \code{PSRAM\_DATA\_WIDTH} & 16 & Physical PSRAM data bus width, unchanged from V1. \\ \rowa \code{CLK\_FREQ\_MHZ} & 64 & Real system clock, generated by \code{ecp5\_pll\_sys\_clk.v} from the 16\,MHz oscillator; 80\,MHz confirmed NO-GO (ch.~\ref{ch:hw} \S\ref{sec:clock-closure-current}). \\
\code{CLK\_FREQ\_MHZ} & 80 & Frequency used in \code{psram\_controller.v}'s own timing formulas (unmodified V1 module). \\
\bottomrule \bottomrule
\end{tabularx} \end{tabularx}
\begin{fnwarn}[\texttt{N\_SLOTS} is a real, measured trade-off, not a free parameter] \begin{fnwarn}[\texttt{N\_SLOTS} is a real, measured trade-off, not a free parameter]
Unlike V1's \code{PARALLEL} (a pure resource/frequency trade-off), \code{N\_SLOTS}=4 is the production default: the largest slot count
\code{N\_SLOTS} interacts with a real, measured system bottleneck (the that reliably closes real timing at 64\,MHz on every tested placement
one physical PSRAM port). \code{N\_SLOTS}=1 and \code{N\_SLOTS}=2 both seed (8/8). \code{N\_SLOTS}=8 is functionally correct (bit-exact) but
show a real net wall-clock win over the pre-optimization baseline; only 3/8 seeds close timing --- deferred, not production-frozen. Do
\code{N\_SLOTS}=4 shows \emph{no} additional real throughput and, with not simply raise \code{N\_SLOTS} without re-running the real 8-seed
the activation cache active, \textbf{fails the 80\,MHz timing target \code{nextpnr-ecp5} matrix.
outright} (ch.~\ref{ch:impl2}). Do not simply raise \code{N\_SLOTS} for
more perceived parallelism without re-running the real benchmark suite.
\end{fnwarn} \end{fnwarn}
\section{Word-alignment constraint (post word-burst rewrite)} \section{Word-alignment constraint (post word-burst rewrite)}
@@ -40,7 +45,9 @@ land on an even byte address. \code{P\_IN} even and \code{x\_base}/
every job --- true of every address this project's own testbenches use, every job --- true of every address this project's own testbenches use,
and a trivial constraint for any real loader/host to satisfy. and a trivial constraint for any real loader/host to satisfy.
\section{Characterized configurations} \section{Characterized configurations (PSRAM-era; see ch.~\ref{ch:hw}
\S\ref{sec:clock-closure-current} for the current SDRAM-architecture
numbers)}
\begin{tabularx}{\textwidth}{C{1.6cm} C{2.6cm} C{2.6cm} Y} \begin{tabularx}{\textwidth}{C{1.6cm} C{2.6cm} C{2.6cm} Y}
\toprule \toprule
\rowh \thd{N\_SLOTS} & \thd{Fmax, word-burst only} & \thd{Fmax, $+$activation cache} & \thd{Notes} \\ \rowh \thd{N\_SLOTS} & \thd{Fmax, word-burst only} & \thd{Fmax, $+$activation cache} & \thd{Notes} \\
@@ -169,17 +169,21 @@ two copies of the same real data).
\subsection{Real, measured clock closure} \subsection{Real, measured clock closure}
\textbf{N\_SLOTS=4 @ 64\,MHz is the frozen production configuration}: \textbf{N\_SLOTS=4 @ 64\,MHz is the frozen production configuration}:
real \code{nextpnr-ecp5} P\&R, 8/8 tested seeds PASS (worst 66.58\,MHz, real \code{nextpnr-ecp5} P\&R, 8/8 tested seeds PASS. \textbf{N\_SLOTS=8
worst WNS $+0.605$\,ns). \textbf{N\_SLOTS=8 @ 64\,MHz remains an open @ 64\,MHz is deferred}, not production-frozen: 3/8 seeds PASS in the
item}: 5/8 seeds PASS (worst 60.12\,MHz, worst WNS $-1.009$\,ns) after final, current RTL state. 80\,MHz was tested with a genuinely
a real critical-path optimization (\code{sdram\_unified\_backend.v}'s
weight-cache hit-index encoder, rewritten from a serially-dependent
priority scan to a flat, parallel one-hot compare --- real errors.log
ERR-0029/decisions.log DEC-0040). 80\,MHz was tested with a genuinely
regenerated PLL (not merely a \code{--freq} flag) and is \textbf{not regenerated PLL (not merely a \code{--freq} flag) and is \textbf{not
achievable} at either processor count (0/8 seeds pass, both before and achievable} at either processor count --- the achievable Fmax is a
after the ERR-0029 optimization) --- the achievable Fmax is a property property of the routed fabric, confirmed identical between the
of the routed fabric, confirmed identical between the 64\,MHz- and 64\,MHz- and 80\,MHz-targeted netlists. Bit-exact functional
80\,MHz-targeted netlists. Bit-exact functional correctness (D-Stress, correctness (D-Stress, 256/256 neurons vs.\ golden model) is unaffected
256/256 neurons vs.\ golden model) is unaffected at every configuration at every configuration tested.
tested, including through this optimization.
\begin{fnnote}[Single source of truth for exact numbers]
The exact per-seed Fmax/WNS table, its full revision history (three
successive real critical-path fixes: ERR-0027, ERR-0028, ERR-0029, plus
a later fan-out fix, DEC-0042), and the SDRAM directed boundary-test
result (21/21 PASS, both 64\,MHz and 166\,MHz) are kept in one place to
avoid two copies of the same real data --- see ch.~\ref{ch:hw}
\S\ref{sec:clock-closure-current} and \S\ref{sec:sdram-addendum}.
\end{fnnote}
@@ -71,7 +71,10 @@ real ball assignments (\code{spi\_sclk}/\code{spi\_mosi}/
ch.~\ref{ch:hw}). WRITE\_JOB carries the full table from ch.~\ref{ch:hw}). WRITE\_JOB carries the full table from
\S\ref{ch:host} above as an 18-byte payload (grew from 15 after the \S\ref{ch:host} above as an 18-byte payload (grew from 15 after the
64MB memory upgrade widened every address field from 3 to 4 bytes --- 64MB memory upgrade widened every address field from 3 to 4 bytes ---
\code{decisions.log} DEC-0039). \code{decisions.log} DEC-0039). \textbf{Maximum verified operating
clock: 12\,MHz recommended} (exact deterministic CDC edge at
12.8\,MHz $=$ 64\,MHz/5, triple-flop synchronizer) --- see
ch.~\ref{ch:hw} \S\ref{sec:spi-max-verified} for the full sweep.
\textbf{Completion notification}: \code{FPGA\_DATA\_READY}, a real \textbf{Completion notification}: \code{FPGA\_DATA\_READY}, a real
output pin (ball \code{G3}, bank~7), closes the exact gap this output pin (ball \code{G3}, bank~7), closes the exact gap this
+64
View File
@@ -0,0 +1,64 @@
\chapter{Top-level module}
\label{ch:toplevel}
\begin{fnwarn}[Real, board-level top --- not the PSRAM-era compute core]
This chapter describes \code{fpga\_neural\_v2\_top.v}, the module that
is actually placed\&routed against real balls
(\code{hardware/v2/constraints/v2\_board\_top.lpf}) and whose Fmax
numbers appear throughout this datasheet. It supersedes an earlier
milestone's \code{neural\_multiprocessor.v} top level, which drove
V1's own PSRAM chain directly and is retained in the repository for
regression purposes (\code{tb\_nms\_dstress\_sdram\_unified.v}'s own
wrapper, \S\ref{sec:sdram-mem-addendum}) but is not the physical top.
\end{fnwarn}
\section{\texttt{fpga\_neural\_v2\_top.v}}
The real, board-level top: a PLL/reset front-end, a real SPI host
bridge, the compute/scheduling core, and a single unified SDRAM
backend --- 18 physical ports, every one ball-assigned.
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.2cm} C{1.6cm} Y}
\toprule
\rowh \thd{Port} & \thd{Dir} & \thd{Width} & \thd{Function} \\
\midrule
\code{osc\_clk} & IN & 1 & 16\,MHz board oscillator (ball H5). \\
\rowa \code{ext\_rst\_n} & IN & 1 & External POR/supervisor, active-low (ball B4). \\
\code{spi\_sclk}, \code{spi\_mosi}, \code{spi\_cs\_n} & IN & 1 each & Physical SPI host transport (ch.~\ref{ch:host}). \\
\rowa \code{spi\_miso} & OUT & 1 & SPI host transport, response direction. \\
\code{sdram\_clk} & OUT & 1 & SDRAM chip's own \code{CLK} pin --- a real board-level output, not internal-only routing (found missing during this session's own schematic review; ball J4). \\
\rowa \code{sdram\_cke}, \code{sdram\_cs\_n}, \code{sdram\_ras\_n}, \code{sdram\_cas\_n}, \code{sdram\_we\_n} & OUT & 1 each & SDRAM control lines. \\
\code{sdram\_ba} & OUT & 2 & SDRAM bank address. \\
\rowa \code{sdram\_a} & OUT & 13 & SDRAM row/column address (widened 12$\to$13 bits for the 64\,MB device, DEC-0039). \\
\code{sdram\_dq} & INOUT & 16 & SDRAM bidirectional data bus. \\
\rowa \code{sdram\_dqm} & OUT & 2 & SDRAM byte mask. \\
\code{data\_ready} & OUT & 1 & \code{FPGA\_DATA\_READY}, system-idle completion flag (ball G3, \S\ref{sec:host-addendum}). \\
\rowa \code{pll\_locked} & OUT & 1 & PLL lock status, bring-up/debug (ball L1). \\
\bottomrule
\end{tabularx}
\section{Internal hierarchy}
\noindent\code{fpga\_neural\_v2\_top.v}
\begin{itemize}[leftmargin=2.4em]
\footnotesize
\item \code{u\_pll} : \code{ecp5\_pll\_sys\_clk.v} (real \code{EHXPLLL} primitive, 16$\to$64\,MHz)
\item \code{u\_reset\_sync} : \code{reset\_sync.v} (async assert, sync deassert, gated by \code{ext\_rst\_n} AND \code{pll\_locked})
\item \code{u\_spi\_bridge} : \code{spi\_host\_bridge.v} (real SPI Mode~0 slave, triple-flop CDC)
\item \code{u\_dataflow\_core} : \code{nms\_dataflow\_core\_sdram.v}
\begin{itemize}
\item \code{u\_dep\_mgr} : \code{dependency\_manager.v}
\item \code{u\_director} : \code{neural\_director.v}
\item \code{GEN\_SLOT[0..N\_SLOTS-1]}: \code{nms\_memory\_manager\_stream\_wide.v} $+$ \code{neural\_processor.v}
\end{itemize}
\item \code{u\_arbiter\_w}, \code{u\_arbiter\_ar} : \code{slot\_mem\_arbiter.v} (one per logical SDRAM port, W and AR)
\item \code{u\_sdram\_backend} : \code{sdram\_unified\_backend.v} $\to$ \code{sdram\_controller.v} (single physical SDRAM)
\end{itemize}
\begin{fnnote}[No shared activation cache in this datapath]
The PSRAM-era shared activation cache (\code{activation\_cache.v},
ch.~\ref{ch:mem} \S\ref{sec:cache}) is not part of the current SDRAM
top-level's instantiation tree --- \code{nms\_memory\_manager\_stream\_wide.v}
handles per-slot activation/weight/result streaming directly against
the unified SDRAM backend. The PSRAM-era module remains real, correct,
and documented for the architecture it was measured on
(ch.~\ref{ch:mem}), but is not reused here.
\end{fnnote}
@@ -169,7 +169,7 @@ DSP/LUT/FF availability & Not the bottleneck at N\_SLOTS$\le$4 --- all well unde
\bottomrule \bottomrule
\end{tabularx} \end{tabularx}
\section{Limitations, honestly stated} \section{Limitations, honestly stated (PSRAM-era campaign above)}
\begin{itemize} \begin{itemize}
\item V1's own memory-utilization/stall figures were not re-measured \item V1's own memory-utilization/stall figures were not re-measured
this session (V1 is frozen); only its already-certified numbers this session (V1 is frozen); only its already-certified numbers
@@ -185,3 +185,73 @@ DSP/LUT/FF availability & Not the bottleneck at N\_SLOTS$\le$4 --- all well unde
chain) after either memory optimization --- only chain) after either memory optimization --- only
\code{dataflow\_core.v} alone, pre-optimization (92.63~MHz). \code{dataflow\_core.v} alone, pre-optimization (92.63~MHz).
\end{itemize} \end{itemize}
\section{SDRAM-era benchmark addendum (2026-09-07) --- current,
authoritative results}
\label{sec:impl-sdram-addendum}
\begin{fnwarn}[Supersedes the PSRAM/\code{N\_SLOTS}$\le$2-era campaign
above for the current hardware baseline]
Every section above (V1 vs.\ V2 comparison, \code{N\_SLOTS} sweep,
parallel scaling, memory optimizations \#1/\#2, bottleneck analysis)
describes an earlier V2 milestone built on V1's own PSRAM chain,
recommending \code{N\_SLOTS}=2. The project has since replaced external
memory with a single SDR SDRAM device (ch.~\ref{ch:mem}
\S\ref{sec:sdram-mem-addendum}) and closed on
\textbf{\code{N\_SLOTS}=4 as the production configuration}. This
section is the current, real, measured state; the PSRAM-era numbers
above remain real and correctly measured for the architecture they
describe, but do not apply to the current board.
\end{fnwarn}
\subsection{Real resource utilization (\code{N\_SLOTS}=4, SDRAM
architecture, post real critical-path fixes)}
\begin{tabularx}{\textwidth}{L{4.2cm} C{2.4cm} Y}
\toprule
\rowh \thd{Resource} & \thd{Count} & \thd{Notes} \\
\midrule
TRELLIS\_COMB (LUT4-equiv) & 7,175 / 43,848 (16.4\%) & Real Yosys synthesis, most recent measurement (post-ERR-0029) \\
\rowa MULT18X18D & 32 / 72 (44.4\%) & Exactly $4\times8$ (\code{N\_SLOTS}$\times$\code{P\_IN}), confirmed --- the ERR-0027 fix removed a spurious 33rd multiplier \\
DP16KD (block RAM) & 0 / 108 & All small SRAMs synthesize to distributed RAM \\
\rowa EHXPLLL & 1 & Real \code{EHXPLLL} primitive, \code{ecppll}-derived parameters \\
TRELLIS\_FF & $\ge$6,322 (last individually re-quoted figure) & Real, same SDRAM architecture, pre-dates the ERR-0027/0028/0029 restructuring; not independently re-synthesized standalone since --- disclosed as a lower-bound reference, not re-invented as exact \\
\bottomrule
\end{tabularx}
\subsection{Real, current clock closure and functional regression}
See ch.~\ref{ch:hw} \S\ref{sec:clock-closure-current} for the complete
per-seed Fmax/WNS table (single source of truth, not duplicated here):
\textbf{\code{N\_SLOTS}=4 @ 64\,MHz, 8/8 seeds PASS} (worst 64.55\,MHz,
best 72.37\,MHz); \code{N\_SLOTS}=8 deferred (3/8); 80\,MHz confirmed
NO-GO at either processor count with a genuinely regenerated PLL.
D-Stress functional regression (256 neurons, 256/256 bit-exact vs.\
golden model): \textbf{49,927 cycles} at \code{N\_SLOTS}=4 ---
\textbf{780\,\textmu s} real wall-clock at the P\&R-verified 64\,MHz
system clock ($49{,}927 / 64{,}000{,}000$, \textsc{Derived}). SDRAM
directed boundary verification (ch.~\ref{ch:hw}
\S\ref{sec:sdram-addendum}): 21/21 PASS, zero bugs found, both 64\,MHz
and 166\,MHz.
\subsection{Real SPI host protocol throughput}
Board-level smoke test (\code{tb\_fpga\_neural\_v2\_top\_smoke.v}, 11/11
PASS): single job 99--100 cycles/job; back-to-back 88--100 cycles/job;
steady-state throughput unaffected by inter-job gap (100\,ns/5\,\textmu
s/50\,\textmu s tested). Maximum verified SPI host clock: \textbf{12\,MHz
recommended} (exact deterministic CDC edge at 12.8\,MHz $=$ 64\,MHz/5)
--- see ch.~\ref{ch:hw} \S\ref{sec:spi-max-verified} for the full sweep.
\subsection{Limitations, honestly stated (current SDRAM architecture)}
\begin{itemize}
\item Power/energy: \textbf{NOT MEASURED} --- no ECP5 power estimator
is available in this toolchain (unchanged from the PSRAM-era
disclosure above).
\item Hold-time closure: \textbf{OPEN --- tool-chain limitation}, not a
real defect; see ch.~\ref{ch:hw} \S\ref{sec:hw-open-items} for
the complete, consolidated open-items list.
\item \code{N\_SLOTS}=8 is functionally correct but not
timing-closed on every tested seed --- deferred by explicit
project direction, not attempted further this pass.
\item No embedded-host (ESP32-class) physical baseline exists; all
host-side numbers above are protocol-level simulation, not
measured on real silicon.
\end{itemize}
@@ -1,64 +1,39 @@
\chapter{Hardware and board} \chapter{Hardware and board}
\label{ch:hw} \label{ch:hw}
\section{Unchanged from V1} \section{Board summary}
V2 targets the identical board and component set as V1: Lattice ECP5 V2 targets Lattice ECP5 \code{LFE5U-45F-8BG381C} ($-8$, commercial
\code{LFE5U-45F-8BG381C} ($-8$, CABGA381), ISSI grade, 381-ball caBGA, 0.8\,mm pitch, real package geometry
\code{IS66WVE4M16EBLL-70BLI} PSRAM (64\,Mb, 4M$\times$16), same 16\,MHz 17$\times$17$\times$1.76\,mm) --- the same die/package family as V1,
reference oscillator. The real PSRAM controller but the board around it has diverged substantially: V2 replaces V1's
(\code{psram\_controller.v}) and its byte$\leftrightarrow$word adapter PSRAM with a single external SDR SDRAM device (\S\ref{sec:sdram-addendum}),
(\code{memory\_interface.v}) are reused byte-for-byte, unmodified, from adds a real, placed SPI host transport and \code{FPGA\_DATA\_READY}
\code{hardware/v1/} throughout every V2 milestone --- their real, completion pin (ch.~\ref{ch:host}), and has a real, exported KiCad
already-verified electrical/timing requirements and page-mode behavior schematic capture and BOM (\S\ref{sec:schematic-capture}--\ref{sec:bom}).
are unchanged, because the controller itself was never touched. Every top-level signal of \code{fpga\_neural\_v2\_top.v} carries a real
ball assignment in \code{hardware/v2/constraints/v2\_board\_top.lpf} ---
no unconstrained/placeholder pins remain in this revision.
\begin{fnnote}[Real ball assignment: defer to V1's own chapter] \begin{fnnote}[V1's own PSRAM chain: retained in RTL, not on this board]
V1's own hardware chapter documents a real, \code{iodb.json}-verified, \code{psram\_controller.v}/\code{memory\_interface.v} remain byte-for-byte
place\&route-confirmed ball assignment for every PSRAM signal identical to V1's own copies in the repository (frozen golden reference),
(\code{psram\_a}, \code{psram\_dq}, \code{psram\_ce\_n/oe\_n/we\_n/ but are \textbf{not instantiated anywhere in V2's real physical top}
lb\_n/ub\_n/zz\_n}). Since V2's own \code{neural\_multiprocessor.v} --- confirmed by inspection (\code{grep -ri psram hardware/v2/} returns
drives these signals through the identical, unmodified controller, that nothing outside historical commentary). V1's own PSRAM ball assignment
same real ball assignment applies unchanged if V2 is deployed on the therefore does not apply to this board.
same physical board --- it is not repeated here to avoid maintaining two
copies of the same real data; see the V1 datasheet directly.
\end{fnnote} \end{fnnote}
\section{What V2 has not yet placed on real hardware}
As stated in ch.~\ref{ch:host}, V2's own node-registration bus has no
physical pin assignment in this revision --- every V2 characterization
to date used either a Verilator testbench or an unconstrained
(\code{--lpf-allow-unconstrained}) synthesis top-level. A real deployment
would need:
\begin{itemize}
\item A physical host transport for the registration bus (ch.~\ref{ch:host}).
\item A real, constrained \code{nextpnr-ecp5} place\&route run
producing a genuine \code{.lpf}/ball assignment for
\code{neural\_multiprocessor.v}'s own top-level pins, analogous to
V1's own \code{tools/pinout/gen\_lpf.py} flow.
\item Re-verification that the real Fmax numbers in ch.~\ref{ch:impl2}
(obtained unconstrained) hold once real pin locations are fixed ---
pin placement can itself affect routing and therefore Fmax.
\end{itemize}
\section{Power supply, oscillator, configuration}
Unchanged from V1: same board-level power sequencing, same oscillator,
same JTAG/config-SPI boot path (fixed-function dedicated pins, outside
RTL scope). No V2-specific hardware change was made or is required
beyond the (not yet placed) registration-bus transport above.
\section{SDRAM upgrade addendum (2026-09-07) --- current, authoritative \section{SDRAM upgrade addendum (2026-09-07) --- current, authoritative
board state} board state}
\label{sec:sdram-addendum} \label{sec:sdram-addendum}
\begin{fnwarn}[This section supersedes the PSRAM description above for \begin{fnwarn}[Real, closed architectural decision]
the current hardware baseline] An earlier V2 milestone reused V1's own PSRAM chain, placed
The sections above describe an earlier V2 milestone that still reused unconstrained. The project has since made a closed architectural
V1's own PSRAM chain unconstrained. The project has since made a decision (real \code{decisions.log} DEC-0034) to replace external
closed architectural decision (real \code{decisions.log} DEC-0034) to memory with a single SDR SDRAM device, and has since upgraded that
replace external memory with a single SDR SDRAM device, and has since device's capacity (8\,MB $\to$ 64\,MB) and re-verified real,
upgraded that device's capacity and re-verified real, constrained constrained place\&route timing end to end. This section is the
place\&route timing. This section is the current, real, measured state current, real, measured state.
--- see \code{hardware/v2/docs/MEMORY\_UPGRADE\_64MB\_N8.md} in the
repository for the full investigation.
\end{fnwarn} \end{fnwarn}
\subsection{Memory device} \subsection{Memory device}
@@ -141,15 +116,48 @@ See \code{decisions.log} DEC-0042 for full detail. A further
pipelining fix on the same broadcast path is a real, identified, pipelining fix on the same broadcast path is a real, identified,
not-yet-attempted option if more margin is ever needed. not-yet-attempted option if more margin is ever needed.
\subsection{Directed SDRAM boundary verification}
A dedicated directed testbench (\code{tb\_sdram\_boundary.v}, 21 checks)
covers every address/row/bank boundary the randomized D-Stress
regression does not directly target: exact first/last address
(\code{0x000000}/\code{0x3FFFFF}), the row-10/row-11 column boundary,
all three inter-bank crossings, the real V2 memory-map boundaries
(weights/activations/results base and last-word-before-next-region),
and all four byte-mask combinations with distinct deterministic
patterns. All 21 addresses are written first, then read back in
\textbf{reversed} order with address-derived patterns, proving no
write corrupts any neighbouring address. \textbf{Result: 21/21 PASS at
both 64\,MHz and 166\,MHz --- no bug found}, closing the one directed
boundary-test gap disclosed earlier in the project's own verification
history.
\subsection{Verified SPI host operating clock}
\label{sec:spi-max-verified}
A dedicated sweep testbench (\code{tb\_spi\_freq\_sweep.v}) drives the
real \code{fpga\_neural\_v2\_top} (not \code{spi\_host\_bridge} in
isolation) at the real 64\,MHz system clock and sweeps the SPI bit
rate across single-job, back-to-back, gapped, and raw
\code{WRITE\_MEM}/\code{READ\_MEM} traffic. The breakpoint is
\textbf{exact and deterministic}: PASS at every rate up to
\textbf{12.8\,MHz (precisely 64\,MHz/5)}, FAIL (data corruption, then
protocol FSM hang) at every rate at or above it --- the triple-flop CDC
synchronizer plus edge-detect/FSM reaction in \code{spi\_host\_bridge.v}
requires at least 5 full system-clock cycles per SPI bit period to
reliably track \code{sclk}/\code{mosi}/\code{cs\_n} transitions, a real
property of the CDC design (correct, standard practice), not a bug.
\textbf{SPI\_MAX\_VERIFIED = 12\,MHz} is the recommended host operating
point (real margin below the hard 12.8\,MHz edge, $\approx$6.7\%
headroom). Board-level electrical limits (trace length, driver
rise/fall time, ground bounce, real metastability risk) are
\textbf{not} modeled by this deterministic simulation and remain to be
confirmed empirically at bring-up.
\section{Power supply design (2026-09-07) --- verified against the real \section{Power supply design (2026-09-07) --- verified against the real
Lattice hardware checklist} Lattice hardware checklist}
\label{sec:power-addendum} \label{sec:power-addendum}
\begin{fnwarn}[Supersedes the generic \S3 stub above] \begin{fnwarn}[Real design data, not estimated]
The ``Power supply, oscillator, configuration'' section earlier in The actual rail topology, sized against the real, primary-source
this chapter only said ``unchanged from V1'' without real design data. Lattice and TI documents below.
This section replaces that stub with the actual rail topology, sized
against the real, primary-source Lattice and TI documents below --- not
estimated.
\end{fnwarn} \end{fnwarn}
\subsection{Rail topology} \subsection{Rail topology}
@@ -236,11 +244,18 @@ without external caps at the regulator itself); the 10\,\textmu F$+$
100\,nF on \code{VCCAUX} above are the FPGA-side filter from 100\,nF on \code{VCCAUX} above are the FPGA-side filter from
FPGA-TN-02038, not regulator-stability caps, and are still required. FPGA-TN-02038, not regulator-stability caps, and are still required.
\begin{fnnote}[Open item carried from \S3 above] \begin{fnnote}[16\,MHz oscillator: frozen]
The 16\,MHz reference oscillator's exact manufacturer part number is \textbf{ECS Inc. International \code{ECS-3225MV-160-BN-TR}} --- a
not yet specified in this document (only ``16\,MHz'' as a frequency quartz crystal oscillator (XO, not a bare crystal; direct digital clock
requirement) --- flagged, not invented, pending finalization of the output, no external oscillator circuit needed), 3225 SMD package
schematic capture. (3.2$\times$2.5\,mm, 4-pad, matching the real KiCad footprint for U5),
3.3\,V supply (matches \code{osc\_clk}'s real \code{IO\_TYPE=LVCMOS33}
ball H5 exactly, no level-shifting needed), $\pm$50\,ppm stability,
$-40$ to $+85^{\circ}$C. One 100\,nF decoupling capacitor across
\code{VDD}/\code{GND}, placed close to the supply pin. The exact
terminal order-code suffix (stability/output-enable option letters)
should be cross-checked against ECS's current published datasheet at
BOM lock --- normal due diligence, not an open architectural question.
\end{fnnote} \end{fnnote}
\subsection{Power tree} \subsection{Power tree}
@@ -290,9 +305,11 @@ timing recovery that followed) for the complete history.
\end{fnwarn} \end{fnwarn}
\subsection{One physical flash chip: boot bitstream only} \subsection{One physical flash chip: boot bitstream only}
Connects exclusively to the ECP5's own dedicated sysCONFIG pins, \textbf{Winbond \code{W25Q128JVPIM}} (128\,Mbit, WSON-8, 6$\times$5\,mm
Master SPI mode, auto-boots every power-up, zero ESP32 involvement in --- real BOM entry U9, \S\ref{sec:bom}). Connects exclusively to the
normal operation. No second flash device, no on-board neural-network ECP5's own dedicated sysCONFIG pins, Master SPI mode, auto-boots every
power-up, zero ESP32 involvement in normal operation. No second flash
device, no on-board neural-network
weight persistence in the current design --- the host (ESP32) is weight persistence in the current design --- the host (ESP32) is
responsible for pushing weight/activation data into SDRAM fresh each responsible for pushing weight/activation data into SDRAM fresh each
session via the real SPI application protocol session via the real SPI application protocol
@@ -527,3 +544,21 @@ Target: a castellated-edge SMD module, approximately
dimensions and pin-out placeholder, real layout pending. This section dimensions and pin-out placeholder, real layout pending. This section
will be filled in with the actual module outline, castellation pin will be filled in with the actual module outline, castellation pin
map, and mechanical drawing once available. map, and mechanical drawing once available.
\section{Verification status --- real, disclosed open items}
\label{sec:hw-open-items}
Everything above is real (simulated, synthesized, and/or place\&route
measured); this section lists what is genuinely \textbf{not yet}
verified, honestly, rather than silently omitted.
\begin{tabularx}{\textwidth}{L{4.4cm} Y}
\toprule
\rowh \thd{Item} & \thd{Status} \\
\midrule
Hold-time closure & \textbf{OPEN --- tool-chain limitation.} \code{nextpnr-ecp5}'s own timing report contains setup-side (posedge$\to$posedge max-delay) data only; no hold/min-delay analysis. No \code{pytrellis}-based min-delay pass or vendor (Lattice Diamond/Radiant) static timing analysis is available in this environment. Setup timing is fully verified (\S\ref{sec:clock-closure-current}). \\
\rowa FPGA dynamic power/current draw & \textbf{OPEN --- not computable without post-implementation tools.} No ECP5 power estimator (\code{ecppower} or equivalent) is available in this toolchain. Regulator current ratings (\S\ref{sec:power-addendum}) are real, datasheet-supported engineering margin against this unknown, not a computed budget. \\
N\_SLOTS=8 @ 64\,MHz & \textbf{Deferred, not production-frozen} --- functionally correct (bit-exact), 3/8 seeds pass timing closure. See \S\ref{sec:clock-closure-current}. \\
\rowa Board-level SPI electrical limit & \textbf{OPEN --- requires real hardware.} \S\ref{sec:spi-max-verified}'s 12\,MHz recommendation is a simulation-verified logical limit; real trace length, driver rise/fall time, and metastability risk are not modeled by simulation. \\
Embedded-host (ESP32-class) benchmark baseline & \textbf{OPEN --- no hardware available.} No comparison against a real ESP32 host exists; all host-side timing is protocol-level (ch.~\ref{ch:host}), not measured on real silicon. \\
\bottomrule
\end{tabularx}
@@ -1,13 +1,22 @@
\chapter{Register-level interface \& internal state encodings} \chapter{Register-level interface \& internal state encodings}
\label{ch:regs} \label{ch:regs}
\begin{fnwarn}[No SPI register map in this revision] \begin{fnwarn}[Real SPI opcode map exists; state encodings below are
V1's own quick-reference chapter documents a real SPI opcode/register per-module reference]
map (\code{STATUS}, \code{SET\_BASE}, \code{READ\_CONFIG}, \ldots). V2 has Ch.~\ref{ch:host} now documents V2's real, physical SPI opcode map
no equivalent yet (ch.~\ref{ch:host}) --- this chapter instead documents (\code{WRITE\_JOB}/\code{WRITE\_MEM}/\code{READ\_MEM}/\code{STATUS}/
the \textbf{node-registration field layout} (repeated here for quick \code{RESET}) --- this chapter's own node-registration field layout
reference) and the \textbf{internal FSM state encodings} exposed by each below remains the logical field reference (repeated here for quick
module, useful for simulation-level debug and for a future host driver. reference). The \textbf{internal FSM state encodings} below are useful
for simulation-level debug; \S\S\ref{ch:regs}'s Dependency
Manager/Neural Director tables are shared by every V2 architecture
(unchanged between the PSRAM-era and current SDRAM boards). The Memory
Manager and Neural Processor tables were captured from the PSRAM-era
\code{memory\_manager.v}/\code{neural\_processor.v} pairing (ch.~\ref{ch:arch})
--- the current SDRAM board's \code{nms\_memory\_manager\_stream\_wide.v}
implements the same functional handshake (prefetch $\to$ stream $\to$
write-back $\to$ done) against the SDRAM backend instead of PSRAM, but
its own internal state encoding was not re-transcribed into this table.
\end{fnwarn} \end{fnwarn}
\section{Node registration fields (quick reference)} \section{Node registration fields (quick reference)}
@@ -17,6 +17,26 @@ data shows the choice between them is configuration-dependent, not a
strict win for either. strict win for either.
\end{fnnote} \end{fnnote}
\begin{fnwarn}[This is the direct ancestor of the current, real board
--- read this before the rest of the chapter]
The \code{nms\_*}-prefixed modules introduced in this chapter
(\code{nms\_dataflow\_core.v}, \code{nms\_neural\_multiprocessor.v},
\ldots) are the \textbf{direct code ancestors} of the real, current
board-level RTL documented in ch.~\ref{ch:hw}/\ref{ch:toplevel}
(\code{nms\_dataflow\_core\_sdram.v}, \code{fpga\_neural\_v2\_top.v}).
The project's own path was: Current V2 (PSRAM, ch.~\ref{ch:arch}) $\to$
NMS (this chapter, still PSRAM, replicated on-chip SRAM) $\to$
\textbf{single unified SDRAM} (ch.~\ref{ch:hw}
\S\ref{sec:sdram-mem-addendum}, the current, real, shipped board). This
chapter's own STEP9/10 recommendation below (``adopt NMS at
\code{N\_SLOTS}$\le$2'') was itself superseded by that final SDRAM
step, which changed the backing memory device and re-closed timing at
\code{N\_SLOTS}=4 (ch.~\ref{ch:hw} \S\ref{sec:clock-closure-current}).
Read this chapter as \textbf{real history explaining how the current
architecture was reached}, not as a currently-open choice between three
systems.
\end{fnwarn}
\section{Design goal} \section{Design goal}
Current V2's own memory path is fundamentally an on-demand, Current V2's own memory path is fundamentally an on-demand,
per-request architecture: every tile fetch is a fresh transaction, per-request architecture: every tile fetch is a fresh transaction,
Binary file not shown.
@@ -1,119 +0,0 @@
% ======================================================================
% FPGA-Neural -- INT8 Neural Network Engine
% Datasheet / Manuale di riferimento tecnico
% Repository: github.com/manvalan/FPGA-Neural
% ======================================================================
\documentclass[11pt,a4paper,openany]{report}
\newcommand{\datasheetrev}{A1}
\newcommand{\datasheetdate}{Settembre 2026}
\input{preamble}
\begin{document}
\sloppy
% ======================================================================
% FRONTESPIZIO
% ======================================================================
\begin{titlepage}
\thispagestyle{empty}
\begin{tikzpicture}[remember picture,overlay]
\fill[fnDark] (current page.north west) rectangle
([yshift=-4.3cm]current page.north east);
\fill[fnTeal] ([yshift=-4.3cm]current page.north west) rectangle
([yshift=-4.55cm]current page.north east);
\node[anchor=north west,text=white,font=\Huge\bfseries]
at ([xshift=2.2cm,yshift=-1.15cm]current page.north west)
{FPGA\,--\,Neural};
\node[anchor=north west,text=fnLight,font=\large]
at ([xshift=2.25cm,yshift=-2.15cm]current page.north west)
{INT8 Neural Network Engine per FPGA};
\node[anchor=north west,text=fnLight2,font=\normalsize]
at ([xshift=2.25cm,yshift=-2.85cm]current page.north west)
{Acceleratore hardware parametrico -- Datasheet e manuale di riferimento};
\node[anchor=north east,text=white,font=\ttfamily\small]
at ([xshift=-2.2cm,yshift=-3.55cm]current page.north east)
{Rev.~\datasheetrev~~\textbullet~~\datasheetdate};
\end{tikzpicture}
\vspace*{5.0cm}
% --- diagramma a blocchi sintetico sul frontespizio ---
\begin{center}
\begin{tikzpicture}[node distance=7mm and 12mm]
\node[fnblockD,minimum width=30mm] (host) {HOST\\{\scriptsize Linux / ESP32 / MCU / PC}};
\node[fnblockT,right=18mm of host,minimum width=34mm] (fpga)
{FPGA\\{\scriptsize Neural Network Engine}};
\node[fnblock,right=18mm of fpga,minimum width=26mm] (ram)
{PSRAM\\{\scriptsize 8\,MB dedicata}};
\draw[fnbus] (host) -- node[fnlbl,above]{SPI Mode 0} (fpga);
\draw[fnbus] (fpga) -- node[fnlbl,above]{async 16-bit} (ram);
\node[below=1mm of fpga,font=\scriptsize\itshape,text=fnGrey]
{calcolo interamente on-chip};
\end{tikzpicture}
\end{center}
\vfill
\begin{center}
\begin{tikzpicture}
\node[draw=fnRule,rounded corners=3pt,inner sep=10pt,fill=fnLight,text width=15.5cm]{
\footnotesize
\textbf{\color{fnDark}Dispositivo target di riferimento:} Lattice ECP5 \code{LFE5U-45F-8BG381C}
(speed grade $-8$, CABGA381, 72$\times$MULT18X18D, $\approx$44k LUT).\\[2pt]
\textbf{\color{fnDark}Configurazione baseline:} INT8/INT32, \code{N\_INPUTS}=256, \code{N\_NEURONS}=4,
\code{PARALLEL} parametrico, memoria di lavoro PSRAM ISSI \code{IS66WVE4M16EBLL-70BLI}.\\[2pt]
\textbf{\color{fnDark}Stato:} RTL verificato in simulazione (Icarus) e sintesi reale
(Yosys + nextpnr-ecp5). Documento descrittivo del progetto allo stato del \datasheetdate.
};
\end{tikzpicture}
\end{center}
\vspace{0.6cm}
{\footnotesize\color{fnGrey}\raggedright
Autore del progetto: Michele Bigi \textbullet{} MIKILAB / manvalan.\\
Questo datasheet documenta il codice RTL, la documentazione e i benchmark
presenti nella repository \texttt{github.com/manvalan/FPGA-Neural}.\par}
\end{titlepage}
% ======================================================================
% PAGINA "FEATURES" (stile datasheet)
% ======================================================================
\input{chapters/00-features}
% ======================================================================
% SUNTO PINOUT (pagine 2-3, pin per pin -- non a bus)
% ======================================================================
\newpage
\input{chapters/00b-pinout}
% ======================================================================
% INDICE
% ======================================================================
\newpage
\pagenumbering{roman}
{\color{fnDark}\tableofcontents}
\newpage
\pagenumbering{arabic}
% ======================================================================
% CAPITOLI
% ======================================================================
\include{chapters/01-overview}
\include{chapters/02-architettura}
\include{chapters/03-datapath}
\include{chapters/04-parametri}
\include{chapters/05-memoria}
\include{chapters/06-sequencer}
\include{chapters/06b-grafo}
\include{chapters/07-spi}
\include{chapters/07b-programmazione}
\include{chapters/08-toplevel}
\include{chapters/09-implementazione}
\include{chapters/10-hardware}
\include{chapters/11-registri}
\include{chapters/12-roadmap}
\appendix
\include{chapters/A-moduli}
\end{document}
-45
View File
@@ -1,45 +0,0 @@
# FPGA-Neural — Datasheet
Datasheet tecnico multicapitolo dell'engine FPGA-Neural, in italiano e inglese.
Ricostruito a partire dal codice RTL, dalla documentazione e dai benchmark presenti
nella repository (revisione A1, settembre 2026).
## Struttura
```
docs/datasheet/
├── FPGA-Neural-Datasheet.pdf ← PDF italiano (36 pagine)
├── FPGA-Neural-Datasheet.tex ← sorgente principale (IT)
├── preamble.tex ← stili, palette, box, TikZ
├── chapters/ ← 14 capitoli (IT)
└── en/
├── FPGA-Neural-Datasheet-EN.pdf ← PDF inglese (36 pagine)
├── FPGA-Neural-Datasheet-EN.tex ← sorgente principale (EN)
├── preamble.tex ← stili (EN)
└── chapters/ ← 14 capitoli (EN)
```
## Compilazione
Serve una distribuzione LaTeX con `pgfplots`, `tikz-timing`, `tcolorbox`,
`ltablex`, `listings`, `babel`.
```sh
# Italiano
cd docs/datasheet
pdflatex FPGA-Neural-Datasheet.tex
pdflatex FPGA-Neural-Datasheet.tex # 2ª passata per indice e riferimenti
# Inglese
cd docs/datasheet/en
pdflatex FPGA-Neural-Datasheet-EN.tex
pdflatex FPGA-Neural-Datasheet-EN.tex
```
## Nota sul pinout
Il capitolo *Progetto hardware e mappa dei segnali* riporta l'analisi completa
segnale-per-segnale del top-level `spi_neuron_top`, con la colonna **Ball**
compilata con assegnazioni CABGA381 reali (53 segnali, `.lpf` reale in
`synth/`) e verificata da un place\&route reale (`nextpnr-ecp5`, 0 errori di
vincolo, `Program finished normally`) — non più auto-piazzate.
@@ -1,119 +0,0 @@
\thispagestyle{plain}
\noindent
\begin{tikzpicture}
\node[fill=fnDark,text=white,rounded corners=2pt,inner sep=6pt,
minimum width=\textwidth,anchor=west]
{\large\bfseries FPGA-Neural --- Descrizione generale e caratteristiche};
\end{tikzpicture}
\vspace{6pt}
\noindent
{\small FPGA-Neural è un \textbf{acceleratore hardware parametrico per reti neurali}
feed-forward completamente contenuto nell'FPGA. Il calcolo (moltiplicazione,
accumulo, bias, attivazione, saturazione) avviene interamente on-chip in aritmetica
intera INT8/INT32; il sistema host fornisce solo configurazione, pesi, dati di
ingresso e controllo attraverso una semplice interfaccia SPI, senza mai far parte
del datapath computazionale. Un unico bitstream serve qualunque topologia fino al
massimo di build.}
\vspace{8pt}
\begin{multicols}{2}
{\color{fnDark}\large\bfseries Caratteristiche}\\[2pt]
{\footnotesize
\begin{itemize}[leftmargin=1.1em]
\item Datapath \textbf{INT8 $\times$ INT8 $\to$ INT16 $\to$ INT32}, accumulo a 32~bit
con estensione di segno.
\item \textbf{Balanced binary adder tree} ($O(\log_2 \text{PARALLEL})$) al posto della
riduzione lineare.
\item MAC parallelo configurabile: \code{PARALLEL} MAC hardware simultanei per neurone,
mappati su DSP \code{MULT18X18D}.
\item Architettura completamente \textbf{parametrica}: \code{N\_INPUTS}, \code{N\_NEURONS},
\code{PARALLEL}, \code{DATA\_WIDTH}, \code{ACC\_WIDTH}, \code{N\_LAYERS}.
\item \textbf{Larghezza di rete a runtime}: \code{n\_inputs\_real}/\code{n\_neurons\_real}
per-layer, un solo bitstream per ogni topologia fino al massimo.
\item Attivazioni configurabili: \code{ACT\_RELU} (default) e \code{ACT\_NONE} (lineare
con saturazione bilaterale), con saturazione INT8.
\item \textbf{Due tipi di rete}: classica multi-layer dense (\code{layer\_sequencer},
buffer ping-pong) e \textbf{grafo arbitrario sparse} (\code{graph\_engine} +
buffer di attivazione in block RAM \code{DP16KD}), selezionabili a runtime.
\item Sottosistema di \textbf{memoria dedicata}: interfaccia byte$\leftrightarrow$word,
controller PSRAM parallelo asincrono con \textbf{page mode} (70~ns accesso
casuale, 20~ns burst di pagina), 8~MB indirizzabili (23~bit).
\item Interfaccia host \textbf{SPI Mode 0} MSB-first, \code{SET\_NET\_TYPE}+dispatch, \code{STATUS.done}
sticky/clear-on-read, \code{READ\_CONFIG} runtime.
\item \textbf{Sottosistema flash} boot/persistenza: accesso esclusivo della FPGA a una
\code{W25Q128JV} SPI NOR (16~MB) via SPI master dedicato, copy engine
flash$\leftrightarrow$PSRAM e catalogo a 16 slot con CRC32, 8 opcode host.
\item Verificato in \textbf{simulazione} (Icarus Verilog) e \textbf{sintesi reale}
(Yosys + nextpnr-ecp5 + ecppack).
\end{itemize}}
\columnbreak
{\color{fnDark}\large\bfseries Applicazioni}\\[2pt]
{\footnotesize
\begin{itemize}[leftmargin=1.1em]
\item Inferenza a bassa latenza deterministica come periferica di
SoC Linux, Raspberry-Pi-like, ESP32, microcontrollori.
\item Blocco hardware riusabile integrabile in progetti eterogenei
(piattaforma, non singola rete).
\item Edge AI su reti dense compatte quantizzate INT8.
\item Off-loading del carico neurale dalla CPU host verso hardware
dedicato con throughput prevedibile.
\end{itemize}}
\vspace{4pt}
{\color{fnDark}\large\bfseries Target \& toolchain}\\[2pt]
{\footnotesize
\begin{itemize}[leftmargin=1.1em]
\item FPGA: Lattice ECP5 \code{LFE5U-45F-8BG381C} ($-8$, CABGA381).
\item Sintesi: Yosys; place\&route: nextpnr-ecp5; bitstream: Project~Trellis
(\code{ecppack}).
\item Simulazione: Icarus Verilog (\code{-g2012}).
\item PSRAM: ISSI \code{IS66WVE4M16EBLL-70BLI} (64\,Mb, 4M$\times$16).
\end{itemize}}
\end{multicols}
\vspace{2pt}
% --- tabella parametri chiave ---
\noindent
{\small\color{fnDark}\bfseries Parametri chiave (configurazione baseline caratterizzata)}
\vspace{2pt}
\noindent
\begin{tabularx}{\textwidth}{L{3.2cm}L{3.6cm}Y}
\toprule
\rowh \thd{Grandezza} & \thd{Valore} & \thd{Note} \\
\midrule
Precisione dati & INT8 (signed) & \code{DATA\_WIDTH}=8 \\
\rowa Accumulatore & INT32 (signed) & \code{ACC\_WIDTH}=32 \\
Ingressi / neuroni & 256 / 4 & baseline benchmark datapath \\
\rowa MAC simultanei & $2\ldots64$ & $=$\code{PARALLEL}$\times$\code{N\_NEURONS} \\
Attivazioni & ReLU, lineare & \code{ACT\_RELU} / \code{ACT\_NONE} \\
\rowa Fmax (P=2, datapath) & 87.88~MHz & benchmark datapath isolato \\
Fmax (P=2, sistema integrato) & 67.91~MHz & sistema completo incl. sottosistema flash, place\&route reale \\
Throughput MAC (P=16) & $\approx$3.34~G\,MAC/s & teorico, solo datapath \\
\rowa Memoria di lavoro & 8~MB PSRAM & bus parallelo 16-bit, 70~ns / 20~ns page mode \\
Spazio indirizzi & 23~bit (byte) & \code{ADDR\_WIDTH}=23 \\
\bottomrule
\end{tabularx}
\vspace{8pt}
\noindent
{\small\color{fnDark}\bfseries Diagramma a blocchi del sistema}
\begin{center}
\begin{tikzpicture}[node distance=6mm and 10mm,font=\footnotesize]
\node[fnblockD,minimum width=26mm,minimum height=13mm] (host){HOST\\{\scriptsize configura / addestra / controlla}};
\node[fnblockT,right=16mm of host,minimum width=52mm,minimum height=22mm] (eng){};
\node[anchor=north,font=\footnotesize\bfseries,text=fnDark] at (eng.north){FPGA -- Neural Network Engine};
\node[fnreg,fill=white] (spi) at ([yshift=-2mm]eng.center){\code{spi\_slave} + \code{spi\_engine}};
\node[fnreg,fill=white,below=2.5mm of spi] (arb){\code{mem\_arbiter} + \code{layer\_sequencer}};
\node[fnreg,fill=white,above=2.5mm of spi] (core){\code{neuron\_memory} $\to$ \code{neuron\_parallel} $\to$ \code{mac8}};
\node[fnblock,right=16mm of eng,minimum width=24mm,minimum height=13mm] (ram){PSRAM 8\,MB\\{\scriptsize \code{psram\_controller}}};
\draw[fnbus] (host) -- node[fnlbl,above]{SPI} (eng.west|-host);
\draw[fnbus] (eng.east|-ram) -- node[fnlbl,above]{16-bit async} (ram);
\end{tikzpicture}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
Il datapath neurale è interamente nell'FPGA; l'host non partecipa alle singole
operazioni MAC.\end{center}
@@ -1,104 +0,0 @@
\thispagestyle{plain}
\noindent
\begin{tikzpicture}
\node[fill=fnDark,text=white,rounded corners=2pt,inner sep=6pt,
minimum width=\textwidth,anchor=west]
{\large\bfseries Sunto del pinout --- collegamento pin per pin};
\end{tikzpicture}
\vspace{6pt}
\noindent
{\footnotesize
Tabella di riferimento rapido: i \textbf{57 segnali reali} del top-level
\code{spi\_neuron\_top}, ciascuno con la propria ball \code{CABGA381}
individuale (\textbf{non} un intervallo di bus) --- dati reali dal database
di dispositivo di Project~Trellis (\code{iodb.json}), \textbf{verificati da
un place\&route \code{nextpnr-ecp5} completo a 0 errori} (non un pinout
pianificato). Descrizione completa, razionale di collocazione per banco e
schema di collegamento pin-per-pin verso la PSRAM ISSI: cap.~\ref{ch:hw}.
}
\vspace{4pt}
\noindent
\renewcommand{\arraystretch}{1.08}
\begin{tabularx}{\textwidth}{L{2.7cm} C{1.0cm} C{1.0cm} C{0.9cm} Y}
\toprule
\rowh \thd{Segnale} & \thd{Ball} & \thd{Banco} & \thd{Dir} & \thd{Pin corrispondente / funzione} \\
\midrule
\multicolumn{5}{l}{\textit{\color{fnDark}Clock e reset}}\\
\code{clk} & H5 & 7 & IN & Clock di sistema, pad \code{GR\_PCLK7\_0} (clock globale dedicato). \\
\rowa \code{rst} & B4 & 7 & IN & Reset globale sincrono, attivo alto. \\
\multicolumn{5}{l}{\textit{\color{fnDark}SPI applicativo (host $\leftrightarrow$ FPGA, Mode~0)}}\\
\code{sclk} & B5 & 7 & IN & SPI clock (CPOL=0, CPHA=0). \\
\rowa \code{mosi} & C5 & 7 & IN & Master-Out Slave-In. \\
\code{miso} & A3 & 7 & OUT & Master-In Slave-Out. \\
\rowa \code{cs\_n} & B3 & 7 & IN & Chip-select, attivo basso. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Attenzione host (attivi bassi, di livello)}}\\
\code{data\_ready\_n} & C3 & 7 & OUT & Basso finché un risultato attende lettura. \\
\rowa \code{irq\_n} & C4 & 7 & OUT & Basso se il guard load-time del grafo è scattato. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Flash subsystem --- SPI verso W25Q128JV (boot/persistenza)}}\\
\code{flash\_sclk} & E3 & 7 & OUT & SPI clock verso la flash --- GPIO ordinario, indipendente (Fase F7, cap.~\ref{ch:hw}). \\
\rowa \code{flash\_mosi} & D3 & 7 & OUT & Master-Out Slave-In verso la flash NOR onboard. \\
\code{flash\_miso} & D5 & 7 & IN & Master-In Slave-Out dalla flash. \\
\rowa \code{flash\_cs\_n} & E4 & 7 & OUT & Chip-select flash, attivo basso. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Bus indirizzi PSRAM --- \code{psram\_a[21:0]} (22 linee reali)}}\\
\code{psram\_a[0]} & E16 & 2 & OUT & PSRAM A0 \\
\rowa \code{psram\_a[1]} & F16 & 2 & OUT & PSRAM A1 \\
\code{psram\_a[2]} & D18 & 2 & OUT & PSRAM A2 \\
\rowa \code{psram\_a[3]} & E17 & 2 & OUT & PSRAM A3 \\
\code{psram\_a[4]} & E18 & 2 & OUT & PSRAM A4 \\
\rowa \code{psram\_a[5]} & F18 & 2 & OUT & PSRAM A5 \\
\code{psram\_a[6]} & F17 & 2 & OUT & PSRAM A6 \\
\rowa \code{psram\_a[7]} & G16 & 2 & OUT & PSRAM A7 \\
\code{psram\_a[8]} & G18 & 2 & OUT & PSRAM A8 \\
\rowa \code{psram\_a[9]} & H16 & 2 & OUT & PSRAM A9 \\
\code{psram\_a[10]} & H17 & 2 & OUT & PSRAM A10 \\
\rowa \code{psram\_a[11]} & H18 & 2 & OUT & PSRAM A11 \\
\code{psram\_a[12]} & J16 & 2 & OUT & PSRAM A12 \\
\rowa \code{psram\_a[13]} & J17 & 2 & OUT & PSRAM A13 \\
\code{psram\_a[14]} & C20 & 2 & OUT & PSRAM A14 \\
\rowa \code{psram\_a[15]} & D19 & 2 & OUT & PSRAM A15 \\
\code{psram\_a[16]} & E19 & 2 & OUT & PSRAM A16 \\
\rowa \code{psram\_a[17]} & E20 & 2 & OUT & PSRAM A17 \\
\code{psram\_a[18]} & F19 & 2 & OUT & PSRAM A18 \\
\rowa \code{psram\_a[19]} & F20 & 2 & OUT & PSRAM A19 \\
\code{psram\_a[20]} & G20 & 2 & OUT & PSRAM A20 \\
\rowa \code{psram\_a[21]} & H20 & 2 & OUT & PSRAM A21 \\
\code{psram\_a[22]} & P18 & 3 & OUT & Sempre 0 (shift byte$\to$word) --- NC su scheda. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Bus dati PSRAM --- \code{psram\_dq[15:0]} (bidirezionale)}}\\
\rowa \code{psram\_dq[0]} & K18 & 2 & IO & PSRAM DQ0 \\
\code{psram\_dq[1]} & C18 & 2 & IO & PSRAM DQ1 \\
\rowa \code{psram\_dq[2]} & D17 & 2 & IO & PSRAM DQ2 \\
\code{psram\_dq[3]} & D20 & 2 & IO & PSRAM DQ3 \\
\rowa \code{psram\_dq[4]} & G19 & 2 & IO & PSRAM DQ4 \\
\code{psram\_dq[5]} & J18 & 2 & IO & PSRAM DQ5 \\
\rowa \code{psram\_dq[6]} & J19 & 2 & IO & PSRAM DQ6 \\
\code{psram\_dq[7]} & J20 & 2 & IO & PSRAM DQ7 \\
\rowa \code{psram\_dq[8]} & K19 & 2 & IO & PSRAM DQ8 \\
\code{psram\_dq[9]} & K20 & 2 & IO & PSRAM DQ9 \\
\rowa \code{psram\_dq[10]} & L17 & 3 & IO & PSRAM DQ10 \\
\code{psram\_dq[11]} & M18 & 3 & IO & PSRAM DQ11 \\
\rowa \code{psram\_dq[12]} & M17 & 3 & IO & PSRAM DQ12 \\
\code{psram\_dq[13]} & N16 & 3 & IO & PSRAM DQ13 \\
\rowa \code{psram\_dq[14]} & N18 & 3 & IO & PSRAM DQ14 \\
\code{psram\_dq[15]} & P17 & 3 & IO & PSRAM DQ15 \\
\multicolumn{5}{l}{\textit{\color{fnDark}Controllo PSRAM}}\\
\rowa \code{psram\_ce\_n} & N17 & 3 & OUT & PSRAM CE\# --- chip enable, attivo basso. \\
\code{psram\_oe\_n} & R16 & 3 & OUT & PSRAM OE\# --- output enable (lettura). \\
\rowa \code{psram\_we\_n} & R17 & 3 & OUT & PSRAM WE\# --- write enable (scrittura). \\
\code{psram\_lb\_n} & T16 & 3 & OUT & PSRAM LB\# --- lower-byte enable (DQ[7:0]). \\
\rowa \code{psram\_ub\_n} & N19 & 3 & OUT & PSRAM UB\# --- upper-byte enable (DQ[15:8]). \\
\code{psram\_zz\_n} & N20 & 3 & OUT & PSRAM ZZ\# --- sleep/snooze (alto in funzionamento). \\
\bottomrule
\end{tabularx}
\renewcommand{\arraystretch}{1.25}
\vspace{4pt}
\noindent
{\footnotesize\color{fnGrey}
Standard I/O: LVCMOS33 su tutti i 57 segnali. Ball di JTAG e config-SPI di
boot (pin dedicati a funzione fissa, senza porta RTL) non compaiono in
questa tabella --- vedi cap.~\ref{ch:hw} §``Configurazione e
programmazione''. Sorgente: \code{synth/ecp5/spi\_neuron\_top.lpf},
generato da \code{tools/pinout/gen\_lpf.py} contro
\code{iodb.json} di Project~Trellis.\par}
@@ -1,94 +0,0 @@
\chapter{Panoramica del sistema}
\label{ch:overview}
\section{Obiettivo del progetto}
FPGA-Neural implementa un \textbf{Neural Network Engine riusabile in hardware FPGA}.
L'insieme è composto da tre elementi: l'FPGA, che è il vero acceleratore; una RAM
dedicata fisicamente associata all'FPGA e non condivisa con l'host; e un'interfaccia
host indipendente dal sistema operativo, inizialmente SPI (con possibile estensione
futura a Dual~SPI).
Il principio fondante è la separazione fra chi \emph{esegue} il calcolo e chi lo
\emph{usa}: il calcolo della rete neurale avviene interamente dentro l'FPGA, mentre
il sistema host fornisce solo configurazione, parametri di rete, dati di ingresso,
controllo e lettura dei risultati. L'host non fa parte del datapath computazionale.
Sistemi host possibili includono SoC Linux, sistemi tipo Raspberry~Pi, ESP32,
microcontrollori e PC di sviluppo: la stessa architettura di engine deve poter essere
usata in sistemi completamente diversi.
\begin{center}
\begin{tikzpicture}[font=\footnotesize,node distance=8mm]
\node[fnblockD,minimum width=42mm,minimum height=20mm] (host){\textbf{HOST}\\[2pt]
{\scriptsize Configurazione}\\{\scriptsize Addestramento}\\{\scriptsize Controllo}};
\node[fnblockT,below=14mm of host,minimum width=42mm,minimum height=20mm] (fpga)
{\textbf{FPGA}\\[2pt]{\scriptsize Neural Network Engine}\\{\scriptsize Compute / Control}};
\node[fnblock,below=14mm of fpga,minimum width=42mm,minimum height=13mm] (ram)
{\textbf{RAM dedicata}\\{\scriptsize pesi / bias / buffer}};
\draw[fnbus] (host) -- node[fnlbl,right]{SPI / Dual SPI} (fpga);
\draw[fnbus] (fpga) -- node[fnlbl,right]{bus parallelo} (ram);
\end{tikzpicture}
\end{center}
\section{Configurazione hardware contro configurazione di rete}
Il progetto distingue con precisione fra l'\textbf{architettura hardware}
dell'acceleratore e i \textbf{parametri della rete neurale}.
L'architettura fisica dell'engine è definita al momento della sintesi e
dell'implementazione dell'FPGA. I parametri hardware tipici sono \code{N\_INPUTS},
\code{N\_NEURONS}, \code{N\_LAYERS}, \code{PARALLEL}, \code{DATA\_WIDTH},
\code{ACC\_WIDTH}: sono parametri Verilog risolti in fase di sintesi e determinano il
datapath contenuto nel bitstream. I parametri della rete --- pesi, bias, parametri di
attivazione e di quantizzazione, costanti specifiche --- vengono invece caricati a
runtime attraverso l'interfaccia host e memorizzati nella RAM associata all'FPGA.
\begin{fnnote}[Principio architetturale centrale]
Una build fissa il \emph{soffitto} della macchina (numero massimo di layer, larghezza
massima, \code{PARALLEL}); l'host configura la rete \emph{reale} --- numero di layer,
larghezza ingressi/uscite per-layer, attivazione per-layer e parametri addestrati ---
interamente a runtime, via SPI, nella memoria locale dell'FPGA. Un solo bitstream
serve qualunque topologia fino a quel soffitto.
\end{fnnote}
\section{Boot e inizializzazione}
L'FPGA viene configurato all'accensione tramite il consueto meccanismo di
configurazione (caricamento del bitstream da flash SPI). Il bitstream definisce
l'architettura hardware dell'engine; l'host non costruisce dinamicamente il datapath
durante il funzionamento normale, ma configura i dati di rete su cui il datapath già
esistente opera.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=4.5mm,start chain=going below,
every node/.style={on chain}]
\node[fnblockA,minimum width=60mm](p){Power-on};
\node[fnblock,minimum width=60mm]{Configurazione FPGA (bitstream da flash)};
\node[fnblockT,minimum width=60mm]{Neural Network Engine disponibile};
\node[fnblock,minimum width=60mm]{Inizializzazione host (SPI)};
\node[fnblock,minimum width=60mm]{Caricamento parametri di rete / pesi / bias};
\node[fnblockD,minimum width=60mm]{Engine pronto};
\begin{scope}[every path/.style={fnarrow}]
\foreach \a/\b in {1/2,2/3,3/4,4/5,5/6}{}
\end{scope}
\foreach \i [count=\j from 2] in {1,...,5}{
\draw[fnarrow] (chain-\i) -- (chain-\j);}
\end{tikzpicture}
\end{center}
\section{Addestramento e inferenza}
Addestramento e inferenza sono concettualmente separati. La prima implementazione non
richiede che l'FPGA esegua l'addestramento: i pesi possono essere calcolati
esternamente (PC/Linux/altro host) e trasferiti via SPI nella RAM dell'FPGA, che poi
esegue l'inferenza. Questo riduce drasticamente la complessità dell'hardware iniziale,
senza precludere una futura implementazione di training assistito o interamente
hardware (Fase~8 della roadmap, cap.~\ref{ch:roadmap}). Durante l'inferenza l'host
fornisce solo i dati di ingresso e recupera il risultato, ottenendo calcolo
deterministico, carico ridotto sull'host, parallelismo hardware, latenza prevedibile e
indipendenza dall'architettura della CPU host.
\section{Filosofia di progetto e riuso}
Il progetto va inteso come una \emph{piattaforma di accelerazione neurale FPGA
riusabile} più che come una singola rete. L'applicazione determina dimensione degli
ingressi, topologia, numero di layer e neuroni, parallelismo, precisione numerica,
funzioni di attivazione, requisiti di memoria e prestazioni; il processo di
generazione hardware produce l'implementazione FPGA corrispondente. La stessa
architettura HDL rimane concettualmente invariata mentre i parametri di sintesi
generano implementazioni appropriate ai diversi target applicativi.
@@ -1,80 +0,0 @@
\chapter[Architettura RTL]{Architettura RTL e gerarchia dei moduli}
\label{ch:arch}
\section{Organizzazione gerarchica}
Il design è organizzato per livelli, dal moltiplicatore-accumulatore elementare fino
al top-level integrato con interfaccia SPI e PSRAM. Ogni livello incapsula il
precedente e ne astrae i dettagli: il datapath validato (\code{mac\_unit},
\code{mac8}, \code{neuron\_parallel}) non viene mai modificato dai livelli di
orchestrazione superiori.
\begin{center}
\begin{tikzpicture}[font=\footnotesize,every node/.style={fnblock,minimum width=40mm},
level distance=13mm,sibling distance=0mm]
\node[fnblockD,minimum width=62mm](top){\code{spi\_neuron\_top} \\ {\scriptsize top-level integrato}};
\node[fnblockT,minimum width=62mm,below=8mm of top](arb){\code{mem\_arbiter} \;/\; \code{layer\_sequencer} \\ {\scriptsize arbitraggio 3 porte + sequenza layer}};
\node[fnblock,minimum width=62mm,below=8mm of arb](nm){\code{neuron\_memory} \\ {\scriptsize ponte memoria $\leftrightarrow$ neurone, loop neuroni}};
\node[fnblock,minimum width=62mm,below=8mm of nm](np){\code{neuron\_parallel} \\ {\scriptsize FSM neurone: gruppi, bias, attivazione, saturazione}};
\node[fnblockT,minimum width=62mm,below=8mm of np](m8){\code{mac8} \\ {\scriptsize \code{PARALLEL} MAC + balanced adder tree}};
\node[fnblock,minimum width=62mm,below=8mm of m8](mu){\code{mac\_unit} \\ {\scriptsize $x\cdot w$ + estensione segno + accumulo}};
\foreach \a/\b in {top/arb,arb/nm,nm/np,np/m8,m8/mu}
\draw[fnarrow] (\a) -- (\b);
% rami memoria a destra
\node[fnblockA,minimum width=34mm,right=14mm of nm](ma){\code{int8\_memory\_access}\\{\scriptsize byte $\leftrightarrow$ word 16-bit}};
\node[fnblockA,minimum width=34mm,below=6mm of ma](mi){\code{memory\_interface}\\{\scriptsize handshake req/ready}};
\node[fnblockA,minimum width=34mm,below=6mm of mi](pc){\code{psram\_controller}\\{\scriptsize bus fisico PSRAM}};
\draw[fnarrowT] (ma)--(mi); \draw[fnarrowT] (mi)--(pc);
\draw[fnarrowT,dashed] (nm.east) -- (ma.west);
% rami SPI a sinistra
\node[fnblockA,minimum width=30mm,left=14mm of arb,yshift=6mm](ss){\code{spi\_slave}\\{\scriptsize layer fisico Mode 0}};
\node[fnblockA,minimum width=30mm,below=6mm of ss](se){\code{spi\_engine}\\{\scriptsize FSM opcode + registri}};
\draw[fnarrowT] (ss)--(se);
\draw[fnarrowT,dashed] (se.east) -- (arb.west);
\end{tikzpicture}
\end{center}
\section{Ruolo di ciascun modulo}
\begin{tabularx}{\textwidth}{L{3.4cm}Y}
\toprule
\rowh \thd{Modulo} & \thd{Funzione} \\
\midrule
\code{mac\_unit} & Singolo prodotto-accumulatore: $\mathrm{acc\_out}=\mathrm{acc\_in}+(x\cdot w)$, con estensione di segno del prodotto ad \code{ACC\_WIDTH}. Parametrico su \code{DATA\_WIDTH}/\code{ACC\_WIDTH}. \\
\rowa \code{mac8} & \code{PARALLEL} istanze di \code{mac\_unit} i cui prodotti vengono sommati da un \emph{balanced binary adder tree} di profondità $\log_2(\text{PARALLEL})$; il risultato è aggiunto all'accumulatore in ingresso. \\
\code{neuron\_parallel} & FSM di un singolo neurone: elabora \code{N\_INPUTS} ingressi in gruppi di \code{PARALLEL}, accumula tra i gruppi, somma il bias, applica l'attivazione e satura a INT8. Include il guard di elaborazione su \code{N\_INPUTS \% PARALLEL} e la larghezza runtime \code{n\_inputs\_real}. \\
\rowa \code{layer} & Istanzia \code{N\_NEURONS} neuroni \emph{in parallelo} sullo stesso vettore di ingresso; \code{busy}=OR, \code{done}=AND dei neuroni. Percorso puramente combinatorio-di-dati usato nei benchmark del datapath. \\
\code{neuron\_memory} & Integra il calcolo con la memoria: legge $X$ (condiviso) una volta, poi per ogni neurone rilegge $W$ e bias dalla RAM e riusa una singola istanza \code{neuron\_parallel} (memory-bound, un neurone per volta). Uscita \code{y\_bus} packed neuron-major. \\
\rowa \code{layer\_sequencer} & Concatena fino a \code{N\_LAYERS} esecuzioni di \code{neuron\_memory} leggendo una tabella descrittori scritta dall'host e alternando i buffer ping-pong in RAM (Fase~5). \\
\code{act\_buffer} & Buffer di attivazione globale in block RAM \code{DP16KD}, indicizzato per id di segnale (Tipo \#2). \\
\rowa \code{graph\_engine} & Motore della rete a grafo (Tipo \#2): gather da \code{act\_buffer}, riusa \code{neuron\_parallel}, scrive le uscite per id (cap.~\ref{ch:grafo}). \\
\code{int8\_memory\_access} & Converte l'interfaccia byte/INT8 (indirizzo di byte) nell'interfaccia a parola 16-bit, selezionando il byte basso/alto tramite \code{lb\_n}/\code{ub\_n} e \code{addr>>1}. \\
\rowa \code{memory\_interface} & FSM di handshake a 2 stati (IDLE/WAIT) che serializza la singola transazione verso il controller. \\
\code{psram\_controller} & Controller del bus PSRAM parallelo asincrono con \textbf{page mode} di lettura: accesso casuale a 70~ns (\code{tAA}), burst nella stessa pagina a 20~ns (\code{tAPA}) con CE\#/OE\# tenuti attivi; abilita il page mode sul chip all'avvio via registro di configurazione (cap.~\ref{ch:mem}, \S~5.5). Pilota \code{ce\_n/oe\_n/we\_n/lb\_n/ub\_n/zz\_n} e il bus dati tri-state. \\
\rowa \code{mem\_arbiter} & Arbitro a priorità fissa (B$>$C$>$A) fra tre master byte-level: \code{spi\_engine} (A), \code{neuron\_memory} (B), \code{layer\_sequencer} (C). \\
\code{spi\_slave} & Layer fisico SPI Mode 0, MSB-first, sincronizzatore CDC a 3 stadi su SCLK/MOSI/CS\_N, shift-register e framing di CS. \\
\rowa \code{spi\_engine} & FSM di protocollo/opcode e banco registri (\code{x\_base}, \code{w\_base}, \code{bias\_addr}, base ping-pong, attivazione, larghezze runtime\ldots), con \code{STATUS.done} sticky/clear-on-read. \\
\code{spi\_neuron\_top} & Top-level: collega SPI, arbitro, sequencer, \code{neuron\_memory} e catena PSRAM; multiplexa il controllo di \code{neuron\_memory} fra sequencer e percorso diretto single-layer. \\
\bottomrule
\end{tabularx}
\vspace{6pt}
\begin{fnnote}[Modelli di simulazione]
\code{psram\_model.v} (in \code{sim/}) e \code{memory\_model.v} sono modelli
comportamentali della memoria usati nei testbench; non fanno parte del design
sintetizzabile ma riproducono la latenza reale per la verifica end-to-end.
\end{fnnote}
\section{Due percorsi di esecuzione}
Il top-level espone due modalità mutuamente esclusive verso lo stesso motore di
calcolo \code{neuron\_memory}:
\begin{itemize}
\item \textbf{Percorso single-layer / manuale}: l'host imposta le basi con
\op{SET\_BASE}, avvia con \op{START} e legge con \op{READ\_OUTPUT}. \code{spi\_engine}
pilota direttamente \code{neuron\_memory}.
\item \textbf{Percorso multi-layer}: l'host scrive la tabella descrittori e avvia con
\op{RUN\_NETWORK}; \code{layer\_sequencer} prende possesso del controllo di
\code{neuron\_memory} (mentre \code{seq\_busy} è alto) e concatena i layer.
\end{itemize}
Il multiplexer del top-level commuta le linee di controllo di \code{neuron\_memory}
in base a \code{seq\_busy}, restituendo il motore al percorso diretto a fine sequenza.
@@ -1,165 +0,0 @@
\chapter{Datapath di calcolo}
\label{ch:datapath}
\section{Catena aritmetica INT8/INT32}
Il datapath elementare implementa la sequenza tipica di un neurone quantizzato:
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=3mm,start chain=going right,
every node/.style={fnblock,minimum width=15mm,minimum height=8mm,on chain}]
\node[fnblockT]{INT8\\$\times$\,INT8};
\node{INT16\\prodotto};
\node{sign-ext\\INT32};
\node[fnblockD]{accumulo\\INT32};
\node{$+$ bias};
\node[fnblockA]{attivazione};
\node[fnblockT]{sat. INT8};
\foreach \i [count=\j from 2] in {1,...,6}
\draw[fnarrow] (chain-\i) -- (chain-\j);
\end{tikzpicture}
\end{center}
Ogni prodotto INT8$\times$INT8 sta in 16~bit; viene esteso con segno a 32~bit prima
dell'accumulo, così l'accumulatore non trabocca su vettori lunghi. Bias e attivazione
operano a 32~bit; solo l'uscita finale viene saturata a INT8.
\section{\texttt{mac\_unit} --- moltiplicatore-accumulatore}
Il modulo \code{mac\_unit} è puramente combinatorio e parametrico su \code{DATA\_WIDTH}
e \code{ACC\_WIDTH}. Calcola:
\[
\mathrm{acc\_out} = \mathrm{acc\_in} + \mathrm{signext}_{ACC}(x \cdot w)
\]
Il prodotto ha larghezza $2\times$\code{DATA\_WIDTH} e viene esteso con segno replicando
il bit più significativo. Su ECP5 la moltiplicazione mappa su un blocco DSP
\code{MULT18X18D}.
\begin{lstlisting}[caption={\texttt{rtl/mac\_unit.v} --- nucleo aritmetico},label={lst:macunit}]
localparam PROD_WIDTH = 2 * DATA_WIDTH;
wire signed [PROD_WIDTH-1:0] product = x * w;
wire signed [ACC_WIDTH-1:0] product_ext =
{{(ACC_WIDTH-PROD_WIDTH){product[PROD_WIDTH-1]}}, product};
assign acc_out = acc_in + product_ext;
\end{lstlisting}
\section{\texttt{mac8} --- MAC parallelo e balanced adder tree}
\code{mac8} istanzia \code{PARALLEL} unità \code{mac\_unit} che generano
\code{PARALLEL} prodotti indipendenti, poi li somma con un \emph{albero di addizione
binario bilanciato}. Rispetto alla riduzione lineare
$((((p_0{+}p_1){+}p_2){+}p_3){+}\dots)$, di profondità $O(\text{PARALLEL})$, l'albero
ha profondità $O(\log_2 \text{PARALLEL})$, riducendo drasticamente il percorso
combinatorio.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,level distance=11mm,
every node/.style={fnreg,minimum width=8mm},
level 1/.style={sibling distance=30mm},
level 2/.style={sibling distance=15mm},
level 3/.style={sibling distance=8mm},
edge from parent/.style={fnarrowT,draw}]
\node[fnblockD]{sum}
child {node[fnblockT]{$+$}
child {node[fnblockT]{$+$}
child {node{$p_0$}} child {node{$p_1$}}}
child {node[fnblockT]{$+$}
child {node{$p_2$}} child {node{$p_3$}}}}
child {node[fnblockT]{$+$}
child {node[fnblockT]{$+$}
child {node{$p_4$}} child {node{$p_5$}}}
child {node[fnblockT]{$+$}
child {node{$p_6$}} child {node{$p_7$}}}};
\end{tikzpicture}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
Esempio con PARALLEL=8: 3 livelli. PARALLEL=16 $\to$ 4 livelli; PARALLEL=32 $\to$ 5
livelli.\end{center}
\begin{fnnote}[PARALLEL potenza di due]
L'albero è pensato per \code{PARALLEL} potenza di due (8, 16, 32\ldots). Questo è anche
il valore usato in tutte le configurazioni del progetto.
\end{fnnote}
\section{\texttt{neuron\_parallel} --- FSM del neurone}
\code{neuron\_parallel} elabora \code{N\_INPUTS} ingressi in gruppi di \code{PARALLEL},
mantenendo l'accumulatore tra un gruppo e il successivo. Alla fine somma il bias,
applica l'attivazione e satura a INT8. Il numero di gruppi è
$\text{GROUPS}=\text{N\_INPUTS}/\text{PARALLEL}$.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=4mm,start chain=going below,
every node/.style={on chain,fnblock,minimum width=46mm}]
\node[fnblockA]{\code{start}};
\node{gruppo 0 $\to$ accumulo};
\node{gruppo 1 $\to$ accumulo};
\node[draw=none,fill=none]{\vdots};
\node{gruppo GROUPS$-$1 $\to$ accumulo};
\node{$+$ bias};
\node[fnblockA]{attivazione (ACT\_RELU / ACT\_NONE)};
\node[fnblockT]{saturazione INT8};
\node[fnblockD]{\code{done}, \code{y}};
\foreach \i [count=\j from 2] in {1,...,8}
\draw[fnarrow] (chain-\i) -- (chain-\j);
\end{tikzpicture}
\end{center}
\subsection{Guard di parametri (elaboration-time)}
Se \code{PARALLEL} non divide esattamente \code{N\_INPUTS} si verificano due guasti,
entrambi confermati empiricamente in \code{sim/parameter\_sweep\_tb.v}:
\begin{itemize}
\item la divisione intera tronca \code{GROUPS} e gli ingressi in eccesso non vengono
mai letti $\to$ risultato \textbf{errato}, senza errore né avviso;
\item se \code{PARALLEL > N\_INPUTS}, \code{GROUPS=0} e la condizione terminale non è mai
soddisfatta $\to$ il neurone \textbf{si blocca} (busy alto, done mai asserito).
\end{itemize}
La soluzione non modifica il datapath validato: un blocco \code{generate} istanzia un
modulo deliberatamente indefinito quando $\text{N\_INPUTS} \bmod \text{PARALLEL}\neq0$,
forzando un errore in \emph{elaborazione} sia in simulazione sia in sintesi. Per le
configurazioni valide il ramo non viene mai elaborato.
\begin{lstlisting}[caption={\texttt{rtl/neuron\_parallel.v} --- guard di parametri}]
generate
if (N_INPUTS == 0 || N_INPUTS % PARALLEL != 0) begin : PARAMETER_ERROR
neuron_parallel_requires_N_INPUTS_multiple_of_PARALLEL
invalid_parameter_combination();
end
endgenerate
\end{lstlisting}
\begin{fnnote}[Caso limite \texttt{N\_INPUTS=0} (corretto 2026-09-04)]
La condizione originale (\code{N\_INPUTS \% PARALLEL != 0}) non intercetta
\code{N\_INPUTS=0}, poiché $0 \bmod \text{PARALLEL}=0$ per ogni \code{PARALLEL}: il modulo
elaborava con successo (sia in simulazione sia in sintesi reale Yosys) lasciando
\code{x\_bus}/\code{w\_bus} non pilotati e \code{start} silenziosamente inefficace. Trovato
durante la campagna di ri-certificazione (\code{docs/validation/bugs.md}, BUG-002) e
corretto estendendo il guard come sopra --- \code{N\_INPUTS=0} ora fallisce l'elaborazione
esattamente come gli altri casi degeneri.
\end{fnnote}
\section{Funzioni di attivazione}
\code{neuron\_parallel} accetta una porta \code{activation} a 2~bit. Il default è
\code{ACT\_RELU}, l'unico comportamento esistente prima dell'introduzione della porta,
così ogni chiamante preesistente resta invariato.
\begin{tabularx}{\textwidth}{L{2.6cm} C{1.4cm} Y}
\toprule
\rowh \thd{Codifica} & \thd{Valore} & \thd{Comportamento} \\
\midrule
\code{ACT\_NONE} & \code{2'd0} & Lineare: nessun clamp a zero, saturazione bilaterale al range INT8 $[-128,+127]$. \\
\rowa \code{ACT\_RELU} & \code{2'd1} & $\max(0,x)$, poi saturazione positiva a $+127$ (default; fallback anche per codifiche riservate). \\
\bottomrule
\end{tabularx}
\section{Saturazione INT8}
Dopo bias e attivazione, l'accumulatore a 32~bit viene ridotto a INT8:
\[
y=\begin{cases}
+127 & \text{se } \mathrm{final\_acc} > 127\\
-128 & \text{se } \mathrm{final\_acc} < -128 \ \text{(solo ACT\_NONE)}\\
0 & \text{se } \mathrm{final\_acc}\le 0 \ \text{(solo ACT\_RELU)}\\
\mathrm{final\_acc}[7:0] & \text{altrimenti}
\end{cases}
\]
\section{\texttt{layer} --- neuroni in parallelo}
\code{layer} istanzia \code{N\_NEURONS} neuroni che condividono il vettore di ingresso
\code{x\_bus} ma hanno pesi e bias distinti; \code{busy} è l'OR e \code{done} l'AND dei
segnali dei neuroni. È il modulo usato nei benchmark del datapath (cap.~\ref{ch:impl}),
dove tutti i neuroni lavorano simultaneamente. La convenzione di indirizzamento è
neuron-major: i pesi del neurone $n$ occupano \code{weights\_bus[n*N\_INPUTS*DATA\_WIDTH +: N\_INPUTS*DATA\_WIDTH]}.
@@ -1,87 +0,0 @@
\chapter{Parametri e configurabilità}
\label{ch:param}
\section{Parametri di build (synthesis-time)}
L'architettura hardware è fissata alla sintesi tramite i parametri Verilog seguenti.
Determinano il datapath contenuto nel bitstream e il suo \emph{soffitto} di capacità.
\begin{tabularx}{\textwidth}{L{3.0cm} C{1.8cm} Y}
\toprule
\rowh \thd{Parametro} & \thd{Default} & \thd{Significato} \\
\midrule
\code{DATA\_WIDTH} & 8 & Larghezza dei dati (INT8). \\
\rowa \code{ACC\_WIDTH} & 32 & Larghezza dell'accumulatore (INT32). \\
\code{N\_INPUTS} & 32 / 256 & Numero massimo di ingressi per neurone (baseline benchmark: 256). \\
\rowa \code{N\_NEURONS} & 1 / 4 & Numero massimo di neuroni per layer. \\
\code{PARALLEL} & 8 & MAC hardware simultanei per neurone; deve dividere \code{N\_INPUTS} e conviene sia potenza di due. \\
\rowa \code{N\_LAYERS} & 4 & Numero massimo di layer concatenabili da \code{layer\_sequencer}. \\
\code{ADDR\_WIDTH} & 23 & Larghezza dell'indirizzo di byte (8~MB). \\
\rowa \code{MEM\_DATA\_WIDTH} & 16 & Larghezza del bus dati fisico PSRAM. \\
\code{CLK\_FREQ\_MHZ} & 80 & Frequenza usata per le formule di temporizzazione PSRAM (va allineata all'oscillatore reale). \\
\bottomrule
\end{tabularx}
\begin{fnwarn}[Vincolo \texttt{N\_INPUTS} \% \texttt{PARALLEL}]
\code{PARALLEL} deve dividere esattamente \code{N\_INPUTS}, altrimenti scatta il guard
di elaborazione (§\ref{ch:datapath}). Lo stesso vincolo vale a runtime su
\code{n\_inputs\_real}.
\end{fnwarn}
\section{Larghezza di rete a runtime}
Un singolo bitstream serve qualunque topologia \emph{fino} al massimo di build. La
larghezza reale di ciascuna esecuzione è un valore separato, impostato dall'host:
\begin{itemize}
\item \code{n\_inputs\_real} --- ingressi realmente usati in questa esecuzione (deve
essere multiplo di \code{PARALLEL});
\item \code{n\_neurons\_real} --- neuroni realmente calcolati in questa esecuzione.
\end{itemize}
Entrambi hanno default pari al massimo di build, così ogni chiamante che li lascia
scollegati elabora l'intera larghezza come prima dell'introduzione delle porte.
\begin{fnnote}[Terminazione anticipata reale]
Non si tratta di semplice contabilità di indirizzi: i due valori limitano
direttamente i loop hardware (letture X/W di \code{neuron\_memory}, conteggio gruppi
MAC di \code{neuron\_parallel} e lunghezza della copia ping-pong per \code{RUN\_NETWORK}).
Un layer più stretto \emph{calcola} e \emph{copia} davvero più in fretta e non richiede
zero-padding della RAM per la coda non usata: i dati oltre
\code{n\_inputs\_real}/\code{n\_neurons\_real} non vengono mai letti.
\end{fnnote}
Questo permette a una rete di rastremarsi dentro una sola esecuzione concatenata, ad
esempio $256\to64\to16\to4$, con ogni layer che dichiara la propria larghezza reale
nella tabella descrittori (cap.~\ref{ch:seq}).
\subsection{Risparmio misurato}
La terminazione anticipata è stata misurata end-to-end:
\begin{tabularx}{\textwidth}{L{5.5cm} C{3.0cm} Y}
\toprule
\rowh \thd{Test} & \thd{Cicli} & \thd{Confronto} \\
\midrule
\code{neuron\_parallel\_tb.v} (T7) & 3 vs 6 & ridotto vs pieno, con dati ``spazzatura'' nelle corsie saltate (prova che non vengono lette). \\
\rowa \code{neuron\_memory\_tb.v} (T5) & 209 vs 788 & 8-di-32 vs 32 pieni, attraverso lo stack PSRAM reale. \\
\bottomrule
\end{tabularx}
\section{Configurazioni caratterizzate}
Alcune combinazioni convalidate in simulazione e/o sintesi:
\begin{tabularx}{\textwidth}{C{2.0cm} C{2.0cm} C{2.0cm} Y}
\toprule
\rowh \thd{N\_INPUTS} & \thd{N\_NEURONS} & \thd{PARALLEL} & \thd{Note} \\
\midrule
32 & 4 & 8 & Primo test parametrico funzionale (Fase~1). \\
\rowa 256 & 4 & 2/4/8/16 & Sweep di benchmark del datapath (Fase~7). \\
32 & 1..3 & 8 & Integrazione memoria mono/multi-neurone (Fase~3). \\
\rowa 4 & 4 & 2 & Test end-to-end \code{RUN\_NETWORK} a 2 layer su SPI reale. \\
\bottomrule
\end{tabularx}
\section{Riepilogo build contro runtime}
\begin{center}
\begin{tikzpicture}[font=\footnotesize,node distance=6mm]
\node[fnblockD,minimum width=54mm,minimum height=15mm](b){\textbf{BUILD (sintesi)}\\[2pt]
{\scriptsize N\_INPUTS, N\_NEURONS, N\_LAYERS,}\\{\scriptsize PARALLEL, DATA\_WIDTH, ACC\_WIDTH}\\{\scriptsize $\Rightarrow$ soffitto della macchina}};
\node[fnblockT,right=16mm of b,minimum width=54mm,minimum height=15mm](r){\textbf{RUNTIME (host, SPI)}\\[2pt]
{\scriptsize n\_inputs\_real, n\_neurons\_real,}\\{\scriptsize attivazione, num\_layers, pesi/bias}\\{\scriptsize $\Rightarrow$ rete effettiva}};
\draw[fnbus] (b) -- node[fnlbl,above]{$\le$} (r);
\end{tikzpicture}
\end{center}
@@ -1,193 +0,0 @@
\chapter{Sottosistema di memoria}
\label{ch:mem}
\section{Catena di memoria}
Il motore di calcolo lavora con indirizzi e dati a livello di \emph{byte} (INT8), mentre
la PSRAM è un dispositivo a parola da 16~bit. Tre moduli in cascata realizzano la
conversione e l'accesso fisico:
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=8mm]
\node[fnblockD,minimum width=30mm,minimum height=12mm](nm){master byte-level\\{\scriptsize \code{neuron\_memory} / \code{spi\_engine} / \code{layer\_sequencer}}};
\node[fnblockT,right=10mm of nm,minimum width=28mm,minimum height=12mm](ia){\code{int8\_memory\_access}\\{\scriptsize byte $\leftrightarrow$ word 16-bit}};
\node[fnblock,right=10mm of ia,minimum width=26mm,minimum height=12mm](mi){\code{memory\_interface}\\{\scriptsize FSM IDLE/WAIT}};
\node[fnblockA,below=9mm of mi,minimum width=26mm,minimum height=12mm](pc){\code{psram\_controller}\\{\scriptsize bus fisico async 70\,ns}};
\node[fnblock,left=10mm of pc,minimum width=26mm,minimum height=12mm](ps){PSRAM\\{\scriptsize 8\,MB 4M$\times$16}};
\draw[fnbus] (nm)--node[fnlbl,above]{req/wr/addr}(ia);
\draw[fnbus] (ia)--node[fnlbl,above]{16-bit}(mi);
\draw[fnbus] (mi)--(pc);
\draw[fnbus] (pc)--node[fnlbl,above]{DQ/A/ctrl}(ps);
\end{tikzpicture}
\end{center}
\section{\texttt{int8\_memory\_access} --- conversione byte/word}
Converte l'interfaccia INT8 (indirizzo di byte) nell'interfaccia a parola. L'indirizzo
di byte viene diviso per due (\code{addr>>1}) per ottenere l'indirizzo di parola; il
bit meno significativo seleziona il byte:
\begin{itemize}
\item \code{addr[0]=0} $\to$ byte basso: \code{lb\_n=0}, \code{ub\_n=1}, dato su DQ[7:0];
\item \code{addr[0]=1} $\to$ byte alto: \code{lb\_n=1}, \code{ub\_n=0}, dato su DQ[15:8].
\end{itemize}
In lettura estrae il byte corretto da \code{mem\_rdata}. La FSM ha due stati (IDLE,
WAIT) e restituisce \code{ready} come impulso di un ciclo.
\section{\texttt{memory\_interface} --- handshake}
FSM a due stati che serializza una singola transazione: in IDLE, alla richiesta
\code{req}, latcha \code{wr/addr/wdata/lb\_n/ub\_n} ed emette un impulso \code{mem\_req}
di un ciclo verso il controller; in WAIT attende \code{mem\_ready}, cattura
\code{rdata} in lettura e asserisce \code{ready}. Garantisce il contratto ``una
transazione per volta''.
\section{\texttt{psram\_controller} --- bus fisico}
Controller del bus PSRAM parallelo asincrono, con supporto al \textbf{page mode}
di lettura del chip (\S~\ref{sec:pagemode}). La macchina a stati principale è:
\begin{center}
\begin{tikzpicture}[font=\scriptsize]
\node[fnstate](init) at (0,0){INIT};
\node[fnstate](idle) at (3.2,0){IDLE};
\node[fnstate](read) at (7,2.7){READ};
\node[fnstate](popen) at (11,2.7){PAGE\\OPEN};
\node[fnstate](write) at (7,-2.7){WRITE};
\node[fnstate](ww) at (11,-2.7){WRITE\\WAIT};
\draw[fnarrow] (init)--node[fnlbl,above]{INIT\_CYCLES + CR load}(idle);
\draw[fnarrow] (idle)--node[fnlbl,above,sloped]{req \& !wr}(read);
\draw[fnarrow] (idle)--node[fnlbl,below,sloped]{req \& wr}(write);
\draw[fnarrow] (read)--node[fnlbl,above]{ready}(popen);
\draw[fnarrowT] (popen) to[bend left=25] node[fnlbl,below]{req \& !wr}(read);
\draw[fnarrow] (popen) to[bend right=20] node[fnlbl,above,sloped]{req \& wr}(write);
\draw[fnarrow] (popen) to[out=-100,in=15,looseness=1.15] node[fnlbl,pos=0.55]{timeout tCEM}(idle);
\draw[fnarrow] (write)--node[fnlbl,above]{ACCESS\_CYCLES}(ww);
\draw[fnarrow] (ww) to[out=160,in=-70] node[fnlbl,pos=0.5,left]{ready}(idle);
\end{tikzpicture}
\end{center}
Da INIT il controller passa automaticamente per una sotto-sequenza di caricamento del
registro di configurazione (\code{STATE\_CR\_INIT}, 4 passi) prima di raggiungere IDLE
per la prima volta --- vedi \S~\ref{sec:pagemode}. La transizione PAGE~OPEN
$\to$~WRITE (freccia in basso a destra) passa internamente per due micro-stati di
transito, \code{STATE\_PAGE\_CLOSE} e \code{STATE\_PAGE\_REOPEN} (un ciclo ciascuno):
il primo forza CE\#/OE\# alti per almeno un ciclo prima che il controller inizi a
pilotare il bus dati, evitando contesa con l'uscita ancora attiva della PSRAM
($\geq t_{HZ}$); il secondo riavvia la transazione già latchata esattamente come
farebbe IDLE. Non sono disegnati come nodi separati per non appesantire la figura.
\subsection{Temporizzazione}
\begin{fnspec}[Formule di temporizzazione]
$\text{ACCESS\_CYCLES}=\lceil (70\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(latenza di accesso casuale, $t_{AA}$/$t_{RC}$ = 70~ns)\\[3pt]
$\text{PAGE\_CYCLES}=\lceil (20\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(continuazione nella stessa pagina, $t_{APA}$/$t_{PC}$ = 20~ns)\\[3pt]
$\text{INIT\_CYCLES}=150\times \text{CLK\_FREQ\_MHZ}$ \quad
(inizializzazione di power-up, $t_{PU}$ = 150~\textmu s)\\[3pt]
$\text{PAGE\_TIMEOUT\_CYCLES}=\lceil (6000\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(chiusura automatica della pagina, margine di sicurezza sotto $t_{CEM}$ = 8~\textmu s)
\end{fnspec}
Il bus dati è pilotato in tri-state: \code{psram\_dq = dq\_oe ? dq\_out : Z}. In lettura
\code{dq\_oe=0}; in scrittura \code{dq\_oe=1} durante l'impulso di \code{we\_n}. Uno
stato di WRITE\_WAIT mantiene attivi \code{ce\_n/lb\_n/ub\_n} per l'hold finale prima
del rilascio.
\begin{fnwarn}[Non è QSPI]
Questa è un'interfaccia SRAM-asincrona classica, \textbf{non} QSPI: la maggior parte
delle ``PSRAM'' serie/QSPI in commercio non è compatibile con questo controller senza
riscrittura. Vedere il cap.~\ref{ch:hw} per la parte raccomandata (ISSI parallela).
\end{fnwarn}
\section{Page mode di lettura}
\label{sec:pagemode}
Il chip raccomandato (cap.~\ref{ch:hw}) è ``asynchronous/\textbf{page mode}'': una
volta fatto un primo accesso casuale a $t_{AA}$~=~70~ns, letture successive
all'interno della stessa pagina da 16 word (bit di indirizzo sopra \code{A[3]}
invariati) costano solo $t_{APA}$/$t_{PC}$~=~20~ns, perché CE\#/OE\# restano attivi
e cambia solo il bus indirizzi. Il page mode è \textbf{disabilitato di default}
all'accensione (bit~7 del registro di configurazione, CR~=~\texttt{0x0070} di
default) e va abilitato esplicitamente.
\begin{itemize}
\item \textbf{Abilitazione all'avvio}: subito dopo INIT, il controller esegue la
``software-access sequence'' del datasheet (2 letture dummy + 2 scritture,
\texttt{0x0000} di sblocco poi CR reale \texttt{0x00F0} = default con il bit
Page attivo) all'indirizzo più alto del chip --- riusa esattamente la stessa
logica READ/WRITE di ogni altra transazione, quindi passa dagli stessi controlli
di temporizzazione.
\item \textbf{Burst di pagina}: dopo una READ il controller non chiude più
CE\#/OE\# (stato PAGE~OPEN). Una READ successiva nella stessa pagina aspetta solo
PAGE\_CYCLES; una READ che attraversa pagina resta comunque senza toggle di CE\#
ma paga un ACCESS\_CYCLES pieno per quella parola (qualunque cambio a
\code{A[4]} o superiore richiede un nuovo $t_{AA}$). Un contatore chiude la
pagina prima del limite $t_{CEM}$ con margine di sicurezza.
\item \textbf{Solo una WRITE chiude la pagina.} I cambi di \code{lb\_n}/\code{ub\_n}
\emph{non} la chiudono: \code{int8\_memory\_access} alterna questi segnali a quasi
ogni accesso (accesso byte-granulare su bus a 16~bit), quindi trattarli come
condizione di chiusura --- primo tentativo di implementazione --- rendeva il
workload reale \emph{più lento}, non più veloce (misurato: 53.25$\to$61.25
cicli/edge sul gather di \code{graph\_engine}); rimosso, corretto a
53.25$\to$37.53 cicli/edge (banda +42\%, \S~\ref{sec:bandwidth}).
\end{itemize}
\begin{fnwarn}[Nessun beneficio senza pattern sequenziale]
Il page mode accelera solo accessi che restano nella stessa pagina (o quasi) mentre
il controller resta in attesa di una nuova richiesta con la pagina ancora aperta.
Accessi isolati e sparsi (indirizzo casuale ogni volta) pagano comunque
ACCESS\_CYCLES pieno, più un piccolo overhead di chiusura/riapertura se preceduti
da una WRITE o da un timeout $t_{CEM}$: non è un guadagno universale, dipende dal
pattern di accesso del chiamante.
\end{fnwarn}
Fmax reale (\code{nextpnr-ecp5}, cap.~\ref{ch:impl}) sul sistema integrato
\code{spi\_neuron\_top} con Tipo~\#2 abilitato: \textbf{75.73~MHz} a
\code{PARALLEL}=2 (era 55.59~MHz prima dell'aggiunta del page mode) e
\textbf{65.13~MHz} a \code{PARALLEL}=8, entrambe ancora FAIL all'obiettivo di
80~MHz ma non regredite. Il percorso critico resta, in entrambi i casi,
interamente dentro \code{u\_graph\_engine.u\_neuron} (catena di accumulo
\code{mac8}/\code{neuron\_parallel}, cap.~\ref{ch:impl}) --- \code{psram\_controller}
non compare mai nel percorso critico nonostante la crescita di risorse del page
mode.
\section{Mappa degli indirizzi e convenzioni}
Lo spazio di indirizzamento è di \code{ADDR\_WIDTH}=23~bit (indirizzo di \emph{byte}),
per 8~MB pieni. Le regioni non hanno indirizzi cablati: le loro basi sono registri
impostati dall'host via \op{SET\_BASE} (percorso single-layer) o lette dalla tabella
descrittori (percorso multi-layer).
\begin{tabularx}{\textwidth}{L{3.2cm} L{3.4cm} Y}
\toprule
\rowh \thd{Regione} & \thd{Base} & \thd{Contenuto / convenzione} \\
\midrule
Ingresso $X$ & \code{x\_base} & Vettore di ingresso condiviso, letto una volta per invocazione. \\
\rowa Pesi $W$ & \code{w\_base} & Neuron-major: i pesi del neurone $n$ a \code{w\_base + n*N\_INPUTS} byte. \\
Bias & \code{bias\_addr} & Un byte per neurone: bias del neurone $n$ a \code{bias\_addr + n}. \\
\rowa Tabella descrittori & \code{table\_base} & \code{N\_LAYERS} voci da 11 byte (cap.~\ref{ch:seq}). \\
Buffer ping-pong A/B & \code{buf\_a\_base} / \code{buf\_b\_base} & Uscite intermedie tra layer. \\
\bottomrule
\end{tabularx}
\subsection{Indirizzamento fisico della PSRAM}
La PSRAM raccomandata è 4M$\times$16 (8~MB), che richiede un indirizzo di parola a
22~bit (A0--A21). \code{int8\_memory\_access} calcola \code{addr>>1} portando l'indirizzo
di byte a 23~bit in un indirizzo di parola a 22~bit che mappa esattamente su A0--A21; il
bit~22 di \code{psram\_a} è quindi sempre 0 e sul PCB restano 22 linee di indirizzo
reali.
\section{Larghezza di banda}
\label{sec:bandwidth}
Misurata sul gather della lista di edge di \code{graph\_engine} (cap.~\ref{ch:grafo}),
per differenza tra due dimensioni di grafo per isolare il costo per-edge dall'overhead
fisso per-neurone (\code{sim/graph\_engine\_bandwidth\_tb.v}):
\begin{tabularx}{\textwidth}{L{5.2cm} Y Y Y}
\toprule
\rowh \thd{} & \thd{Prima (no page mode)} & \thd{Dopo (page mode)} & \thd{$\Delta$} \\
\midrule
Cicli/edge & 53.25 & 37.53 & $-29.5\%$ \\
\rowa Banda @80\,MHz & 6.01\,MB/s & 8.53\,MB/s & $+41.9\%$ \\
Banda @16\,MHz\textsuperscript{*} & 1.20\,MB/s & 1.71\,MB/s & $+41.9\%$ \\
\bottomrule
\end{tabularx}
\textsuperscript{*}oscillatore reale raccomandato (cap.~\ref{ch:hw}).
Il modello resta comunque memory-bound per costruzione: \code{neuron\_memory} legge
$X$ una volta e rilegge $W$/bias per ciascun neurone (cap.~\ref{ch:seq}), un neurone
per volta; il page mode riduce il costo per-byte dell'accesso sequenziale, non elimina
il pattern di accesso stesso.
@@ -1,111 +0,0 @@
\chapter[Memoria, multi-neurone e multi-layer]{Integrazione memoria, multi-neurone e multi-layer}
\label{ch:seq}
\section{\texttt{neuron\_memory} --- ponte memoria/neurone}
\code{neuron\_memory} collega il datapath di calcolo alla memoria e gestisce il loop sui
neuroni. Legge il vettore $X$ una sola volta (ingresso condiviso), poi per ciascun
neurone rilegge $W$ e bias dalla RAM e li invia a una singola istanza riusata di
\code{neuron\_parallel}: il progetto è memory-bound, un neurone calcolato per volta,
senza duplicare il datapath. L'uscita è \code{y\_bus}, packed neuron-major
(\code{DATA\_WIDTH*N\_NEURONS} bit).
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=13mm]
\node[fnstate](idle){IDLE};
\node[fnstate,right=of idle](rx){READ\_X};
\node[fnstate,right=of rx](rw){READ\_W};
\node[fnstate,below=10mm of rw](rb){READ\_BIAS};
\node[fnstate,left=of rb](sn){START\_N};
\node[fnstate,left=of sn](wn){WAIT\_N};
\draw[fnarrow] (idle)--node[fnlbl,above]{start}(rx);
\draw[fnarrow] (rx)--node[fnlbl,above]{X letto}(rw);
\draw[fnarrow] (rw)--(rb);
\draw[fnarrow] (rb)--(sn);
\draw[fnarrow] (sn)--(wn);
\draw[fnarrow] (wn) to[bend left=18] node[fnlbl,above]{neurone succ.}(rw);
\draw[fnarrow] (wn) to[bend right=28] node[fnlbl,below]{ultimo neurone: done}(idle);
\end{tikzpicture}
\end{center}
Gli stati sono IDLE, READ\_X, READ\_W, READ\_BIAS, START\_N, WAIT\_N. Dopo l'ultimo
neurone la FSM torna in IDLE e asserisce \code{done}. Il conteggio di neuroni e ingressi
realmente elaborati è dato da \code{n\_neurons\_real}/\code{n\_inputs\_real}
(cap.~\ref{ch:param}).
\section{\texttt{layer\_sequencer} --- rete multi-layer}
\code{layer\_sequencer} concatena fino a \code{N\_LAYERS} esecuzioni della stessa
istanza \code{neuron\_memory}, realizzando una rete densa feed-forward \emph{senza}
toccare il core di calcolo validato. Legge una tabella descrittori scritta dall'host e
alterna i due buffer di uscita in RAM (ping-pong).
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=13mm]
\node[fnstate](i){IDLE};
\node[fnstate,right=of i](rd){READ\\DESC};
\node[fnstate,right=of rd](rw){READ\\WAIT};
\node[fnstate,below=10mm of rw](sl){START\\LAYER};
\node[fnstate,left=of sl](wl){WAIT\\LAYER};
\node[fnstate,left=of wl](ci){COPY\\ISSUE};
\node[fnstate,below=9mm of ci](cw){COPY\\WAIT};
\draw[fnarrow] (i)--node[fnlbl,above]{run\_start}(rd);
\draw[fnarrow] (rd)--(rw);
\draw[fnarrow] (rw)--(sl);
\draw[fnarrow] (sl)--(wl);
\draw[fnarrow] (wl)--(ci);
\draw[fnarrow] (ci)--(cw);
\draw[fnarrow] (cw) to[bend left=15] node[fnlbl,left]{layer succ.}(rd);
\draw[fnarrow] (cw) to[bend right=12] node[fnlbl,below]{ultimo: seq\_done}(i);
\end{tikzpicture}
\end{center}
\subsection{Buffer ping-pong}
Il layer~0 legge l'ingresso esterno \code{x\_base}. Il layer $k>0$ legge dal buffer
scritto dal layer $k-1$; l'uscita di ciascun layer viene copiata nell'altro buffer,
alternando A e B. L'uscita finale resta sia in \code{y\_bus} (leggibile con
\op{READ\_OUTPUT}) sia nel buffer ping-pong su cui è stata copiata.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=7mm]
\node[fnblockA,minimum width=18mm](x){X\\\code{x\_base}};
\node[fnblockD,right=10mm of x,minimum width=20mm](l0){Layer 0};
\node[fnblock,right=10mm of l0,minimum width=18mm](ba){buf A};
\node[fnblockD,right=10mm of ba,minimum width=20mm](l1){Layer 1};
\node[fnblock,right=10mm of l1,minimum width=18mm](bb){buf B};
\node[fnblockD,right=10mm of bb,minimum width=20mm](l2){Layer 2};
\draw[fnarrow] (x)--(l0); \draw[fnarrow] (l0)--(ba);
\draw[fnarrow] (ba)--(l1); \draw[fnarrow] (l1)--(bb);
\draw[fnarrow] (bb)--(l2);
\draw[fnarrowT,dashed] (l2.south) to[bend left=25] node[fnlbl,below]{copia in buf A} (ba.south);
\end{tikzpicture}
\end{center}
\subsection{Tabella descrittori}
Scritta dall'host in RAM a \code{table\_base} con \op{WRITE\_RAM}; \code{N\_LAYERS} voci
da 11 byte ciascuna, MSB-first:
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.6cm} Y}
\toprule
\rowh \thd{Campo} & \thd{Byte} & \thd{Significato} \\
\midrule
\code{w\_base} & 3 & Base dei pesi del layer. \\
\rowa \code{bias\_addr} & 3 & Base dei bias del layer. \\
\code{activation} & 1 & Attivazione del layer (2 bit bassi, cfr. \code{ACT\_*}). \\
\rowa \code{n\_inputs\_real} & 2 & Ingressi reali del layer (multiplo di \code{PARALLEL}). \\
\code{n\_neurons\_real} & 2 & Neuroni reali del layer. \\
\midrule
\rowh \thd{Totale} & \thd{11} & per voce/layer \\
\bottomrule
\end{tabularx}
\begin{fnnote}[Copia proporzionale alla larghezza reale]
Il sequencer copia esattamente \code{n\_neurons\_real} byte di \code{y\_bus} nel buffer
ping-pong (non l'intera larghezza di build): un layer più stretto viene copiato più in
fretta, senza zero-padding in RAM. Ogni attivazione è letta per-layer dalla tabella,
indipendente dal registro \code{activation} del percorso single-layer.
\end{fnnote}
\section{Gerarchia dei segnali \texttt{busy}/\texttt{done}}
Nel percorso multi-layer, \code{STATUS.busy} è l'OR dei busy single-layer e sequencer,
mentre \code{STATUS.done} latcha solo al completamento dell'\emph{ultimo} layer, non a
ogni layer intermedio (cap.~\ref{ch:spi}). Il top-level restituisce il controllo di
\code{neuron\_memory} al percorso diretto \op{START} al termine della sequenza.
-191
View File
@@ -1,191 +0,0 @@
\chapter[Rete a grafo (Tipo \#2)]{Configurazione a due livelli: rete a grafo (Tipo \#2)}
\label{ch:grafo}
\section{Due tipi di rete}
L'engine espone due \emph{tipi di rete} selezionabili dall'host, con lo stesso comando di
avvio che instrada verso il motore corretto:
\begin{itemize}
\item \textbf{Tipo \#1 --- rete classica (dense).} Layer con neuroni per layer, fully
connected tra layer consecutivi. È il percorso di \code{layer\_sequencer}
(cap.~\ref{ch:seq}), avviato da \op{RUN\_NETWORK}. Connessioni \emph{implicite per
posizione}: non si enumera nulla, si definiscono solo i pesi indirizzati come
\code{w\_base + k*n\_inputs + j}.
\item \textbf{Tipo \#2 --- grafo arbitrario (sparse).} A partire dagli id dei neuroni di
ingresso si definiscono le connessioni di ogni neurone fino all'uscita, tramite una
\emph{edge-list sparsa} per-neurone. Connessioni \emph{esplicite per enumerazione}: ogni
connessione è un edge \code{(src\_id, peso)}; se non è nella lista, non esiste.
\end{itemize}
\begin{fnnote}[La differenza in una riga]
Dense: definisci i \emph{pesi} per posizione in una matrice. Graph: definisci ogni
\emph{connessione} come edge \code{(src\_id, peso)} in una lista per-neurone. Le due
tabelle descrittori hanno lo stesso formato di 11~byte ma campi diversi; il registro
\code{net\_type} dice al motore quale interpretazione usare.
\end{fnnote}
\section{Buffer di attivazione globale}
Il Tipo \#2 introduce un \textbf{buffer di attivazione} indicizzato per \emph{id di
segnale}, un byte INT8 per id, realizzato in \textbf{block RAM on-chip \code{DP16KD}}
(\code{rtl/act\_buffer.v}). Gli id \code{0..N\_in-1} sono gli ingressi; ogni neurone
scrive la propria uscita nel proprio id. Il gather delle sorgenti legge da qui a
\emph{accesso random a un ciclo}: è ciò che rende economico il grafo, perché è l'accesso
che la PSRAM (70~ns, sequenziale) non potrebbe accelerare.
\begin{fnspec}[Dimensionamento V1]
\code{N\_TOTAL}=4096 segnali, id a 16~bit (spazio fino a 65.536 senza cambiare formato).
Buffer = 4~KB, cioè 2 blocchi \code{DP16KD} su 108. Il vincolo reale diventa la capacità
PSRAM per gli edge ($\approx$2\,M edge a 4~B), non la block RAM.
\end{fnspec}
\section{DAG feed-forward e vincolo \texttt{src\_id < out\_id}}
Il grafo è un DAG feed-forward: ogni connessione punta a un id \textbf{già calcolato}
(\code{src\_id < out\_id}). I neuroni si elaborano in ordine di id crescente, così quando
si calcola un neurone tutte le sue sorgenti sono pronte nel buffer. Cicli e ricorrenza
sono fuori scope per la V1. Il vincolo è verificato a due livelli: dall'assemblatore host
(a compile time) e da un guard a runtime in \code{graph\_engine} (\code{STATUS.err}),
nella stessa filosofia del guard di elaborazione su \code{N\_INPUTS \% PARALLEL}.
\section{Formati dati}
Entrambi i descrittori sono da 11~byte/voce, MSB-first, a \code{table\_base}.
\subsection{Descrittore Tipo \#2 (grafo)}
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.6cm} Y}
\toprule
\rowh \thd{Campo} & \thd{Byte} & \thd{Significato} \\
\midrule
\code{conn\_ptr} & 3 & Indirizzo byte in PSRAM del blocco edge del neurone. \\
\rowa \code{n\_conn} & 2 & Connessioni reali (pre-padding). \\
\code{out\_id} & 2 & Id in cui scrivere l'uscita del neurone. \\
\rowa \code{activation} & 1 & \code{ACT\_RELU} / \code{ACT\_NONE} (2 bit bassi). \\
\code{bias} & 1 & Bias del neurone (INT8). \\
\rowa \code{reserved} & 2 & 0. \\
\midrule
\rowh \thd{Totale} & \thd{11} & voci in ordine di \code{out\_id} crescente \\
\bottomrule
\end{tabularx}
\subsection{Edge del grafo (4~byte, allineato)}
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.6cm} Y}
\toprule
\rowh \thd{Campo} & \thd{Byte} & \thd{Significato} \\
\midrule
\code{src\_id} & 2 & Id sorgente (uint16 BE). \\
\rowa \code{weight} & 1 & Peso (INT8). \\
\code{reserved} & 1 & 0 (allineamento a 4~byte). \\
\bottomrule
\end{tabularx}
\begin{fnnote}[Padding a \texttt{PARALLEL}]
\code{n\_conn} arbitrario non è multiplo di \code{PARALLEL}: la edge-list del neurone è
riempita fino al multiplo con edge a \textbf{peso zero} (spreco $\le$\code{PARALLEL}$-1$
per neurone). Così il datapath e il suo guard restano intatti.
\end{fnnote}
\section{\texttt{graph\_engine} --- motore del grafo}
\code{rtl/graph\_engine.v} orchestra il Tipo \#2 \textbf{riusando \code{neuron\_parallel}
senza modificarlo}, come fa \code{neuron\_memory} per il caso denso. Differenza chiave: tra
i due modi cambia \emph{solo l'indirizzamento di X}. In Tipo \#1 l'input è contiguo
(\code{x\_base + i}); in Tipo \#2 è un gather (\code{act\_buf[src\_id]}). Il core aritmetico
non si tocca.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=4mm,start chain=going below,
every node/.style={on chain,fnblock,minimum width=52mm}]
\node[fnblockA]{\code{COPY\_INPUTS}: PSRAM \code{x\_base} $\to$ \code{act\_buf[0..N\_in-1]}};
\node{\code{READ\_DESC}: descrittore del neurone k};
\node{\code{READ\_EDGES}: stream edge + gather \code{act\_buf[src\_id]}};
\node{\code{START\_N} / \code{WAIT\_N}: gruppo da \code{PARALLEL} $\to$ \code{neuron\_parallel}};
\node[fnblockT]{\code{WRITE\_ACT}: y $\to$ \code{act\_buf[out\_id]}};
\node{neurone successivo (ordine di id)};
\node[fnblockD]{\code{WRITE\_OUTPUTS}: ultimi \code{n\_out} $\to$ PSRAM \code{out\_base}};
\foreach \i [count=\j from 2] in {1,...,6}
\draw[fnarrow] (chain-\i) -- (chain-\j);
\end{tikzpicture}
\end{center}
Le uscite sono gli \textbf{ultimi \code{n\_out}} id: nei DAG con l'ordinamento
\code{src\_id < out\_id} i neuroni di uscita (sink, non riusati come sorgente) finiscono
naturalmente con gli id più alti. A fine esecuzione \code{graph\_engine} copia questi
\code{n\_out} byte in una regione PSRAM a \code{out\_base}, che l'host rilegge con
\op{READ\_RAM}.
\section{Opcode e registri del Tipo \#2}
La selezione del tipo avviene con un nuovo opcode; \op{RUN\_NETWORK} fa il dispatch sul
registro \code{net\_type} (dettagli in cap.~\ref{ch:spi}).
\begin{tabularx}{\textwidth}{L{2.6cm} L{3.4cm} Y}
\toprule
\rowh \thd{Opcode / sel} & \thd{Nome} & \thd{Funzione} \\
\midrule
\op{0x11} & SET\_NET\_TYPE & \code{type(1B)}: \code{0x01}=dense (\#1), \code{0x02}=graph (\#2). Default dopo \op{RESET}=dense. \\
\rowa \code{SET\_BASE sel 9} & num\_neurons\_graph & Numero di neuroni del grafo (uint16). \\
\code{SET\_BASE sel 10} & n\_out & Numero di id di uscita (uint16). \\
\bottomrule
\end{tabularx}
\begin{fnnote}[Zero regressioni sul Tipo \#1]
Con \code{net\_type=dense} (valore di default dopo \op{RESET}) il percorso \#1 è
bit-identico a prima: \op{RUN\_NETWORK} mantiene il payload \code{num\_layers(1B)} e il
framing degli opcode esistenti non cambia.
\end{fnnote}
\section{Occupazione (Tipo \#2 abilitato)}
Sintesi Yosys del sistema completo \code{spi\_neuron\_top} con Tipo \#2 abilitato
(\code{PARALLEL}=2):
\begin{tabularx}{\textwidth}{L{4.6cm} Y}
\toprule
\rowh \thd{Risorsa} & \thd{Uso} \\
\midrule
\code{DP16KD} (block RAM) & 2 (buffer di attivazione) \\
\rowa \code{MULT18X18D} (DSP) & 4 (2 \code{neuron\_memory} + 2 \code{graph\_engine}) \\
LUT4 & 2619 \\
\rowa TRELLIS\_FF & 2467 \\
\code{\$\_TBUF\_} (bus PSRAM) & 16 \\
\bottomrule
\end{tabularx}
Il device (108 \code{DP16KD}, 72 DSP, $\approx$44k LUT/FF) resta ampiamente sotto la
saturazione: il Tipo \#2 aggiunge una modalità completa a costo di risorse contenuto.
LUT4/TRELLIS\_FF sono cresciuti rispetto a una misura precedente (2367/2406) per via del
page mode PSRAM aggiunto al controller (cap.~\ref{ch:mem}, \S~5.5) --- sotto il 6\% di
utilizzo, nessun impatto pratico.
\section{Banda del gather (misurata)}
Il costo per-edge del gather è stato \textbf{isolato} costruendo due grafi identici per
struttura ma con conteggio edge diverso e differenziando i cicli: la sottrazione cancella
l'overhead fisso per-neurone e lascia il solo costo dell'edge.
\begin{fnspec}[Costo per-edge]
\textbf{37.53 cicli/edge} con il page mode PSRAM abilitato (cap.~\ref{ch:mem},
\S~5.5) --- \textbf{53.25 cicli/edge} senza (baseline pre-page-mode, coerente con la
teoria: 4~byte/edge $\times$ $\approx$13 cicli/byte via PSRAM asincrona
$\approx$52). A 80~MHz: $\approx$2.13\,M edge/s ($\approx$8.5~MB/s, +42\% vs
baseline); al clock reale di 16~MHz: $\approx$426\,k edge/s ($\approx$1.71~MB/s).
\end{fnspec}
Il page-mode read (roadmap G7, cap.~\ref{ch:roadmap}) è stato implementato e misurato:
l'accesso sequenziale del gather ne beneficia direttamente, riducendo il costo per-edge
del 29.5\% (53.25$\to$37.53 cicli/edge). Ogni edge continua comunque a pagare l'accesso
byte-granulare di \code{int8\_memory\_access} (4 byte/edge); il page mode riduce il costo
di ciascun byte sequenziale, non il numero di accessi.
\section{Assemblatore host \texttt{netasm}}
La configurazione leggibile della rete non richiede logica dedicata in FPGA: uno
pseudo-assembly viene compilato \emph{sull'host} (\code{tools/netasm/}) nei byte esatti
delle tabelle e degli edge, poi caricati con \op{WRITE\_RAM}. L'assemblatore valida a
compile time (\code{src\_id < out\_id}, limiti \code{N\_TOTAL}, padding a \code{PARALLEL}),
complementando il guard runtime.
\begin{lstlisting}[language=,caption={Esempio di pseudo-assembly (grafo)},basicstyle=\ttfamily\scriptsize]
NET graph
INPUTS 4 ; id 0..3
NEURON n4 relu bias=2
CONN 0 w=5
CONN 1 w=-3
NEURON n5 none bias=0
CONN n4 w=2 ; riferimento simbolico all'uscita di n4
CONN 2 w=7
OUTPUT n5
END
\end{lstlisting}
-297
View File
@@ -1,297 +0,0 @@
\chapter{Interfaccia host SPI}
\label{ch:spi}
\section{Livello fisico}
L'FPGA è sempre \textbf{slave} SPI. Il protocollo v1 usa SPI \textbf{Mode~0}
(CPOL=0, CPHA=0), MSB-first, single-SPI. Un comando per periodo di CS basso; il byte~0
di ogni transazione è l'opcode. I campi multi-byte sono big-endian.
\begin{fnspec}[Campionamento Mode 0]
\code{mosi} è campionato sul fronte di \textbf{salita} di \code{sclk}; \code{miso} è
pilotato sul fronte di \textbf{discesa} (stabile prima del successivo campionamento del
master). \code{spi\_slave} sincronizza \code{sclk/mosi/cs\_n} con un doppio flip-flop
(CDC a 3 stadi) prima di ogni rilevazione di fronte.
\end{fnspec}
\begin{center}
\begin{tikztimingtable}[timing/dslope=0.1,timing/.style={x=3.4ex,y=2.2ex},
xscale=1.0,font=\scriptsize]
\sig{CS\_N} & H 1L 16L 1H \\
\sig{SCLK} & L 1L {2C(2)}8{2C(2)} 6L \\
\sig{MOSI} & U 1U 2D{b7} 2D{b6} 2D{b5} 2D{b4} 2D{b3} 2D{b2} 2D{b1} 2D{b0} 2U \\
\sig{MISO} & Z 1Z 16D{dato} 1Z \\
\end{tikztimingtable}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
Framing di un byte: CS scende, 8 colpi di SCLK, MSB per primo; MISO in tri-state fuori
transazione.\end{center}
\begin{fnnote}[Contratto \texttt{tx\_byte\_req}]
\code{tx\_byte\_req} è un \emph{prefetch hint}, non un evento ``byte consumato'': un
consumatore deve avanzare i puntatori (indirizzo RAM, indice byte di risposta) su
\code{rx\_valid}, che pulsa esattamente una volta per byte reale trasferito.
\end{fnnote}
\section{Framing e lunghezza esplicita}
La lunghezza dei trasferimenti RAM è \textbf{esplicita}, non delimitata dal fronte di
CS: \op{WRITE\_RAM}/\op{READ\_RAM} portano un campo lunghezza a 2~byte, così il
controller SPI necessita solo di un contatore di byte. Gli indirizzi di byte sono a
23~bit, trasportati in un campo di 3~byte con il bit più alto riservato a 0.
\section{Tabella degli opcode}
\renewcommand{\arraystretch}{1.16}
\begin{longtable}{C{1.1cm} L{2.4cm} L{3.9cm} L{2.4cm} L{4.0cm}}
\toprule
\rowh \thd{Op} & \thd{Nome} & \thd{Payload (host$\to$FPGA)} & \thd{Risposta} & \thd{Funzione} \\
\midrule
\endfirsthead
\rowh \thd{Op} & \thd{Nome} & \thd{Payload} & \thd{Risposta} & \thd{Funzione} \\ \midrule
\endhead
\bottomrule
\endfoot
\op{0x00} & NOP & --- & --- & Nessuna operazione (idle/dummy clocking). \\
\rowa \op{0x01} & WRITE\_RAM & addr(3B)+len(2B)+dati & --- & Scrive un blocco in PSRAM (X, pesi, bias, parametri). \\
\op{0x02} & READ\_RAM & addr(3B)+len(2B) & \code{len} byte & Rilegge un blocco da PSRAM. \\
\rowa \op{0x0F} & RESET & --- & --- & Reset sincrono del motore e azzeramento del latch STATUS; non cancella la PSRAM. \\
\op{0x10} & SET\_BASE & sel(1B)+addr(3B) & --- & Imposta le basi/registri (vedi §\ref{sec:setbase}). \\
\rowa \op{0x11} & SET\_NET\_TYPE & type(1B) & --- & Tipo di rete: \code{0x01}=dense (\#1), \code{0x02}=graph (\#2). Default dopo RESET=dense. \\
\rowa \op{0x20} & START & --- & --- & Avvia \code{neuron\_memory} (percorso single-layer); ignorato se busy. \\
\op{0x21} & STATUS & --- & 1 byte & bit0=\code{busy} (live), bit1=\code{done} (sticky, clear-on-read), bit2=\code{err} (guard grafo), bit3=\code{flash\_err} (sticky, clear-on-read), bit4=\code{flash\_busy} (live); bit7:5=0. \\
\rowa \op{0x22} & READ\_OUTPUT & --- & \code{N\_NEURONS} byte & \code{y\_bus} neuron-major (byte~0 = neurone~0); solo percorso dense (Tipo \#1). \\
\op{0x23} & RUN\_NETWORK & num\_layers(1B) & --- & Avvia l'esecuzione: dispatch su \code{net\_type} verso \code{layer\_sequencer} (\#1) o \code{graph\_engine} (\#2); ignorato se busy. \\
\rowa \op{0x30} & READ\_CONFIG & --- & 11 byte & Record di configurazione hardware (§\ref{sec:readcfg}). \\
\op{0x40} & FLASH\_READ\_BLOCK & flash\_addr(3B)+psram\_addr(3B)+len(3B) & --- & Lettura raw flash$\to$PSRAM, bypassa il catalogo. \\
\rowa \op{0x41} & FLASH\_WRITE\_BLOCK & psram\_addr(3B)+flash\_addr(3B)+len(3B) & --- & Scrittura raw PSRAM$\to$flash (erase-before-write interno + loop Page Program $\leq$256B + poll WIP, trasparente all'host), bypassa il catalogo. \\
\op{0x42} & FLASH\_ERASE & sector\_addr(3B) & --- & Erase di un settore da 4~KB (deve essere sector-aligned), bypassa il catalogo. \\
\rowa \op{0x43} & CAT\_READ & --- & --- & Ricarica il catalogo a 16 slot (registri on-chip) dal settore riservato in flash. \\
\op{0x44} & CAT\_WRITE\_SLOT & slot\_id(1B)+offset(3B)+len(3B)+tipo(1B) & --- & Registra/aggiorna (offset, lunghezza, tipo) dello slot nel catalogo on-chip e lo persiste in flash; marca lo slot \emph{non valido} finché \op{SAVE\_SLOT} non lo conferma. \\
\rowa \op{0x45} & LOAD\_SLOT & slot\_id(1B)+psram\_addr(3B) & --- & Flash$\to$PSRAM per lo slot (offset/lunghezza dal catalogo), verifica CRC32 live; \code{STATUS.flash\_err} se lo slot non è valido o il CRC non torna. \\
\op{0x46} & SAVE\_SLOT & slot\_id(1B)+psram\_addr(3B)+len(3B) & --- & PSRAM$\to$flash all'offset già registrato dello slot, calcola il CRC32 live; a esito positivo aggiorna e persiste la entry di catalogo (lunghezza, CRC, valid=1). \\
\rowa \op{0x47} & CAT\_INSPECT & slot\_id(1B) & 16 byte & Lettura sincrona di una entry di catalogo già caricata: offset[3]+len[3]+tipo[1]+valid[1]+CRC32[4]+riservato[4], MSB-first. \\
\end{longtable}
Tutti gli opcode flash sono \emph{fire-and-forget}: l'host fa polling su \op{STATUS}
(bit4=\code{flash\_busy}, bit3=\code{flash\_err}) o sui pin \code{irq\_n}/\code{data\_ready\_n}
per l'esito, eccetto \op{CAT\_INSPECT} che risponde in modo sincrono.
Gli 8 opcode flash (\op{0x40}--\op{0x47}) sono descritti in dettaglio, con
razionale di progetto e latenze reali misurate, in §\ref{sec:flashspi} sotto.
\section{Selettori \texttt{SET\_BASE}}
\label{sec:setbase}
\begin{tabularx}{\textwidth}{C{1.2cm} L{3.2cm} Y}
\toprule
\rowh \thd{sel} & \thd{Registro} & \thd{Uso} \\
\midrule
0 & \code{x\_base} & Base ingresso $X$. \\
\rowa 1 & \code{w\_base} & Base pesi. \\
2 & \code{bias\_addr} & Base bias. \\
\rowa 3 & \code{table\_base} & Base tabella descrittori (multi-layer). \\
4 & \code{buf\_a\_base} & Buffer ping-pong A. \\
\rowa 5 & \code{buf\_b\_base} & Buffer ping-pong B. \\
6 & \code{activation} & Attivazione (2 bit bassi) --- solo percorso single-layer. \\
\rowa 7 & \code{n\_inputs\_real} & Larghezza ingressi runtime (16-bit BE) --- single-layer. \\
8 & \code{n\_neurons\_real} & Larghezza neuroni runtime (16-bit BE) --- single-layer. \\
\rowa 9 & \code{num\_neurons\_graph} & Numero neuroni del grafo (16-bit BE) --- Tipo \#2. \\
10 & \code{n\_out} & Numero id di uscita (16-bit BE) --- Tipo \#2. \\
\bottomrule
\end{tabularx}
I selettori 6--8 riguardano solo il percorso single-layer/manuale; con \op{RUN\_NETWORK}
i valori equivalenti sono letti per-layer dalla tabella descrittori.
\begin{fnwarn}[Casi limite ``reale=0'' corretti (2026-09-04)]
La campagna di ri-certificazione (\code{docs/validation/bugs.md}) ha trovato che diversi
valori runtime pari a zero non erano protetti da alcun guard, con esiti che andavano da un
risultato silenziosamente ignorato fino a hang o scritture PSRAM a indirizzi arbitrari.
Tutti e cinque i casi seguenti sono ora no-op sicuri, verificati indipendentemente:
\begin{itemize}
\item \code{n\_inputs\_real=0} (selettore 7): completa in 1 ciclo con
$y=\text{activation}(\text{bias})$ (BUG-003).
\item \code{n\_neurons\_real=0} (selettore 8): completa senza eseguire alcun calcolo
per-neurone, molto più rapido di un run a piena larghezza (BUG-004).
\item \code{num\_neurons\_graph=0} (selettore 9): completa immediatamente dopo la copia
degli ingressi, senza mai entrare nel loop dei descrittori (BUG-006).
\item \op{RUN\_NETWORK} con \code{num\_layers=0} (percorso dense): no-op immediato ---
\textbf{prima del fix eseguiva 256 layer fasulli leggendo dati PSRAM arbitrari come
descrittori} (BUG-005, CRITICO, vedi \S\ref{sec:run-network} sotto).
\item \op{SET\_NET\_TYPE} ricevuto mentre un run è in corso: ora rifiutato silenziosamente
(nessun effetto, nessun errore SPI) invece di rimappare il multiplexer dell'arbitro a metà
esecuzione --- \textbf{prima del fix causava un hang permanente del motore in corso}
(BUG-007, CRITICO).
\end{itemize}
Dettagli, evidenza e verifica di ciascun fix in \code{docs/validation/bugs.md}.
\end{fnwarn}
\section{\texttt{STATUS.done} sticky / clear-on-read}
In \code{neuron\_memory} il segnale \code{done} è un impulso di un solo ciclo. Un host
che effettua polling via SPI (molto più lento del clock FPGA) mancherebbe quasi
certamente un impulso grezzo di un ciclo. Il banco registri SPI latcha quindi
\code{done} in un bit sticky sull'impulso e lo azzera quando l'host legge \op{STATUS}
(o \op{RESET}). Il bit \code{busy} è invece mantenuto a livello per tutta la
computazione e si legge live.
\begin{fnwarn}[Race corretto (2026-09-02)]
Una race reale nel meccanismo sticky (presente dalla Fase~4) è stata corretta latchando
uno \code{status\_snapshot} all'accettazione dell'opcode \op{STATUS} e condizionando la
pulizia del bit sticky a \code{status\_snapshot[1]} (si azzera solo se il byte
effettivamente trasmesso mostrava \code{done=1}). Un \code{done} che arriva troppo tardi
per uno snapshot viene riportato al polling successivo invece di essere perso.
\end{fnwarn}
\section{Pin di attenzione host (\texttt{data\_ready\_n}, \texttt{irq\_n})}
Oltre al polling di \op{STATUS}, il top-level espone due pin fisici attivi bassi (banco 7,
cap.~\ref{ch:hw}) che rispecchiano i bit sticky senza richiedere una transazione SPI,
utili per pilotare un GPIO/IRQ dell'host:
\begin{itemize}
\item \code{data\_ready\_n} = $\sim$\code{STATUS.done} (sticky): basso quando un risultato è
pronto da leggere, torna alto alla lettura di \op{STATUS} (clear-on-read).
\item \code{irq\_n} = $\sim$\code{STATUS.err} (guard grafo): basso quando il guard load-time
di \code{graph\_engine} è scattato. \textbf{Non} è clear-on-read: si azzera solo con
\op{RESET} o un nuovo avvio di grafo, così un errore non passa inosservato tra un polling e
l'altro.
\end{itemize}
Sono porte aggiuntive: non toccano gli opcode né i registri esistenti.
\begin{fnwarn}[\code{flash\_err} non ha un pin dedicato]
\code{STATUS.flash\_err} (bit3) è riportato \textbf{solo} nel byte \op{STATUS}, per scelta
di progetto: riusare \code{irq\_n} lo avrebbe confuso con gli errori del guard grafo (due
domini di errore indipendenti sullo stesso pin), mentre un'operazione flash è sempre
avviata dall'host con un opcode appena emesso, quindi il polling di \op{STATUS} subito dopo
--- già implicito nella convenzione ``fire-and-forget, poi polling \op{STATUS}/
\code{data\_ready\_n}'' --- è già naturale, senza bisogno di un pin asincrono in più.
\code{data\_ready\_n} invece \emph{si azzera anche al termine di un'operazione flash}: lo
specchia \code{STATUS.done} (bit1), che ora latcha anche sul completamento di un op flash,
non solo su \op{RUN\_NETWORK}/\op{START}.
\end{fnwarn}
\section{\texttt{READ\_CONFIG}}
\label{sec:readcfg}
Payload fisso di \textbf{11 byte}: permette a un unico firmware host di funzionare con
bitstream diversi senza ricompilare. I valori \code{N\_INPUTS}/\code{N\_NEURONS} riportano
il \emph{massimo} di build (il soffitto), non necessariamente la rete correntemente
caricata.
\begin{tabularx}{\textwidth}{C{1.6cm} L{3.6cm} Y}
\toprule
\rowh \thd{Byte} & \thd{Campo} & \thd{Sorgente} \\
\midrule
0 & \code{ADDR\_WIDTH} (bit) & \code{neuron\_memory.ADDR\_WIDTH} \\
\rowa 1--2 & \code{N\_INPUTS} (16-bit BE) & massimo di build \\
3 & \code{N\_NEURONS} & massimo di build \\
\rowa 4 & \code{PARALLEL} & parametro di build \\
5 & \code{DATA\_WIDTH} (bit) & parametro di build \\
\rowa 6--7 & versione protocollo (BE) & \code{0x0001} \\
8--9 & \code{N\_TOTAL} (16-bit BE) & massimo segnali grafo (Tipo \#2) \\
\rowa 10 & flag di capacità & bit0=\code{GRAPH\_SUPPORTED}=1 \\
\bottomrule
\end{tabularx}
\section{Sottosistema flash (opcode 0x40--0x47, completato 2026-09-04)}
\label{sec:flashspi}
La FPGA ha accesso \textbf{esclusivo} alla flash di boot/persistenza onboard (Winbond
\code{W25Q128JV}, 16~MB SPI NOR, cap.~\ref{ch:hw} §6/§7) tramite un SPI master dedicato e
fisicamente separato (\code{rtl/spi\_flash\_master.v}), mai per accesso diretto dell'host ai
pin della flash. \textbf{Non} è un filesystem: un catalogo a dimensione fissa (16 slot,
\code{rtl/flash\_slot\_manager.v}) mappa \code{slot\_id}~$\to$~(offset, lunghezza, tipo,
valid, CRC32) in un settore riservato della flash (settore 0) --- nessuna allocazione
dinamica, nessun garbage collection.
\begin{fnnote}[Stratificazione (ogni livello testabile a sé)]
\begin{itemize}
\item \code{rtl/spi\_flash\_master.v} --- SPI master grezzo verso il chip flash
(RDID/READ/WREN/PP/SE/RDSR-1). Bus a 4 fili completamente indipendente
(\code{sclk}/\code{mosi}/\code{miso}/\code{cs\_n}, tutti GPIO ordinario ---
Fase F7, 2026-09-04): una versione precedente riusava il pad \code{CCLK} di boot
via la primitiva ECP5 \code{USRMCLK} per risparmiare un pin, abbandonato perché
rendeva fuorviante l'affermazione di ``bus esclusivo'' (elettricamente dipendeva
comunque dal motore di configurazione) e comportava un gap di verifica mai chiuso
(timing di \code{USRMCLKTS} mai verificato contro la guida Lattice primaria).
\item \code{rtl/flash\_copy\_engine.v} --- motore di streaming a blocchi: flash$\to$PSRAM
(\code{DIR\_LOAD}), PSRAM$\to$flash con erase-before-write interno + loop Page
Program $\leq$256B + poll WIP (\code{DIR\_SAVE}), erase di settore standalone
(\code{DIR\_ERASE}). Master a bassa priorità (Porta D) su \code{rtl/mem\_arbiter.v}:
le operazioni flash sono su scala dei ms e non bloccano mai l'inferenza.
\item \code{rtl/flash\_slot\_manager.v} --- il catalogo a slot sopra, più un CRC32
(\code{rtl/crc32.v}, IEEE~802.3/zlib) calcolato live sul flusso di byte reale durante
\op{LOAD\_SLOT}/\op{SAVE\_SLOT}, così uno slot corrotto o scritto a metà (es.
alimentazione persa durante l'erase) è rilevato anche quando l'operazione flash
sottostante ha riportato successo.
\end{itemize}
\end{fnnote}
\begin{fnwarn}[Allineamento a settore obbligatorio]
\op{SAVE\_SLOT} (e i raw \op{FLASH\_WRITE\_BLOCK}/\op{FLASH\_ERASE}) richiedono che
l'indirizzo flash target sia allineato a settore da 4~KB --- rifiutato come errore
altrimenti, invece di un silenzioso read-modify-erase-write parziale del settore (non
esiste un buffer di scratch abbastanza grande per farlo, e ogni \op{SAVE\_SLOT} reale scrive
già uno slot intero e allineato per costruzione).
\end{fnwarn}
Razionale completo, ogni citazione da datasheet, ogni test avversariale (CRC non
corrispondente, slot mai salvato, attraversamento di confine pagina, simulazione di perdita
di alimentazione, contesa sull'arbitro) e i due bug reali trovati e corretti durante il
bring-up (uno pre-esistente in \code{psram\_controller.v}, uno nel nuovo handshake di
richiesta dell'arbitro) sono in \code{WORKLOG.md} (voci Fasi F1-F6) e
\code{docs/FPGA-Neural-Flash-Subsystem-Verification.md} (sunto di copertura per modulo, non
ripetuto qui).
\begin{tabularx}{\textwidth}{L{3.4cm}Y}
\toprule
\rowh \thd{Operazione} & \thd{Latenza reale misurata} \\
\midrule
ERASE (settore 4~KB) & $\approx$400~ms (dominata dal tSE interno del chip flash, indipendente dal clock host) \\
\rowa SAVE (pagina 256~B, incl. erase interno) & $\approx$403~ms (idem, tSE+tPP) \\
LOAD (4096~B) & 1.74~ms (2.35~MB/s) @80~MHz; 8.71~ms (0.47~MB/s) @16~MHz (solo SPI-clock-bound) \\
\bottomrule
\end{tabularx}
Metodologia di misura completa in \code{docs/FPGA-Neural-Flash-Subsystem-Verification.md}.
\section{Sequenze di sessione}
\subsection{Percorso single-layer}
\begin{lstlisting}[language=,caption={Sessione single-layer},basicstyle=\ttfamily\scriptsize]
RESET -> 0x0F
READ_CONFIG -> 0x30 (l'host apprende N_INPUTS/N_NEURONS/...)
WRITE_RAM (pesi) -> 0x01 ...
WRITE_RAM (bias) -> 0x01 ...
SET_BASE (X/W/BIAS) -> 0x10 x3
WRITE_RAM (input X) -> 0x01 ...
START -> 0x20
poll STATUS -> 0x21 (finche' done=1; si azzera a questa lettura)
READ_OUTPUT -> 0x22
\end{lstlisting}
\subsection{Percorso multi-layer (RUN\_NETWORK)}
\label{sec:run-network}
\begin{lstlisting}[language=,caption={Sessione multi-layer},basicstyle=\ttfamily\scriptsize]
WRITE_RAM (tabella descrittori) -> 0x01 ...
WRITE_RAM (pesi/bias per layer, X layer0)-> 0x01 ...
SET_BASE (X/TABLE/BUF_A/BUF_B) -> 0x10 x4
RUN_NETWORK(num_layers) -> 0x23 <num_layers>
poll STATUS -> 0x21 (finche' done=1)
READ_OUTPUT -> 0x22 (y_bus del layer finale)
\end{lstlisting}
\begin{fnnote}[Fuori ambito per v1]
Dual~SPI e CRC/checksum sui trasferimenti host (SPI assunto affidabile su traccia di
scheda --- da non confondere con il CRC32 del catalogo flash, §\ref{sec:flashspi}, che
protegge un dominio diverso: la persistenza flash$\leftrightarrow$PSRAM, non il link SPI
host).
\end{fnnote}
\begin{fnwarn}[\op{WRITE\_RAM}/\op{READ\_RAM} senza backpressure verso l'host --- rischio reale, non teorico]
Ogni byte ricevuto/prodotto deve essere completamente processato da \code{spi\_engine}
prima che arrivi il successivo confine di byte scandito da SCLK --- ragionevole per il
bulk-loading iniziale di pesi/ingressi, non un percorso real-time. Il rischio concreto: se
un host emette \op{WRITE\_RAM}/\op{READ\_RAM} prima che la sequenza di power-up di
\code{psram\_controller.v} sia completata ($\sim$150~\textmu s dopo il reset,
\code{STATE\_INIT}+\code{STATE\_CR\_INIT}), \code{spi\_engine} si blocca in attesa che il
primo accesso PSRAM completi, mentre l'host --- non rallentato da alcun handshake ---
continua a scandire byte. I byte ricevuti durante quello stallo vengono \textbf{scartati
silenziosamente}, senza errore e senza hang: solo dati sbagliati in PSRAM. Trovato durante
il lavoro sul sottosistema flash (\code{WORKLOG.md}, Fase~F5) con una riproduzione minimale
solo-\op{WRITE\_RAM}, senza alcun opcode flash coinvolto: è un rischio generale per
qualunque host, non specifico agli opcode flash. \textbf{Mitigazione attuale: l'host deve
attendere il power-up della PSRAM (o assicurarsi che la FPGA sia fuori reset da
$>$150~\textmu s) prima del suo primo \op{WRITE\_RAM}/\op{READ\_RAM}.} Non risolto a livello
di protocollo (richiederebbe una vera backpressure, una modifica più ampia) ---
dichiarato qui come rischio aperto, non aggirato silenziosamente.
\end{fnwarn}
@@ -1,233 +0,0 @@
\chapter[Programmazione della rete]{Programmazione della rete neurale}
\label{ch:prog}
Questo capitolo è la guida pratica alla codifica di una rete per FPGA-Neural: come si
dispone in memoria, quali registri si impostano e come si avvia, per entrambe le
topologie. Presuppone gli opcode SPI (cap.~\ref{ch:spi}) e i formati descrittore
(cap.~\ref{ch:seq}, \ref{ch:grafo}).
\section{Flusso generale}
Qualunque sia il tipo, il ciclo è lo stesso: l'host \emph{costruisce le strutture dati in
RAM}, imposta i \emph{registri base}, dichiara il \emph{tipo di rete}, \emph{avvia} e
\emph{rilegge} il risultato.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=3mm,start chain=going below,
every node/.style={on chain,fnblock,minimum width=64mm}]
\node[fnblockA]{1. \op{RESET} --- azzera il motore e il latch STATUS};
\node{2. \op{SET\_NET\_TYPE} --- dense (\#1) o graph (\#2)};
\node{3. \op{WRITE\_RAM} --- tabelle, pesi/edge, bias, input X};
\node{4. \op{SET\_BASE} --- registri base (x, table, \ldots)};
\node[fnblockT]{5. \op{RUN\_NETWORK} --- dispatch su \code{net\_type}};
\node{6. \op{STATUS} in polling --- attende \code{done}};
\node[fnblockD]{7. \op{READ\_OUTPUT} / \op{READ\_RAM} --- risultato};
\foreach \i [count=\j from 2] in {1,...,6} \draw[fnarrow] (chain-\i)--(chain-\j);
\end{tikzpicture}
\end{center}
\section{Registri e opcode coinvolti}
Tutti i valori base si impostano con \op{SET\_BASE} \code{sel(1B)+addr(3B)}. Selettori:
\begin{tabularx}{\textwidth}{C{1.0cm} L{3.4cm} C{1.4cm} C{1.4cm} Y}
\toprule
\rowh \thd{sel} & \thd{Registro} & \thd{Tipo \#1} & \thd{Tipo \#2} & \thd{Uso} \\
\midrule
0 & \code{x\_base} & \checkmark & \checkmark & Base input $X$. \\
\rowa 3 & \code{table\_base} & \checkmark & \checkmark & Tabella descrittori. \\
4 & \code{buf\_a\_base} & \checkmark & \checkmark\textsuperscript{$\ast$} & Ping-pong A (\#1) / \code{out\_base} riuso (\#2). \\
\rowa 5 & \code{buf\_b\_base} & \checkmark & --- & Ping-pong B (\#1). \\
9 & \code{num\_neurons\_graph} & --- & \checkmark & Numero neuroni del grafo. \\
\rowa 10 & \code{n\_out} & --- & \checkmark & Numero id di uscita. \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
$\ast$ In Tipo \#2 i buffer ping-pong non servono: il selettore 4 è riusato come
\code{out\_base} (regione dove copiare le uscite). I selettori 1/2/6/7/8 riguardano solo
il percorso single-layer manuale (\op{START}), non \op{RUN\_NETWORK}.\end{center}
Per il Tipo \#1, i \code{w\_base}/\code{bias\_addr} \emph{per-layer} \textbf{non} si
impostano con \op{SET\_BASE}: sono campi della tabella descrittori. \op{SET\_NET\_TYPE}
default dopo \op{RESET} è \emph{dense}, quindi una rete \#1 funziona anche senza emetterlo.
% ======================================================================
\section{Tipo \#1 --- rete densa}
\subsection{Layout in memoria}
\begin{tabularx}{\textwidth}{L{3.4cm} Y}
\toprule
\rowh \thd{Struttura} & \thd{Formato} \\
\midrule
Input $X$ & \code{n\_inputs\_real} byte INT8 a \code{x\_base}. \\
\rowa Pesi (per layer) & Neuron-major: neurone $k$ a \code{w\_base + k*n\_inputs\_real}, \code{n\_neurons*n\_inputs} byte. \\
Bias (per layer) & Un byte INT8 per neurone a \code{bias\_addr}. \\
\rowa Tabella descrittori & \code{num\_layers} voci da 11 byte a \code{table\_base}. \\
Buffer A/B & Uscite intermedie ping-pong. \\
\bottomrule
\end{tabularx}
Descrittore (11 byte, MSB-first): \code{w\_base}(3) $|$ \code{bias\_addr}(3) $|$
\code{activation}(1) $|$ \code{n\_inputs\_real}(2) $|$ \code{n\_neurons\_real}(2).
\subsection{Esempio completo: rete $4\to4\to2$}
Layer~0: 4 input, 4 neuroni, ReLU. Layer~1: 4 input, 2 neuroni, lineare
(\code{PARALLEL}=2, quindi ogni \code{n\_inputs\_real} è multiplo di 2). Indirizzi scelti:
\code{table\_base}=\code{0x000000}, \code{x\_base}=\code{0x001000}, pesi/bias L0 a
\code{0x002000}/\code{0x002100}, L1 a \code{0x002200}/\code{0x002300}, buffer a
\code{0x003000}/\code{0x003100}.
\begin{lstlisting}[language=,caption={Tabella descrittori dense (22 byte)},basicstyle=\ttfamily\scriptsize]
Layer 0: 00 20 00 | 00 21 00 | 01 | 00 04 | 00 04
w_base bias_addr ReLU n_in=4 n_neu=4
Layer 1: 00 22 00 | 00 23 00 | 00 | 00 04 | 00 02
w_base bias_addr NONE n_in=4 n_neu=2
\end{lstlisting}
\begin{lstlisting}[language=,caption={Sessione SPI (dense)},basicstyle=\ttfamily\scriptsize]
0x0F RESET
0x11 01 SET_NET_TYPE = dense
0x01 000000 0016 <22 byte tabella> WRITE_RAM tabella
0x01 002000 0010 <16 byte pesi L0> WRITE_RAM pesi L0 (neuron-major)
0x01 002100 0004 <4 byte bias L0>
0x01 002200 0008 <8 byte pesi L1>
0x01 002300 0002 <2 byte bias L1>
0x01 001000 0004 <x0 x1 x2 x3> WRITE_RAM input X
0x10 00 001000 SET_BASE x_base
0x10 03 000000 SET_BASE table_base
0x10 04 003000 SET_BASE buf_a
0x10 05 003100 SET_BASE buf_b
0x23 02 RUN_NETWORK num_layers=2
0x21 ... poll STATUS finche' done=1
0x22 READ_OUTPUT -> 2 byte (layer finale)
\end{lstlisting}
\subsection{Pseudocodice host (dense)}
\begin{lstlisting}[language=,caption={Codifica e caricamento di una rete densa},basicstyle=\ttfamily\scriptsize]
def load_dense(layers, X): # layers in ordine di esecuzione
spi(RESET); spi(SET_NET_TYPE, DENSE)
table = b""
for L in layers: # L: pesi[n][k], bias[n], act, n_in, n_out
assert L.n_in % PARALLEL == 0
w = alloc(L.weights_neuron_major) # k lento, input veloce
b = alloc(L.bias)
table += u24(w)+u24(b)+u8(L.act)+u16(L.n_in)+u16(L.n_out)
write_ram(TABLE_BASE, table)
write_ram(X_BASE, X)
set_base(0, X_BASE); set_base(3, TABLE_BASE)
set_base(4, BUF_A); set_base(5, BUF_B)
spi(RUN_NETWORK, len(layers))
wait_status_done()
return read_output(layers[-1].n_out)
\end{lstlisting}
% ======================================================================
\section{Tipo \#2 --- rete a grafo}
\subsection{Layout in memoria}
\begin{tabularx}{\textwidth}{L{3.4cm} Y}
\toprule
\rowh \thd{Struttura} & \thd{Formato} \\
\midrule
Input $X$ & \code{N\_in} byte a \code{x\_base}; copiati in \code{act\_buf[0..N\_in-1]} all'avvio. \\
\rowa Tabella descrittori & \code{num\_neurons\_graph} voci da 11 byte a \code{table\_base}, in ordine di \code{out\_id} crescente. \\
Blocchi edge & Per neurone: \code{n\_conn} edge da 4 byte a \code{conn\_ptr}, con padding a multiplo di \code{PARALLEL} (edge peso 0). \\
\rowa Uscite & \code{n\_out} byte scritti a \code{out\_base} (=selettore 4). \\
\bottomrule
\end{tabularx}
Descrittore graph (11 byte): \code{conn\_ptr}(3) $|$ \code{n\_conn}(2) $|$ \code{out\_id}(2)
$|$ \code{activation}(1) $|$ \code{bias}(1) $|$ \code{reserved}(2). \quad
Edge (4 byte): \code{src\_id}(2) $|$ \code{weight}(1) $|$ \code{reserved}(1). \quad
Vincolo: \code{src\_id < out\_id} (DAG feed-forward).
\subsection{Esempio completo}
4 ingressi (id 0--3). Neurone n4 (\code{out\_id}=4, ReLU, bias=2) connesso agli id 0 e 1;
neurone n5 (\code{out\_id}=5, lineare, bias=0) connesso a n4 (id~4) e all'id~2; uscita = n5
(\code{n\_out}=1). \code{PARALLEL}=2, entrambi hanno 2 connessioni (nessun padding).
Indirizzi: \code{table\_base}=\code{0x000000}, edge a \code{0x000100}, \code{x\_base}=
\code{0x001000}, \code{out\_base}=\code{0x002000}.
\begin{lstlisting}[language=,caption={Descrittori + edge grafo},basicstyle=\ttfamily\scriptsize]
Descrittori (a 0x000000, 22 byte):
n4: 00 01 00 | 00 02 | 00 04 | 01 | 02 | 00 00
conn_ptr n_conn out_id ReLU bias rsv
n5: 00 01 08 | 00 02 | 00 05 | 00 | 00 | 00 00
conn_ptr n_conn out_id NONE bias rsv
Blocchi edge (a 0x000100, 4 byte/edge: src_id, weight, rsv):
n4 @0x000100: 00 00 05 00 (src=0, w=+5)
00 01 FD 00 (src=1, w=-3) ; -3 = 0xFD
n5 @0x000108: 00 04 02 00 (src=4, w=+2) ; id4 = uscita di n4
00 02 07 00 (src=2, w=+7)
\end{lstlisting}
\begin{lstlisting}[language=,caption={Sessione SPI (graph)},basicstyle=\ttfamily\scriptsize]
0x0F RESET
0x11 02 SET_NET_TYPE = graph
0x01 000000 0016 <22 byte tabella> WRITE_RAM descrittori
0x01 000100 0010 <16 byte edge> WRITE_RAM blocchi edge
0x01 001000 0004 <x0 x1 x2 x3> WRITE_RAM input X
0x10 00 001000 SET_BASE x_base
0x10 03 000000 SET_BASE table_base
0x10 04 002000 SET_BASE out_base (riuso sel 4)
0x10 09 000002 SET_BASE num_neurons_graph = 2
0x10 0A 000001 SET_BASE n_out = 1
0x23 00 RUN_NETWORK (dispatch a graph_engine)
0x21 ... poll STATUS (bit2=err se src_id>=out_id)
0x02 002000 0001 READ_RAM out_base -> 1 byte (uscita n5)
\end{lstlisting}
\subsection{Pseudocodice host (graph)}
\begin{lstlisting}[language=,caption={Codifica e caricamento di un grafo},basicstyle=\ttfamily\scriptsize]
def load_graph(neurons, X, n_out): # neurons ordinati per out_id crescente
spi(RESET); spi(SET_NET_TYPE, GRAPH)
edges = b""; table = b""
for N in neurons: # N: out_id, conns=[(src_id,w)...], act, bias
for (src,_) in N.conns:
assert src < N.out_id and src < N_TOTAL # regola DAG
conn_ptr = EDGE_BASE + len(edges)
padded = pad(N.conns, PARALLEL, fill=(0,0)) # edge peso 0
for (src,w) in padded:
edges += u16(src)+i8(w)+u8(0)
table += u24(conn_ptr)+u16(len(N.conns))+u16(N.out_id) \
+ u8(N.act)+i8(N.bias)+u16(0)
write_ram(TABLE_BASE, table); write_ram(EDGE_BASE, edges)
write_ram(X_BASE, X)
set_base(0, X_BASE); set_base(3, TABLE_BASE); set_base(4, OUT_BASE)
set_base(9, len(neurons)); set_base(10, n_out)
spi(RUN_NETWORK, 0) # payload ignorato in graph
wait_status_done()
return read_ram(OUT_BASE, n_out)
\end{lstlisting}
\subsection{Pseudo-assembly \texttt{netasm}}
La descrizione leggibile viene compilata dall'assemblatore host (\code{tools/netasm/})
esattamente nei byte delle tabelle e degli edge sopra. Esempio equivalente al grafo
dell'esempio:
\begin{lstlisting}[language=,caption={netasm: sorgente e byte generati},basicstyle=\ttfamily\scriptsize]
; --- sorgente ---
NET graph
INPUTS 4 ; id 0..3
NEURON n4 relu bias=2
CONN 0 w=5
CONN 1 w=-3
NEURON n5 none bias=0
CONN n4 w=2 ; riferimento simbolico -> id 4
CONN 2 w=7
OUTPUT n5
END
; --- l'assemblatore emette ---
; id assegnati: n4=4, n5=5 (garantito src_id < out_id)
; descrittori: 00 01 00 00 02 00 04 01 02 00 00
; 00 01 08 00 02 00 05 00 00 00 00
; edge: 00 00 05 00 00 01 FD 00 (n4)
; 00 04 02 00 00 02 07 00 (n5)
; registri: table_base, x_base, out_base, num_neurons=2, n_out=1
; validato a compile-time: src_id<out_id, N_TOTAL, padding a PARALLEL
\end{lstlisting}
\begin{fnnote}[Perche' due livelli di codifica]
Lo pseudocodice host e il \code{netasm} producono gli \emph{stessi byte}. Il primo è utile
quando la rete è generata a runtime (es. pesi da training); il secondo quando la topologia
è scritta a mano o versionata come sorgente. In entrambi i casi l'FPGA riceve solo tabelle
e dati via \op{WRITE\_RAM}: nessun interprete a bordo.
\end{fnnote}
@@ -1,78 +0,0 @@
\chapter[Arbitraggio e top-level]{Arbitraggio e integrazione top-level}
\label{ch:top}
\section{\texttt{mem\_arbiter} --- arbitro a tre porte}
Un unico master di memoria byte-level (che alimenta la catena condivisa
\code{int8\_memory\_access} $\to$ \code{memory\_interface} $\to$ \code{psram\_controller})
è arbitrato tra tre richiedenti:
\begin{tabularx}{\textwidth}{C{1.3cm} L{3.4cm} Y}
\toprule
\rowh \thd{Porta} & \thd{Master} & \thd{Accessi} \\
\midrule
A & \code{spi\_engine} & \op{WRITE\_RAM} / \op{READ\_RAM}. \\
\rowa B & \code{neuron\_memory} & Letture X/W/bias durante un'esecuzione. \\
C & \code{layer\_sequencer} & Letture descrittori + scritture buffer tra layer. \\
\bottomrule
\end{tabularx}
Priorità fissa \textbf{B $>$ C $>$ A}: un'inferenza in corso è più critica della
contabilità del sequencer, che a sua volta è più critica di un accesso SPI manuale
appena arrivato. In funzionamento normale B e C sono comunque temporalmente disgiunti
(\code{neuron\_memory} richiede solo durante un'esecuzione, \code{layer\_sequencer} solo
nelle pause tra layer), quindi la priorità conta soprattutto per il caso limite di un
\op{WRITE\_RAM}/\op{READ\_RAM} manuale che arriva durante un'esecuzione multi-layer.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=6mm]
\node[fnblock,minimum width=30mm](a){Port A --- \code{spi\_engine}};
\node[fnblock,below=4mm of a,minimum width=30mm](b){Port B --- \code{neuron\_memory}};
\node[fnblock,below=4mm of b,minimum width=30mm](c){Port C --- \code{layer\_sequencer}};
\node[fnblockD,right=16mm of b,minimum width=26mm,minimum height=16mm](arb){\code{mem\_arbiter}\\{\scriptsize B$>$C$>$A}};
\node[fnblockT,right=14mm of arb,minimum width=26mm](m){catena memoria\\{\scriptsize condivisa}};
\draw[fnarrow] (a)-|(arb.west|-a); \draw[fnarrow] (b)--(arb.west);
\draw[fnarrow] (c)-|(arb.west|-c);
\draw[fnbus] (arb)--(m);
\end{tikzpicture}
\end{center}
Concesso l'accesso, l'arbitro mantiene la proprietà fino all'impulso \code{m\_ready}
della singola transazione, poi rilascia: tutti e tre i master emettono \code{req} come
impulso pulito di un ciclo, quindi è sufficiente un design grant-and-forward senza code.
\section{\texttt{spi\_neuron\_top} --- integrazione completa}
Il top-level collega SPI (\code{spi\_slave}+\code{spi\_engine}), l'arbitro, il sequencer,
\code{neuron\_memory} e la catena PSRAM. Il reset di \code{neuron\_memory} è l'OR del
reset globale con l'impulso di soft-reset dell'opcode \op{RESET}, così l'host può
recuperare il motore via SPI senza reset fisico (la RAM resta intatta).
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=7mm]
\node[fnblockA,minimum width=22mm](ss){\code{spi\_slave}};
\node[fnblockA,right=8mm of ss,minimum width=22mm](se){\code{spi\_engine}};
\node[fnblockT,below=8mm of se,minimum width=26mm](sq){\code{layer\_sequencer}};
\node[fnblockD,right=10mm of se,minimum width=24mm](mux){MUX ctrl\\{\scriptsize su \code{seq\_busy}}};
\node[fnblock,below=8mm of mux,minimum width=26mm](nm){\code{neuron\_memory}};
\node[fnblockD,right=10mm of mux,minimum width=22mm](arb){\code{mem\_arbiter}};
\node[fnblockA,right=8mm of arb,minimum width=26mm](mem){catena PSRAM};
\draw[fnarrow] (ss)--(se);
\draw[fnarrow] (se)--(mux);
\draw[fnarrow] (sq)--(mux);
\draw[fnarrow] (mux)--(nm);
\draw[fnarrow] (se.south) to[bend right=10] (arb.north west);
\draw[fnarrow] (nm)--(arb);
\draw[fnarrow] (sq.east) to[bend right=20] (arb.south west);
\draw[fnbus] (arb)--(mem);
\end{tikzpicture}
\end{center}
Il multiplexer commuta le linee di controllo di \code{neuron\_memory} tra il sequencer
(mentre \code{seq\_busy} è alto) e il percorso diretto di \code{spi\_engine} (modalità
single-layer legacy), restituendo il motore al percorso diretto a fine sequenza.
\begin{fnnote}[Verifica end-to-end]
\code{spi\_neuron\_top} è verificato in simulazione con PSRAM reale
(\code{psram\_model.v}, nessun mock): RESET/READ\_CONFIG/WRITE\_RAM/READ\_RAM/SET\_BASE/
START/STATUS/READ\_OUTPUT e \op{RUN\_NETWORK} sono esercitati puramente su SPI simulato
(cap.~\ref{ch:impl}).
\end{fnnote}
@@ -1,172 +0,0 @@
\chapter[Implementazione ECP5]{Implementazione e caratterizzazione ECP5}
\label{ch:impl}
\section{Flusso e verifica}
Il progetto è verificato su due piani complementari: \textbf{simulazione} funzionale con
Icarus Verilog (algebra signed, prodotti, accumulo, gruppi, bias, ReLU, saturazione,
segnali busy/done) e \textbf{implementazione} reale con Yosys (sintesi) $+$
nextpnr-ecp5 (place\&route, timing) $+$ Project~Trellis (\code{ecppack}).
\begin{tabularx}{\textwidth}{L{5.0cm} C{3.0cm} Y}
\toprule
\rowh \thd{Fase di verifica} & \thd{Esito} & \thd{Copre} \\
\midrule
RTL funzionale & \PASS & correttezza del datapath \\
\rowa Simulazione parametrica & \PASS & sweep di configurazioni \\
Sintesi ECP5 & \PASS & sintetizzabilità, mapping \\
\rowa Placement / Routing & \PASS & LUT/FF/DSP, timing \\
Bitstream (\code{ecppack}) & \PASS & flusso completo, 0 errori (P2 e P8) \\
\bottomrule
\end{tabularx}
\begin{fnnote}[Toolchain end-to-end fino al bitstream]
L'intero flusso RTL $\to$ Yosys $\to$ nextpnr-ecp5 $\to$ \code{ecppack} produce un
bitstream valido per P2 e P8, \textbf{0 errori in ogni stadio}. Header verificato
byte-per-byte: \code{Part: LFE5U-45F-8CABGA381}, il part number reale del target, non un
placeholder. Verificata la sola \emph{generazione}: nessun test su hardware fisico in
questa sessione.
\end{fnnote}
\section{Benchmark del datapath (256$\times$4)}
Configurazione: INT8/INT32, \code{N\_INPUTS}=256, \code{N\_NEURONS}=4, \code{PARALLEL}
variabile, target 80~MHz, dispositivo \code{LFE5U-45F-8BG381C} ($-8$). I bus di test
sono generati \emph{dentro} il wrapper di benchmark per non esporre migliaia di I/O; il
top-level espone solo \code{clk/rst/start/y\_bus/busy/done}.
\begin{tabularx}{\textwidth}{C{1.4cm} C{1.8cm} C{1.4cm} C{1.6cm} C{1.6cm} C{1.5cm} C{1.4cm}}
\toprule
\rowh \thd{PAR} & \thd{MAC tot} & \thd{DSP} & \thd{Fmax} & \thd{Tcrit} & \thd{80\,MHz} & \thd{LUT4} \\
\midrule
16 & 64 & 64/72 & 52.13 & 19.18 & \FAIL & $\approx$2531 \\
\rowa 8 & 32 & 32/72 & 61.71 & 16.20 & \FAIL & --- \\
4 & 16 & 16/72 & 75.01 & 13.33 & \FAIL & 804 \\
\rowa 2 & 8 & 8/72 & 87.88 & 11.38 & \PASS & 481 \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
Fmax e Tcrit in MHz e ns. MAC totali $=$ PARALLEL$\times$4 neuroni.\end{center}
\subsection{Fmax e throughput contro parallelismo}
\begin{center}
\begin{tikzpicture}
\begin{axis}[
width=0.62\textwidth,height=6.0cm,
axis y line*=left, axis x line=bottom,
xlabel={\footnotesize PARALLEL}, ylabel={\footnotesize Fmax [MHz]},
xtick={2,4,8,16}, xmode=log, log basis x=2,
ymin=40,ymax=95, ytick={40,55,70,85},
tick label style={font=\scriptsize}, label style={font=\footnotesize},
grid=major, grid style={fnRule!40},
legend style={font=\scriptsize,at={(0.5,-0.28)},anchor=north,legend columns=2}]
\addplot[fnTeal,mark=*,thick,mark options={fill=fnTeal}]
coordinates {(2,87.88)(4,75.01)(8,61.71)(16,52.13)};
\addlegendentry{Fmax}
\draw[fnAmber,dashed,thick] (axis cs:2,80)--(axis cs:16,80);
\node[font=\scriptsize,text=fnAmber] at (axis cs:11,82.5){target 80 MHz};
\end{axis}
\begin{axis}[
width=0.62\textwidth,height=6.0cm,
axis y line*=right, axis x line=none,
xmode=log, log basis x=2, xmin=2,xmax=16,
ylabel={\footnotesize throughput [G\,MAC/s]},
ymin=0,ymax=3.6, ytick={0,1,2,3},
tick label style={font=\scriptsize}, label style={font=\footnotesize}]
\addplot[fnBlue,mark=square*,thick,mark options={fill=fnBlue}]
coordinates {(2,0.703)(4,1.20)(8,1.97)(16,3.34)};
\label{plt:tp}
\end{axis}
\end{tikzpicture}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
Trade-off fondamentale: al crescere di PARALLEL la Fmax cala (routing/albero più
profondi) ma il throughput teorico sale. La linea blu (quadrati) è il throughput
$\approx$MAC/ciclo$\times$Fmax.\end{center}
\subsection{Interpretazione}
Riducendo \code{PARALLEL} calano MAC simultanei, DSP, profondità dell'adder tree e
congestione di routing, quindi la Fmax sale; ma aumenta il numero di gruppi e quindi la
latenza. La sola frequenza non basta a scegliere: conta il throughput complessivo
$\approx$MAC/ciclo$\times$frequenza.
\begin{fnnote}[Scelte architetturali]
\code{PARALLEL=8} è il candidato per la V1 orientata al throughput: esattamente 32~MAC
simultanei con 4 neuroni, DSP al $\approx$44\%, lasciando risorse per controller,
buffer, SPI e pipeline future. \code{PARALLEL=2} è il riferimento orientato alla
frequenza: 87.88~MHz, unico a superare il target 80~MHz, ma richiede 128 gruppi per un
neurone da 256 ingressi.
\end{fnnote}
\subsection{Percorso critico e limite a 100~MHz}
Il target 100~MHz non è raggiunto (miglior risultato 87.88~MHz con P2). Il limite è
\emph{temporale}, non di occupazione: con P2 l'FPGA è usato pochissimo (DSP $\approx$11\%,
LUT $\approx$1\%). Il percorso critico attraversa FF pesi $\to$ \code{MULT18X18D} $\to$
prodotti $\to$ adder/carry $\to$ \code{acc\_next} $\to$ ReLU/saturazione $\to$ FF uscita.
Superare 100~MHz richiederà una o più pipeline interne, non ancora necessarie per
proseguire.
\section{Sistema integrato completo}
Sintesi reale di \code{spi\_neuron\_top} (SPI + arbitro + \code{neuron\_memory} +
\code{graph\_engine} + catena PSRAM), speed grade $-8$. Prima della timing closure il
sistema integrato mancava il target 80~MHz (P2 $\approx$55~MHz, P8 $\approx$45~MHz), con
un percorso critico interamente interno a \code{neuron\_parallel}.
\subsection{Causa: catena di saturazione/ReLU}
L'utilizzo di risorse non è la causa (device sotto il 10\% ovunque). Il percorso critico
del sistema integrato è la \textbf{catena di riporto \code{CCU2C} del comparatore di
saturazione/ReLU} in \code{neuron\_parallel.v} --- \emph{non} lo SPI, l'arbitro, la PSRAM
né i moduli del Tipo \#2. La saturazione era scritta come confronto aritmetico
(\code{acc > 127}, \code{acc < -128}), mappato dal sintetizzatore su un sottrattore a
32~bit con carry chain lunga.
\subsection{Timing closure (2026-09-03)}
Deroga esplicita al vincolo ``datapath intoccabile'' per un task separato di timing
closure, con l'unico vincolo dell'\textbf{equivalenza bit-esatta} su tutta la regressione.
Due passi:
\begin{itemize}
\item \textbf{Passo 1 --- saturazione/ReLU come bit-test.} Un valore signed a 32~bit sta
in INT8 se e solo se \code{acc[31:7]} sono tutti uguali: riduzione AND/OR su una fetta di
bit invece di 32~bit di riporto. Semplificazione corretta e verificata bit-esatta, guadagno
di logica reale ma da solo sommerso dal rumore di piazzamento.
\item \textbf{Passo 2 --- registro di pipeline} tra accumulo e attivazione (\code{+1}
ciclo di latenza per neurone, assorbito dall'handshake \code{start}/\code{done}, trasparente
per i chiamanti). È il passo decisivo.
\end{itemize}
\begin{tabularx}{\textwidth}{L{4.6cm} C{2.6cm} C{2.4cm} Y}
\toprule
\rowh \thd{Config} & \thd{Prima} & \thd{Dopo} & \thd{$\Delta$} \\
\midrule
P2, \code{.lpf} reale & 54.58 & \textbf{75.30} & $+38\%$ \\
\rowa P2, sweep 5 seed & 55.59 & 73.38--75.55 & robusto \\
P8, unconstrained & 45.47 & \textbf{60.26} & $+33\%$ \\
\rowa P8, sweep 5 seed & 43.15--50.48 & 60.26--68.87 & non sovrapposto \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
Fmax in MHz, place\&route reale (\code{nextpnr-ecp5}). Guadagno robusto su 5 seed, non
attribuibile a fortuna di placement.\end{center}
\begin{fnnote}[Criterio di stop e margine reale]
80~MHz non è raggiunto (75.30~MHz a P2, 94\% del target) ma il guadagno è enorme e reale
($+38\%$/$+33\%$). Il passo successivo (registro di uscita del \code{MULT18X18D}, che
toccherebbe \code{mac\_unit.v}) è stato lasciato: gli 80~MHz sono \emph{headroom} in vista
del \code{.lpf} reale, non un requisito operativo. Con l'oscillatore previsto a 16~MHz,
anche il numero peggiore misurato ($\approx$45~MHz a P8) ha $2.8\times$ di margine.
\textbf{Superato 2026-09-04}: dopo l'aggiunta del sottosistema flash (cap.~\ref{ch:spi}
§\ref{sec:flashspi}, cap.~\ref{ch:roadmap}) la Fmax del sistema completo (P2, stesso
pinout reale + 3 nuovi segnali flash) era 66.68~MHz, percorso critico ancora sulla
stessa catena di accumulo di \code{neuron\_parallel} identificata qui sopra --- non un
nuovo collo di bottiglia, la differenza rispetto a 75.30~MHz rumore di piazzamento/routing
dovuto ai pin/logica aggiuntivi. \textbf{Aggiornato di nuovo lo stesso giorno (Fase F7)}:
reso il bus SPI della flash genuinamente indipendente (rimosso il riuso del pad \code{CCLK}
via \code{USRMCLK}, aggiunto un 4°~pin \code{flash\_sclk} ordinario), Fmax ri-misurata
\textbf{67.91~MHz} (leggero miglioramento, percorso critico confermato ancora identico).
Margine sull'oscillatore 16~MHz: $4.2\times$.
\end{fnnote}
\begin{fnnote}[Ottimizzazione futura separata]
Indipendente dalla timing closure: gli array \code{x\_mem}/\code{w\_mem} di
\code{neuron\_memory} sono ancora inferiti come RAM distribuita su LUT anziché su
\code{DP16KD}. Spostarli su block RAM libererebbe LUT ed è un candidato per la Fase~7 ---
non era però sul percorso critico risolto qui.
\end{fnnote}
@@ -1,325 +0,0 @@
\chapter[Progetto hardware e pinout]{Progetto hardware e mappa dei segnali}
\label{ch:hw}
\begin{fnnote}[Stato del pinout --- assegnato e verificato]
Esiste ora un \code{.lpf} reale (\code{synth/ecp5/spi\_neuron\_top.lpf}) con i \textbf{57
segnali} del top-level assegnati a ball CABGA381 concrete, \textbf{verificato
da un place\&route \code{nextpnr-ecp5} completo a 0 errori} (non più
\code{-{}-lpf-allow-unconstrained}). Le ball derivano dal database di dispositivo di
Project~Trellis (\code{iodb.json}, lo stesso che usa nextpnr) e sono state validate in modo
indipendente contro la §4.3.2 del datasheet Lattice ufficiale (conteggi GPIO per banco:
coincidenza esatta su 6 banchi su 7, scostamento di 1 ball sul banco 3, irrilevante perché
nessun segnale assegnato lo usa). \code{TRELLIS\_IO}: 57/245 (23\%). Fmax del build
corrente (sistema completo incl. sottosistema flash con bus SPI indipendente, Fase F7,
2026-09-04) \textbf{67.91~MHz}, percorso critico confermato ancora sulla catena di accumulo
di \code{neuron\_parallel}, invariato rispetto alle build precedenti (cap.~\ref{ch:impl}).
Sunto pin-per-pin a inizio documento (pagg.~2--3). Le ball di config-SPI di boot e JTAG non
compaiono qui perché sono pin dedicati a funzione fissa, senza porta RTL corrispondente:
nextpnr non le richiede mai (0 errori), contano solo per lo schematic PCB.
\end{fnnote}
\section{Dispositivo target}
\begin{tabularx}{\textwidth}{L{4.2cm}Y}
\toprule
\rowh \thd{Parametro} & \thd{Valore} \\
\midrule
Dispositivo & Lattice ECP5 \code{LFE5U-45F-8BG381C} \\
\rowa Package & CABGA381 (381 ball) \\
Speed grade & $-8$ (il più veloce della famiglia ECP5) \\
\rowa Risorse & $\approx$44k LUT/FF, 72$\times$\code{MULT18X18D}, block RAM \code{DP16KD} \\
I/O utilizzabili & $\approx$232 ball su 381 (resto: alimentazione/massa/NC) \\
\bottomrule
\end{tabularx}
\section{Budget dei pin}
Il progetto richiede circa 60 segnali su $\approx$232 I/O utilizzabili: ampio margine
($>$170 pin liberi), quindi la scheda non è pin-constrained.
\begin{tabularx}{\textwidth}{Y C{2.2cm}}
\toprule
\rowh \thd{Funzione} & \thd{Pin} \\
\midrule
PSRAM (indirizzi 22, dati 16, controllo 6) & fino a 44 \\
\rowa SPI applicativo (\code{sclk/mosi/miso/cs\_n}) & 4 \\
Clock, reset & 2 \\
\rowa Pin attenzione host (\code{irq\_n}, \code{data\_ready\_n}) & 2 \\
Bus SPI flash runtime (\code{flash\_sclk/flash\_mosi/flash\_miso/flash\_cs\_n}, GPIO ordinario, bus indipendente --- Fase F7) & 4 \\
\rowa JTAG (bring-up / debug, consigliato) & 4 \\
\midrule
\rowh \thd{Totale} & \thd{$\approx$60} \\
\bottomrule
\end{tabularx}
\section{Mappa dei segnali (top-level \texttt{spi\_neuron\_top}) --- ball reali}
Assegnazione reale dei 57 segnali del top-level, verificata da place\&route, \textbf{ball
individuale per ogni bit} (mai un intervallo di bus). Standard I/O: LVCMOS33
(alimentazione I/O a 3.3~V). Le ball provengono dal \code{.lpf} reale
place\&route-verified. Sunto compatto della stessa tabella anche a inizio documento
(pagg.~2--3).
\renewcommand{\arraystretch}{1.1}
\begin{tabularx}{\textwidth}{L{3.0cm} C{1.0cm} C{1.9cm} C{1.0cm} Y}
\toprule
\rowh \thd{Segnale} & \thd{Dir} & \thd{Ball} & \thd{Banco} & \thd{Funzione} \\
\midrule
\multicolumn{5}{l}{\textit{\color{fnDark}Clock e reset (banco 7, lato sinistro)}}\\
\code{clk} & IN & H5 & 7 & Clock di sistema su pad \code{GR\_PCLK7\_0} (clock globale dedicato). \\
\rowa \code{rst} & IN & B4 & 7 & Reset globale sincrono, attivo alto. \\
\multicolumn{5}{l}{\textit{\color{fnDark}SPI applicativo (banco 7, opposto al bus PSRAM)}}\\
\code{sclk} & IN & B5 & 7 & SPI clock (CPOL=0, CPHA=0). \\
\rowa \code{mosi} & IN & C5 & 7 & Master-Out Slave-In. \\
\code{miso} & OUT & A3 & 7 & Master-In Slave-Out (pilotato sul fronte di discesa). \\
\rowa \code{cs\_n} & IN & B3 & 7 & Chip-select attivo basso. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Pin di attenzione host (banco 7, attivi bassi, di livello)}}\\
\code{data\_ready\_n} & OUT & C3 & 7 & Basso finché un risultato attende lettura (specchio di \code{STATUS.done}, clear su lettura STATUS). \\
\rowa \code{irq\_n} & OUT & C4 & 7 & Basso se il guard load-time del grafo è scattato (specchio di \code{STATUS.err}); si azzera solo su \code{RESET} o nuovo \code{run\_start}, \emph{non} su lettura STATUS. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Flash subsystem --- SPI verso W25Q128JV onboard, bus indipendente (banco 7, Fasi F1-F7)}}\\
\code{flash\_sclk} & OUT & E3 & 7 & SPI clock verso la flash --- GPIO ordinario, nessuna primitiva di config coinvolta (Fase F7). \\
\rowa \code{flash\_mosi} & OUT & D3 & 7 & Master-Out Slave-In verso la flash. \\
\code{flash\_miso} & IN & D5 & 7 & Master-In Slave-Out dalla flash. \\
\rowa \code{flash\_cs\_n} & OUT & E4 & 7 & Chip-select flash, attivo basso. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Bus PSRAM indirizzi \code{psram\_a[21:0]} --- 22 ball individuali (banco 2)}}\\
\code{psram\_a[0]} & OUT & E16 & 2 & PSRAM A0 \\
\rowa \code{psram\_a[1]} & OUT & F16 & 2 & PSRAM A1 \\
\code{psram\_a[2]} & OUT & D18 & 2 & PSRAM A2 \\
\rowa \code{psram\_a[3]} & OUT & E17 & 2 & PSRAM A3 \\
\code{psram\_a[4]} & OUT & E18 & 2 & PSRAM A4 \\
\rowa \code{psram\_a[5]} & OUT & F18 & 2 & PSRAM A5 \\
\code{psram\_a[6]} & OUT & F17 & 2 & PSRAM A6 \\
\rowa \code{psram\_a[7]} & OUT & G16 & 2 & PSRAM A7 \\
\code{psram\_a[8]} & OUT & G18 & 2 & PSRAM A8 \\
\rowa \code{psram\_a[9]} & OUT & H16 & 2 & PSRAM A9 \\
\code{psram\_a[10]} & OUT & H17 & 2 & PSRAM A10 \\
\rowa \code{psram\_a[11]} & OUT & H18 & 2 & PSRAM A11 \\
\code{psram\_a[12]} & OUT & J16 & 2 & PSRAM A12 \\
\rowa \code{psram\_a[13]} & OUT & J17 & 2 & PSRAM A13 \\
\code{psram\_a[14]} & OUT & C20 & 2 & PSRAM A14 \\
\rowa \code{psram\_a[15]} & OUT & D19 & 2 & PSRAM A15 \\
\code{psram\_a[16]} & OUT & E19 & 2 & PSRAM A16 \\
\rowa \code{psram\_a[17]} & OUT & E20 & 2 & PSRAM A17 \\
\code{psram\_a[18]} & OUT & F19 & 2 & PSRAM A18 \\
\rowa \code{psram\_a[19]} & OUT & F20 & 2 & PSRAM A19 \\
\code{psram\_a[20]} & OUT & G20 & 2 & PSRAM A20 \\
\rowa \code{psram\_a[21]} & OUT & H20 & 2 & PSRAM A21 \\
\code{psram\_a[22]} & OUT & P18 & 3 & Sempre 0 (shift byte$\to$word): NC sulla scheda. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Bus PSRAM dati \code{psram\_dq[15:0]} --- 16 ball individuali (banchi 2 e 3)}}\\
\rowa \code{psram\_dq[0]} & IO & K18 & 2 & PSRAM DQ0 \\
\code{psram\_dq[1]} & IO & C18 & 2 & PSRAM DQ1 (dual-function, usata come GPIO ordinario). \\
\rowa \code{psram\_dq[2]} & IO & D17 & 2 & PSRAM DQ2 \\
\code{psram\_dq[3]} & IO & D20 & 2 & PSRAM DQ3 \\
\rowa \code{psram\_dq[4]} & IO & G19 & 2 & PSRAM DQ4 \\
\code{psram\_dq[5]} & IO & J18 & 2 & PSRAM DQ5 \\
\rowa \code{psram\_dq[6]} & IO & J19 & 2 & PSRAM DQ6 \\
\code{psram\_dq[7]} & IO & J20 & 2 & PSRAM DQ7 \\
\rowa \code{psram\_dq[8]} & IO & K19 & 2 & PSRAM DQ8 \\
\code{psram\_dq[9]} & IO & K20 & 2 & PSRAM DQ9 \\
\rowa \code{psram\_dq[10]} & IO & L17 & 3 & PSRAM DQ10 \\
\code{psram\_dq[11]} & IO & M18 & 3 & PSRAM DQ11 \\
\rowa \code{psram\_dq[12]} & IO & M17 & 3 & PSRAM DQ12 \\
\code{psram\_dq[13]} & IO & N16 & 3 & PSRAM DQ13 \\
\rowa \code{psram\_dq[14]} & IO & N18 & 3 & PSRAM DQ14 \\
\code{psram\_dq[15]} & IO & P17 & 3 & PSRAM DQ15 (bus dati bidirezionale tri-state, \code{dq\_oe} = direzione). \\
\multicolumn{5}{l}{\textit{\color{fnDark}Controllo PSRAM (banco 3)}}\\
\rowa \code{psram\_ce\_n} & OUT & N17 & 3 & Chip enable, attivo basso. \\
\code{psram\_oe\_n} & OUT & R16 & 3 & Output enable (lettura). \\
\rowa \code{psram\_we\_n} & OUT & R17 & 3 & Write enable (scrittura). \\
\code{psram\_lb\_n} & OUT & T16 & 3 & Lower-byte enable (DQ[7:0]). \\
\rowa \code{psram\_ub\_n} & OUT & N19 & 3 & Upper-byte enable (DQ[15:8]). \\
\code{psram\_zz\_n} & OUT & N20 & 3 & Sleep/snooze (inattivo=alto in funzionamento). \\
\bottomrule
\end{tabularx}
\renewcommand{\arraystretch}{1.25}
\begin{fnnote}[Segnali di scheda non esposti come porte RTL]
Non sono porte di \code{spi\_neuron\_top} ma vanno previsti a livello di scheda: le linee
di \textbf{SPI di configurazione} verso la flash NOR onboard (\code{PROGRAMN}/\code{INITN}/
\code{DONE}/\code{CCLK}\ldots, i ``Miscellaneous Dedicated Pins'' del datasheet) e le 4
linee \textbf{JTAG} (\code{TCK}/\code{TMS}/\code{TDI}/\code{TDO}), l'\textbf{oscillatore}
sul pad \code{PCLK}, le \textbf{alimentazioni}. I loro numeri di ball non sono nel datasheet
Lattice (file separato) ma non servono qui: sono pin dedicati senza porta RTL, nextpnr non
li richiede mai (0 errori), contano solo per lo schematic PCB.
\end{fnnote}
\begin{fnwarn}[SPI applicativo separato dallo SPI di configurazione]
L'SPI applicativo (\code{sclk/mosi/miso/cs\_n}) deve cadere su I/O ordinarie,
\textbf{mai} sui pin dell'SPI di configurazione: il pin di clock della config-SPI non è
riutilizzabile come ingresso generico dopo la configurazione senza workaround a livello
di scheda. Tenerli fisicamente separati evita quel problema.
\end{fnwarn}
\section{Allocazione per banchi (geometria reale del die)}
La collocazione segue la geometria dei bordi del die (da \code{globals.json} di Trellis,
ball~$\to$~(col,row)~$\to$~banco): i banchi \textbf{2 e 3} sono contigui lungo il bordo
\textbf{destro} del chip e ospitano insieme l'intero bus PSRAM (44+1 segnali) --- esattamente
gli ``uno o due banchi adiacenti'' raccomandati. Il banco \textbf{7} (bordo \textbf{sinistro},
fisicamente opposto al bus PSRAM) ospita SPI applicativo e clock/reset, deliberatamente sul
lato opposto per non far incrociare i due bus. \code{clk} è sul pad dedicato \code{H5}
(\code{GR\_PCLK7\_0}). Dove un banco esauriva le ball ``plain'' (parte di \code{psram\_dq}),
è stata usata la ball dual-function successiva come GPIO ordinario, confermata utilizzabile
dal place\&route reale.
\begin{tabularx}{\textwidth}{Y C{1.6cm} L{4.4cm}}
\toprule
\rowh \thd{Gruppo di segnali} & \thd{N. pin} & \thd{Banco (reale)} \\
\midrule
Indirizzi PSRAM \code{psram\_a[21:0]} & 22 & banco 2 (bordo destro) \\
\rowa Dati PSRAM \code{psram\_dq[15:0]} & 16 & banchi 2 + 3 (adiacenti) \\
Controllo PSRAM (ce/oe/we/lb/ub/zz) & 6 & banco 3 \\
\rowa SPI applicativo & 4 & banco 7 (bordo sinistro) \\
Pin attenzione host (\code{irq\_n}, \code{data\_ready\_n}) & 2 & banco 7 \\
\rowa Bus SPI flash indipendente (\code{flash\_sclk/flash\_mosi/flash\_miso/flash\_cs\_n}) & 4 & banco 7 \\
Clock / reset & 2 & banco 7, \code{clk} su \code{GR\_PCLK7\_0} \\
\rowa Config SPI boot / JTAG & --- & pin dedicati (fuori RTL, solo PCB) \\
\bottomrule
\end{tabularx}
\section{Sottosistema PSRAM}
Il controller \code{psram\_controller.v} implementa un'interfaccia \textbf{parallela
asincrona} (bus indirizzi, dati 16-bit, \code{ce\_n/oe\_n/we\_n} e byte-lane
\code{lb\_n/ub\_n}, più \code{zz\_n}) con latenza di accesso \textbf{70~ns} cablata come
$\lceil 70\,\text{ns}\times f_{clk}\rceil$. È un bus in stile SRAM asincrona, non QSPI.
\begin{tabularx}{\textwidth}{L{3.0cm}Y}
\toprule
\rowh \thd{Ruolo} & \thd{Componente} \\
\midrule
Memoria di lavoro & ISSI \code{IS66WVE4M16EBLL-70BLI} --- PSRAM parallela 64\,Mbit (4M$\times$16, 8~MB), async, 70~ns, corrispondente esatto alla temporizzazione del controller. \\
\rowa Fallback & ISSI \code{IS61WV6416DBLL} / \code{IS61WV102416BLL} (SRAM async vera, drop-in sugli stessi segnali, \code{zz\_n} inattivo, $\sim$10~ns, densità minore). \\
Storage persistente & Winbond \code{W25Q128JV} --- flash NOR SPI 16~MB per bitstream, pesi, bias, metadati di rete. \\
\bottomrule
\end{tabularx}
\subsection{Collegamento PSRAM (esclusivo della FPGA)}
La PSRAM è pilotata \textbf{esclusivamente dalla FPGA} tramite \code{psram\_controller.v}:
nessun master esterno accede al bus. L'host esterno (RPi/ESP32/MCU) parla solo SPI con la
FPGA e non tocca mai queste linee. Collegamento pin-per-pin FPGA~$\leftrightarrow$~ISSI
\code{IS66WVE4M16EBLL-70BLI}:
\begin{tabularx}{\textwidth}{L{3.6cm} L{3.0cm} Y}
\toprule
\rowh \thd{Segnale FPGA} & \thd{Pin PSRAM} & \thd{Funzione} \\
\midrule
\code{psram\_a[21:0]} & A0--A21 & Bus indirizzi (22 linee, 8~MB word address). \\
\rowa \code{psram\_dq[15:0]} & DQ0--DQ15 & Bus dati bidirezionale (tri-state, \code{dq\_oe}=direzione). \\
\code{psram\_ce\_n} & CE\# & Chip enable (attivo basso). \\
\rowa \code{psram\_oe\_n} & OE\# & Output enable (lettura). \\
\code{psram\_we\_n} & WE\# & Write enable (scrittura). \\
\rowa \code{psram\_lb\_n} & LB\# & Lower-byte enable (DQ[7:0]). \\
\code{psram\_ub\_n} & UB\# & Upper-byte enable (DQ[15:8]). \\
\rowa \code{psram\_zz\_n} & ZZ\# & Sleep/snooze (tenuto alto in funzionamento). \\
\bottomrule
\end{tabularx}
Alimentazione PSRAM: \textbf{3.3~V} (variante BLL), sullo stesso rail I/O dei banchi 2/3
a cui è cablata (cap.~\ref{ch:hw}, ball reali). Disaccoppiamento per pin di alimentazione
secondo il datasheet ISSI.
\section{Clock}
\label{sec:clock}
Non esiste ancora alcun PLL nell'RTL: \code{CLK\_FREQ\_MHZ} è un \emph{parametro di
temporizzazione} (alimenta le formule di accesso PSRAM), non un generatore di clock.
L'oscillatore montato pilota \code{clk} direttamente. Raccomandazione: oscillatore MEMS
16~MHz (famiglia SiTime SiT2001B), ben al di sotto dei 67.91~MHz di Fmax del sistema
integrato completo (incl. sottosistema flash, cap.~\ref{ch:impl}). \code{CLK\_FREQ\_MHZ} deve essere impostato al valore reale
dell'oscillatore montato, altrimenti la temporizzazione PSRAM risulta errata.
\section{Alimentazione}
Albero a \textbf{tre rail} (la sezione SERDES dell'eval board Lattice non serve e va
omessa: niente \code{VCCA}/\code{VCCHTX} a 1.2~V):
\begin{tabularx}{\textwidth}{L{3.4cm} C{2.0cm} Y}
\toprule
\rowh \thd{Rail} & \thd{Tensione} & \thd{Alimenta / regolatore} \\
\midrule
\code{VCC} (core) & 1.1~V & Core logico FPGA. Buck \code{TLV62568}, $\geq$600~mA. \\
\rowa \code{VCCIO0/2/3/6/7} & 3.3~V & I/O di tutti i banchi usati + PSRAM. Buck \code{TLV62568}, 1~A. \\
\code{VCCAUX} & 2.5~V & Ausiliario FPGA. LDO \code{TLV73325}, 10~mA. \\
\bottomrule
\end{tabularx}
Disaccoppiamento: almeno un condensatore per pin di alimentazione + bulk per rail, secondo
la checklist hardware ECP5 Lattice. Ingresso: 12~V esterno (o adatta i buck alla sorgente).
\section{Configurazione e programmazione}
\label{sec:config}
La ``scrittura della mappa'' dell'FPGA (bitstream) avviene tramite pin dedicati del
silicio, \textbf{non} porte del top-level RTL. Modo di default: \textbf{MSPI} --- boot
automatico dalla flash NOR all'accensione (prodotto standalone); JTAG disponibile per lo
sviluppo.
\subsection{JTAG (sviluppo / debug)}
\begin{tabularx}{\textwidth}{L{3.0cm} C{2.2cm} Y}
\toprule
\rowh \thd{Segnale} & \thd{Ball\textsuperscript{$\dagger$}} & \thd{Funzione} \\
\midrule
\code{TCK} & T5 & Test clock. \\
\rowa \code{TDI} & R5 & Test data in. \\
\code{TDO} & V4 & Test data out. \\
\rowa \code{TMS} & U5 & Test mode select. \\
\bottomrule
\end{tabularx}
\subsection{Config-SPI verso boot flash}
La FPGA carica il bitstream dalla \textbf{Winbond \code{W25Q128JV}} (128~Mbit SPI NOR,
Quad read) all'accensione. Il sottosistema flash (\code{rtl/flash\_slot\_manager.v}, Fasi
F1-F7, cap.~\ref{ch:impl}) usa la \textbf{stessa flash fisica} per pesi/bias/metadati di
rete a runtime, accesso esclusivo della FPGA: dopo la configurazione, la FPGA riprende il
controllo del chip via un bus SPI a 4 fili completamente indipendente,
\code{flash\_sclk/flash\_mosi/flash\_miso/flash\_cs\_n} (tutti GPIO ordinario, pagg.~2--3 e
§``Mappa dei segnali'' --- nessuna primitiva di configurazione ECP5 coinvolta, Fase F7) ---
implica comunque un doppio collegamento a livello di scheda (DI/DO/CS/CLK della flash
cablati sia ai pin dedicati di boot sotto sia a queste 4 ball ordinarie, poiché è lo stesso
chip fisico a svolgere entrambi i ruoli), non ancora riportato in uno schematico (nessuno
esiste ancora, vedi checklist sotto).
\begin{tabularx}{\textwidth}{L{3.4cm} C{2.2cm} Y}
\toprule
\rowh \thd{Segnale} & \thd{Ball\textsuperscript{$\dagger$}} & \thd{Funzione} \\
\midrule
\code{CCLK/MCLK/SCK} & U3 & Clock di configurazione. \\
\rowa \code{DQ0\_MOSI} & W2 & Dato config (MOSI). \\
\code{DQ1\_MISO} & V2 & Dato config (MISO). \\
\rowa \code{BUSY\_CSSPIN} & R2 & Chip-select flash. \\
\code{DQ2 / DQ3} & Y2 / W1 & Linee per Quad read. \\
\rowa \code{PROGRAMN} & W3 & Avvia riconfigurazione (pulsante, attivo basso). \\
\code{INITN} & V3 & Init / errore di configurazione (LED). \\
\rowa \code{DONE} & Y3 & Configurazione completata (LED). \\
\code{CFGMDN[2:0]} & R4/T4/U4 & Selezione modo (vedi sotto). \\
\bottomrule
\end{tabularx}
\subsection{Modi di configurazione (\texttt{CFGMDN})}
\begin{tabularx}{\textwidth}{L{4.0cm} C{4.0cm} Y}
\toprule
\rowh \thd{Modo} & \thd{CFGMDN[2:0]} & \thd{Uso} \\
\midrule
MSPI (boot da flash) & \code{010} & \textbf{Default} --- standalone. \\
\rowa SSPI (slave SPI) & \code{001} & Config da host esterno. \\
SCM (slave serial) & \code{101} & Config seriale. \\
\rowa SPCM (slave parallel) & \code{111} & Config parallela 8-bit. \\
\bottomrule
\end{tabularx}
\begin{fnwarn}[Ball di configurazione da verificare sul 45F]
\textsuperscript{$\dagger$}Le ball di JTAG e config-SPI qui riportate sono il
\emph{riferimento} dell'eval board Lattice (device 85F). JTAG e config-SPI sono pin
dedicati e in gran parte fissi nella famiglia ECP5, ma le posizioni esatte sul target
\code{LFE5U-45F-8BG381C} vanno confermate sul file pinout Lattice del 45F (Diamond/Radiant
o database Trellis) prima di committarle nello schematico, come già fatto per i segnali
applicativi (cap.~\ref{ch:hw}).
\end{fnwarn}
\section{Attività aperte prima della cattura schematica}
\begin{itemize}
\item[\OK] \code{ADDR\_WIDTH}=23 (8~MB pieni) su tutti i moduli e testbench.
\item[\OK] \code{.lpf} reale con l'assegnazione ball CABGA381, place\&route-verified a
0 errori (\code{synth/ecp5/spi\_neuron\_top.lpf}, 57 segnali incl. sottosistema flash).
\item[\OK] Sottosistema flash boot/persistenza (Fasi F1-F7): SPI master, copy engine,
catalogo a slot con CRC32, bus SPI a 4 fili indipendente (nessuna primitiva di
configurazione condivisa), sintesi reale a 0 errori, Fmax 67.91~MHz (\code{WORKLOG.md}).
\item[$\square$] Confermare signal integrity PSRAM/SPI al clock effettivamente montato.
\item[$\square$] Schema di doppio collegamento DI/DO/CS/CLK della flash (pin dedicati di
boot + le 4 ball ordinarie del sottosistema flash) --- non ancora catturato a
schematico.
\item[$\square$] Scelta del footprint del connettore JTAG.
\item[$\square$] Cattura schematica (KiCad o altro): nessuno schema esiste ancora per
questa combinazione dispositivo/package.
\end{itemize}
@@ -1,81 +0,0 @@
\chapter{Riferimento rapido}
\label{ch:ref}
\section{Opcode SPI}
\begin{tabularx}{\textwidth}{C{1.4cm} L{3.2cm} C{2.4cm} Y}
\toprule
\rowh \thd{Valore} & \thd{Nome} & \thd{Risposta} & \thd{Sintesi} \\
\midrule
\op{0x00} & NOP & --- & idle \\
\rowa \op{0x01} & WRITE\_RAM & --- & scrittura blocco PSRAM \\
\op{0x02} & READ\_RAM & \code{len} B & lettura blocco PSRAM \\
\rowa \op{0x0F} & RESET & --- & reset motore + latch STATUS \\
\op{0x10} & SET\_BASE & --- & imposta base/registro (sel 0..10) \\
\rowa \op{0x11} & SET\_NET\_TYPE & --- & tipo rete \#1/\#2 \\
\rowa \op{0x20} & START & --- & avvio single-layer \\
\op{0x21} & STATUS & 1 B & busy(live)/done(sticky) \\
\rowa \op{0x22} & READ\_OUTPUT & N\_NEURONS B & \code{y\_bus} \\
\op{0x23} & RUN\_NETWORK & --- & avvio multi-layer \\
\rowa \op{0x30} & READ\_CONFIG & 11 B & record configurazione \\
\bottomrule
\end{tabularx}
\section{Byte di STATUS}
\begin{center}
\begin{tikzpicture}[font=\scriptsize]
\foreach \i/\lbl [count=\x from 0] in {7/0,6/0,5/0,4/0,3/0,2/0,1/{done},0/{busy}}{
\node[fnreg,minimum width=13mm,minimum height=9mm] (b\x) at (\x*13mm,0) {\lbl};
\node[font=\tiny,text=fnGrey,above=0.5mm of b\x] {bit \i};
}
\node[fill=fnAmber,text=white,rounded corners=1pt,inner sep=1.5pt,font=\tiny]
at (b7.center){riservati = 0};
\node[fill=fnTeal,text=white,rounded corners=1pt,inner sep=1.5pt,font=\tiny]
at (b6.center){};
\end{tikzpicture}
\end{center}
\code{done} è sticky, clear-on-read; \code{busy} è live; \code{bit2=err} (guard grafo).
\section{Selettori SET\_BASE}
\begin{multicols}{2}\footnotesize
\begin{itemize}
\item 0 --- \code{x\_base}
\item 1 --- \code{w\_base}
\item 2 --- \code{bias\_addr}
\item 3 --- \code{table\_base}
\item 4 --- \code{buf\_a\_base}
\columnbreak
\item 5 --- \code{buf\_b\_base}
\item 6 --- \code{activation} (single-layer)
\item 7 --- \code{n\_inputs\_real} (single-layer)
\item 8 --- \code{n\_neurons\_real} (single-layer)
\item 9 --- \code{num\_neurons\_graph} (Tipo \#2)
\item 10 --- \code{n\_out} (Tipo \#2)
\end{itemize}
\end{multicols}
\section{Tabella descrittori (11 byte/layer, MSB-first)}
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=0mm]
\node[fnreg,minimum width=20mm,minimum height=8mm](a){\code{w\_base}\\3B};
\node[fnreg,minimum width=20mm,minimum height=8mm,right=0mm of a](b){\code{bias\_addr}\\3B};
\node[fnreg,minimum width=14mm,minimum height=8mm,right=0mm of b](c){\code{act}\\1B};
\node[fnreg,minimum width=22mm,minimum height=8mm,right=0mm of c](d){\code{n\_inputs\_real}\\2B};
\node[fnreg,minimum width=22mm,minimum height=8mm,right=0mm of d](e){\code{n\_neurons\_real}\\2B};
\end{tikzpicture}
\end{center}
\section{Parametri di build}
\begin{multicols}{2}\footnotesize
\begin{itemize}
\item \code{DATA\_WIDTH} --- 8 (INT8)
\item \code{ACC\_WIDTH} --- 32 (INT32)
\item \code{N\_INPUTS} --- max ingressi
\item \code{N\_NEURONS} --- max neuroni
\item \code{PARALLEL} --- MAC simultanei
\columnbreak
\item \code{N\_LAYERS} --- max layer
\item \code{ADDR\_WIDTH} --- 23 (8 MB)
\item \code{MEM\_DATA\_WIDTH} --- 16
\item \code{CLK\_FREQ\_MHZ} --- timing PSRAM
\end{itemize}
\end{multicols}
@@ -1,65 +0,0 @@
\chapter{Roadmap e stato di sviluppo}
\label{ch:roadmap}
\section{Fasi di sviluppo}
\begin{tabularx}{\textwidth}{C{1.2cm} L{4.6cm} C{1.8cm} Y}
\toprule
\rowh \thd{Fase} & \thd{Titolo} & \thd{Stato} & \thd{Contenuto} \\
\midrule
1 & Layer parametrico & \OK & ingressi/neuroni/parallelismo, accumulo, bias, ReLU; test 32$\times$4/P=8. \\
\rowa 2 & Parameter sweep & \OK & configurazioni multiple incl. non-multiple e degeneri; guard di elaborazione aggiunto. \\
3 & Architettura di memoria & \OK & \code{neuron\_memory} mono/multi-neurone, PSRAM reale testata; buffer multi-layer $\to$ Fase~5. \\
\rowa 4 & Interfaccia SPI & \OK & \code{spi\_slave}+\code{spi\_engine}, 17 opcode incl. sottosistema flash, Fmax controllata a livello di sistema completo. \\
5 & Rete multi-layer & \OK$^\dagger$ & \code{layer\_sequencer}, attivazioni configurabili, larghezza runtime; toolchain reale controllata. \\
\rowa 6 & Software host & pianificata & driver Linux ed ESP32 sullo stesso protocollo. \\
7 & Ottimizzazione & in corso & timing closure fatta (55$\to$75~MHz); page-mode PSRAM fatto (banda gather +42\%); resta block RAM per $x$/$w$. \\
\rowa 8 & Training hardware (opz.) & futura & backprop, gradienti, aggiornamento pesi. \\
9 & Sottosistema flash (F1-F7) & \OK & SPI master dedicato, copy engine flash$\leftrightarrow$PSRAM, catalogo a 16 slot con CRC32, bus SPI a 4 fili indipendente (F7), 8 opcode (\op{0x40}--\op{0x47}, cap.~\ref{ch:spi} §\ref{sec:flashspi}); sintesi reale 0 errori. \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
$\dagger$ RTL, unit test ed end-to-end su SPI simulato completi; timing closure eseguita:
75.30~MHz (P2) / 60.26~MHz (P8) al tempo della Fase~5, bit-esatta su tutta la regressione;
Fmax del sistema completo dopo Fase~9 (incl. sottosistema flash indipendente): \textbf{67.91~MHz}
(cap.~\ref{ch:impl}).\end{center}
\section{Stato dei componenti}
\begin{tabularx}{\textwidth}{Y C{4.2cm}}
\toprule
\rowh \thd{Componente} & \thd{Stato} \\
\midrule
Layer neurale parametrico & \OK{} funzionante \\
\rowa Ingressi/neuroni/parallelismo parametrici & \OK \\
Accumulo, bias, ReLU & \OK \\
\rowa Validazione 32$\times$4 / P=8 & \OK \\
RAM dedicata (interfaccia + controller + accesso INT8) & \OK{} testata su PSRAM reale \\
\rowa Interfaccia SPI (17 opcode incl. RUN\_NETWORK + flash) & \OK{} Fmax a livello di sistema completo \\
Dual SPI & futura \\
\rowa Motore multi-layer & \OK{} timing closure 75.30~MHz (P2) al tempo della Fase~5 \\
Attivazioni configurabili (ACT\_NONE/ACT\_RELU) & \OK \\
\rowa Larghezza rete runtime (un bitstream, ogni topologia) & \OK{} risparmio misurato \\
Rete a grafo Tipo \#2 (act\_buffer, graph\_engine, netasm) & \OK{} RTL + test + sintesi \\
\rowa Pinout CABGA381 (\code{.lpf} reale, 57 segnali incl. flash) & \OK{} place\&route-verified 0 errori \\
Page-mode PSRAM (G7) & \OK{} fatto (37.53 cicli/edge, banda +42\%) \\
\rowa Sottosistema flash (SPI master, copy engine, catalogo CRC32, bus indipendente F7) & \OK{} sintesi reale 0 errori, Fmax 67.91~MHz \\
Bitstream reale (\code{ecppack}, P2/P8) & \OK{} 0 errori, part LFE5U-45F-8CABGA381 \\
\rowa Driver host Linux / ESP32 & pianificato \\
Training hardware & futuro \\
\bottomrule
\end{tabularx}
\section{Principio architetturale (sintesi)}
\begin{fnspec}[Fondamento del progetto]
L'FPGA implementa la macchina neurale e possiede la propria RAM; l'host configura e usa
la macchina. Una build fissa il \emph{soffitto} (max layer, max larghezza, PARALLEL);
l'host configura la rete \emph{reale} --- numero di layer, larghezza per-layer,
attivazione per-layer, parametri addestrati --- interamente a runtime, via SPI, nella
memoria locale dell'FPGA. Un solo bitstream serve qualunque topologia fino a quel
soffitto.
\end{fnspec}
\section{Visione a lungo termine}
L'obiettivo finale è un blocco hardware riusabile integrabile in progetti futuri
diversi: la piattaforma host può cambiare (Linux, ESP32, MCU, PC) senza cambiare
l'architettura fondamentale dell'engine. L'FPGA diventa una periferica di computazione
neurale dedicata, ottimizzata per la topologia richiesta da ciascuna applicazione.
@@ -1,97 +0,0 @@
\chapter[Moduli e toolchain]{Moduli, porte e toolchain}
\label{ch:appmod}
\section{Elenco dei moduli RTL}
\begin{tabularx}{\textwidth}{L{3.4cm} C{2.0cm} Y}
\toprule
\rowh \thd{File} & \thd{Tipo} & \thd{Ruolo} \\
\midrule
\code{rtl/mac\_unit.v} & combinatorio & prodotto-accumulatore singolo \\
\rowa \code{rtl/mac8.v} & combinatorio & MAC parallelo + adder tree bilanciato \\
\code{rtl/neuron\_parallel.v} & FSM & neurone: gruppi, bias, attivazione, saturazione \\
\rowa \code{rtl/layer.v} & strutturale & N\_NEURONS neuroni in parallelo \\
\code{rtl/neuron\_memory.v} & FSM & ponte memoria/neurone, loop neuroni \\
\rowa \code{rtl/layer\_sequencer.v} & FSM & sequenza multi-layer, ping-pong \\
\code{rtl/int8\_memory\_access.v} & FSM & conversione byte $\leftrightarrow$ word \\
\rowa \code{rtl/memory\_interface.v} & FSM & handshake req/ready \\
\code{rtl/psram\_controller.v} & FSM & bus fisico PSRAM async, page mode 70/20~ns \\
\rowa \code{rtl/mem\_arbiter.v} & arbitro & 3 porte, priorità B$>$C$>$A \\
\code{rtl/spi\_slave.v} & FSM & layer fisico SPI Mode 0 + CDC \\
\rowa \code{rtl/spi\_engine.v} & FSM & opcode + banco registri \\
\code{rtl/act\_buffer.v} & block RAM & buffer di attivazione DP16KD (Tipo \#2) \\
\rowa \code{rtl/graph\_engine.v} & FSM & motore rete a grafo (Tipo \#2) \\
\code{rtl/spi\_neuron\_top.v} & top & integrazione completa \\
\rowa \code{rtl/memory\_model.v} & modello & RAM comportamentale (sim) \\
\bottomrule
\end{tabularx}
\section{Porte del top-level \texttt{spi\_neuron\_top}}
Vedere la tabella segnale-per-segnale completa nel cap.~\ref{ch:hw}. In sintesi: clock
e reset (\code{clk}, \code{rst}); SPI applicativo (\code{sclk}, \code{mosi},
\code{miso}, \code{cs\_n}); bus PSRAM (\code{psram\_a[22:0]}, \code{psram\_dq[15:0]},
\code{psram\_ce\_n/oe\_n/we\_n/lb\_n/ub\_n/zz\_n}).
\section{Toolchain}
\begin{tabularx}{\textwidth}{L{3.6cm} L{3.4cm} Y}
\toprule
\rowh \thd{Strumento} & \thd{Versione} & \thd{Uso} \\
\midrule
Yosys & 0.68+post & sintesi RTL $\to$ netlist JSON, mapping ECP5 \\
\rowa nextpnr-ecp5 & 0.11.1-19-g8dbcee5 & placement, routing, timing \\
Project Trellis & install & \code{ecppack}/\code{ecppll}/\code{ecpbram} \\
\rowa Icarus Verilog & \code{-g2012} & simulazione funzionale \\
\bottomrule
\end{tabularx}
\subsection{Parametri nextpnr principali}
\begin{lstlisting}[language=,basicstyle=\ttfamily\scriptsize]
--45k seleziona LFE5U-45F
--package CABGA381 package
--speed 8 speed grade -8
--json <netlist> netlist da Yosys
--lpf <vincoli> vincoli di pin (attualmente vuoti)
--lpf-allow-unconstrained permette I/O non vincolate (benchmark)
--freq 80 timing target 80 MHz
\end{lstlisting}
\subsection{Esempio di simulazione}
\begin{lstlisting}[language=,basicstyle=\ttfamily\scriptsize]
iverilog -g2012 -Ptb.PARALLEL=16 -o sim/parametric_256x4_p16 \
sim/parametric_tb.v rtl/mac_unit.v rtl/mac8.v \
rtl/neuron_parallel.v rtl/layer.v
vvp sim/parametric_256x4_p16
\end{lstlisting}
\section{Testbench principali}
\begin{tabularx}{\textwidth}{L{5.4cm} Y}
\toprule
\rowh \thd{Testbench} & \thd{Copertura} \\
\midrule
\code{parametric\_tb.v} & datapath 256$\times$4, casi accumulo/bias/ReLU/saturazione \\
\rowa \code{parameter\_sweep\_tb.v} & sweep configurazioni valide \\
\code{neuron\_parallel\_tb.v} & attivazioni, larghezza runtime (T7) \\
\rowa \code{neuron\_memory\_tb.v} / \code{\_multi\_tb.v} & integrazione memoria mono/multi-neurone, PSRAM reale (T5) \\
\code{psram\_controller\_tb.v} & controller PSRAM \\
\rowa \code{psram\_page\_mode\_tb.v} & burst di pagina, attraversamento pagina, chiusura su WRITE/timeout $t_{CEM}$, cambi di byte-enable (§~5.5) \\
\code{spi\_slave\_tb.v} & layer fisico SPI (4 test) \\
\rowa \code{spi\_engine\_tb.v} & opcode, registri (10+ test) \\
\code{spi\_neuron\_top\_tb.v} & end-to-end, PSRAM reale su SPI simulato \\
\rowa \code{spi\_neuron\_top\_runnetwork\_tb.v} & RUN\_NETWORK 2 layer end-to-end \\
\code{layer\_sequencer\_tb.v} & sequenza 2 layer, ping-pong, copia byte-exact \\
\bottomrule
\end{tabularx}
\vfill
\begin{center}
\begin{tikzpicture}
\node[draw=fnRule,rounded corners=3pt,inner sep=8pt,fill=fnLight,text width=15.5cm]{
\footnotesize\color{fnGrey}
Questo datasheet è generato a partire dal codice RTL, dalla documentazione e dai
benchmark presenti nella repository \texttt{github.com/manvalan/FPGA-Neural} allo stato
del \datasheetdate. I valori di Fmax, utilizzo risorse e throughput sono quelli
riportati nelle misure della repository (\texttt{.lpf} reale già assegnato e
verificato da place\&route, cap.~\ref{ch:hw}) e vanno riverificati ad ogni
variazione sostanziale dell'RTL o della chiusura del timing di Fase~7, tuttora in
corso (cap.~\ref{ch:roadmap}).};
\end{tikzpicture}
\end{center}
Binary file not shown.
@@ -1,119 +0,0 @@
% ======================================================================
% FPGA-Neural -- INT8 Neural Network Engine
% Datasheet / Technical reference manual
% Repository: github.com/manvalan/FPGA-Neural
% ======================================================================
\documentclass[11pt,a4paper,openany]{report}
\newcommand{\datasheetrev}{A1}
\newcommand{\datasheetdate}{September 2026}
\input{preamble}
\begin{document}
\sloppy
% ======================================================================
% TITLE PAGE
% ======================================================================
\begin{titlepage}
\thispagestyle{empty}
\begin{tikzpicture}[remember picture,overlay]
\fill[fnDark] (current page.north west) rectangle
([yshift=-4.3cm]current page.north east);
\fill[fnTeal] ([yshift=-4.3cm]current page.north west) rectangle
([yshift=-4.55cm]current page.north east);
\node[anchor=north west,text=white,font=\Huge\bfseries]
at ([xshift=2.2cm,yshift=-1.15cm]current page.north west)
{FPGA\,--\,Neural};
\node[anchor=north west,text=fnLight,font=\large]
at ([xshift=2.25cm,yshift=-2.15cm]current page.north west)
{INT8 Neural Network Engine for FPGA};
\node[anchor=north west,text=fnLight2,font=\normalsize]
at ([xshift=2.25cm,yshift=-2.85cm]current page.north west)
{Parametric hardware accelerator -- Datasheet and reference manual};
\node[anchor=north east,text=white,font=\ttfamily\small]
at ([xshift=-2.2cm,yshift=-3.55cm]current page.north east)
{Rev.~\datasheetrev~~\textbullet~~\datasheetdate};
\end{tikzpicture}
\vspace*{5.0cm}
% --- compact block diagram on the title page ---
\begin{center}
\begin{tikzpicture}[node distance=7mm and 12mm]
\node[fnblockD,minimum width=30mm] (host) {HOST\\{\scriptsize Linux / ESP32 / MCU / PC}};
\node[fnblockT,right=18mm of host,minimum width=34mm] (fpga)
{FPGA\\{\scriptsize Neural Network Engine}};
\node[fnblock,right=18mm of fpga,minimum width=26mm] (ram)
{PSRAM\\{\scriptsize 8\,MB dedicated}};
\draw[fnbus] (host) -- node[fnlbl,above]{SPI Mode 0} (fpga);
\draw[fnbus] (fpga) -- node[fnlbl,above]{async 16-bit} (ram);
\node[below=1mm of fpga,font=\scriptsize\itshape,text=fnGrey]
{computation entirely on-chip};
\end{tikzpicture}
\end{center}
\vfill
\begin{center}
\begin{tikzpicture}
\node[draw=fnRule,rounded corners=3pt,inner sep=10pt,fill=fnLight,text width=15.5cm]{
\footnotesize
\textbf{\color{fnDark}Reference target device:} Lattice ECP5 \code{LFE5U-45F-8BG381C}
(speed grade $-8$, CABGA381, 72$\times$MULT18X18D, $\approx$44k LUT).\\[2pt]
\textbf{\color{fnDark}Baseline configuration:} INT8/INT32, \code{N\_INPUTS}=256, \code{N\_NEURONS}=4,
parametric \code{PARALLEL}, PSRAM working memory ISSI \code{IS66WVE4M16EBLL-70BLI}.\\[2pt]
\textbf{\color{fnDark}Status:} RTL verified in simulation (Icarus) and real synthesis
(Yosys + nextpnr-ecp5). Document describing the project as of \datasheetdate.
};
\end{tikzpicture}
\end{center}
\vspace{0.6cm}
{\footnotesize\color{fnGrey}\raggedright
Project author: Michele Bigi \textbullet{} MIKILAB / manvalan.\\
This datasheet documents the RTL code, documentation and benchmarks
present in the repository \texttt{github.com/manvalan/FPGA-Neural}.\par}
\end{titlepage}
% ======================================================================
% "FEATURES" PAGE (datasheet style)
% ======================================================================
\input{chapters/00-features}
% ======================================================================
% PINOUT SUMMARY (pages 2-3, pin-by-pin -- not bus ranges)
% ======================================================================
\newpage
\input{chapters/00b-pinout}
% ======================================================================
% TABLE OF CONTENTS
% ======================================================================
\newpage
\pagenumbering{roman}
{\color{fnDark}\tableofcontents}
\newpage
\pagenumbering{arabic}
% ======================================================================
% CHAPTERS
% ======================================================================
\include{chapters/01-overview}
\include{chapters/02-architettura}
\include{chapters/03-datapath}
\include{chapters/04-parametri}
\include{chapters/05-memoria}
\include{chapters/06-sequencer}
\include{chapters/06b-grafo}
\include{chapters/07-spi}
\include{chapters/07b-programmazione}
\include{chapters/08-toplevel}
\include{chapters/09-implementazione}
\include{chapters/10-hardware}
\include{chapters/11-registri}
\include{chapters/12-roadmap}
\appendix
\include{chapters/A-moduli}
\end{document}
@@ -1,120 +0,0 @@
\thispagestyle{plain}
\noindent
\begin{tikzpicture}
\node[fill=fnDark,text=white,rounded corners=2pt,inner sep=6pt,
minimum width=\textwidth,anchor=west]
{\large\bfseries FPGA-Neural --- General description and features};
\end{tikzpicture}
\vspace{6pt}
\noindent
{\small FPGA-Neural is a \textbf{parametric hardware accelerator for feed-forward
neural networks} contained entirely within the FPGA. Computation (multiplication,
accumulation, bias, activation, saturation) takes place entirely on-chip in INT8/INT32
integer arithmetic; the host system only provides configuration, weights, input data
and control through a simple SPI interface, without ever being part of the
computational datapath. A single bitstream serves any topology up to the build
maximum.}
\vspace{8pt}
\begin{multicols}{2}
{\color{fnDark}\large\bfseries Features}\\[2pt]
{\footnotesize
\begin{itemize}[leftmargin=1.1em]
\item \textbf{INT8 $\times$ INT8 $\to$ INT16 $\to$ INT32} datapath, 32-bit accumulation
with sign extension.
\item \textbf{Balanced binary adder tree} ($O(\log_2 \text{PARALLEL})$) instead of
linear reduction.
\item Configurable parallel MAC: \code{PARALLEL} simultaneous hardware MACs per neuron,
mapped onto \code{MULT18X18D} DSPs.
\item Fully \textbf{parametric} architecture: \code{N\_INPUTS}, \code{N\_NEURONS},
\code{PARALLEL}, \code{DATA\_WIDTH}, \code{ACC\_WIDTH}, \code{N\_LAYERS}.
\item \textbf{Runtime network width}: per-layer \code{n\_inputs\_real}/\code{n\_neurons\_real},
a single bitstream for every topology up to the maximum.
\item Configurable activations: \code{ACT\_RELU} (default) and \code{ACT\_NONE} (linear
with bilateral saturation), with INT8 saturation.
\item \textbf{Two network types}: classic multi-layer dense (\code{layer\_sequencer},
ping-pong buffers) and \textbf{arbitrary sparse graph} (\code{graph\_engine} +
activation buffer in \code{DP16KD} block RAM), selectable at runtime.
\item \textbf{Dedicated memory} subsystem: byte$\leftrightarrow$word interface,
asynchronous parallel PSRAM controller with \textbf{page mode} (70~ns random
access, 20~ns page burst), 8~MB addressable (23~bit).
\item \textbf{SPI Mode 0} MSB-first host interface, \code{SET\_NET\_TYPE}+dispatch, \code{STATUS.done}
sticky/clear-on-read, runtime \code{READ\_CONFIG}.
\item \textbf{Flash subsystem} for boot/persistence: FPGA-exclusive access to a
\code{W25Q128JV} SPI NOR (16~MB) via a dedicated SPI master, a
flash$\leftrightarrow$PSRAM copy engine, and a 16-slot catalog with CRC32,
8 host opcodes.
\item Verified in \textbf{simulation} (Icarus Verilog) and \textbf{real synthesis}
(Yosys + nextpnr-ecp5 + ecppack).
\end{itemize}}
\columnbreak
{\color{fnDark}\large\bfseries Applications}\\[2pt]
{\footnotesize
\begin{itemize}[leftmargin=1.1em]
\item Deterministic low-latency inference as a peripheral of a
Linux SoC, Raspberry-Pi-like board, ESP32, microcontrollers.
\item Reusable hardware block integrable into heterogeneous projects
(a platform, not a single network).
\item Edge AI on compact dense INT8-quantized networks.
\item Off-loading the neural workload from the host CPU to dedicated
hardware with predictable throughput.
\end{itemize}}
\vspace{4pt}
{\color{fnDark}\large\bfseries Target \& toolchain}\\[2pt]
{\footnotesize
\begin{itemize}[leftmargin=1.1em]
\item FPGA: Lattice ECP5 \code{LFE5U-45F-8BG381C} ($-8$, CABGA381).
\item Synthesis: Yosys; place\&route: nextpnr-ecp5; bitstream: Project~Trellis
(\code{ecppack}).
\item Simulation: Icarus Verilog (\code{-g2012}).
\item PSRAM: ISSI \code{IS66WVE4M16EBLL-70BLI} (64\,Mb, 4M$\times$16).
\end{itemize}}
\end{multicols}
\vspace{2pt}
% --- key parameter table ---
\noindent
{\small\color{fnDark}\bfseries Key parameters (characterized baseline configuration)}
\vspace{2pt}
\noindent
\begin{tabularx}{\textwidth}{L{3.2cm}L{3.6cm}Y}
\toprule
\rowh \thd{Quantity} & \thd{Value} & \thd{Notes} \\
\midrule
Data precision & INT8 (signed) & \code{DATA\_WIDTH}=8 \\
\rowa Accumulator & INT32 (signed) & \code{ACC\_WIDTH}=32 \\
Inputs / neurons & 256 / 4 & datapath benchmark baseline \\
\rowa Simultaneous MACs & $2\ldots64$ & $=$\code{PARALLEL}$\times$\code{N\_NEURONS} \\
Activations & ReLU, linear & \code{ACT\_RELU} / \code{ACT\_NONE} \\
\rowa Fmax (P=2, datapath) & 87.88~MHz & isolated datapath benchmark \\
Fmax (P=2, integrated system) & 67.91~MHz & full system incl. flash subsystem, real place\&route \\
MAC throughput (P=16) & $\approx$3.34~G\,MAC/s & theoretical, datapath only \\
\rowa Working memory & 8~MB PSRAM & 16-bit parallel bus, 70~ns / 20~ns page mode \\
Address space & 23~bit (byte) & \code{ADDR\_WIDTH}=23 \\
\bottomrule
\end{tabularx}
\vspace{8pt}
\noindent
{\small\color{fnDark}\bfseries System block diagram}
\begin{center}
\begin{tikzpicture}[node distance=6mm and 10mm,font=\footnotesize]
\node[fnblockD,minimum width=26mm,minimum height=13mm] (host){HOST\\{\scriptsize configures / trains / controls}};
\node[fnblockT,right=16mm of host,minimum width=52mm,minimum height=22mm] (eng){};
\node[anchor=north,font=\footnotesize\bfseries,text=fnDark] at (eng.north){FPGA -- Neural Network Engine};
\node[fnreg,fill=white] (spi) at ([yshift=-2mm]eng.center){\code{spi\_slave} + \code{spi\_engine}};
\node[fnreg,fill=white,below=2.5mm of spi] (arb){\code{mem\_arbiter} + \code{layer\_sequencer}};
\node[fnreg,fill=white,above=2.5mm of spi] (core){\code{neuron\_memory} $\to$ \code{neuron\_parallel} $\to$ \code{mac8}};
\node[fnblock,right=16mm of eng,minimum width=24mm,minimum height=13mm] (ram){PSRAM 8\,MB\\{\scriptsize \code{psram\_controller}}};
\draw[fnbus] (host) -- node[fnlbl,above]{SPI} (eng.west|-host);
\draw[fnbus] (eng.east|-ram) -- node[fnlbl,above]{16-bit async} (ram);
\end{tikzpicture}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
The neural datapath is entirely inside the FPGA; the host does not take part in the
individual MAC operations.\end{center}
@@ -1,103 +0,0 @@
\thispagestyle{plain}
\noindent
\begin{tikzpicture}
\node[fill=fnDark,text=white,rounded corners=2pt,inner sep=6pt,
minimum width=\textwidth,anchor=west]
{\large\bfseries Pinout summary --- pin-by-pin connection};
\end{tikzpicture}
\vspace{6pt}
\noindent
{\footnotesize
Quick-reference table: the \textbf{57 real signals} of the top-level
\code{spi\_neuron\_top}, each with its own individual \code{CABGA381} ball
(\textbf{not} a bus range) --- real data from Project~Trellis's device
database (\code{iodb.json}), \textbf{verified by a complete
\code{nextpnr-ecp5} place\&route run at 0 errors} (not a planned pinout).
Full description, per-bank placement rationale and the pin-by-pin
connection to the ISSI PSRAM: ch.~\ref{ch:hw}.
}
\vspace{4pt}
\noindent
\renewcommand{\arraystretch}{1.08}
\begin{tabularx}{\textwidth}{L{2.7cm} C{1.0cm} C{1.0cm} C{0.9cm} Y}
\toprule
\rowh \thd{Signal} & \thd{Ball} & \thd{Bank} & \thd{Dir} & \thd{Corresponding pin / function} \\
\midrule
\multicolumn{5}{l}{\textit{\color{fnDark}Clock and reset}}\\
\code{clk} & H5 & 7 & IN & System clock, pad \code{GR\_PCLK7\_0} (dedicated global clock). \\
\rowa \code{rst} & B4 & 7 & IN & Global synchronous reset, active high. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Application SPI (host $\leftrightarrow$ FPGA, Mode~0)}}\\
\code{sclk} & B5 & 7 & IN & SPI clock (CPOL=0, CPHA=0). \\
\rowa \code{mosi} & C5 & 7 & IN & Master-Out Slave-In. \\
\code{miso} & A3 & 7 & OUT & Master-In Slave-Out. \\
\rowa \code{cs\_n} & B3 & 7 & IN & Chip-select, active low. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Host attention (active-low, level)}}\\
\code{data\_ready\_n} & C3 & 7 & OUT & Low while a result is waiting to be read. \\
\rowa \code{irq\_n} & C4 & 7 & OUT & Low while the graph engine's load-time guard has tripped. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Flash subsystem --- SPI toward W25Q128JV (boot/persistence)}}\\
\code{flash\_sclk} & E3 & 7 & OUT & SPI clock toward the flash --- ordinary GPIO, independent (Phase F7, ch.~\ref{ch:hw}). \\
\rowa \code{flash\_mosi} & D3 & 7 & OUT & Master-Out Slave-In toward the onboard flash. \\
\code{flash\_miso} & D5 & 7 & IN & Master-In Slave-Out from the flash. \\
\rowa \code{flash\_cs\_n} & E4 & 7 & OUT & Flash chip-select, active low. \\
\multicolumn{5}{l}{\textit{\color{fnDark}PSRAM address bus \code{psram\_a[21:0]} --- 22 individual balls (bank 2)}}\\
\code{psram\_a[0]} & E16 & 2 & OUT & PSRAM A0 \\
\rowa \code{psram\_a[1]} & F16 & 2 & OUT & PSRAM A1 \\
\code{psram\_a[2]} & D18 & 2 & OUT & PSRAM A2 \\
\rowa \code{psram\_a[3]} & E17 & 2 & OUT & PSRAM A3 \\
\code{psram\_a[4]} & E18 & 2 & OUT & PSRAM A4 \\
\rowa \code{psram\_a[5]} & F18 & 2 & OUT & PSRAM A5 \\
\code{psram\_a[6]} & F17 & 2 & OUT & PSRAM A6 \\
\rowa \code{psram\_a[7]} & G16 & 2 & OUT & PSRAM A7 \\
\code{psram\_a[8]} & G18 & 2 & OUT & PSRAM A8 \\
\rowa \code{psram\_a[9]} & H16 & 2 & OUT & PSRAM A9 \\
\code{psram\_a[10]} & H17 & 2 & OUT & PSRAM A10 \\
\rowa \code{psram\_a[11]} & H18 & 2 & OUT & PSRAM A11 \\
\code{psram\_a[12]} & J16 & 2 & OUT & PSRAM A12 \\
\rowa \code{psram\_a[13]} & J17 & 2 & OUT & PSRAM A13 \\
\code{psram\_a[14]} & C20 & 2 & OUT & PSRAM A14 \\
\rowa \code{psram\_a[15]} & D19 & 2 & OUT & PSRAM A15 \\
\code{psram\_a[16]} & E19 & 2 & OUT & PSRAM A16 \\
\rowa \code{psram\_a[17]} & E20 & 2 & OUT & PSRAM A17 \\
\code{psram\_a[18]} & F19 & 2 & OUT & PSRAM A18 \\
\rowa \code{psram\_a[19]} & F20 & 2 & OUT & PSRAM A19 \\
\code{psram\_a[20]} & G20 & 2 & OUT & PSRAM A20 \\
\rowa \code{psram\_a[21]} & H20 & 2 & OUT & PSRAM A21 \\
\code{psram\_a[22]} & P18 & 3 & OUT & Always 0 (byte$\to$word shift): NC on the board. \\
\multicolumn{5}{l}{\textit{\color{fnDark}PSRAM data bus \code{psram\_dq[15:0]} --- 16 individual balls (banks 2 and 3)}}\\
\rowa \code{psram\_dq[0]} & K18 & 2 & IO & PSRAM DQ0 \\
\code{psram\_dq[1]} & C18 & 2 & IO & PSRAM DQ1 (dual-function ball, used as ordinary GPIO). \\
\rowa \code{psram\_dq[2]} & D17 & 2 & IO & PSRAM DQ2 \\
\code{psram\_dq[3]} & D20 & 2 & IO & PSRAM DQ3 \\
\rowa \code{psram\_dq[4]} & G19 & 2 & IO & PSRAM DQ4 \\
\code{psram\_dq[5]} & J18 & 2 & IO & PSRAM DQ5 \\
\rowa \code{psram\_dq[6]} & J19 & 2 & IO & PSRAM DQ6 \\
\code{psram\_dq[7]} & J20 & 2 & IO & PSRAM DQ7 \\
\rowa \code{psram\_dq[8]} & K19 & 2 & IO & PSRAM DQ8 \\
\code{psram\_dq[9]} & K20 & 2 & IO & PSRAM DQ9 \\
\rowa \code{psram\_dq[10]} & L17 & 3 & IO & PSRAM DQ10 \\
\code{psram\_dq[11]} & M18 & 3 & IO & PSRAM DQ11 \\
\rowa \code{psram\_dq[12]} & M17 & 3 & IO & PSRAM DQ12 \\
\code{psram\_dq[13]} & N16 & 3 & IO & PSRAM DQ13 \\
\rowa \code{psram\_dq[14]} & N18 & 3 & IO & PSRAM DQ14 \\
\code{psram\_dq[15]} & P17 & 3 & IO & PSRAM DQ15 \\
\multicolumn{5}{l}{\textit{\color{fnDark}PSRAM control}}\\
\rowa \code{psram\_ce\_n} & N17 & 3 & OUT & PSRAM CE\# --- chip enable, active low. \\
\code{psram\_oe\_n} & R16 & 3 & OUT & PSRAM OE\# --- output enable (read). \\
\rowa \code{psram\_we\_n} & R17 & 3 & OUT & PSRAM WE\# --- write enable. \\
\code{psram\_lb\_n} & T16 & 3 & OUT & PSRAM LB\# --- lower-byte enable (DQ[7:0]). \\
\rowa \code{psram\_ub\_n} & N19 & 3 & OUT & PSRAM UB\# --- upper-byte enable (DQ[15:8]). \\
\code{psram\_zz\_n} & N20 & 3 & OUT & PSRAM ZZ\# --- sleep/snooze (high during normal operation). \\
\bottomrule
\end{tabularx}
\renewcommand{\arraystretch}{1.25}
\vspace{4pt}
\noindent
{\footnotesize\color{fnGrey}
Standard I/O: LVCMOS33 on all 57 signals. Boot config-SPI and JTAG balls (fixed-function
dedicated pins, no RTL port) do not appear in this table --- see ch.~\ref{ch:hw}
§``Configuration and programming''. Source: \code{synth/ecp5/spi\_neuron\_top.lpf},
generated by \code{tools/pinout/gen\_lpf.py} against Project~Trellis's
\code{iodb.json}.\par}
@@ -1,93 +0,0 @@
\chapter{System overview}
\label{ch:overview}
\section{Project goal}
FPGA-Neural implements a \textbf{reusable Neural Network Engine in FPGA hardware}.
The whole is made of three elements: the FPGA, which is the actual accelerator; a
dedicated RAM physically associated with the FPGA and not shared with the host; and a
host interface independent of the operating system, initially SPI (with possible
future extension to Dual~SPI).
The founding principle is the separation between who \emph{executes} the computation
and who \emph{uses} it: the neural network computation happens entirely inside the
FPGA, while the host system only provides configuration, network parameters, input
data, control and result readback. The host is not part of the computational datapath.
Possible host systems include Linux SoCs, Raspberry~Pi-like systems, ESP32,
microcontrollers and development PCs: the same engine architecture must be usable in
completely different systems.
\begin{center}
\begin{tikzpicture}[font=\footnotesize,node distance=8mm]
\node[fnblockD,minimum width=42mm,minimum height=20mm] (host){\textbf{HOST}\\[2pt]
{\scriptsize Configuration}\\{\scriptsize Training}\\{\scriptsize Control}};
\node[fnblockT,below=14mm of host,minimum width=42mm,minimum height=20mm] (fpga)
{\textbf{FPGA}\\[2pt]{\scriptsize Neural Network Engine}\\{\scriptsize Compute / Control}};
\node[fnblock,below=14mm of fpga,minimum width=42mm,minimum height=13mm] (ram)
{\textbf{Dedicated RAM}\\{\scriptsize weights / bias / buffers}};
\draw[fnbus] (host) -- node[fnlbl,right]{SPI / Dual SPI} (fpga);
\draw[fnbus] (fpga) -- node[fnlbl,right]{parallel bus} (ram);
\end{tikzpicture}
\end{center}
\section{Hardware configuration versus network configuration}
The project draws a precise distinction between the accelerator's \textbf{hardware
architecture} and the \textbf{neural network parameters}.
The physical architecture of the engine is defined at FPGA synthesis and
implementation time. Typical hardware parameters are \code{N\_INPUTS},
\code{N\_NEURONS}, \code{N\_LAYERS}, \code{PARALLEL}, \code{DATA\_WIDTH},
\code{ACC\_WIDTH}: they are Verilog parameters resolved at synthesis and they
determine the datapath contained in the bitstream. The network parameters --- weights,
bias, activation and quantization parameters, specific constants --- are instead loaded
at runtime through the host interface and stored in the RAM associated with the FPGA.
\begin{fnnote}[Central architectural principle]
A build fixes the \emph{ceiling} of the machine (maximum number of layers, maximum
width, \code{PARALLEL}); the host configures the \emph{actual} network --- number of
layers, per-layer input/output width, per-layer activation and trained parameters ---
entirely at runtime, over SPI, into the FPGA's local memory. A single bitstream serves
any topology up to that ceiling.
\end{fnnote}
\section{Boot and initialization}
The FPGA is configured at power-on through the usual configuration mechanism (bitstream
loading from SPI flash). The bitstream defines the hardware architecture of the engine;
the host does not dynamically build the datapath during normal operation, but rather
configures the network data on which the already existing datapath operates.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=4.5mm,start chain=going below,
every node/.style={on chain}]
\node[fnblockA,minimum width=60mm](p){Power-on};
\node[fnblock,minimum width=60mm]{FPGA configuration (bitstream from flash)};
\node[fnblockT,minimum width=60mm]{Neural Network Engine available};
\node[fnblock,minimum width=60mm]{Host initialization (SPI)};
\node[fnblock,minimum width=60mm]{Loading network parameters / weights / bias};
\node[fnblockD,minimum width=60mm]{Engine ready};
\begin{scope}[every path/.style={fnarrow}]
\foreach \a/\b in {1/2,2/3,3/4,4/5,5/6}{}
\end{scope}
\foreach \i [count=\j from 2] in {1,...,5}{
\draw[fnarrow] (chain-\i) -- (chain-\j);}
\end{tikzpicture}
\end{center}
\section{Training and inference}
Training and inference are conceptually separate. The first implementation does not
require the FPGA to perform training: weights can be computed externally
(PC/Linux/other host) and transferred over SPI into the FPGA's RAM, which then performs
inference. This drastically reduces the complexity of the initial hardware, without
precluding a future implementation of assisted or fully hardware training (roadmap
Phase~8, ch.~\ref{ch:roadmap}). During inference the host only provides the input data
and retrieves the result, obtaining deterministic computation, reduced host load,
hardware parallelism, predictable latency and independence from the host CPU
architecture.
\section{Design philosophy and reuse}
The project should be understood as a \emph{reusable FPGA neural acceleration platform}
rather than a single network. The application determines input size, topology, number
of layers and neurons, parallelism, numeric precision, activation functions, memory and
performance requirements; the hardware generation process produces the corresponding
FPGA implementation. The same HDL architecture remains conceptually unchanged while the
synthesis parameters generate implementations appropriate to the different application
targets.
@@ -1,80 +0,0 @@
\chapter[RTL architecture]{RTL architecture and module hierarchy}
\label{ch:arch}
\section{Hierarchical organization}
The design is organized in layers, from the elementary multiply-accumulator up to the
integrated top-level with SPI interface and PSRAM. Each layer encapsulates the previous
one and abstracts away its details: the validated datapath (\code{mac\_unit},
\code{mac8}, \code{neuron\_parallel}) is never modified by the higher orchestration
layers.
\begin{center}
\begin{tikzpicture}[font=\footnotesize,every node/.style={fnblock,minimum width=40mm},
level distance=13mm,sibling distance=0mm]
\node[fnblockD,minimum width=62mm](top){\code{spi\_neuron\_top} \\ {\scriptsize integrated top-level}};
\node[fnblockT,minimum width=62mm,below=8mm of top](arb){\code{mem\_arbiter} \;/\; \code{layer\_sequencer} \\ {\scriptsize 3-port arbitration + layer sequencing}};
\node[fnblock,minimum width=62mm,below=8mm of arb](nm){\code{neuron\_memory} \\ {\scriptsize memory $\leftrightarrow$ neuron bridge, neuron loop}};
\node[fnblock,minimum width=62mm,below=8mm of nm](np){\code{neuron\_parallel} \\ {\scriptsize neuron FSM: groups, bias, activation, saturation}};
\node[fnblockT,minimum width=62mm,below=8mm of np](m8){\code{mac8} \\ {\scriptsize \code{PARALLEL} MACs + balanced adder tree}};
\node[fnblock,minimum width=62mm,below=8mm of m8](mu){\code{mac\_unit} \\ {\scriptsize $x\cdot w$ + sign extension + accumulate}};
\foreach \a/\b in {top/arb,arb/nm,nm/np,np/m8,m8/mu}
\draw[fnarrow] (\a) -- (\b);
% memory branches on the right
\node[fnblockA,minimum width=34mm,right=14mm of nm](ma){\code{int8\_memory\_access}\\{\scriptsize byte $\leftrightarrow$ 16-bit word}};
\node[fnblockA,minimum width=34mm,below=6mm of ma](mi){\code{memory\_interface}\\{\scriptsize req/ready handshake}};
\node[fnblockA,minimum width=34mm,below=6mm of mi](pc){\code{psram\_controller}\\{\scriptsize physical PSRAM bus}};
\draw[fnarrowT] (ma)--(mi); \draw[fnarrowT] (mi)--(pc);
\draw[fnarrowT,dashed] (nm.east) -- (ma.west);
% SPI branches on the left
\node[fnblockA,minimum width=30mm,left=14mm of arb,yshift=6mm](ss){\code{spi\_slave}\\{\scriptsize Mode 0 physical layer}};
\node[fnblockA,minimum width=30mm,below=6mm of ss](se){\code{spi\_engine}\\{\scriptsize opcode FSM + registers}};
\draw[fnarrowT] (ss)--(se);
\draw[fnarrowT,dashed] (se.east) -- (arb.west);
\end{tikzpicture}
\end{center}
\section{Role of each module}
\begin{tabularx}{\textwidth}{L{3.4cm}Y}
\toprule
\rowh \thd{Module} & \thd{Function} \\
\midrule
\code{mac\_unit} & Single multiply-accumulate: $\mathrm{acc\_out}=\mathrm{acc\_in}+(x\cdot w)$, with sign extension of the product to \code{ACC\_WIDTH}. Parametric on \code{DATA\_WIDTH}/\code{ACC\_WIDTH}. \\
\rowa \code{mac8} & \code{PARALLEL} instances of \code{mac\_unit} whose products are summed by a \emph{balanced binary adder tree} of depth $\log_2(\text{PARALLEL})$; the result is added to the input accumulator. \\
\code{neuron\_parallel} & FSM of a single neuron: processes \code{N\_INPUTS} inputs in groups of \code{PARALLEL}, accumulates across groups, adds the bias, applies the activation and saturates to INT8. Includes the processing guard on \code{N\_INPUTS \% PARALLEL} and the runtime width \code{n\_inputs\_real}. \\
\rowa \code{layer} & Instantiates \code{N\_NEURONS} neurons \emph{in parallel} on the same input vector; \code{busy}=OR, \code{done}=AND of the neurons. A purely data-combinational path used in the datapath benchmarks. \\
\code{neuron\_memory} & Integrates computation with memory: reads $X$ (shared) once, then for each neuron re-reads $W$ and bias from RAM and reuses a single \code{neuron\_parallel} instance (memory-bound, one neuron at a time). Output \code{y\_bus} packed neuron-major. \\
\rowa \code{layer\_sequencer} & Chains up to \code{N\_LAYERS} executions of \code{neuron\_memory} by reading a descriptor table written by the host and alternating the ping-pong buffers in RAM (Phase~5). \\
\code{act\_buffer} & Global activation buffer in \code{DP16KD} block RAM, indexed by signal id (Type \#2). \\
\rowa \code{graph\_engine} & Graph-network engine (Type \#2): gather from \code{act\_buffer}, reuses \code{neuron\_parallel}, writes outputs by id (ch.~\ref{ch:grafo}). \\
\code{int8\_memory\_access} & Converts the byte/INT8 interface (byte address) into the 16-bit word interface, selecting the low/high byte via \code{lb\_n}/\code{ub\_n} and \code{addr>>1}. \\
\rowa \code{memory\_interface} & 2-state handshake FSM (IDLE/WAIT) that serializes the single transaction toward the controller. \\
\code{psram\_controller} & Asynchronous parallel PSRAM bus controller with read \textbf{page mode}: 70~ns random access (\code{tAA}), 20~ns same-page bursts (\code{tAPA}) with CE\#/OE\# held asserted; enables page mode on the chip at boot via the configuration register (ch.~\ref{ch:mem}, \S~5.5). Drives \code{ce\_n/oe\_n/we\_n/lb\_n/ub\_n/zz\_n} and the tri-state data bus. \\
\rowa \code{mem\_arbiter} & Fixed-priority arbiter (B$>$C$>$A) among three byte-level masters: \code{spi\_engine} (A), \code{neuron\_memory} (B), \code{layer\_sequencer} (C). \\
\code{spi\_slave} & SPI Mode 0 physical layer, MSB-first, 3-stage CDC synchronizer on SCLK/MOSI/CS\_N, shift register and CS framing. \\
\rowa \code{spi\_engine} & Protocol/opcode FSM and register bank (\code{x\_base}, \code{w\_base}, \code{bias\_addr}, ping-pong base, activation, runtime widths\ldots), with sticky/clear-on-read \code{STATUS.done}. \\
\code{spi\_neuron\_top} & Top-level: connects SPI, arbiter, sequencer, \code{neuron\_memory} and the PSRAM chain; multiplexes control of \code{neuron\_memory} between the sequencer and the direct single-layer path. \\
\bottomrule
\end{tabularx}
\vspace{6pt}
\begin{fnnote}[Simulation models]
\code{psram\_model.v} (in \code{sim/}) and \code{memory\_model.v} are behavioral memory
models used in the testbenches; they are not part of the synthesizable design but they
reproduce the real latency for end-to-end verification.
\end{fnnote}
\section{Two execution paths}
The top-level exposes two mutually exclusive modes toward the same \code{neuron\_memory}
compute engine:
\begin{itemize}
\item \textbf{Single-layer / manual path}: the host sets the bases with
\op{SET\_BASE}, starts with \op{START} and reads with \op{READ\_OUTPUT}.
\code{spi\_engine} drives \code{neuron\_memory} directly.
\item \textbf{Multi-layer path}: the host writes the descriptor table and starts with
\op{RUN\_NETWORK}; \code{layer\_sequencer} takes over control of \code{neuron\_memory}
(while \code{seq\_busy} is high) and chains the layers.
\end{itemize}
The top-level multiplexer switches the control lines of \code{neuron\_memory} based on
\code{seq\_busy}, returning the engine to the direct path at the end of the sequence.
@@ -1,166 +0,0 @@
\chapter{Compute datapath}
\label{ch:datapath}
\section{INT8/INT32 arithmetic chain}
The elementary datapath implements the typical sequence of a quantized neuron:
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=3mm,start chain=going right,
every node/.style={fnblock,minimum width=15mm,minimum height=8mm,on chain}]
\node[fnblockT]{INT8\\$\times$\,INT8};
\node{INT16\\product};
\node{sign-ext\\INT32};
\node[fnblockD]{accumulate\\INT32};
\node{$+$ bias};
\node[fnblockA]{activation};
\node[fnblockT]{sat. INT8};
\foreach \i [count=\j from 2] in {1,...,6}
\draw[fnarrow] (chain-\i) -- (chain-\j);
\end{tikzpicture}
\end{center}
Each INT8$\times$INT8 product fits in 16~bits; it is sign-extended to 32~bits before
accumulation, so the accumulator does not overflow on long vectors. Bias and activation
operate at 32~bits; only the final output is saturated to INT8.
\section{\texttt{mac\_unit} --- multiply-accumulator}
The \code{mac\_unit} module is purely combinational and parametric on \code{DATA\_WIDTH}
and \code{ACC\_WIDTH}. It computes:
\[
\mathrm{acc\_out} = \mathrm{acc\_in} + \mathrm{signext}_{ACC}(x \cdot w)
\]
The product has width $2\times$\code{DATA\_WIDTH} and is sign-extended by replicating
the most significant bit. On ECP5 the multiplication maps onto a \code{MULT18X18D} DSP
block.
\begin{lstlisting}[caption={\texttt{rtl/mac\_unit.v} --- arithmetic core},label={lst:macunit}]
localparam PROD_WIDTH = 2 * DATA_WIDTH;
wire signed [PROD_WIDTH-1:0] product = x * w;
wire signed [ACC_WIDTH-1:0] product_ext =
{{(ACC_WIDTH-PROD_WIDTH){product[PROD_WIDTH-1]}}, product};
assign acc_out = acc_in + product_ext;
\end{lstlisting}
\section{\texttt{mac8} --- parallel MAC and balanced adder tree}
\code{mac8} instantiates \code{PARALLEL} \code{mac\_unit} units that generate
\code{PARALLEL} independent products, then sums them with a \emph{balanced binary
adder tree}. Compared to the linear reduction
$((((p_0{+}p_1){+}p_2){+}p_3){+}\dots)$, of depth $O(\text{PARALLEL})$, the tree has
depth $O(\log_2 \text{PARALLEL})$, drastically reducing the combinational path.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,level distance=11mm,
every node/.style={fnreg,minimum width=8mm},
level 1/.style={sibling distance=30mm},
level 2/.style={sibling distance=15mm},
level 3/.style={sibling distance=8mm},
edge from parent/.style={fnarrowT,draw}]
\node[fnblockD]{sum}
child {node[fnblockT]{$+$}
child {node[fnblockT]{$+$}
child {node{$p_0$}} child {node{$p_1$}}}
child {node[fnblockT]{$+$}
child {node{$p_2$}} child {node{$p_3$}}}}
child {node[fnblockT]{$+$}
child {node[fnblockT]{$+$}
child {node{$p_4$}} child {node{$p_5$}}}
child {node[fnblockT]{$+$}
child {node{$p_6$}} child {node{$p_7$}}}};
\end{tikzpicture}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
Example with PARALLEL=8: 3 levels. PARALLEL=16 $\to$ 4 levels; PARALLEL=32 $\to$ 5
levels.\end{center}
\begin{fnnote}[PARALLEL as a power of two]
The tree is designed for \code{PARALLEL} as a power of two (8, 16, 32\ldots). This is
also the value used in all project configurations.
\end{fnnote}
\section{\texttt{neuron\_parallel} --- neuron FSM}
\code{neuron\_parallel} processes \code{N\_INPUTS} inputs in groups of \code{PARALLEL},
maintaining the accumulator from one group to the next. At the end it adds the bias,
applies the activation and saturates to INT8. The number of groups is
$\text{GROUPS}=\text{N\_INPUTS}/\text{PARALLEL}$.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=4mm,start chain=going below,
every node/.style={on chain,fnblock,minimum width=46mm}]
\node[fnblockA]{\code{start}};
\node{group 0 $\to$ accumulate};
\node{group 1 $\to$ accumulate};
\node[draw=none,fill=none]{\vdots};
\node{group GROUPS$-$1 $\to$ accumulate};
\node{$+$ bias};
\node[fnblockA]{activation (ACT\_RELU / ACT\_NONE)};
\node[fnblockT]{INT8 saturation};
\node[fnblockD]{\code{done}, \code{y}};
\foreach \i [count=\j from 2] in {1,...,8}
\draw[fnarrow] (chain-\i) -- (chain-\j);
\end{tikzpicture}
\end{center}
\subsection{Parameter guard (elaboration-time)}
If \code{PARALLEL} does not exactly divide \code{N\_INPUTS} two failures occur, both
confirmed empirically in \code{sim/parameter\_sweep\_tb.v}:
\begin{itemize}
\item integer division truncates \code{GROUPS} and the excess inputs are never read
$\to$ \textbf{wrong} result, with no error and no warning;
\item if \code{PARALLEL > N\_INPUTS}, \code{GROUPS=0} and the terminal condition is
never satisfied $\to$ the neuron \textbf{hangs} (busy high, done never asserted).
\end{itemize}
The solution does not modify the validated datapath: a \code{generate} block
instantiates a deliberately undefined module when
$\text{N\_INPUTS} \bmod \text{PARALLEL}\neq0$, forcing an error at \emph{elaboration}
both in simulation and in synthesis. For valid configurations the branch is never
elaborated.
\begin{lstlisting}[caption={\texttt{rtl/neuron\_parallel.v} --- parameter guard}]
generate
if (N_INPUTS == 0 || N_INPUTS % PARALLEL != 0) begin : PARAMETER_ERROR
neuron_parallel_requires_N_INPUTS_multiple_of_PARALLEL
invalid_parameter_combination();
end
endgenerate
\end{lstlisting}
\begin{fnnote}[Edge case \texttt{N\_INPUTS=0} (fixed 2026-09-04)]
The original condition (\code{N\_INPUTS \% PARALLEL != 0}) does not catch
\code{N\_INPUTS=0}, since $0 \bmod \text{PARALLEL}=0$ for any \code{PARALLEL}: the module
elaborated successfully (both in simulation and in real Yosys synthesis) while leaving
\code{x\_bus}/\code{w\_bus} undriven and \code{start} silently ineffective. Found during
the re-certification campaign (\code{docs/validation/bugs.md}, BUG-002) and fixed by
extending the guard as above --- \code{N\_INPUTS=0} now fails elaboration exactly like the
other degenerate cases.
\end{fnnote}
\section{Activation functions}
\code{neuron\_parallel} accepts a 2-bit \code{activation} port. The default is
\code{ACT\_RELU}, the only behavior that existed before the port was introduced, so
every pre-existing caller remains unchanged.
\begin{tabularx}{\textwidth}{L{2.6cm} C{1.4cm} Y}
\toprule
\rowh \thd{Encoding} & \thd{Value} & \thd{Behavior} \\
\midrule
\code{ACT\_NONE} & \code{2'd0} & Linear: no clamp to zero, bilateral saturation to the INT8 range $[-128,+127]$. \\
\rowa \code{ACT\_RELU} & \code{2'd1} & $\max(0,x)$, then positive saturation to $+127$ (default; also the fallback for reserved encodings). \\
\bottomrule
\end{tabularx}
\section{INT8 saturation}
After bias and activation, the 32-bit accumulator is reduced to INT8:
\[
y=\begin{cases}
+127 & \text{if } \mathrm{final\_acc} > 127\\
-128 & \text{if } \mathrm{final\_acc} < -128 \ \text{(ACT\_NONE only)}\\
0 & \text{if } \mathrm{final\_acc}\le 0 \ \text{(ACT\_RELU only)}\\
\mathrm{final\_acc}[7:0] & \text{otherwise}
\end{cases}
\]
\section{\texttt{layer} --- neurons in parallel}
\code{layer} instantiates \code{N\_NEURONS} neurons that share the input vector
\code{x\_bus} but have distinct weights and bias; \code{busy} is the OR and \code{done}
the AND of the neurons' signals. It is the module used in the datapath benchmarks
(ch.~\ref{ch:impl}), where all neurons work simultaneously. The addressing convention
is neuron-major: the weights of neuron $n$ occupy
\code{weights\_bus[n*N\_INPUTS*DATA\_WIDTH +: N\_INPUTS*DATA\_WIDTH]}.
@@ -1,88 +0,0 @@
\chapter{Parameters and configurability}
\label{ch:param}
\section{Build parameters (synthesis-time)}
The hardware architecture is fixed at synthesis through the following Verilog
parameters. They determine the datapath contained in the bitstream and its capacity
\emph{ceiling}.
\begin{tabularx}{\textwidth}{L{3.0cm} C{1.8cm} Y}
\toprule
\rowh \thd{Parameter} & \thd{Default} & \thd{Meaning} \\
\midrule
\code{DATA\_WIDTH} & 8 & Data width (INT8). \\
\rowa \code{ACC\_WIDTH} & 32 & Accumulator width (INT32). \\
\code{N\_INPUTS} & 32 / 256 & Maximum number of inputs per neuron (benchmark baseline: 256). \\
\rowa \code{N\_NEURONS} & 1 / 4 & Maximum number of neurons per layer. \\
\code{PARALLEL} & 8 & Simultaneous hardware MACs per neuron; must divide \code{N\_INPUTS} and should be a power of two. \\
\rowa \code{N\_LAYERS} & 4 & Maximum number of layers chainable by \code{layer\_sequencer}. \\
\code{ADDR\_WIDTH} & 23 & Byte-address width (8~MB). \\
\rowa \code{MEM\_DATA\_WIDTH} & 16 & Width of the physical PSRAM data bus. \\
\code{CLK\_FREQ\_MHZ} & 80 & Frequency used in the PSRAM timing formulas (must be aligned to the real oscillator). \\
\bottomrule
\end{tabularx}
\begin{fnwarn}[\texttt{N\_INPUTS} \% \texttt{PARALLEL} constraint]
\code{PARALLEL} must divide \code{N\_INPUTS} exactly, otherwise the elaboration guard
fires (§\ref{ch:datapath}). The same constraint applies at runtime to
\code{n\_inputs\_real}.
\end{fnwarn}
\section{Runtime network width}
A single bitstream serves any topology \emph{up to} the build maximum. The actual width
of each execution is a separate value, set by the host:
\begin{itemize}
\item \code{n\_inputs\_real} --- inputs actually used in this execution (must be a
multiple of \code{PARALLEL});
\item \code{n\_neurons\_real} --- neurons actually computed in this execution.
\end{itemize}
Both default to the build maximum, so any caller that leaves them unconnected processes
the full width as before the ports were introduced.
\begin{fnnote}[Real early termination]
This is not mere address bookkeeping: the two values directly bound the hardware loops
(X/W reads of \code{neuron\_memory}, MAC group count of \code{neuron\_parallel} and the
length of the ping-pong copy for \code{RUN\_NETWORK}). A narrower layer actually
\emph{computes} and \emph{copies} faster and does not require zero-padding of the RAM
for the unused tail: data beyond \code{n\_inputs\_real}/\code{n\_neurons\_real} is never
read.
\end{fnnote}
This lets a network taper within a single chained execution, for example
$256\to64\to16\to4$, with each layer declaring its own actual width in the descriptor
table (ch.~\ref{ch:seq}).
\subsection{Measured savings}
Early termination was measured end-to-end:
\begin{tabularx}{\textwidth}{L{5.5cm} C{3.0cm} Y}
\toprule
\rowh \thd{Test} & \thd{Cycles} & \thd{Comparison} \\
\midrule
\code{neuron\_parallel\_tb.v} (T7) & 3 vs 6 & reduced vs full, with ``garbage'' data in the skipped lanes (proof that they are not read). \\
\rowa \code{neuron\_memory\_tb.v} (T5) & 209 vs 788 & 8-of-32 vs full 32, through the real PSRAM stack. \\
\bottomrule
\end{tabularx}
\section{Characterized configurations}
Some combinations validated in simulation and/or synthesis:
\begin{tabularx}{\textwidth}{C{2.0cm} C{2.0cm} C{2.0cm} Y}
\toprule
\rowh \thd{N\_INPUTS} & \thd{N\_NEURONS} & \thd{PARALLEL} & \thd{Notes} \\
\midrule
32 & 4 & 8 & First functional parametric test (Phase~1). \\
\rowa 256 & 4 & 2/4/8/16 & Datapath benchmark sweep (Phase~7). \\
32 & 1..3 & 8 & Single/multi-neuron memory integration (Phase~3). \\
\rowa 4 & 4 & 2 & End-to-end 2-layer \code{RUN\_NETWORK} test over real SPI. \\
\bottomrule
\end{tabularx}
\section{Build versus runtime summary}
\begin{center}
\begin{tikzpicture}[font=\footnotesize,node distance=6mm]
\node[fnblockD,minimum width=54mm,minimum height=15mm](b){\textbf{BUILD (synthesis)}\\[2pt]
{\scriptsize N\_INPUTS, N\_NEURONS, N\_LAYERS,}\\{\scriptsize PARALLEL, DATA\_WIDTH, ACC\_WIDTH}\\{\scriptsize $\Rightarrow$ machine ceiling}};
\node[fnblockT,right=16mm of b,minimum width=54mm,minimum height=15mm](r){\textbf{RUNTIME (host, SPI)}\\[2pt]
{\scriptsize n\_inputs\_real, n\_neurons\_real,}\\{\scriptsize activation, num\_layers, weights/bias}\\{\scriptsize $\Rightarrow$ actual network}};
\draw[fnbus] (b) -- node[fnlbl,above]{$\le$} (r);
\end{tikzpicture}
\end{center}
@@ -1,187 +0,0 @@
\chapter{Memory subsystem}
\label{ch:mem}
\section{Memory chain}
The compute engine works with addresses and data at the \emph{byte} level (INT8), while
the PSRAM is a 16-bit word device. Three cascaded modules realize the conversion and
the physical access:
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=8mm]
\node[fnblockD,minimum width=30mm,minimum height=12mm](nm){byte-level master\\{\scriptsize \code{neuron\_memory} / \code{spi\_engine} / \code{layer\_sequencer}}};
\node[fnblockT,right=10mm of nm,minimum width=28mm,minimum height=12mm](ia){\code{int8\_memory\_access}\\{\scriptsize byte $\leftrightarrow$ 16-bit word}};
\node[fnblock,right=10mm of ia,minimum width=26mm,minimum height=12mm](mi){\code{memory\_interface}\\{\scriptsize IDLE/WAIT FSM}};
\node[fnblockA,below=9mm of mi,minimum width=26mm,minimum height=12mm](pc){\code{psram\_controller}\\{\scriptsize async 70\,ns physical bus}};
\node[fnblock,left=10mm of pc,minimum width=26mm,minimum height=12mm](ps){PSRAM\\{\scriptsize 8\,MB 4M$\times$16}};
\draw[fnbus] (nm)--node[fnlbl,above]{req/wr/addr}(ia);
\draw[fnbus] (ia)--node[fnlbl,above]{16-bit}(mi);
\draw[fnbus] (mi)--(pc);
\draw[fnbus] (pc)--node[fnlbl,above]{DQ/A/ctrl}(ps);
\end{tikzpicture}
\end{center}
\section{\texttt{int8\_memory\_access} --- byte/word conversion}
Converts the INT8 interface (byte address) into the word interface. The byte address is
divided by two (\code{addr>>1}) to obtain the word address; the least significant bit
selects the byte:
\begin{itemize}
\item \code{addr[0]=0} $\to$ low byte: \code{lb\_n=0}, \code{ub\_n=1}, data on DQ[7:0];
\item \code{addr[0]=1} $\to$ high byte: \code{lb\_n=1}, \code{ub\_n=0}, data on DQ[15:8].
\end{itemize}
On read it extracts the correct byte from \code{mem\_rdata}. The FSM has two states
(IDLE, WAIT) and returns \code{ready} as a one-cycle pulse.
\section{\texttt{memory\_interface} --- handshake}
Two-state FSM that serializes a single transaction: in IDLE, on the \code{req} request,
it latches \code{wr/addr/wdata/lb\_n/ub\_n} and emits a one-cycle \code{mem\_req} pulse
toward the controller; in WAIT it waits for \code{mem\_ready}, captures \code{rdata} on
read and asserts \code{ready}. It guarantees the ``one transaction at a time'' contract.
\section{\texttt{psram\_controller} --- physical bus}
Asynchronous parallel PSRAM bus controller, with support for the chip's read
\textbf{page mode} (\S~\ref{sec:pagemode}). The main state machine is:
\begin{center}
\begin{tikzpicture}[font=\scriptsize]
\node[fnstate](init) at (0,0){INIT};
\node[fnstate](idle) at (3.2,0){IDLE};
\node[fnstate](read) at (7,2.7){READ};
\node[fnstate](popen) at (11,2.7){PAGE\\OPEN};
\node[fnstate](write) at (7,-2.7){WRITE};
\node[fnstate](ww) at (11,-2.7){WRITE\\WAIT};
\draw[fnarrow] (init)--node[fnlbl,above]{INIT\_CYCLES + CR load}(idle);
\draw[fnarrow] (idle)--node[fnlbl,above,sloped]{req \& !wr}(read);
\draw[fnarrow] (idle)--node[fnlbl,below,sloped]{req \& wr}(write);
\draw[fnarrow] (read)--node[fnlbl,above]{ready}(popen);
\draw[fnarrowT] (popen) to[bend left=25] node[fnlbl,below]{req \& !wr}(read);
\draw[fnarrow] (popen) to[bend right=20] node[fnlbl,above,sloped]{req \& wr}(write);
\draw[fnarrow] (popen) to[out=-100,in=15,looseness=1.15] node[fnlbl,pos=0.55]{tCEM timeout}(idle);
\draw[fnarrow] (write)--node[fnlbl,above]{ACCESS\_CYCLES}(ww);
\draw[fnarrow] (ww) to[out=160,in=-70] node[fnlbl,pos=0.5,left]{ready}(idle);
\end{tikzpicture}
\end{center}
From INIT the controller automatically goes through a configuration-register load
sub-sequence (\code{STATE\_CR\_INIT}, 4 steps) before reaching IDLE for the first
time --- see \S~\ref{sec:pagemode}. The PAGE~OPEN~$\to$~WRITE transition
(bottom-right arrow) internally passes through two transit micro-states,
\code{STATE\_PAGE\_CLOSE} and \code{STATE\_PAGE\_REOPEN} (one cycle each): the
first forces CE\#/OE\# high for at least one cycle before the controller starts
driving the data bus, avoiding contention with the PSRAM's still-active output
($\geq t_{HZ}$); the second restarts the already-latched transaction exactly as
IDLE would. They are not drawn as separate nodes to keep the figure readable.
\subsection{Timing}
\begin{fnspec}[Timing formulas]
$\text{ACCESS\_CYCLES}=\lceil (70\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(random-access latency, $t_{AA}$/$t_{RC}$ = 70~ns)\\[3pt]
$\text{PAGE\_CYCLES}=\lceil (20\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(same-page continuation, $t_{APA}$/$t_{PC}$ = 20~ns)\\[3pt]
$\text{INIT\_CYCLES}=150\times \text{CLK\_FREQ\_MHZ}$ \quad
(power-up initialization, $t_{PU}$ = 150~\textmu s)\\[3pt]
$\text{PAGE\_TIMEOUT\_CYCLES}=\lceil (6000\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(automatic page close, safety margin under $t_{CEM}$ = 8~\textmu s)
\end{fnspec}
The data bus is tri-state driven: \code{psram\_dq = dq\_oe ? dq\_out : Z}. On read
\code{dq\_oe=0}; on write \code{dq\_oe=1} during the \code{we\_n} pulse. A WRITE\_WAIT
state keeps \code{ce\_n/lb\_n/ub\_n} active for the final hold before release.
\begin{fnwarn}[This is not QSPI]
This is a classic asynchronous-SRAM interface, \textbf{not} QSPI: most commercial
serial/QSPI ``PSRAM'' parts are not compatible with this controller without a rewrite.
See ch.~\ref{ch:hw} for the recommended part (parallel ISSI).
\end{fnwarn}
\section{Read page mode}
\label{sec:pagemode}
The recommended chip (ch.~\ref{ch:hw}) is ``asynchronous/\textbf{page mode}'': once
an initial random access at $t_{AA}$~=~70~ns has been done, further reads inside the
same 16-word page (address bits above \code{A[3]} unchanged) only cost
$t_{APA}$/$t_{PC}$~=~20~ns, because CE\#/OE\# stay asserted and only the address bus
changes. Page mode is \textbf{disabled by default} at power-up (bit~7 of the
configuration register, CR~=~\texttt{0x0070} by default) and must be explicitly
enabled.
\begin{itemize}
\item \textbf{Enable at boot}: right after INIT, the controller runs the
datasheet's ``software-access sequence'' (2 dummy reads + 2 writes, \texttt{0x0000}
unlock then real CR \texttt{0x00F0} = default with the Page bit set) at the
chip's highest address --- it reuses exactly the same READ/WRITE logic as every
other transaction, so it goes through the same timing checks.
\item \textbf{Page bursts}: after a READ the controller no longer closes CE\#/OE\#
(PAGE~OPEN state). A following read in the same page only waits PAGE\_CYCLES; a
read crossing into a different page still avoids a CE\# toggle but pays a full
ACCESS\_CYCLES for that one word (any change at \code{A[4]} or above requires a
new $t_{AA}$). A counter closes the page before the $t_{CEM}$ limit with a
safety margin.
\item \textbf{Only a WRITE closes the page.} Changes to \code{lb\_n}/\code{ub\_n}
do \emph{not} close it: \code{int8\_memory\_access} alternates these signals on
nearly every access (byte-granular access over the 16-bit bus), so treating them
as a close condition --- the first implementation attempt --- made the real
workload \emph{slower}, not faster (measured: 53.25$\to$61.25 cycles/edge on
\code{graph\_engine}'s gather); removed, corrected to 53.25$\to$37.53
cycles/edge (bandwidth +42\%, \S~\ref{sec:bandwidth}).
\end{itemize}
\begin{fnwarn}[No benefit without a sequential pattern]
Page mode only speeds up accesses that stay in the same page (or nearly) while the
controller is waiting for a new request with the page still open. Isolated,
scattered accesses (a random address every time) still pay a full ACCESS\_CYCLES,
plus a small close/reopen overhead if preceded by a WRITE or a $t_{CEM}$ timeout:
it is not a universal win, it depends on the caller's access pattern.
\end{fnwarn}
Real Fmax (\code{nextpnr-ecp5}, ch.~\ref{ch:impl}) on the integrated
\code{spi\_neuron\_top} system with Type~\#2 enabled: \textbf{75.73~MHz} at
\code{PARALLEL}=2 (was 55.59~MHz before page mode was added) and
\textbf{65.13~MHz} at \code{PARALLEL}=8, both still FAIL against the 80~MHz
target but not regressed. The critical path stays, in both cases, entirely
inside \code{u\_graph\_engine.u\_neuron} (the \code{mac8}/\code{neuron\_parallel}
accumulate chain, ch.~\ref{ch:impl}) --- \code{psram\_controller} never appears
in the critical path despite page mode's resource growth.
\section{Address map and conventions}
The addressing space is \code{ADDR\_WIDTH}=23~bits (\emph{byte} address), for a full
8~MB. The regions do not have hardwired addresses: their bases are registers set by the
host via \op{SET\_BASE} (single-layer path) or read from the descriptor table
(multi-layer path).
\begin{tabularx}{\textwidth}{L{3.2cm} L{3.4cm} Y}
\toprule
\rowh \thd{Region} & \thd{Base} & \thd{Content / convention} \\
\midrule
Input $X$ & \code{x\_base} & Shared input vector, read once per invocation. \\
\rowa Weights $W$ & \code{w\_base} & Neuron-major: weights of neuron $n$ at \code{w\_base + n*N\_INPUTS} bytes. \\
Bias & \code{bias\_addr} & One byte per neuron: bias of neuron $n$ at \code{bias\_addr + n}. \\
\rowa Descriptor table & \code{table\_base} & \code{N\_LAYERS} 11-byte entries (ch.~\ref{ch:seq}). \\
Ping-pong buffers A/B & \code{buf\_a\_base} / \code{buf\_b\_base} & Intermediate outputs between layers. \\
\bottomrule
\end{tabularx}
\subsection{PSRAM physical addressing}
The recommended PSRAM is 4M$\times$16 (8~MB), which requires a 22-bit word address
(A0--A21). \code{int8\_memory\_access} computes \code{addr>>1}, turning the 23-bit byte
address into a 22-bit word address that maps exactly onto A0--A21; bit~22 of
\code{psram\_a} is therefore always 0 and 22 real address lines remain on the PCB.
\section{Bandwidth}
\label{sec:bandwidth}
Measured on \code{graph\_engine}'s edge-list gather (ch.~\ref{ch:grafo}), by
difference between two graph sizes to isolate the per-edge cost from the fixed
per-neuron overhead (\code{sim/graph\_engine\_bandwidth\_tb.v}):
\begin{tabularx}{\textwidth}{L{5.2cm} Y Y Y}
\toprule
\rowh \thd{} & \thd{Before (no page mode)} & \thd{After (page mode)} & \thd{$\Delta$} \\
\midrule
Cycles/edge & 53.25 & 37.53 & $-29.5\%$ \\
\rowa Bandwidth @80\,MHz & 6.01\,MB/s & 8.53\,MB/s & $+41.9\%$ \\
Bandwidth @16\,MHz\textsuperscript{*} & 1.20\,MB/s & 1.71\,MB/s & $+41.9\%$ \\
\bottomrule
\end{tabularx}
\textsuperscript{*}recommended real oscillator (ch.~\ref{ch:hw}).
The model still remains memory-bound by construction: \code{neuron\_memory} reads
$X$ once and re-reads $W$/bias for each neuron (ch.~\ref{ch:seq}), one neuron at a
time; page mode reduces the per-byte cost of a sequential access, it does not
eliminate the access pattern itself.
@@ -1,111 +0,0 @@
\chapter[Memory, multi-neuron and multi-layer]{Memory integration, multi-neuron and multi-layer}
\label{ch:seq}
\section{\texttt{neuron\_memory} --- memory/neuron bridge}
\code{neuron\_memory} connects the compute datapath to memory and manages the loop over
the neurons. It reads the $X$ vector only once (shared input), then for each neuron
re-reads $W$ and bias from RAM and feeds them to a single reused instance of
\code{neuron\_parallel}: the design is memory-bound, one neuron computed at a time,
without duplicating the datapath. The output is \code{y\_bus}, packed neuron-major
(\code{DATA\_WIDTH*N\_NEURONS} bits).
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=13mm]
\node[fnstate](idle){IDLE};
\node[fnstate,right=of idle](rx){READ\_X};
\node[fnstate,right=of rx](rw){READ\_W};
\node[fnstate,below=10mm of rw](rb){READ\_BIAS};
\node[fnstate,left=of rb](sn){START\_N};
\node[fnstate,left=of sn](wn){WAIT\_N};
\draw[fnarrow] (idle)--node[fnlbl,above]{start}(rx);
\draw[fnarrow] (rx)--node[fnlbl,above]{X read}(rw);
\draw[fnarrow] (rw)--(rb);
\draw[fnarrow] (rb)--(sn);
\draw[fnarrow] (sn)--(wn);
\draw[fnarrow] (wn) to[bend left=18] node[fnlbl,above]{next neuron}(rw);
\draw[fnarrow] (wn) to[bend right=28] node[fnlbl,below]{last neuron: done}(idle);
\end{tikzpicture}
\end{center}
The states are IDLE, READ\_X, READ\_W, READ\_BIAS, START\_N, WAIT\_N. After the last
neuron the FSM returns to IDLE and asserts \code{done}. The count of neurons and inputs
actually processed is given by \code{n\_neurons\_real}/\code{n\_inputs\_real}
(ch.~\ref{ch:param}).
\section{\texttt{layer\_sequencer} --- multi-layer network}
\code{layer\_sequencer} chains up to \code{N\_LAYERS} executions of the same
\code{neuron\_memory} instance, realizing a dense feed-forward network \emph{without}
touching the validated compute core. It reads a descriptor table written by the host and
alternates the two output buffers in RAM (ping-pong).
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=13mm]
\node[fnstate](i){IDLE};
\node[fnstate,right=of i](rd){READ\\DESC};
\node[fnstate,right=of rd](rw){READ\\WAIT};
\node[fnstate,below=10mm of rw](sl){START\\LAYER};
\node[fnstate,left=of sl](wl){WAIT\\LAYER};
\node[fnstate,left=of wl](ci){COPY\\ISSUE};
\node[fnstate,below=9mm of ci](cw){COPY\\WAIT};
\draw[fnarrow] (i)--node[fnlbl,above]{run\_start}(rd);
\draw[fnarrow] (rd)--(rw);
\draw[fnarrow] (rw)--(sl);
\draw[fnarrow] (sl)--(wl);
\draw[fnarrow] (wl)--(ci);
\draw[fnarrow] (ci)--(cw);
\draw[fnarrow] (cw) to[bend left=15] node[fnlbl,left]{next layer}(rd);
\draw[fnarrow] (cw) to[bend right=12] node[fnlbl,below]{last: seq\_done}(i);
\end{tikzpicture}
\end{center}
\subsection{Ping-pong buffers}
Layer~0 reads the external input \code{x\_base}. Layer $k>0$ reads from the buffer
written by layer $k-1$; the output of each layer is copied into the other buffer,
alternating A and B. The final output remains both in \code{y\_bus} (readable with
\op{READ\_OUTPUT}) and in the ping-pong buffer into which it was copied.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=7mm]
\node[fnblockA,minimum width=18mm](x){X\\\code{x\_base}};
\node[fnblockD,right=10mm of x,minimum width=20mm](l0){Layer 0};
\node[fnblock,right=10mm of l0,minimum width=18mm](ba){buf A};
\node[fnblockD,right=10mm of ba,minimum width=20mm](l1){Layer 1};
\node[fnblock,right=10mm of l1,minimum width=18mm](bb){buf B};
\node[fnblockD,right=10mm of bb,minimum width=20mm](l2){Layer 2};
\draw[fnarrow] (x)--(l0); \draw[fnarrow] (l0)--(ba);
\draw[fnarrow] (ba)--(l1); \draw[fnarrow] (l1)--(bb);
\draw[fnarrow] (bb)--(l2);
\draw[fnarrowT,dashed] (l2.south) to[bend left=25] node[fnlbl,below]{copy into buf A} (ba.south);
\end{tikzpicture}
\end{center}
\subsection{Descriptor table}
Written by the host into RAM at \code{table\_base} with \op{WRITE\_RAM}; \code{N\_LAYERS}
entries of 11 bytes each, MSB-first:
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.6cm} Y}
\toprule
\rowh \thd{Field} & \thd{Bytes} & \thd{Meaning} \\
\midrule
\code{w\_base} & 3 & Weight base of the layer. \\
\rowa \code{bias\_addr} & 3 & Bias base of the layer. \\
\code{activation} & 1 & Layer activation (low 2 bits, cf. \code{ACT\_*}). \\
\rowa \code{n\_inputs\_real} & 2 & Actual inputs of the layer (multiple of \code{PARALLEL}). \\
\code{n\_neurons\_real} & 2 & Actual neurons of the layer. \\
\midrule
\rowh \thd{Total} & \thd{11} & per entry/layer \\
\bottomrule
\end{tabularx}
\begin{fnnote}[Copy proportional to the actual width]
The sequencer copies exactly \code{n\_neurons\_real} bytes of \code{y\_bus} into the
ping-pong buffer (not the full build width): a narrower layer is copied faster, without
zero-padding in RAM. Each activation is read per-layer from the table, independent of the
\code{activation} register of the single-layer path.
\end{fnnote}
\section{Hierarchy of the \texttt{busy}/\texttt{done} signals}
In the multi-layer path, \code{STATUS.busy} is the OR of the single-layer and sequencer
busy signals, while \code{STATUS.done} latches only at completion of the \emph{last}
layer, not at each intermediate layer (ch.~\ref{ch:spi}). The top-level returns control
of \code{neuron\_memory} to the direct \op{START} path at the end of the sequence.
@@ -1,190 +0,0 @@
\chapter[Graph network (Type \#2)]{Two-level configuration: graph network (Type \#2)}
\label{ch:grafo}
\section{Two network types}
The engine exposes two \emph{network types} selectable by the host, with the same start
command dispatching to the correct engine:
\begin{itemize}
\item \textbf{Type \#1 --- classic network (dense).} Layers with neurons per layer, fully
connected between consecutive layers. It is the \code{layer\_sequencer} path
(ch.~\ref{ch:seq}), started by \op{RUN\_NETWORK}. Connections are \emph{implicit by
position}: nothing is enumerated, only the weights are defined, addressed as
\code{w\_base + k*n\_inputs + j}.
\item \textbf{Type \#2 --- arbitrary graph (sparse).} Starting from the input neuron ids,
each neuron's connections up to the output are defined through a per-neuron \emph{sparse
edge-list}. Connections are \emph{explicit by enumeration}: each connection is an edge
\code{(src\_id, weight)}; if it is not in the list, it does not exist.
\end{itemize}
\begin{fnnote}[The difference in one line]
Dense: you define the \emph{weights} by position in a matrix. Graph: you define each
\emph{connection} as an edge \code{(src\_id, weight)} in a per-neuron list. The two
descriptor tables share the same 11-byte format but different fields; the
\code{net\_type} register tells the engine which interpretation to use.
\end{fnnote}
\section{Global activation buffer}
Type \#2 introduces an \textbf{activation buffer} indexed by \emph{signal id}, one INT8
byte per id, implemented in \textbf{on-chip \code{DP16KD} block RAM}
(\code{rtl/act\_buffer.v}). Ids \code{0..N\_in-1} are the inputs; each neuron writes its
own output into its own id. The source gather reads from here with \emph{single-cycle
random access}: this is what makes the graph cheap, because it is the access that PSRAM
(70~ns, sequential) could not accelerate.
\begin{fnspec}[V1 sizing]
\code{N\_TOTAL}=4096 signals, 16-bit id (room to 65\,536 without changing the format).
Buffer = 4~KB, i.e. 2 \code{DP16KD} blocks out of 108. The real constraint becomes the
PSRAM edge capacity ($\approx$2\,M edges at 4~B), not block RAM.
\end{fnspec}
\section{Feed-forward DAG and the \texttt{src\_id < out\_id} rule}
The graph is a feed-forward DAG: every connection points to an \textbf{already-computed}
id (\code{src\_id < out\_id}). Neurons are processed in ascending id order, so that when a
neuron is computed all its sources are ready in the buffer. Cycles and recurrence are out
of scope for V1. The rule is checked at two levels: by the host assembler (compile time)
and by a runtime guard in \code{graph\_engine} (\code{STATUS.err}), in the same philosophy
as the elaboration guard on \code{N\_INPUTS \% PARALLEL}.
\section{Data formats}
Both descriptors are 11~bytes/entry, MSB-first, at \code{table\_base}.
\subsection{Type \#2 descriptor (graph)}
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.6cm} Y}
\toprule
\rowh \thd{Field} & \thd{Bytes} & \thd{Meaning} \\
\midrule
\code{conn\_ptr} & 3 & Byte address in PSRAM of the neuron's edge block. \\
\rowa \code{n\_conn} & 2 & Real connections (pre-padding). \\
\code{out\_id} & 2 & Id into which the neuron's output is written. \\
\rowa \code{activation} & 1 & \code{ACT\_RELU} / \code{ACT\_NONE} (low 2 bits). \\
\code{bias} & 1 & Neuron bias (INT8). \\
\rowa \code{reserved} & 2 & 0. \\
\midrule
\rowh \thd{Total} & \thd{11} & entries in ascending \code{out\_id} order \\
\bottomrule
\end{tabularx}
\subsection{Graph edge (4~bytes, aligned)}
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.6cm} Y}
\toprule
\rowh \thd{Field} & \thd{Bytes} & \thd{Meaning} \\
\midrule
\code{src\_id} & 2 & Source id (uint16 BE). \\
\rowa \code{weight} & 1 & Weight (INT8). \\
\code{reserved} & 1 & 0 (4-byte alignment). \\
\bottomrule
\end{tabularx}
\begin{fnnote}[Padding to \texttt{PARALLEL}]
An arbitrary \code{n\_conn} is not a multiple of \code{PARALLEL}: the neuron's edge-list
is padded up to the multiple with \textbf{zero-weight} edges (waste
$\le$\code{PARALLEL}$-1$ per neuron). This keeps the datapath and its guard intact.
\end{fnnote}
\section{\texttt{graph\_engine} --- graph engine}
\code{rtl/graph\_engine.v} orchestrates Type \#2 \textbf{reusing \code{neuron\_parallel}
unmodified}, as \code{neuron\_memory} does for the dense case. Key difference: between the
two modes only the \emph{X addressing} changes. In Type \#1 the input is contiguous
(\code{x\_base + i}); in Type \#2 it is a gather (\code{act\_buf[src\_id]}). The arithmetic
core is untouched.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=4mm,start chain=going below,
every node/.style={on chain,fnblock,minimum width=52mm}]
\node[fnblockA]{\code{COPY\_INPUTS}: PSRAM \code{x\_base} $\to$ \code{act\_buf[0..N\_in-1]}};
\node{\code{READ\_DESC}: descriptor of neuron k};
\node{\code{READ\_EDGES}: stream edges + gather \code{act\_buf[src\_id]}};
\node{\code{START\_N} / \code{WAIT\_N}: group of \code{PARALLEL} $\to$ \code{neuron\_parallel}};
\node[fnblockT]{\code{WRITE\_ACT}: y $\to$ \code{act\_buf[out\_id]}};
\node{next neuron (id order)};
\node[fnblockD]{\code{WRITE\_OUTPUTS}: last \code{n\_out} $\to$ PSRAM \code{out\_base}};
\foreach \i [count=\j from 2] in {1,...,6}
\draw[fnarrow] (chain-\i) -- (chain-\j);
\end{tikzpicture}
\end{center}
The outputs are the \textbf{last \code{n\_out}} ids: in a DAG with the
\code{src\_id < out\_id} ordering the output neurons (sinks, not reused as sources)
naturally end up with the highest ids. At the end \code{graph\_engine} copies these
\code{n\_out} bytes into a PSRAM region at \code{out\_base}, which the host reads back with
\op{READ\_RAM}.
\section{Type \#2 opcodes and registers}
The type is selected with a new opcode; \op{RUN\_NETWORK} dispatches on the
\code{net\_type} register (details in ch.~\ref{ch:spi}).
\begin{tabularx}{\textwidth}{L{2.6cm} L{3.4cm} Y}
\toprule
\rowh \thd{Opcode / sel} & \thd{Name} & \thd{Function} \\
\midrule
\op{0x11} & SET\_NET\_TYPE & \code{type(1B)}: \code{0x01}=dense (\#1), \code{0x02}=graph (\#2). Default after \op{RESET}=dense. \\
\rowa \code{SET\_BASE sel 9} & num\_neurons\_graph & Number of graph neurons (uint16). \\
\code{SET\_BASE sel 10} & n\_out & Number of output ids (uint16). \\
\bottomrule
\end{tabularx}
\begin{fnnote}[Zero regression on Type \#1]
With \code{net\_type=dense} (the default value after \op{RESET}) the \#1 path is
bit-identical to before: \op{RUN\_NETWORK} keeps its \code{num\_layers(1B)} payload and the
framing of the existing opcodes does not change.
\end{fnnote}
\section{Occupancy (Type \#2 enabled)}
Yosys synthesis of the full \code{spi\_neuron\_top} system with Type \#2 enabled
(\code{PARALLEL}=2):
\begin{tabularx}{\textwidth}{L{4.6cm} Y}
\toprule
\rowh \thd{Resource} & \thd{Use} \\
\midrule
\code{DP16KD} (block RAM) & 2 (activation buffer) \\
\rowa \code{MULT18X18D} (DSP) & 4 (2 \code{neuron\_memory} + 2 \code{graph\_engine}) \\
LUT4 & 2619 \\
\rowa TRELLIS\_FF & 2467 \\
\code{\$\_TBUF\_} (PSRAM bus) & 16 \\
\bottomrule
\end{tabularx}
The device (108 \code{DP16KD}, 72 DSP, $\approx$44k LUT/FF) stays well below saturation:
Type \#2 adds a complete mode at a contained resource cost. LUT4/TRELLIS\_FF grew from an
earlier measurement (2367/2406) because of the PSRAM page mode added to the controller
(ch.~\ref{ch:mem}, \S~5.5) --- under 6\% utilization, no practical impact.
\section{Gather bandwidth (measured)}
The per-edge gather cost was \textbf{isolated} by building two structurally-identical
graphs with different edge counts and differencing the cycles: the subtraction cancels the
fixed per-neuron overhead and leaves the edge cost alone.
\begin{fnspec}[Per-edge cost]
\textbf{37.53 cycles/edge} with PSRAM page mode enabled (ch.~\ref{ch:mem}, \S~5.5) ---
\textbf{53.25 cycles/edge} without it (pre-page-mode baseline, consistent with theory:
4~bytes/edge $\times$ $\approx$13 cycles/byte over async PSRAM $\approx$52). At 80~MHz:
$\approx$2.13\,M edges/s ($\approx$8.5~MB/s, +42\% vs. baseline); at the real 16~MHz
clock: $\approx$426\,k edges/s ($\approx$1.71~MB/s).
\end{fnspec}
Page-mode read (roadmap G7, ch.~\ref{ch:roadmap}) has been implemented and measured: the
gather's sequential access benefits directly, cutting the per-edge cost by 29.5\%
(53.25$\to$37.53 cycles/edge). Each edge still pays \code{int8\_memory\_access}'s
byte-granular access (4 bytes/edge); page mode reduces the cost of each sequential byte,
not the number of accesses.
\section{\texttt{netasm} host assembler}
Readable network configuration needs no dedicated FPGA logic: a pseudo-assembly is
compiled \emph{on the host} (\code{tools/netasm/}) into the exact bytes of the tables and
edges, then loaded with \op{WRITE\_RAM}. The assembler validates at compile time
(\code{src\_id < out\_id}, \code{N\_TOTAL} bounds, padding to \code{PARALLEL}),
complementing the runtime guard.
\begin{lstlisting}[language=,caption={Pseudo-assembly example (graph)},basicstyle=\ttfamily\scriptsize]
NET graph
INPUTS 4 ; ids 0..3
NEURON n4 relu bias=2
CONN 0 w=5
CONN 1 w=-3
NEURON n5 none bias=0
CONN n4 w=2 ; symbolic reference to n4's output
CONN 2 w=7
OUTPUT n5
END
\end{lstlisting}
-290
View File
@@ -1,290 +0,0 @@
\chapter{SPI host interface}
\label{ch:spi}
\section{Physical layer}
The FPGA is always an SPI \textbf{slave}. The v1 protocol uses SPI \textbf{Mode~0}
(CPOL=0, CPHA=0), MSB-first, single-SPI. One command per low-CS period; byte~0 of each
transaction is the opcode. Multi-byte fields are big-endian.
\begin{fnspec}[Mode 0 sampling]
\code{mosi} is sampled on the \textbf{rising} edge of \code{sclk}; \code{miso} is driven
on the \textbf{falling} edge (stable before the master's next sampling). \code{spi\_slave}
synchronizes \code{sclk/mosi/cs\_n} with a double flip-flop (3-stage CDC) before every
edge detection.
\end{fnspec}
\begin{center}
\begin{tikztimingtable}[timing/dslope=0.1,timing/.style={x=3.4ex,y=2.2ex},
xscale=1.0,font=\scriptsize]
\sig{CS\_N} & H 1L 16L 1H \\
\sig{SCLK} & L 1L {2C(2)}8{2C(2)} 6L \\
\sig{MOSI} & U 1U 2D{b7} 2D{b6} 2D{b5} 2D{b4} 2D{b3} 2D{b2} 2D{b1} 2D{b0} 2U \\
\sig{MISO} & Z 1Z 16D{data} 1Z \\
\end{tikztimingtable}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
Framing of one byte: CS falls, 8 SCLK pulses, MSB first; MISO in tri-state outside a
transaction.\end{center}
\begin{fnnote}[\texttt{tx\_byte\_req} contract]
\code{tx\_byte\_req} is a \emph{prefetch hint}, not a ``byte consumed'' event: a consumer
must advance its pointers (RAM address, response byte index) on \code{rx\_valid}, which
pulses exactly once per real byte transferred.
\end{fnnote}
\section{Framing and explicit length}
The length of RAM transfers is \textbf{explicit}, not delimited by the CS edge:
\op{WRITE\_RAM}/\op{READ\_RAM} carry a 2-byte length field, so the SPI controller only
needs a byte counter. Byte addresses are 23-bit, carried in a 3-byte field with the most
significant bit reserved to 0.
\section{Opcode table}
\renewcommand{\arraystretch}{1.16}
\begin{longtable}{C{1.1cm} L{2.4cm} L{3.9cm} L{2.4cm} L{4.0cm}}
\toprule
\rowh \thd{Op} & \thd{Name} & \thd{Payload (host$\to$FPGA)} & \thd{Response} & \thd{Function} \\
\midrule
\endfirsthead
\rowh \thd{Op} & \thd{Name} & \thd{Payload} & \thd{Response} & \thd{Function} \\ \midrule
\endhead
\bottomrule
\endfoot
\op{0x00} & NOP & --- & --- & No operation (idle/dummy clocking). \\
\rowa \op{0x01} & WRITE\_RAM & addr(3B)+len(2B)+data & --- & Writes a block into PSRAM (X, weights, bias, parameters). \\
\op{0x02} & READ\_RAM & addr(3B)+len(2B) & \code{len} bytes & Reads a block back from PSRAM. \\
\rowa \op{0x0F} & RESET & --- & --- & Synchronous reset of the engine and clearing of the STATUS latch; does not erase PSRAM. \\
\op{0x10} & SET\_BASE & sel(1B)+addr(3B) & --- & Sets the bases/registers (see §\ref{sec:setbase}). \\
\rowa \op{0x11} & SET\_NET\_TYPE & type(1B) & --- & Network type: \code{0x01}=dense (\#1), \code{0x02}=graph (\#2). Default after RESET=dense. \\
\rowa \op{0x20} & START & --- & --- & Starts \code{neuron\_memory} (single-layer path); ignored if busy. \\
\op{0x21} & STATUS & --- & 1 byte & bit0=\code{busy} (live), bit1=\code{done} (sticky, clear-on-read), bit2=\code{err} (graph guard), bit3=\code{flash\_err} (sticky, clear-on-read), bit4=\code{flash\_busy} (live); bit7:5=0. \\
\rowa \op{0x22} & READ\_OUTPUT & --- & \code{N\_NEURONS} bytes & \code{y\_bus} neuron-major (byte~0 = neuron~0); dense path only (Type \#1). \\
\op{0x23} & RUN\_NETWORK & num\_layers(1B) & --- & Starts execution: dispatches on \code{net\_type} to \code{layer\_sequencer} (\#1) or \code{graph\_engine} (\#2); ignored if busy. \\
\rowa \op{0x30} & READ\_CONFIG & --- & 11 bytes & Hardware configuration record (§\ref{sec:readcfg}). \\
\op{0x40} & FLASH\_READ\_BLOCK & flash\_addr(3B)+psram\_addr(3B)+len(3B) & --- & Raw flash$\to$PSRAM read, bypasses the catalog. \\
\rowa \op{0x41} & FLASH\_WRITE\_BLOCK & psram\_addr(3B)+flash\_addr(3B)+len(3B) & --- & Raw PSRAM$\to$flash write (internal erase-before-write + $\leq$256B Page Program loop + WIP poll, transparent to the host), bypasses the catalog. \\
\op{0x42} & FLASH\_ERASE & sector\_addr(3B) & --- & Standalone 4~KB sector erase (must be sector-aligned), bypasses the catalog. \\
\rowa \op{0x43} & CAT\_READ & --- & --- & Reloads the 16-slot catalog (on-chip registers) from the flash's reserved sector. \\
\op{0x44} & CAT\_WRITE\_SLOT & slot\_id(1B)+offset(3B)+len(3B)+type(1B) & --- & Registers/updates the slot's (offset, length, type) in the on-chip catalog and persists it to flash; marks the slot \emph{invalid} until \op{SAVE\_SLOT} confirms it. \\
\rowa \op{0x45} & LOAD\_SLOT & slot\_id(1B)+psram\_addr(3B) & --- & Flash$\to$PSRAM for the slot (offset/length from the catalog), verifies the CRC32 live; \code{STATUS.flash\_err} if the slot is invalid or the CRC does not match. \\
\op{0x46} & SAVE\_SLOT & slot\_id(1B)+psram\_addr(3B)+len(3B) & --- & PSRAM$\to$flash at the slot's already-registered offset, computes the CRC32 live; on success updates and persists the catalog entry (length, CRC, valid=1). \\
\rowa \op{0x47} & CAT\_INSPECT & slot\_id(1B) & 16 bytes & Synchronous read of an already-loaded catalog entry: offset[3]+len[3]+type[1]+valid[1]+CRC32[4]+reserved[4], MSB-first. \\
\end{longtable}
All flash opcodes are \emph{fire-and-forget}: the host polls \op{STATUS} (bit4=
\code{flash\_busy}, bit3=\code{flash\_err}) or the \code{irq\_n}/\code{data\_ready\_n} pins
for the outcome, except \op{CAT\_INSPECT}, which responds synchronously.
The 8 flash opcodes (\op{0x40}--\op{0x47}) are described in full, with design rationale and
measured real latencies, in §\ref{sec:flashspi} below.
\section{\texttt{SET\_BASE} selectors}
\label{sec:setbase}
\begin{tabularx}{\textwidth}{C{1.2cm} L{3.2cm} Y}
\toprule
\rowh \thd{sel} & \thd{Register} & \thd{Use} \\
\midrule
0 & \code{x\_base} & Input base $X$. \\
\rowa 1 & \code{w\_base} & Weight base. \\
2 & \code{bias\_addr} & Bias base. \\
\rowa 3 & \code{table\_base} & Descriptor table base (multi-layer). \\
4 & \code{buf\_a\_base} & Ping-pong buffer A. \\
\rowa 5 & \code{buf\_b\_base} & Ping-pong buffer B. \\
6 & \code{activation} & Activation (low 2 bits) --- single-layer path only. \\
\rowa 7 & \code{n\_inputs\_real} & Runtime input width (16-bit BE) --- single-layer. \\
8 & \code{n\_neurons\_real} & Runtime neuron width (16-bit BE) --- single-layer. \\
\rowa 9 & \code{num\_neurons\_graph} & Number of graph neurons (16-bit BE) --- Type \#2. \\
10 & \code{n\_out} & Number of output ids (16-bit BE) --- Type \#2. \\
\bottomrule
\end{tabularx}
Selectors 6--8 concern only the single-layer/manual path; with \op{RUN\_NETWORK} the
equivalent values are read per-layer from the descriptor table.
\begin{fnwarn}[``real=0'' edge cases fixed (2026-09-04)]
The re-certification campaign (\code{docs/validation/bugs.md}) found that several
runtime values equal to zero were unguarded, with outcomes ranging from a silently
ignored limit to a hang or arbitrary-address PSRAM writes. All five cases below are now
safe no-ops, independently verified:
\begin{itemize}
\item \code{n\_inputs\_real=0} (selector 7): completes in 1 cycle with
$y=\text{activation}(\text{bias})$ (BUG-003).
\item \code{n\_neurons\_real=0} (selector 8): completes without performing any
per-neuron computation, far faster than a full-width run (BUG-004).
\item \code{num\_neurons\_graph=0} (selector 9): completes immediately after the input
copy, without ever entering the descriptor loop (BUG-006).
\item \op{RUN\_NETWORK} with \code{num\_layers=0} (dense path): an immediate no-op ---
\textbf{before the fix it executed 256 fabricated layers, reading arbitrary PSRAM data as
descriptors} (BUG-005, CRITICAL, see \S\ref{sec:run-network} below).
\item \op{SET\_NET\_TYPE} received while a run is in progress: now silently rejected
(no effect, no SPI error) instead of remapping the arbiter's multiplexer mid-execution
--- \textbf{before the fix it caused a permanent hang of the in-progress engine}
(BUG-007, CRITICAL).
\end{itemize}
Details, evidence, and per-fix verification are in \code{docs/validation/bugs.md}.
\end{fnwarn}
\section{\texttt{STATUS.done} sticky / clear-on-read}
In \code{neuron\_memory} the \code{done} signal is a single-cycle pulse. A host polling
over SPI (much slower than the FPGA clock) would almost certainly miss a raw one-cycle
pulse. The SPI register bank therefore latches \code{done} into a sticky bit on the pulse
and clears it when the host reads \op{STATUS} (or \op{RESET}). The \code{busy} bit is
instead held at level for the whole computation and is read live.
\begin{fnwarn}[Race corrected (2026-09-02)]
A real race in the sticky mechanism (present since Phase~4) was corrected by latching a
\code{status\_snapshot} on acceptance of the \op{STATUS} opcode and conditioning the
clearing of the sticky bit on \code{status\_snapshot[1]} (it clears only if the byte
actually transmitted showed \code{done=1}). A \code{done} that arrives too late for a
snapshot is reported on the next poll instead of being lost.
\end{fnwarn}
\section{Host attention pins (\texttt{data\_ready\_n}, \texttt{irq\_n})}
Besides \op{STATUS} polling, the top-level exposes two active-low physical pins (bank 7,
ch.~\ref{ch:hw}) that mirror the sticky bits without an SPI transaction, handy for driving
a host GPIO/IRQ:
\begin{itemize}
\item \code{data\_ready\_n} = $\sim$\code{STATUS.done} (sticky): low when a result is ready
to read, returns high on the \op{STATUS} read (clear-on-read).
\item \code{irq\_n} = $\sim$\code{STATUS.err} (graph guard): low when \code{graph\_engine}'s
load-time guard has tripped. It is \textbf{not} clear-on-read: it clears only on \op{RESET}
or a fresh graph start, so an error is not missed between polls.
\end{itemize}
These are additive ports: they touch neither the existing opcodes nor the registers.
\begin{fnwarn}[\code{flash\_err} has no dedicated pin]
\code{STATUS.flash\_err} (bit3) is reported \textbf{only} in the \op{STATUS} byte, by
design: reusing \code{irq\_n} would have conflated it with graph-guard errors (two
independent error domains on one pin), while a flash operation is always host-initiated
with an opcode just issued, so polling \op{STATUS} right after --- already implicit in the
``fire-and-forget, then poll \op{STATUS}/\code{data\_ready\_n}'' convention --- is already a
natural fit, no extra async pin needed. \code{data\_ready\_n}, on the other hand,
\emph{also clears at the end of a flash operation}: it mirrors \code{STATUS.done} (bit1),
which now latches on a completed flash op too, not only on \op{RUN\_NETWORK}/\op{START}.
\end{fnwarn}
\section{\texttt{READ\_CONFIG}}
\label{sec:readcfg}
Fixed \textbf{11-byte} payload: it lets a single host firmware work with different
bitstreams without recompiling. The \code{N\_INPUTS}/\code{N\_NEURONS} values report the
build \emph{maximum} (the ceiling), not necessarily the currently loaded network.
\begin{tabularx}{\textwidth}{C{1.6cm} L{3.6cm} Y}
\toprule
\rowh \thd{Byte} & \thd{Field} & \thd{Source} \\
\midrule
0 & \code{ADDR\_WIDTH} (bit) & \code{neuron\_memory.ADDR\_WIDTH} \\
\rowa 1--2 & \code{N\_INPUTS} (16-bit BE) & build maximum \\
3 & \code{N\_NEURONS} & build maximum \\
\rowa 4 & \code{PARALLEL} & build parameter \\
5 & \code{DATA\_WIDTH} (bit) & build parameter \\
\rowa 6--7 & protocol version (BE) & \code{0x0001} \\
8--9 & \code{N\_TOTAL} (16-bit BE) & max graph signals (Type \#2) \\
\rowa 10 & capability flag & bit0=\code{GRAPH\_SUPPORTED}=1 \\
\bottomrule
\end{tabularx}
\section{Flash subsystem (opcodes 0x40--0x47, completed 2026-09-04)}
\label{sec:flashspi}
The FPGA has \textbf{exclusive} access to the onboard boot/persistence flash (Winbond
\code{W25Q128JV}, 16~MB SPI NOR, ch.~\ref{ch:hw} §6/§7) through a dedicated, physically
separate SPI master (\code{rtl/spi\_flash\_master.v}), never through direct host access to
the flash pins. This is \textbf{not} a filesystem: a fixed-size catalog (16 slots,
\code{rtl/flash\_slot\_manager.v}) maps \code{slot\_id}~$\to$~(offset, length, type, valid,
CRC32) in a reserved flash sector (sector 0) --- no dynamic allocation, no garbage
collection.
\begin{fnnote}[Layering (each level independently testable)]
\begin{itemize}
\item \code{rtl/spi\_flash\_master.v} --- raw SPI master toward the flash chip
(RDID/READ/WREN/PP/SE/RDSR-1). Fully independent 4-wire bus (\code{sclk}/\code{mosi}/
\code{miso}/\code{cs\_n}, all ordinary GPIO --- Phase F7, 2026-09-04): an earlier
version reused the boot \code{CCLK} pad via the ECP5 \code{USRMCLK} primitive to save
one pin, dropped because it made the ``exclusive flash bus'' claim electrically
misleading (SCLK still depended on the same pad as the config engine) and carried an
unresolved verification gap (\code{USRMCLKTS} timing never checked against the
primary Lattice sysCONFIG Usage Guide).
\item \code{rtl/flash\_copy\_engine.v} --- block-streaming engine on top: flash$\to$PSRAM
(\code{DIR\_LOAD}), PSRAM$\to$flash with internal erase-before-write + $\leq$256B
Page Program loop + WIP polling (\code{DIR\_SAVE}), standalone sector erase
(\code{DIR\_ERASE}). A low-priority master (Port D) on \code{rtl/mem\_arbiter.v}:
flash operations are ms-scale and never block inference.
\item \code{rtl/flash\_slot\_manager.v} --- the slot catalog on top of that, plus a CRC32
(\code{rtl/crc32.v}, IEEE~802.3/zlib) computed live over the real byte stream during
\op{LOAD\_SLOT}/\op{SAVE\_SLOT}, so a corrupted or partially-written slot (e.g. power
lost mid-erase) is detected even when the underlying flash operation itself reported
success.
\end{itemize}
\end{fnnote}
\begin{fnwarn}[Sector alignment is mandatory]
\op{SAVE\_SLOT} (and the raw \op{FLASH\_WRITE\_BLOCK}/\op{FLASH\_ERASE}) require the target
flash address to be 4~KB-sector-aligned --- rejected as an error otherwise, rather than a
silent partial-sector read-modify-erase-write (no scratch buffer large enough exists for
that, and every real \op{SAVE\_SLOT} already writes a whole, sector-aligned slot by
construction).
\end{fnwarn}
Full rationale, every datasheet citation, every adversarial test (CRC mismatch, never-saved
slot, page-boundary crossing, simulated power loss, arbiter contention), and the two real
bugs found and fixed during bring-up (one pre-existing in \code{psram\_controller.v}, one in
the new arbiter request handshake) are in \code{WORKLOG.md} (Phases F1-F6 entries) and
\code{docs/FPGA-Neural-Flash-Subsystem-Verification.md} (per-module coverage summary, not
repeated here).
\begin{tabularx}{\textwidth}{L{3.4cm}Y}
\toprule
\rowh \thd{Operation} & \thd{Measured real latency} \\
\midrule
ERASE (4~KB sector) & $\approx$400~ms (dominated by the flash chip's own internal tSE, independent of the host clock) \\
\rowa SAVE (256~B page, incl. its own erase) & $\approx$403~ms (same, tSE+tPP) \\
LOAD (4096~B) & 1.74~ms (2.35~MB/s) @80~MHz; 8.71~ms (0.47~MB/s) @16~MHz (purely SPI-clock-bound) \\
\bottomrule
\end{tabularx}
Full measurement methodology in \code{docs/FPGA-Neural-Flash-Subsystem-Verification.md}.
\section{Session sequences}
\subsection{Single-layer path}
\begin{lstlisting}[language=,caption={Single-layer session},basicstyle=\ttfamily\scriptsize]
RESET -> 0x0F
READ_CONFIG -> 0x30 (host learns N_INPUTS/N_NEURONS/...)
WRITE_RAM (weights) -> 0x01 ...
WRITE_RAM (bias) -> 0x01 ...
SET_BASE (X/W/BIAS) -> 0x10 x3
WRITE_RAM (input X) -> 0x01 ...
START -> 0x20
poll STATUS -> 0x21 (until done=1; cleared by this read)
READ_OUTPUT -> 0x22
\end{lstlisting}
\subsection{Multi-layer path (RUN\_NETWORK)}
\label{sec:run-network}
\begin{lstlisting}[language=,caption={Multi-layer session},basicstyle=\ttfamily\scriptsize]
WRITE_RAM (descriptor table) -> 0x01 ...
WRITE_RAM (weights/bias per layer, X L0) -> 0x01 ...
SET_BASE (X/TABLE/BUF_A/BUF_B) -> 0x10 x4
RUN_NETWORK(num_layers) -> 0x23 <num_layers>
poll STATUS -> 0x21 (until done=1)
READ_OUTPUT -> 0x22 (y_bus of the final layer)
\end{lstlisting}
\begin{fnnote}[Out of scope for v1]
Dual~SPI and CRC/checksum on host transfers (SPI assumed reliable on a board trace --- not
to be confused with the flash catalog's CRC32, §\ref{sec:flashspi}, which protects a
different domain: flash$\leftrightarrow$PSRAM persistence, not the host SPI link).
\end{fnnote}
\begin{fnwarn}[\op{WRITE\_RAM}/\op{READ\_RAM} have no backpressure to the host --- a real risk, not a theoretical one]
Every received/produced byte must be fully processed by \code{spi\_engine} before the next
SCLK-driven byte boundary arrives --- reasonable for the initial bulk-loading of weights/
inputs, not a real-time path. The concrete risk: if a host issues \op{WRITE\_RAM}/
\op{READ\_RAM} before \code{psram\_controller.v}'s power-up sequence has completed
($\sim$150~\textmu s after reset, \code{STATE\_INIT}+\code{STATE\_CR\_INIT}),
\code{spi\_engine} stalls waiting for the very first PSRAM access to complete, while the
host --- not slowed by any handshake --- keeps clocking bytes. Bytes received during that
stall are \textbf{silently dropped}, with no error and no hang: just wrong data in PSRAM.
Found during the flash-subsystem work (\code{WORKLOG.md}, Phase~F5) via a minimal
\op{WRITE\_RAM}-only reproduction with no flash opcodes involved at all: it is a general
hazard for any host, not specific to the flash opcodes. \textbf{Current mitigation: a host
must wait for PSRAM power-up (or otherwise ensure the FPGA has been out of reset for
$>$150~\textmu s) before its first \op{WRITE\_RAM}/\op{READ\_RAM}.} Not fixed at the
protocol level (would need real backpressure, a larger change) --- declared here as an open
risk, not silently worked around.
\end{fnwarn}
@@ -1,232 +0,0 @@
\chapter[Network programming]{Neural network programming}
\label{ch:prog}
This chapter is the practical guide to encoding a network for FPGA-Neural: how it is laid
out in memory, which registers are set and how it is started, for both topologies. It
assumes the SPI opcodes (ch.~\ref{ch:spi}) and the descriptor formats (ch.~\ref{ch:seq},
\ref{ch:grafo}).
\section{General flow}
Whatever the type, the cycle is the same: the host \emph{builds the data structures in
RAM}, sets the \emph{base registers}, declares the \emph{network type}, \emph{starts} and
\emph{reads back} the result.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=3mm,start chain=going below,
every node/.style={on chain,fnblock,minimum width=64mm}]
\node[fnblockA]{1. \op{RESET} --- clears the engine and the STATUS latch};
\node{2. \op{SET\_NET\_TYPE} --- dense (\#1) or graph (\#2)};
\node{3. \op{WRITE\_RAM} --- tables, weights/edges, bias, input X};
\node{4. \op{SET\_BASE} --- base registers (x, table, \ldots)};
\node[fnblockT]{5. \op{RUN\_NETWORK} --- dispatch on \code{net\_type}};
\node{6. \op{STATUS} polling --- waits for \code{done}};
\node[fnblockD]{7. \op{READ\_OUTPUT} / \op{READ\_RAM} --- result};
\foreach \i [count=\j from 2] in {1,...,6} \draw[fnarrow] (chain-\i)--(chain-\j);
\end{tikzpicture}
\end{center}
\section{Registers and opcodes involved}
All base values are set with \op{SET\_BASE} \code{sel(1B)+addr(3B)}. Selectors:
\begin{tabularx}{\textwidth}{C{1.0cm} L{3.4cm} C{1.4cm} C{1.4cm} Y}
\toprule
\rowh \thd{sel} & \thd{Register} & \thd{Type \#1} & \thd{Type \#2} & \thd{Use} \\
\midrule
0 & \code{x\_base} & \checkmark & \checkmark & Input base $X$. \\
\rowa 3 & \code{table\_base} & \checkmark & \checkmark & Descriptor table. \\
4 & \code{buf\_a\_base} & \checkmark & \checkmark\textsuperscript{$\ast$} & Ping-pong A (\#1) / \code{out\_base} reuse (\#2). \\
\rowa 5 & \code{buf\_b\_base} & \checkmark & --- & Ping-pong B (\#1). \\
9 & \code{num\_neurons\_graph} & --- & \checkmark & Number of graph neurons. \\
\rowa 10 & \code{n\_out} & --- & \checkmark & Number of output ids. \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
$\ast$ In Type \#2 the ping-pong buffers are unused: selector 4 is reused as
\code{out\_base} (region into which outputs are copied). Selectors 1/2/6/7/8 concern only
the manual single-layer path (\op{START}), not \op{RUN\_NETWORK}.\end{center}
For Type \#1, the \emph{per-layer} \code{w\_base}/\code{bias\_addr} are \textbf{not} set
with \op{SET\_BASE}: they are fields of the descriptor table. \op{SET\_NET\_TYPE} defaults
to \emph{dense} after \op{RESET}, so a \#1 network works even without issuing it.
% ======================================================================
\section{Type \#1 --- dense network}
\subsection{Memory layout}
\begin{tabularx}{\textwidth}{L{3.4cm} Y}
\toprule
\rowh \thd{Structure} & \thd{Format} \\
\midrule
Input $X$ & \code{n\_inputs\_real} INT8 bytes at \code{x\_base}. \\
\rowa Weights (per layer) & Neuron-major: neuron $k$ at \code{w\_base + k*n\_inputs\_real}, \code{n\_neurons*n\_inputs} bytes. \\
Bias (per layer) & One INT8 byte per neuron at \code{bias\_addr}. \\
\rowa Descriptor table & \code{num\_layers} 11-byte entries at \code{table\_base}. \\
Buffers A/B & Ping-pong intermediate outputs. \\
\bottomrule
\end{tabularx}
Descriptor (11 bytes, MSB-first): \code{w\_base}(3) $|$ \code{bias\_addr}(3) $|$
\code{activation}(1) $|$ \code{n\_inputs\_real}(2) $|$ \code{n\_neurons\_real}(2).
\subsection{Worked example: a $4\to4\to2$ network}
Layer~0: 4 inputs, 4 neurons, ReLU. Layer~1: 4 inputs, 2 neurons, linear
(\code{PARALLEL}=2, so each \code{n\_inputs\_real} is a multiple of 2). Chosen addresses:
\code{table\_base}=\code{0x000000}, \code{x\_base}=\code{0x001000}, L0 weights/bias at
\code{0x002000}/\code{0x002100}, L1 at \code{0x002200}/\code{0x002300}, buffers at
\code{0x003000}/\code{0x003100}.
\begin{lstlisting}[language=,caption={Dense descriptor table (22 bytes)},basicstyle=\ttfamily\scriptsize]
Layer 0: 00 20 00 | 00 21 00 | 01 | 00 04 | 00 04
w_base bias_addr ReLU n_in=4 n_neu=4
Layer 1: 00 22 00 | 00 23 00 | 00 | 00 04 | 00 02
w_base bias_addr NONE n_in=4 n_neu=2
\end{lstlisting}
\begin{lstlisting}[language=,caption={SPI session (dense)},basicstyle=\ttfamily\scriptsize]
0x0F RESET
0x11 01 SET_NET_TYPE = dense
0x01 000000 0016 <22-byte table> WRITE_RAM table
0x01 002000 0010 <16-byte L0 wts> WRITE_RAM L0 weights (neuron-major)
0x01 002100 0004 <4-byte L0 bias>
0x01 002200 0008 <8-byte L1 wts>
0x01 002300 0002 <2-byte L1 bias>
0x01 001000 0004 <x0 x1 x2 x3> WRITE_RAM input X
0x10 00 001000 SET_BASE x_base
0x10 03 000000 SET_BASE table_base
0x10 04 003000 SET_BASE buf_a
0x10 05 003100 SET_BASE buf_b
0x23 02 RUN_NETWORK num_layers=2
0x21 ... poll STATUS until done=1
0x22 READ_OUTPUT -> 2 bytes (final layer)
\end{lstlisting}
\subsection{Host pseudocode (dense)}
\begin{lstlisting}[language=,caption={Encoding and loading a dense network},basicstyle=\ttfamily\scriptsize]
def load_dense(layers, X): # layers in execution order
spi(RESET); spi(SET_NET_TYPE, DENSE)
table = b""
for L in layers: # L: weights[n][k], bias[n], act, n_in, n_out
assert L.n_in % PARALLEL == 0
w = alloc(L.weights_neuron_major) # k slow, input fast
b = alloc(L.bias)
table += u24(w)+u24(b)+u8(L.act)+u16(L.n_in)+u16(L.n_out)
write_ram(TABLE_BASE, table)
write_ram(X_BASE, X)
set_base(0, X_BASE); set_base(3, TABLE_BASE)
set_base(4, BUF_A); set_base(5, BUF_B)
spi(RUN_NETWORK, len(layers))
wait_status_done()
return read_output(layers[-1].n_out)
\end{lstlisting}
% ======================================================================
\section{Type \#2 --- graph network}
\subsection{Memory layout}
\begin{tabularx}{\textwidth}{L{3.4cm} Y}
\toprule
\rowh \thd{Structure} & \thd{Format} \\
\midrule
Input $X$ & \code{N\_in} bytes at \code{x\_base}; copied into \code{act\_buf[0..N\_in-1]} at start. \\
\rowa Descriptor table & \code{num\_neurons\_graph} 11-byte entries at \code{table\_base}, in ascending \code{out\_id} order. \\
Edge blocks & Per neuron: \code{n\_conn} 4-byte edges at \code{conn\_ptr}, padded to a multiple of \code{PARALLEL} (zero-weight edges). \\
\rowa Outputs & \code{n\_out} bytes written to \code{out\_base} (=selector 4). \\
\bottomrule
\end{tabularx}
Graph descriptor (11 bytes): \code{conn\_ptr}(3) $|$ \code{n\_conn}(2) $|$ \code{out\_id}(2)
$|$ \code{activation}(1) $|$ \code{bias}(1) $|$ \code{reserved}(2). \quad
Edge (4 bytes): \code{src\_id}(2) $|$ \code{weight}(1) $|$ \code{reserved}(1). \quad
Rule: \code{src\_id < out\_id} (feed-forward DAG).
\subsection{Worked example}
4 inputs (ids 0--3). Neuron n4 (\code{out\_id}=4, ReLU, bias=2) connected to ids 0 and 1;
neuron n5 (\code{out\_id}=5, linear, bias=0) connected to n4 (id~4) and id~2; output = n5
(\code{n\_out}=1). \code{PARALLEL}=2, both have 2 connections (no padding). Addresses:
\code{table\_base}=\code{0x000000}, edges at \code{0x000100}, \code{x\_base}=
\code{0x001000}, \code{out\_base}=\code{0x002000}.
\begin{lstlisting}[language=,caption={Graph descriptors + edges},basicstyle=\ttfamily\scriptsize]
Descriptors (at 0x000000, 22 bytes):
n4: 00 01 00 | 00 02 | 00 04 | 01 | 02 | 00 00
conn_ptr n_conn out_id ReLU bias rsv
n5: 00 01 08 | 00 02 | 00 05 | 00 | 00 | 00 00
conn_ptr n_conn out_id NONE bias rsv
Edge blocks (at 0x000100, 4 bytes/edge: src_id, weight, rsv):
n4 @0x000100: 00 00 05 00 (src=0, w=+5)
00 01 FD 00 (src=1, w=-3) ; -3 = 0xFD
n5 @0x000108: 00 04 02 00 (src=4, w=+2) ; id4 = n4's output
00 02 07 00 (src=2, w=+7)
\end{lstlisting}
\begin{lstlisting}[language=,caption={SPI session (graph)},basicstyle=\ttfamily\scriptsize]
0x0F RESET
0x11 02 SET_NET_TYPE = graph
0x01 000000 0016 <22-byte table> WRITE_RAM descriptors
0x01 000100 0010 <16-byte edges> WRITE_RAM edge blocks
0x01 001000 0004 <x0 x1 x2 x3> WRITE_RAM input X
0x10 00 001000 SET_BASE x_base
0x10 03 000000 SET_BASE table_base
0x10 04 002000 SET_BASE out_base (sel 4 reuse)
0x10 09 000002 SET_BASE num_neurons_graph = 2
0x10 0A 000001 SET_BASE n_out = 1
0x23 00 RUN_NETWORK (dispatch to graph_engine)
0x21 ... poll STATUS (bit2=err if src_id>=out_id)
0x02 002000 0001 READ_RAM out_base -> 1 byte (n5 output)
\end{lstlisting}
\subsection{Host pseudocode (graph)}
\begin{lstlisting}[language=,caption={Encoding and loading a graph},basicstyle=\ttfamily\scriptsize]
def load_graph(neurons, X, n_out): # neurons sorted by ascending out_id
spi(RESET); spi(SET_NET_TYPE, GRAPH)
edges = b""; table = b""
for N in neurons: # N: out_id, conns=[(src_id,w)...], act, bias
for (src,_) in N.conns:
assert src < N.out_id and src < N_TOTAL # DAG rule
conn_ptr = EDGE_BASE + len(edges)
padded = pad(N.conns, PARALLEL, fill=(0,0)) # zero-weight edges
for (src,w) in padded:
edges += u16(src)+i8(w)+u8(0)
table += u24(conn_ptr)+u16(len(N.conns))+u16(N.out_id) \
+ u8(N.act)+i8(N.bias)+u16(0)
write_ram(TABLE_BASE, table); write_ram(EDGE_BASE, edges)
write_ram(X_BASE, X)
set_base(0, X_BASE); set_base(3, TABLE_BASE); set_base(4, OUT_BASE)
set_base(9, len(neurons)); set_base(10, n_out)
spi(RUN_NETWORK, 0) # payload ignored in graph
wait_status_done()
return read_ram(OUT_BASE, n_out)
\end{lstlisting}
\subsection{\texttt{netasm} pseudo-assembly}
The readable description is compiled by the host assembler (\code{tools/netasm/}) into
exactly the table and edge bytes above. Example equivalent to the worked graph:
\begin{lstlisting}[language=,caption={netasm: source and generated bytes},basicstyle=\ttfamily\scriptsize]
; --- source ---
NET graph
INPUTS 4 ; ids 0..3
NEURON n4 relu bias=2
CONN 0 w=5
CONN 1 w=-3
NEURON n5 none bias=0
CONN n4 w=2 ; symbolic reference -> id 4
CONN 2 w=7
OUTPUT n5
END
; --- the assembler emits ---
; assigned ids: n4=4, n5=5 (guarantees src_id < out_id)
; descriptors: 00 01 00 00 02 00 04 01 02 00 00
; 00 01 08 00 02 00 05 00 00 00 00
; edges: 00 00 05 00 00 01 FD 00 (n4)
; 00 04 02 00 00 02 07 00 (n5)
; registers: table_base, x_base, out_base, num_neurons=2, n_out=1
; compile-time checks: src_id<out_id, N_TOTAL, padding to PARALLEL
\end{lstlisting}
\begin{fnnote}[Why two encoding levels]
The host pseudocode and \code{netasm} produce the \emph{same bytes}. The former is useful
when the network is generated at runtime (e.g. trained weights); the latter when the
topology is hand-written or version-controlled as source. In both cases the FPGA receives
only tables and data via \op{WRITE\_RAM}: no on-board interpreter.
\end{fnnote}
@@ -1,78 +0,0 @@
\chapter[Arbitration and top-level]{Arbitration and top-level integration}
\label{ch:top}
\section{\texttt{mem\_arbiter} --- three-port arbiter}
A single byte-level memory master (which feeds the shared chain
\code{int8\_memory\_access} $\to$ \code{memory\_interface} $\to$ \code{psram\_controller})
is arbitrated among three requesters:
\begin{tabularx}{\textwidth}{C{1.3cm} L{3.4cm} Y}
\toprule
\rowh \thd{Port} & \thd{Master} & \thd{Accesses} \\
\midrule
A & \code{spi\_engine} & \op{WRITE\_RAM} / \op{READ\_RAM}. \\
\rowa B & \code{neuron\_memory} & X/W/bias reads during an execution. \\
C & \code{layer\_sequencer} & Descriptor reads + buffer writes between layers. \\
\bottomrule
\end{tabularx}
Fixed priority \textbf{B $>$ C $>$ A}: an inference in progress is more critical than the
sequencer's bookkeeping, which in turn is more critical than a manual SPI access that has
just arrived. In normal operation B and C are anyway temporally disjoint
(\code{neuron\_memory} requests only during an execution, \code{layer\_sequencer} only in
the pauses between layers), so the priority matters mostly for the corner case of a
manual \op{WRITE\_RAM}/\op{READ\_RAM} arriving during a multi-layer execution.
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=6mm]
\node[fnblock,minimum width=30mm](a){Port A --- \code{spi\_engine}};
\node[fnblock,below=4mm of a,minimum width=30mm](b){Port B --- \code{neuron\_memory}};
\node[fnblock,below=4mm of b,minimum width=30mm](c){Port C --- \code{layer\_sequencer}};
\node[fnblockD,right=16mm of b,minimum width=26mm,minimum height=16mm](arb){\code{mem\_arbiter}\\{\scriptsize B$>$C$>$A}};
\node[fnblockT,right=14mm of arb,minimum width=26mm](m){shared memory\\{\scriptsize chain}};
\draw[fnarrow] (a)-|(arb.west|-a); \draw[fnarrow] (b)--(arb.west);
\draw[fnarrow] (c)-|(arb.west|-c);
\draw[fnbus] (arb)--(m);
\end{tikzpicture}
\end{center}
Once access is granted, the arbiter retains ownership until the single transaction's
\code{m\_ready} pulse, then releases: all three masters emit \code{req} as a clean
one-cycle pulse, so a queue-less grant-and-forward design suffices.
\section{\texttt{spi\_neuron\_top} --- full integration}
The top-level connects SPI (\code{spi\_slave}+\code{spi\_engine}), the arbiter, the
sequencer, \code{neuron\_memory} and the PSRAM chain. The reset of \code{neuron\_memory}
is the OR of the global reset with the soft-reset pulse of the \op{RESET} opcode, so the
host can recover the engine over SPI without a physical reset (the RAM stays intact).
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=7mm]
\node[fnblockA,minimum width=22mm](ss){\code{spi\_slave}};
\node[fnblockA,right=8mm of ss,minimum width=22mm](se){\code{spi\_engine}};
\node[fnblockT,below=8mm of se,minimum width=26mm](sq){\code{layer\_sequencer}};
\node[fnblockD,right=10mm of se,minimum width=24mm](mux){ctrl MUX\\{\scriptsize on \code{seq\_busy}}};
\node[fnblock,below=8mm of mux,minimum width=26mm](nm){\code{neuron\_memory}};
\node[fnblockD,right=10mm of mux,minimum width=22mm](arb){\code{mem\_arbiter}};
\node[fnblockA,right=8mm of arb,minimum width=26mm](mem){PSRAM chain};
\draw[fnarrow] (ss)--(se);
\draw[fnarrow] (se)--(mux);
\draw[fnarrow] (sq)--(mux);
\draw[fnarrow] (mux)--(nm);
\draw[fnarrow] (se.south) to[bend right=10] (arb.north west);
\draw[fnarrow] (nm)--(arb);
\draw[fnarrow] (sq.east) to[bend right=20] (arb.south west);
\draw[fnbus] (arb)--(mem);
\end{tikzpicture}
\end{center}
The multiplexer switches the control lines of \code{neuron\_memory} between the sequencer
(while \code{seq\_busy} is high) and the direct path of \code{spi\_engine} (legacy
single-layer mode), returning the engine to the direct path at the end of the sequence.
\begin{fnnote}[End-to-end verification]
\code{spi\_neuron\_top} is verified in simulation with real PSRAM
(\code{psram\_model.v}, no mock): RESET/READ\_CONFIG/WRITE\_RAM/READ\_RAM/SET\_BASE/
START/STATUS/READ\_OUTPUT and \op{RUN\_NETWORK} are exercised purely over simulated SPI
(ch.~\ref{ch:impl}).
\end{fnnote}
@@ -1,169 +0,0 @@
\chapter[ECP5 implementation]{ECP5 implementation and characterization}
\label{ch:impl}
\section{Flow and verification}
The project is verified on two complementary planes: functional \textbf{simulation} with
Icarus Verilog (signed algebra, products, accumulation, groups, bias, ReLU, saturation,
busy/done signals) and real \textbf{implementation} with Yosys (synthesis) $+$
nextpnr-ecp5 (place\&route, timing) $+$ Project~Trellis (\code{ecppack}).
\begin{tabularx}{\textwidth}{L{5.0cm} C{3.0cm} Y}
\toprule
\rowh \thd{Verification stage} & \thd{Outcome} & \thd{Covers} \\
\midrule
Functional RTL & \PASS & datapath correctness \\
\rowa Parametric simulation & \PASS & configuration sweep \\
ECP5 synthesis & \PASS & synthesizability, mapping \\
\rowa Placement / Routing & \PASS & LUT/FF/DSP, timing \\
Bitstream (\code{ecppack}) & \PASS & full flow, 0 errors (P2 and P8) \\
\bottomrule
\end{tabularx}
\begin{fnnote}[End-to-end toolchain through the bitstream]
The full flow RTL $\to$ Yosys $\to$ nextpnr-ecp5 $\to$ \code{ecppack} produces a valid
bitstream for P2 and P8, \textbf{0 errors at every stage}. Header verified byte-by-byte:
\code{Part: LFE5U-45F-8CABGA381}, the target's real part number, not a placeholder. Only
\emph{generation} is verified: no physical-hardware test in this session.
\end{fnnote}
\section{Datapath benchmark (256$\times$4)}
Configuration: INT8/INT32, \code{N\_INPUTS}=256, \code{N\_NEURONS}=4, variable
\code{PARALLEL}, 80~MHz target, device \code{LFE5U-45F-8BG381C} ($-8$). The test buses
are generated \emph{inside} the benchmark wrapper so as not to expose thousands of I/Os;
the top-level exposes only \code{clk/rst/start/y\_bus/busy/done}.
\begin{tabularx}{\textwidth}{C{1.4cm} C{1.8cm} C{1.4cm} C{1.6cm} C{1.6cm} C{1.5cm} C{1.4cm}}
\toprule
\rowh \thd{PAR} & \thd{tot MAC} & \thd{DSP} & \thd{Fmax} & \thd{Tcrit} & \thd{80\,MHz} & \thd{LUT4} \\
\midrule
16 & 64 & 64/72 & 52.13 & 19.18 & \FAIL & $\approx$2531 \\
\rowa 8 & 32 & 32/72 & 61.71 & 16.20 & \FAIL & --- \\
4 & 16 & 16/72 & 75.01 & 13.33 & \FAIL & 804 \\
\rowa 2 & 8 & 8/72 & 87.88 & 11.38 & \PASS & 481 \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
Fmax and Tcrit in MHz and ns. Total MACs $=$ PARALLEL$\times$4 neurons.\end{center}
\subsection{Fmax and throughput versus parallelism}
\begin{center}
\begin{tikzpicture}
\begin{axis}[
width=0.62\textwidth,height=6.0cm,
axis y line*=left, axis x line=bottom,
xlabel={\footnotesize PARALLEL}, ylabel={\footnotesize Fmax [MHz]},
xtick={2,4,8,16}, xmode=log, log basis x=2,
ymin=40,ymax=95, ytick={40,55,70,85},
tick label style={font=\scriptsize}, label style={font=\footnotesize},
grid=major, grid style={fnRule!40},
legend style={font=\scriptsize,at={(0.5,-0.28)},anchor=north,legend columns=2}]
\addplot[fnTeal,mark=*,thick,mark options={fill=fnTeal}]
coordinates {(2,87.88)(4,75.01)(8,61.71)(16,52.13)};
\addlegendentry{Fmax}
\draw[fnAmber,dashed,thick] (axis cs:2,80)--(axis cs:16,80);
\node[font=\scriptsize,text=fnAmber] at (axis cs:11,82.5){80 MHz target};
\end{axis}
\begin{axis}[
width=0.62\textwidth,height=6.0cm,
axis y line*=right, axis x line=none,
xmode=log, log basis x=2, xmin=2,xmax=16,
ylabel={\footnotesize throughput [G\,MAC/s]},
ymin=0,ymax=3.6, ytick={0,1,2,3},
tick label style={font=\scriptsize}, label style={font=\footnotesize}]
\addplot[fnBlue,mark=square*,thick,mark options={fill=fnBlue}]
coordinates {(2,0.703)(4,1.20)(8,1.97)(16,3.34)};
\label{plt:tp}
\end{axis}
\end{tikzpicture}
\end{center}
\begin{center}\footnotesize\itshape\color{fnGrey}
Fundamental trade-off: as PARALLEL grows, Fmax drops (deeper routing/tree) but the
theoretical throughput rises. The blue line (squares) is the throughput
$\approx$MAC/cycle$\times$Fmax.\end{center}
\subsection{Interpretation}
Reducing \code{PARALLEL} lowers simultaneous MACs, DSPs, adder-tree depth and routing
congestion, so Fmax rises; but the number of groups increases and hence the latency.
Frequency alone is not enough to choose: what matters is the overall throughput
$\approx$MAC/cycle$\times$frequency.
\begin{fnnote}[Architectural choices]
\code{PARALLEL=8} is the candidate for the throughput-oriented V1: exactly 32~simultaneous
MACs with 4 neurons, DSP at $\approx$44\%, leaving resources for controller, buffers, SPI
and future pipelines. \code{PARALLEL=2} is the frequency-oriented reference: 87.88~MHz,
the only one to exceed the 80~MHz target, but it requires 128 groups for a 256-input
neuron.
\end{fnnote}
\subsection{Critical path and the 100~MHz limit}
The 100~MHz target is not met (best result 87.88~MHz with P2). The limit is
\emph{temporal}, not one of occupancy: with P2 the FPGA is barely used (DSP $\approx$11\%,
LUT $\approx$1\%). The critical path runs through weight FF $\to$ \code{MULT18X18D} $\to$
products $\to$ adder/carry $\to$ \code{acc\_next} $\to$ ReLU/saturation $\to$ output FF.
Exceeding 100~MHz will require one or more internal pipelines, not yet necessary to
proceed.
\section{Full integrated system}
Real synthesis of \code{spi\_neuron\_top} (SPI + arbiter + \code{neuron\_memory} +
\code{graph\_engine} + PSRAM chain), speed grade $-8$. Before timing closure the integrated
system missed the 80~MHz target (P2 $\approx$55~MHz, P8 $\approx$45~MHz), with a critical
path entirely inside \code{neuron\_parallel}.
\subsection{Cause: the saturation/ReLU carry chain}
Resource usage is not the cause (device below 10\% everywhere). The integrated system's
critical path is the \textbf{\code{CCU2C} carry chain of the saturation/ReLU comparator} in
\code{neuron\_parallel.v} --- \emph{not} SPI, arbiter, PSRAM, nor the Type~\#2 modules. The
saturation was written as an arithmetic comparison (\code{acc > 127}, \code{acc < -128}),
mapped by the synthesizer onto a 32-bit subtractor with a long carry chain.
\subsection{Timing closure (2026-09-03)}
An explicit waiver of the ``datapath untouchable'' rule for a separate timing-closure task,
with the single constraint of \textbf{bit-exact equivalence} across the whole regression.
Two steps:
\begin{itemize}
\item \textbf{Step 1 --- saturation/ReLU as bit-test.} A signed 32-bit value fits INT8 iff
\code{acc[31:7]} are all equal: an AND/OR reduction over a bit slice instead of a 32-bit
carry chain. A correct, bit-exact-verified simplification; real logic gain, but on its own
submerged by placement noise.
\item \textbf{Step 2 --- pipeline register} between accumulate and activation (\code{+1}
cycle of latency per neuron, absorbed by the \code{start}/\code{done} handshake, transparent
to callers). This is the decisive step.
\end{itemize}
\begin{tabularx}{\textwidth}{L{4.6cm} C{2.6cm} C{2.4cm} Y}
\toprule
\rowh \thd{Config} & \thd{Before} & \thd{After} & \thd{$\Delta$} \\
\midrule
P2, real \code{.lpf} & 54.58 & \textbf{75.30} & $+38\%$ \\
\rowa P2, 5-seed sweep & 55.59 & 73.38--75.55 & robust \\
P8, unconstrained & 45.47 & \textbf{60.26} & $+33\%$ \\
\rowa P8, 5-seed sweep & 43.15--50.48 & 60.26--68.87 & non-overlapping \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
Fmax in MHz, real place\&route (\code{nextpnr-ecp5}). Robust across 5 seeds, not attributable
to placement luck.\end{center}
\begin{fnnote}[Stop criterion and real margin]
80~MHz is not reached (75.30~MHz at P2, 94\% of target) but the gain is large and real
($+38\%$/$+33\%$). The next step (the \code{MULT18X18D} output register, which would touch
\code{mac\_unit.v}) was left out: 80~MHz is \emph{headroom} toward the real \code{.lpf}, not
an operating requirement. With the planned 16~MHz oscillator, even the worst measured number
($\approx$45~MHz at P8) has $2.8\times$ of margin.
\textbf{Superseded 2026-09-04}: after adding the flash subsystem (ch.~\ref{ch:spi}
§\ref{sec:flashspi}, ch.~\ref{ch:roadmap}), Fmax for the full system (P2, same real pinout +
3 new flash signals) was 66.68~MHz, critical path still on the same
\code{neuron\_parallel} accumulator chain identified above --- not a new bottleneck, the
difference from 75.30~MHz was placement/routing noise from the added pins/logic.
\textbf{Updated again the same day (Phase F7)}: made the flash SPI bus genuinely
independent (dropped the \code{CCLK}/\code{USRMCLK} reuse, added a 4th ordinary
\code{flash\_sclk} pin), Fmax re-measured \textbf{67.91~MHz} (slight improvement, critical
path confirmed still identical). Margin on the 16~MHz oscillator: $4.2\times$.
\end{fnnote}
\begin{fnnote}[Separate future optimization]
Independent of timing closure: the \code{x\_mem}/\code{w\_mem} arrays of \code{neuron\_memory}
are still inferred as distributed RAM on LUTs instead of \code{DP16KD}. Moving them to block
RAM would free LUTs and is a Phase~7 candidate --- but it was not on the critical path
resolved here.
\end{fnnote}
@@ -1,326 +0,0 @@
\chapter[Hardware design and pinout]{Hardware design and signal map}
\label{ch:hw}
\begin{fnnote}[Pinout status --- assigned and verified]
A real \code{.lpf} now exists (\code{synth/ecp5/spi\_neuron\_top.lpf}) with the top-level's
\textbf{57 signals} assigned to concrete CABGA381 balls, \textbf{verified by a full
0-error \code{nextpnr-ecp5} place\&route} (no longer \code{-{}-lpf-allow-unconstrained}). The
balls come from Project~Trellis's device database (\code{iodb.json}, the same nextpnr uses)
and were independently validated against §4.3.2 of the official Lattice datasheet (per-bank
GPIO counts: exact match on 6 of 7 banks, off by 1 ball on bank~3, immaterial since no
assigned signal uses it). \code{TRELLIS\_IO}: 57/245 (23\%). Current-build Fmax (full system
incl. flash subsystem with an independent SPI bus, Phase F7, 2026-09-04) \textbf{67.91~MHz},
critical path confirmed still on \code{neuron\_parallel}'s accumulator chain, unchanged from
earlier builds (ch.~\ref{ch:impl}). Pin-by-pin summary at the front of the document
(pp.~2--3). The boot config-SPI and JTAG balls do not appear here: they are dedicated
fixed-function pins with no corresponding RTL port, nextpnr never requires them (0 errors),
they matter only for the PCB schematic.
\end{fnnote}
\section{Target device}
\begin{tabularx}{\textwidth}{L{4.2cm}Y}
\toprule
\rowh \thd{Parameter} & \thd{Value} \\
\midrule
Device & Lattice ECP5 \code{LFE5U-45F-8BG381C} \\
\rowa Package & CABGA381 (381 balls) \\
Speed grade & $-8$ (the fastest of the ECP5 family) \\
\rowa Resources & $\approx$44k LUT/FF, 72$\times$\code{MULT18X18D}, \code{DP16KD} block RAM \\
Usable I/O & $\approx$232 balls out of 381 (rest: power/ground/NC) \\
\bottomrule
\end{tabularx}
\section{Pin budget}
The project requires about 60 signals out of $\approx$232 usable I/Os: ample margin
($>$170 free pins), so the board is not pin-constrained.
\begin{tabularx}{\textwidth}{Y C{2.2cm}}
\toprule
\rowh \thd{Function} & \thd{Pins} \\
\midrule
PSRAM (22 address, 16 data, 6 control) & up to 44 \\
\rowa Application SPI (\code{sclk/mosi/miso/cs\_n}) & 4 \\
Clock, reset & 2 \\
\rowa Host attention pins (\code{irq\_n}, \code{data\_ready\_n}) & 2 \\
Flash runtime SPI bus (\code{flash\_sclk/flash\_mosi/flash\_miso/flash\_cs\_n}, ordinary GPIO, fully independent bus --- Phase F7) & 4 \\
\rowa JTAG (bring-up / debug, recommended) & 4 \\
\midrule
\rowh \thd{Total} & \thd{$\approx$60} \\
\bottomrule
\end{tabularx}
\section{Signal map (top-level \texttt{spi\_neuron\_top}) --- real balls}
Real assignment of the top-level's 57 signals, verified by place\&route, \textbf{an
individual ball for every bit} (never a bus range). I/O standard: LVCMOS33 (3.3~V I/O
supply). The balls come from the real place\&route-verified \code{.lpf}. A compact summary
of the same table also appears at the front of the document (pp.~2--3).
\renewcommand{\arraystretch}{1.1}
\begin{tabularx}{\textwidth}{L{3.0cm} C{1.0cm} C{1.9cm} C{1.0cm} Y}
\toprule
\rowh \thd{Signal} & \thd{Dir} & \thd{Ball} & \thd{Bank} & \thd{Function} \\
\midrule
\multicolumn{5}{l}{\textit{\color{fnDark}Clock and reset (bank 7, left edge)}}\\
\code{clk} & IN & H5 & 7 & System clock on pad \code{GR\_PCLK7\_0} (dedicated global clock). \\
\rowa \code{rst} & IN & B4 & 7 & Global synchronous reset, active high. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Application SPI (bank 7, opposite the PSRAM bus)}}\\
\code{sclk} & IN & B5 & 7 & SPI clock (CPOL=0, CPHA=0). \\
\rowa \code{mosi} & IN & C5 & 7 & Master-Out Slave-In. \\
\code{miso} & OUT & A3 & 7 & Master-In Slave-Out (driven on the falling edge). \\
\rowa \code{cs\_n} & IN & B3 & 7 & Active-low chip-select. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Host attention pins (bank 7, active-low, level)}}\\
\code{data\_ready\_n} & OUT & C3 & 7 & Low while a result awaits reading (mirrors \code{STATUS.done}, clear on STATUS read). \\
\rowa \code{irq\_n} & OUT & C4 & 7 & Low if the graph load-time guard has tripped (mirrors \code{STATUS.err}); clears only on \code{RESET} or a fresh \code{run\_start}, \emph{not} on a STATUS read. \\
\multicolumn{5}{l}{\textit{\color{fnDark}Flash subsystem --- independent SPI bus toward the onboard W25Q128JV (bank 7, Phases F1-F7)}}\\
\code{flash\_sclk} & OUT & E3 & 7 & SPI clock toward the flash --- ordinary GPIO, no config primitive involved (Phase F7). \\
\rowa \code{flash\_mosi} & OUT & D3 & 7 & Master-Out Slave-In toward the flash. \\
\code{flash\_miso} & IN & D5 & 7 & Master-In Slave-Out from the flash. \\
\rowa \code{flash\_cs\_n} & OUT & E4 & 7 & Flash chip-select, active low. \\
\multicolumn{5}{l}{\textit{\color{fnDark}PSRAM address bus \code{psram\_a[21:0]} --- 22 individual balls (bank 2)}}\\
\code{psram\_a[0]} & OUT & E16 & 2 & PSRAM A0 \\
\rowa \code{psram\_a[1]} & OUT & F16 & 2 & PSRAM A1 \\
\code{psram\_a[2]} & OUT & D18 & 2 & PSRAM A2 \\
\rowa \code{psram\_a[3]} & OUT & E17 & 2 & PSRAM A3 \\
\code{psram\_a[4]} & OUT & E18 & 2 & PSRAM A4 \\
\rowa \code{psram\_a[5]} & OUT & F18 & 2 & PSRAM A5 \\
\code{psram\_a[6]} & OUT & F17 & 2 & PSRAM A6 \\
\rowa \code{psram\_a[7]} & OUT & G16 & 2 & PSRAM A7 \\
\code{psram\_a[8]} & OUT & G18 & 2 & PSRAM A8 \\
\rowa \code{psram\_a[9]} & OUT & H16 & 2 & PSRAM A9 \\
\code{psram\_a[10]} & OUT & H17 & 2 & PSRAM A10 \\
\rowa \code{psram\_a[11]} & OUT & H18 & 2 & PSRAM A11 \\
\code{psram\_a[12]} & OUT & J16 & 2 & PSRAM A12 \\
\rowa \code{psram\_a[13]} & OUT & J17 & 2 & PSRAM A13 \\
\code{psram\_a[14]} & OUT & C20 & 2 & PSRAM A14 \\
\rowa \code{psram\_a[15]} & OUT & D19 & 2 & PSRAM A15 \\
\code{psram\_a[16]} & OUT & E19 & 2 & PSRAM A16 \\
\rowa \code{psram\_a[17]} & OUT & E20 & 2 & PSRAM A17 \\
\code{psram\_a[18]} & OUT & F19 & 2 & PSRAM A18 \\
\rowa \code{psram\_a[19]} & OUT & F20 & 2 & PSRAM A19 \\
\code{psram\_a[20]} & OUT & G20 & 2 & PSRAM A20 \\
\rowa \code{psram\_a[21]} & OUT & H20 & 2 & PSRAM A21 \\
\code{psram\_a[22]} & OUT & P18 & 3 & Always 0 (byte$\to$word shift): NC on the board. \\
\multicolumn{5}{l}{\textit{\color{fnDark}PSRAM data bus \code{psram\_dq[15:0]} --- 16 individual balls (banks 2 and 3)}}\\
\rowa \code{psram\_dq[0]} & IO & K18 & 2 & PSRAM DQ0 \\
\code{psram\_dq[1]} & IO & C18 & 2 & PSRAM DQ1 (dual-function ball, used as ordinary GPIO). \\
\rowa \code{psram\_dq[2]} & IO & D17 & 2 & PSRAM DQ2 \\
\code{psram\_dq[3]} & IO & D20 & 2 & PSRAM DQ3 \\
\rowa \code{psram\_dq[4]} & IO & G19 & 2 & PSRAM DQ4 \\
\code{psram\_dq[5]} & IO & J18 & 2 & PSRAM DQ5 \\
\rowa \code{psram\_dq[6]} & IO & J19 & 2 & PSRAM DQ6 \\
\code{psram\_dq[7]} & IO & J20 & 2 & PSRAM DQ7 \\
\rowa \code{psram\_dq[8]} & IO & K19 & 2 & PSRAM DQ8 \\
\code{psram\_dq[9]} & IO & K20 & 2 & PSRAM DQ9 \\
\rowa \code{psram\_dq[10]} & IO & L17 & 3 & PSRAM DQ10 \\
\code{psram\_dq[11]} & IO & M18 & 3 & PSRAM DQ11 \\
\rowa \code{psram\_dq[12]} & IO & M17 & 3 & PSRAM DQ12 \\
\code{psram\_dq[13]} & IO & N16 & 3 & PSRAM DQ13 \\
\rowa \code{psram\_dq[14]} & IO & N18 & 3 & PSRAM DQ14 \\
\code{psram\_dq[15]} & IO & P17 & 3 & PSRAM DQ15 (bidirectional tri-state data bus, \code{dq\_oe} = direction). \\
\multicolumn{5}{l}{\textit{\color{fnDark}PSRAM control (bank 3)}}\\
\rowa \code{psram\_ce\_n} & OUT & N17 & 3 & Chip enable, active low. \\
\code{psram\_oe\_n} & OUT & R16 & 3 & Output enable (read). \\
\rowa \code{psram\_we\_n} & OUT & R17 & 3 & Write enable (write). \\
\code{psram\_lb\_n} & OUT & T16 & 3 & Lower-byte enable (DQ[7:0]). \\
\rowa \code{psram\_ub\_n} & OUT & N19 & 3 & Upper-byte enable (DQ[15:8]). \\
\code{psram\_zz\_n} & OUT & N20 & 3 & Sleep/snooze (inactive=high in operation). \\
\bottomrule
\end{tabularx}
\renewcommand{\arraystretch}{1.25}
\begin{fnnote}[Board signals not exposed as RTL ports]
Not ports of \code{spi\_neuron\_top} but required at board level: the \textbf{configuration
SPI} lines to the onboard NOR flash (\code{PROGRAMN}/\code{INITN}/\code{DONE}/\code{CCLK}\ldots,
the datasheet's ``Miscellaneous Dedicated Pins'') and the 4 \textbf{JTAG} lines
(\code{TCK}/\code{TMS}/\code{TDI}/\code{TDO}), the \textbf{oscillator} on the \code{PCLK}
pad, the \textbf{power supplies}. Their ball numbers are not in the Lattice datasheet
(separate file) but are not needed here: dedicated pins with no RTL port, nextpnr never
requires them (0 errors), they matter only for the PCB schematic.
\end{fnnote}
\begin{fnwarn}[Application SPI separate from configuration SPI]
The application SPI (\code{sclk/mosi/miso/cs\_n}) must land on ordinary I/Os,
\textbf{never} on the configuration-SPI pins: the config-SPI clock pin is not reusable as
a general-purpose input after configuration without a board-level workaround. Keeping them
physically separate avoids that problem.
\end{fnwarn}
\section{Per-bank allocation (real die geometry)}
The placement follows the die-edge geometry (from Trellis's \code{globals.json},
ball~$\to$~(col,row)~$\to$~bank): banks \textbf{2 and 3} sit contiguously along the chip's
\textbf{right} edge and together hold the entire PSRAM bus (44+1 signals) --- exactly the
``one or two adjacent banks'' recommended. Bank \textbf{7} (\textbf{left} edge, physically
opposite the PSRAM bus) holds the application SPI and clock/reset, deliberately on the far
side so the two buses do not cross. \code{clk} is on the dedicated pad \code{H5}
(\code{GR\_PCLK7\_0}). Where a bank ran out of plain balls (part of \code{psram\_dq}), the
next dual-function ball was used as ordinary GPIO, confirmed usable by the real place\&route.
\begin{tabularx}{\textwidth}{Y C{1.6cm} L{4.4cm}}
\toprule
\rowh \thd{Signal group} & \thd{\# pins} & \thd{Bank (real)} \\
\midrule
PSRAM addresses \code{psram\_a[21:0]} & 22 & bank 2 (right edge) \\
\rowa PSRAM data \code{psram\_dq[15:0]} & 16 & banks 2 + 3 (adjacent) \\
PSRAM control (ce/oe/we/lb/ub/zz) & 6 & bank 3 \\
\rowa Application SPI & 4 & bank 7 (left edge) \\
Host attention pins (\code{irq\_n}, \code{data\_ready\_n}) & 2 & bank 7 \\
\rowa Independent flash SPI bus (\code{flash\_sclk/flash\_mosi/flash\_miso/flash\_cs\_n}) & 4 & bank 7 \\
Clock / reset & 2 & bank 7, \code{clk} on \code{GR\_PCLK7\_0} \\
\rowa Boot config SPI / JTAG & --- & dedicated pins (outside RTL, PCB only) \\
\bottomrule
\end{tabularx}
\section{PSRAM subsystem}
The \code{psram\_controller.v} controller implements an \textbf{asynchronous parallel}
interface (address bus, 16-bit data, \code{ce\_n/oe\_n/we\_n} and byte-lanes
\code{lb\_n/ub\_n}, plus \code{zz\_n}) with an access latency of \textbf{70~ns} wired as
$\lceil 70\,\text{ns}\times f_{clk}\rceil$. It is an asynchronous-SRAM-style bus, not QSPI.
\begin{tabularx}{\textwidth}{L{3.0cm}Y}
\toprule
\rowh \thd{Role} & \thd{Component} \\
\midrule
Working memory & ISSI \code{IS66WVE4M16EBLL-70BLI} --- 64\,Mbit parallel PSRAM (4M$\times$16, 8~MB), async, 70~ns, an exact match to the controller timing. \\
\rowa Fallback & ISSI \code{IS61WV6416DBLL} / \code{IS61WV102416BLL} (true async SRAM, drop-in on the same signals, \code{zz\_n} inactive, $\sim$10~ns, lower density). \\
Persistent storage & Winbond \code{W25Q128JV} --- 16~MB SPI NOR flash for bitstream, weights, bias, network metadata. \\
\bottomrule
\end{tabularx}
\subsection{PSRAM connection (FPGA-exclusive)}
The PSRAM is driven \textbf{exclusively by the FPGA} through \code{psram\_controller.v}: no
external master touches the bus. The external host (RPi/ESP32/MCU) only speaks SPI to the
FPGA and never touches these lines. Pin-by-pin connection FPGA~$\leftrightarrow$~ISSI
\code{IS66WVE4M16EBLL-70BLI}:
\begin{tabularx}{\textwidth}{L{3.6cm} L{3.0cm} Y}
\toprule
\rowh \thd{FPGA signal} & \thd{PSRAM pin} & \thd{Function} \\
\midrule
\code{psram\_a[21:0]} & A0--A21 & Address bus (22 lines, 8~MB word address). \\
\rowa \code{psram\_dq[15:0]} & DQ0--DQ15 & Bidirectional data bus (tri-state, \code{dq\_oe}=direction). \\
\code{psram\_ce\_n} & CE\# & Chip enable (active low). \\
\rowa \code{psram\_oe\_n} & OE\# & Output enable (read). \\
\code{psram\_we\_n} & WE\# & Write enable (write). \\
\rowa \code{psram\_lb\_n} & LB\# & Lower-byte enable (DQ[7:0]). \\
\code{psram\_ub\_n} & UB\# & Upper-byte enable (DQ[15:8]). \\
\rowa \code{psram\_zz\_n} & ZZ\# & Sleep/snooze (held high in operation). \\
\bottomrule
\end{tabularx}
PSRAM supply: \textbf{3.3~V} (BLL variant), on the same I/O rail as banks 2/3 to which it is
wired (ch.~\ref{ch:hw}, real balls). Decoupling per supply pin per the ISSI datasheet.
\section{Clock}
\label{sec:clock}
There is no PLL in the RTL yet: \code{CLK\_FREQ\_MHZ} is a \emph{timing parameter} (it
feeds the PSRAM access formulas), not a clock generator. The mounted oscillator drives
\code{clk} directly. Recommendation: a 16~MHz MEMS oscillator (SiTime SiT2001B family),
well below the 67.91~MHz Fmax of the full integrated system (incl. flash subsystem,
ch.~\ref{ch:impl}). \code{CLK\_FREQ\_MHZ} must
be set to the real value of the mounted oscillator, otherwise the PSRAM timing comes out
wrong.
\section{Power}
A \textbf{three-rail} tree (the Lattice eval board's SERDES section is not needed and is
omitted: no 1.2~V \code{VCCA}/\code{VCCHTX}):
\begin{tabularx}{\textwidth}{L{3.4cm} C{2.0cm} Y}
\toprule
\rowh \thd{Rail} & \thd{Voltage} & \thd{Feeds / regulator} \\
\midrule
\code{VCC} (core) & 1.1~V & FPGA core logic. Buck \code{TLV62568}, $\geq$600~mA. \\
\rowa \code{VCCIO0/2/3/6/7} & 3.3~V & I/O of all used banks + PSRAM. Buck \code{TLV62568}, 1~A. \\
\code{VCCAUX} & 2.5~V & FPGA auxiliary. LDO \code{TLV73325}, 10~mA. \\
\bottomrule
\end{tabularx}
Decoupling: at least one capacitor per supply pin + bulk per rail, per the Lattice ECP5
hardware checklist. Input: external 12~V (or match the bucks to the source).
\section{Configuration and programming}
\label{sec:config}
Writing the FPGA ``map'' (bitstream) happens through dedicated silicon pins, \textbf{not}
RTL top-level ports. Default mode: \textbf{MSPI} --- automatic boot from the NOR flash at
power-on (standalone product); JTAG available for development.
\subsection{JTAG (development / debug)}
\begin{tabularx}{\textwidth}{L{3.0cm} C{2.2cm} Y}
\toprule
\rowh \thd{Signal} & \thd{Ball\textsuperscript{$\dagger$}} & \thd{Function} \\
\midrule
\code{TCK} & T5 & Test clock. \\
\rowa \code{TDI} & R5 & Test data in. \\
\code{TDO} & V4 & Test data out. \\
\rowa \code{TMS} & U5 & Test mode select. \\
\bottomrule
\end{tabularx}
\subsection{Config-SPI to boot flash}
The FPGA loads the bitstream from the \textbf{Winbond \code{W25Q128JV}} (128~Mbit SPI NOR,
Quad read) at power-on. The flash subsystem (\code{rtl/flash\_slot\_manager.v}, Phases
F1-F7, ch.~\ref{ch:impl}) uses the \textbf{same physical flash} for network weights/bias/
metadata at runtime, FPGA-exclusive access: after configuration, the FPGA regains control
of the chip through a fully independent 4-wire SPI bus, \code{flash\_sclk/flash\_mosi/
flash\_miso/flash\_cs\_n} (all ordinary GPIO, pp.~2--3 and §``Signal map'' --- no ECP5
config primitive involved, Phase F7) --- this still implies a board-level dual connection
(the flash's DI/DO/CS/CLK pins wired both to the dedicated boot pins below and to these 4
ordinary balls, since it is the same physical chip serving both roles), not yet captured in
a schematic (none exists yet, see the checklist below).
\begin{tabularx}{\textwidth}{L{3.4cm} C{2.2cm} Y}
\toprule
\rowh \thd{Signal} & \thd{Ball\textsuperscript{$\dagger$}} & \thd{Function} \\
\midrule
\code{CCLK/MCLK/SCK} & U3 & Configuration clock. \\
\rowa \code{DQ0\_MOSI} & W2 & Config data (MOSI). \\
\code{DQ1\_MISO} & V2 & Config data (MISO). \\
\rowa \code{BUSY\_CSSPIN} & R2 & Flash chip-select. \\
\code{DQ2 / DQ3} & Y2 / W1 & Quad-read lines. \\
\rowa \code{PROGRAMN} & W3 & Start reconfiguration (button, active low). \\
\code{INITN} & V3 & Init / configuration error (LED). \\
\rowa \code{DONE} & Y3 & Configuration complete (LED). \\
\code{CFGMDN[2:0]} & R4/T4/U4 & Mode select (see below). \\
\bottomrule
\end{tabularx}
\subsection{Configuration modes (\texttt{CFGMDN})}
\begin{tabularx}{\textwidth}{L{4.0cm} C{4.0cm} Y}
\toprule
\rowh \thd{Mode} & \thd{CFGMDN[2:0]} & \thd{Use} \\
\midrule
MSPI (boot from flash) & \code{010} & \textbf{Default} --- standalone. \\
\rowa SSPI (slave SPI) & \code{001} & Config from external host. \\
SCM (slave serial) & \code{101} & Serial config. \\
\rowa SPCM (slave parallel) & \code{111} & 8-bit parallel config. \\
\bottomrule
\end{tabularx}
\begin{fnwarn}[Configuration balls to verify on the 45F]
\textsuperscript{$\dagger$}The JTAG and config-SPI balls listed here are the \emph{reference}
from the Lattice eval board (85F device). JTAG and config-SPI are dedicated, largely fixed
pins in the ECP5 family, but the exact positions on the \code{LFE5U-45F-8BG381C} target must
be confirmed against the Lattice 45F pinout file (Diamond/Radiant or the Trellis database)
before committing them to the schematic, as already done for the application signals
(ch.~\ref{ch:hw}).
\textbf{Distinct from this open item} (do not conflate the two): the flash subsystem's own
runtime SPI pins (\code{flash\_sclk}, \code{flash\_mosi}, \code{flash\_miso},
\code{flash\_cs\_n} --- Phases F1-F6, made fully independent in Phase F7) \textbf{are} real,
pinned, place\&route-verified ordinary GPIO on bank 7 --- \textbf{no pin shared with any
ECP5 config primitive}: an earlier version reused the boot \code{CCLK} pad for SCLK via
\code{USRMCLK}, dropped in Phase F7 (\code{USRMCLK} utilisation in the current full-system
synthesis is 0/1, confirming it is no longer used at all).
\end{fnwarn}
\section{Open tasks before schematic capture}
\begin{itemize}
\item[\OK] \code{ADDR\_WIDTH}=23 (full 8~MB) across all modules and testbenches.
\item[\OK] Real \code{.lpf} with the CABGA381 ball assignment, place\&route-verified with
0 errors (\code{synth/ecp5/spi\_neuron\_top.lpf}, 57 signals incl. flash subsystem).
\item[\OK] Boot/persistence flash subsystem (Phases F1-F7): SPI master, copy engine,
CRC32 slot catalog, fully independent 4-wire SPI bus, real synthesis at 0 errors, Fmax
67.91~MHz (\code{WORKLOG.md}).
\item[$\square$] Confirm PSRAM/SPI signal integrity at the actually mounted clock.
\item[$\square$] Board-level dual-wiring diagram for the flash's DI/DO/CS/CLK pins (dedicated
boot pins + the flash subsystem's 4 ordinary balls) --- not yet captured in a schematic.
\item[$\square$] Choice of the JTAG connector footprint.
\item[$\square$] Schematic capture (KiCad or other): no schematic exists yet for this
device/package combination.
\end{itemize}
@@ -1,80 +0,0 @@
\chapter{Quick reference}
\label{ch:ref}
\section{SPI opcodes}
\begin{tabularx}{\textwidth}{C{1.4cm} L{3.2cm} C{2.4cm} Y}
\toprule
\rowh \thd{Value} & \thd{Name} & \thd{Response} & \thd{Summary} \\
\midrule
\op{0x00} & NOP & --- & idle \\
\rowa \op{0x01} & WRITE\_RAM & --- & PSRAM block write \\
\op{0x02} & READ\_RAM & \code{len} B & PSRAM block read \\
\rowa \op{0x0F} & RESET & --- & engine reset + STATUS latch \\
\op{0x10} & SET\_BASE & --- & set base/register (sel 0..10) \\
\rowa \op{0x20} & START & --- & single-layer start \\
\op{0x21} & STATUS & 1 B & busy(live)/done(sticky) \\
\rowa \op{0x22} & READ\_OUTPUT & N\_NEURONS B & \code{y\_bus} \\
\op{0x23} & RUN\_NETWORK & --- & multi-layer start \\
\rowa \op{0x30} & READ\_CONFIG & 11 B & configuration record \\
\bottomrule
\end{tabularx}
\section{STATUS byte}
\begin{center}
\begin{tikzpicture}[font=\scriptsize]
\foreach \i/\lbl [count=\x from 0] in {7/0,6/0,5/0,4/0,3/0,2/0,1/{done},0/{busy}}{
\node[fnreg,minimum width=13mm,minimum height=9mm] (b\x) at (\x*13mm,0) {\lbl};
\node[font=\tiny,text=fnGrey,above=0.5mm of b\x] {bit \i};
}
\node[fill=fnAmber,text=white,rounded corners=1pt,inner sep=1.5pt,font=\tiny]
at (b7.center){reserved = 0};
\node[fill=fnTeal,text=white,rounded corners=1pt,inner sep=1.5pt,font=\tiny]
at (b6.center){};
\end{tikzpicture}
\end{center}
\code{done} is sticky, clear-on-read; \code{busy} is live.
\section{SET\_BASE selectors}
\begin{multicols}{2}\footnotesize
\begin{itemize}
\item 0 --- \code{x\_base}
\item 1 --- \code{w\_base}
\item 2 --- \code{bias\_addr}
\item 3 --- \code{table\_base}
\item 4 --- \code{buf\_a\_base}
\columnbreak
\item 5 --- \code{buf\_b\_base}
\item 6 --- \code{activation} (single-layer)
\item 7 --- \code{n\_inputs\_real} (single-layer)
\item 8 --- \code{n\_neurons\_real} (single-layer)
\item 9 --- \code{num\_neurons\_graph} (Type \#2)
\item 10 --- \code{n\_out} (Type \#2)
\end{itemize}
\end{multicols}
\section{Descriptor table (11 bytes/layer, MSB-first)}
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=0mm]
\node[fnreg,minimum width=20mm,minimum height=8mm](a){\code{w\_base}\\3B};
\node[fnreg,minimum width=20mm,minimum height=8mm,right=0mm of a](b){\code{bias\_addr}\\3B};
\node[fnreg,minimum width=14mm,minimum height=8mm,right=0mm of b](c){\code{act}\\1B};
\node[fnreg,minimum width=22mm,minimum height=8mm,right=0mm of c](d){\code{n\_inputs\_real}\\2B};
\node[fnreg,minimum width=22mm,minimum height=8mm,right=0mm of d](e){\code{n\_neurons\_real}\\2B};
\end{tikzpicture}
\end{center}
\section{Build parameters}
\begin{multicols}{2}\footnotesize
\begin{itemize}
\item \code{DATA\_WIDTH} --- 8 (INT8)
\item \code{ACC\_WIDTH} --- 32 (INT32)
\item \code{N\_INPUTS} --- max inputs
\item \code{N\_NEURONS} --- max neurons
\item \code{PARALLEL} --- simultaneous MACs
\columnbreak
\item \code{N\_LAYERS} --- max layers
\item \code{ADDR\_WIDTH} --- 23 (8 MB)
\item \code{MEM\_DATA\_WIDTH} --- 16
\item \code{CLK\_FREQ\_MHZ} --- PSRAM timing
\end{itemize}
\end{multicols}
@@ -1,64 +0,0 @@
\chapter{Roadmap and development status}
\label{ch:roadmap}
\section{Development phases}
\begin{tabularx}{\textwidth}{C{1.2cm} L{4.6cm} C{1.8cm} Y}
\toprule
\rowh \thd{Phase} & \thd{Title} & \thd{Status} & \thd{Content} \\
\midrule
1 & Parametric layer & \OK & inputs/neurons/parallelism, accumulation, bias, ReLU; test 32$\times$4/P=8. \\
\rowa 2 & Parameter sweep & \OK & multiple configurations incl. non-multiple and degenerate; elaboration guard added. \\
3 & Memory architecture & \OK & \code{neuron\_memory} single/multi-neuron, real PSRAM tested; multi-layer buffers $\to$ Phase~5. \\
\rowa 4 & SPI interface & \OK & \code{spi\_slave}+\code{spi\_engine}, 17 opcodes incl. flash subsystem, Fmax checked at full-system level. \\
5 & Multi-layer network & \OK$^\dagger$ & \code{layer\_sequencer}, configurable activations, runtime width; real toolchain checked. \\
\rowa 6 & Host software & planned & Linux and ESP32 drivers on the same protocol. \\
7 & Optimization & in progress & timing closure done (55$\to$75~MHz); PSRAM page-mode done (gather bandwidth +42\%); $x$/$w$ block RAM remains. \\
\rowa 8 & Hardware training (opt.) & future & backprop, gradients, weight update. \\
9 & Flash subsystem (F1-F7) & \OK & dedicated SPI master, flash$\leftrightarrow$PSRAM copy engine, 16-slot catalog with CRC32, fully independent 4-wire SPI bus (F7), 8 opcodes (\op{0x40}--\op{0x47}, ch.~\ref{ch:spi} §\ref{sec:flashspi}); real synthesis 0 errors. \\
\bottomrule
\end{tabularx}
\begin{center}\footnotesize\itshape\color{fnGrey}
$\dagger$ RTL, unit tests and end-to-end over simulated SPI complete; timing closure done:
75.30~MHz (P2) / 60.26~MHz (P8) at the time of Phase~5, bit-exact across the whole
regression; Fmax of the full system after Phase~9 (incl. independent flash subsystem):
\textbf{67.91~MHz} (ch.~\ref{ch:impl}).\end{center}
\section{Component status}
\begin{tabularx}{\textwidth}{Y C{4.2cm}}
\toprule
\rowh \thd{Component} & \thd{Status} \\
\midrule
Parametric neural layer & \OK{} working \\
\rowa Parametric inputs/neurons/parallelism & \OK \\
Accumulation, bias, ReLU & \OK \\
\rowa 32$\times$4 / P=8 validation & \OK \\
Dedicated RAM (interface + controller + INT8 access) & \OK{} tested on real PSRAM \\
\rowa SPI interface (17 opcodes incl. RUN\_NETWORK + flash) & \OK{} Fmax at full-system level \\
Dual SPI & future \\
\rowa Multi-layer engine & \OK{} timing closure 75.30~MHz (P2) at the time of Phase~5 \\
Configurable activations (ACT\_NONE/ACT\_RELU) & \OK \\
\rowa Runtime network width (one bitstream, any topology) & \OK{} measured savings \\
Type \#2 graph network (act\_buffer, graph\_engine, netasm) & \OK{} RTL + tests + synthesis \\
\rowa CABGA381 pinout (real \code{.lpf}, 57 signals incl. flash) & \OK{} place\&route-verified, 0 errors \\
PSRAM page-mode (G7) & \OK{} done (37.53 cycles/edge, bandwidth +42\%) \\
\rowa Flash subsystem (SPI master, copy engine, CRC32 catalog, independent bus F7) & \OK{} real synthesis 0 errors, Fmax 67.91~MHz \\
Real bitstream (\code{ecppack}, P2/P8) & \OK{} 0 errors, part LFE5U-45F-8CABGA381 \\
\rowa Linux / ESP32 host driver & planned \\
Hardware training & future \\
\bottomrule
\end{tabularx}
\section{Architectural principle (summary)}
\begin{fnspec}[Foundation of the project]
The FPGA implements the neural machine and owns its own RAM; the host configures and uses
the machine. A build fixes the \emph{ceiling} (max layers, max width, PARALLEL); the host
configures the \emph{actual} network --- number of layers, per-layer width, per-layer
activation, trained parameters --- entirely at runtime, over SPI, into the FPGA's local
memory. A single bitstream serves any topology up to that ceiling.
\end{fnspec}
\section{Long-term vision}
The final goal is a reusable hardware block integrable into different future projects:
the host platform can change (Linux, ESP32, MCU, PC) without changing the fundamental
architecture of the engine. The FPGA becomes a dedicated neural computation peripheral,
optimized for the topology required by each application.
@@ -1,96 +0,0 @@
\chapter[Modules and toolchain]{Modules, ports and toolchain}
\label{ch:appmod}
\section{List of RTL modules}
\begin{tabularx}{\textwidth}{L{3.4cm} C{2.0cm} Y}
\toprule
\rowh \thd{File} & \thd{Type} & \thd{Role} \\
\midrule
\code{rtl/mac\_unit.v} & combinational & single multiply-accumulator \\
\rowa \code{rtl/mac8.v} & combinational & parallel MAC + balanced adder tree \\
\code{rtl/neuron\_parallel.v} & FSM & neuron: groups, bias, activation, saturation \\
\rowa \code{rtl/layer.v} & structural & N\_NEURONS neurons in parallel \\
\code{rtl/neuron\_memory.v} & FSM & memory/neuron bridge, neuron loop \\
\rowa \code{rtl/layer\_sequencer.v} & FSM & multi-layer sequencing, ping-pong \\
\code{rtl/int8\_memory\_access.v} & FSM & byte $\leftrightarrow$ word conversion \\
\rowa \code{rtl/memory\_interface.v} & FSM & req/ready handshake \\
\code{rtl/psram\_controller.v} & FSM & async 70~ns physical PSRAM bus \\
\rowa \code{rtl/mem\_arbiter.v} & arbiter & 3 ports, priority B$>$C$>$A \\
\code{rtl/spi\_slave.v} & FSM & SPI Mode 0 physical layer + CDC \\
\rowa \code{rtl/spi\_engine.v} & FSM & opcode + register bank \\
\code{rtl/act\_buffer.v} & block RAM & DP16KD activation buffer (Type \#2) \\
\rowa \code{rtl/graph\_engine.v} & FSM & graph-network engine (Type \#2) \\
\code{rtl/spi\_neuron\_top.v} & top & full integration \\
\rowa \code{rtl/memory\_model.v} & model & behavioral RAM (sim) \\
\bottomrule
\end{tabularx}
\section{Ports of the top-level \texttt{spi\_neuron\_top}}
See the complete signal-by-signal table in ch.~\ref{ch:hw}. In summary: clock and reset
(\code{clk}, \code{rst}); application SPI (\code{sclk}, \code{mosi}, \code{miso},
\code{cs\_n}); PSRAM bus (\code{psram\_a[22:0]}, \code{psram\_dq[15:0]},
\code{psram\_ce\_n/oe\_n/we\_n/lb\_n/ub\_n/zz\_n}).
\section{Toolchain}
\begin{tabularx}{\textwidth}{L{3.6cm} L{3.4cm} Y}
\toprule
\rowh \thd{Tool} & \thd{Version} & \thd{Use} \\
\midrule
Yosys & 0.68+post & RTL synthesis $\to$ JSON netlist, ECP5 mapping \\
\rowa nextpnr-ecp5 & 0.11.1-19-g8dbcee5 & placement, routing, timing \\
Project Trellis & install & \code{ecppack}/\code{ecppll}/\code{ecpbram} \\
\rowa Icarus Verilog & \code{-g2012} & functional simulation \\
\bottomrule
\end{tabularx}
\subsection{Main nextpnr parameters}
\begin{lstlisting}[language=,basicstyle=\ttfamily\scriptsize]
--45k selects LFE5U-45F
--package CABGA381 package
--speed 8 speed grade -8
--json <netlist> netlist from Yosys
--lpf <constraints> pin constraints (currently empty)
--lpf-allow-unconstrained allows unconstrained I/Os (benchmark)
--freq 80 80 MHz timing target
\end{lstlisting}
\subsection{Simulation example}
\begin{lstlisting}[language=,basicstyle=\ttfamily\scriptsize]
iverilog -g2012 -Ptb.PARALLEL=16 -o sim/parametric_256x4_p16 \
sim/parametric_tb.v rtl/mac_unit.v rtl/mac8.v \
rtl/neuron_parallel.v rtl/layer.v
vvp sim/parametric_256x4_p16
\end{lstlisting}
\section{Main testbenches}
\begin{tabularx}{\textwidth}{L{5.4cm} Y}
\toprule
\rowh \thd{Testbench} & \thd{Coverage} \\
\midrule
\code{parametric\_tb.v} & 256$\times$4 datapath, accumulate/bias/ReLU/saturation cases \\
\rowa \code{parameter\_sweep\_tb.v} & sweep of valid configurations \\
\code{neuron\_parallel\_tb.v} & activations, runtime width (T7) \\
\rowa \code{neuron\_memory\_tb.v} / \code{\_multi\_tb.v} & single/multi-neuron memory integration, real PSRAM (T5) \\
\code{psram\_controller\_tb.v} & PSRAM controller \\
\rowa \code{psram\_page\_mode\_tb.v} & page bursts, page crossing, close on WRITE/$t_{CEM}$ timeout, byte-enable changes (§~5.5) \\
\code{spi\_slave\_tb.v} & SPI physical layer (4 tests) \\
\rowa \code{spi\_engine\_tb.v} & opcodes, registers (10+ tests) \\
\code{spi\_neuron\_top\_tb.v} & end-to-end, real PSRAM over simulated SPI \\
\rowa \code{spi\_neuron\_top\_runnetwork\_tb.v} & RUN\_NETWORK 2-layer end-to-end \\
\code{layer\_sequencer\_tb.v} & 2-layer sequence, ping-pong, byte-exact copy \\
\bottomrule
\end{tabularx}
\vfill
\begin{center}
\begin{tikzpicture}
\node[draw=fnRule,rounded corners=3pt,inner sep=8pt,fill=fnLight,text width=15.5cm]{
\footnotesize\color{fnGrey}
This datasheet is generated from the RTL code, the documentation and the benchmarks
present in the repository \texttt{github.com/manvalan/FPGA-Neural} as of \datasheetdate.
The Fmax, resource usage and throughput values are those reported in the repository
measurements (real \texttt{.lpf} already assigned and place\&route-verified,
ch.~\ref{ch:hw}) and must be re-verified on any substantial RTL change or as the
Phase~7 timing closure, still in progress, continues (ch.~\ref{ch:roadmap}).};
\end{tikzpicture}
\end{center}
-184
View File
@@ -1,184 +0,0 @@
% ======================================================================
% FPGA-Neural Datasheet -- preamble / stile
% ======================================================================
\usepackage[T1]{fontenc}
\usepackage[utf8]{inputenc}
\usepackage[italian,provide=*]{babel}
\usepackage{helvet}
\renewcommand{\familydefault}{\sfdefault}
\usepackage{courier}
\usepackage{microtype}
\usepackage[a4paper,top=2.4cm,bottom=2.3cm,left=2.2cm,right=2.2cm,headheight=15pt]{geometry}
\usepackage[table]{xcolor}
\usepackage{graphicx}
\usepackage{booktabs}
\usepackage{tabularx}
\usepackage{longtable}
\usepackage{array}
\usepackage{ltablex}
\keepXColumns
\usepackage{multirow}
\usepackage{multicol}
\usepackage{enumitem}
\usepackage{amsmath}
\usepackage{amssymb}
\usepackage{ragged2e}
% ---------- Palette ----------------------------------------------------
\definecolor{fnDark}{HTML}{0B2E4F} % blu profondo (primario)
\definecolor{fnBlue}{HTML}{15629B} % blu medio
\definecolor{fnTeal}{HTML}{0E8F8A} % accento teal
\definecolor{fnAmber}{HTML}{C9761B} % accento ambra
\definecolor{fnRed}{HTML}{B22C34} % fail / warning
\definecolor{fnGreen}{HTML}{2E7D32} % pass / ok
\definecolor{fnGrey}{HTML}{5B6B78}
\definecolor{fnLight}{HTML}{EEF3F7} % sfondo chiaro
\definecolor{fnLight2}{HTML}{E2ECF3}
\definecolor{fnRule}{HTML}{9FB4C4}
\definecolor{codebg}{HTML}{F5F7F9}
\definecolor{codekw}{HTML}{15629B}
\definecolor{codecom}{HTML}{5B6B78}
\definecolor{codestr}{HTML}{0E8F8A}
% ---------- Titoli -----------------------------------------------------
\usepackage{titlesec}
\titleformat{\chapter}[display]
{\normalfont\bfseries\color{fnDark}}
{\filright\Large\color{fnTeal}CAPITOLO \thechapter}
{6pt}
{\Huge\filright}
[\vspace{2pt}{\color{fnRule}\titlerule[1.3pt]}]
\titlespacing*{\chapter}{0pt}{6pt}{18pt}
\titleformat{\section}
{\normalfont\large\bfseries\color{fnDark}}{\thesection}{0.6em}{}
\titleformat{\subsection}
{\normalfont\bfseries\color{fnBlue}}{\thesubsection}{0.6em}{}
\titleformat{\subsubsection}
{\normalfont\bfseries\color{fnGrey}}{\thesubsubsection}{0.6em}{}
\titlespacing*{\section}{0pt}{12pt}{4pt}
% ---------- Header / footer -------------------------------------------
\usepackage{fancyhdr}
\pagestyle{fancy}
\fancyhf{}
\renewcommand{\headrulewidth}{0.6pt}
\renewcommand{\footrulewidth}{0.4pt}
\renewcommand{\headrule}{\hbox to\headwidth{\color{fnRule}\leaders\hrule height \headrulewidth\hfill}}
\renewcommand{\footrule}{\hbox to\headwidth{\color{fnRule}\leaders\hrule height \footrulewidth\hfill}}
\renewcommand{\chaptermark}[1]{\markboth{#1}{}}
\fancyhead[L]{\small\color{fnDark}\textbf{FPGA-Neural}}
\fancyhead[R]{\footnotesize\color{fnGrey}\nouppercase{\leftmark}}
\fancyfoot[L]{\small\color{fnGrey}Datasheet~\textbf{Rev.\,\datasheetrev}}
\fancyfoot[C]{\small\color{fnGrey}manvalan/FPGA-Neural}
\fancyfoot[R]{\small\color{fnGrey}\thepage}
\fancypagestyle{plain}{\fancyhf{}%
\fancyfoot[L]{\small\color{fnGrey}Datasheet~\textbf{Rev.\,\datasheetrev}}%
\fancyfoot[C]{\small\color{fnGrey}manvalan/FPGA-Neural}%
\fancyfoot[R]{\small\color{fnGrey}\thepage}%
\renewcommand{\headrulewidth}{0pt}}
% ---------- tcolorbox --------------------------------------------------
\usepackage[most]{tcolorbox}
\tcbuselibrary{skins,breakable}
% Box "nota"
\newtcolorbox{fnnote}[1][Nota]{
enhanced, breakable, colback=fnLight, colframe=fnTeal,
boxrule=0.4pt, left=8pt, right=8pt, top=5pt, bottom=5pt,
fonttitle=\bfseries\color{white}, coltitle=white,
attach boxed title to top left={xshift=6pt,yshift=-3pt},
boxed title style={colback=fnTeal,boxrule=0pt,arc=1pt}, title={#1}}
% Box "attenzione"
\newtcolorbox{fnwarn}[1][Attenzione]{
enhanced, breakable, colback=fnLight, colframe=fnAmber,
boxrule=0.4pt, left=8pt, right=8pt, top=5pt, bottom=5pt,
fonttitle=\bfseries\color{white}, coltitle=white,
attach boxed title to top left={xshift=6pt,yshift=-3pt},
boxed title style={colback=fnAmber,boxrule=0pt,arc=1pt}, title={#1}}
% Box "registro/parametro"
\newtcolorbox{fnspec}[1][Specifica]{
enhanced, breakable, colback=white, colframe=fnBlue,
boxrule=0.7pt, left=8pt, right=8pt, top=5pt, bottom=5pt, arc=1.5pt,
fonttitle=\bfseries\color{white}, coltitle=white,
attach boxed title to top left={xshift=6pt,yshift=-3pt},
boxed title style={colback=fnBlue,boxrule=0pt,arc=1pt}, title={#1}}
% ---------- listings (Verilog) ----------------------------------------
\usepackage{listings}
\lstdefinestyle{verilog}{
language=Verilog,
backgroundcolor=\color{codebg},
basicstyle=\ttfamily\scriptsize,
keywordstyle=\color{codekw}\bfseries,
commentstyle=\color{codecom}\itshape,
stringstyle=\color{codestr},
numbers=left, numberstyle=\tiny\color{fnGrey}, numbersep=7pt,
showstringspaces=false, breaklines=true, frame=leftline,
framerule=1.2pt, rulecolor=\color{fnTeal},
xleftmargin=12pt, framexleftmargin=10pt, tabsize=2,
morekeywords={logic,always_ff,always_comb,localparam,signed,genvar,generate,endgenerate}
}
\lstset{style=verilog}
% ---------- Tabelle ----------------------------------------------------
\newcolumntype{L}[1]{>{\raggedright\arraybackslash}p{#1}}
\newcolumntype{C}[1]{>{\centering\arraybackslash}p{#1}}
\newcolumntype{R}[1]{>{\raggedleft\arraybackslash}p{#1}}
\newcolumntype{Y}{>{\raggedright\arraybackslash}X}
\renewcommand{\arraystretch}{1.25}
\arrayrulecolor{fnRule}
% intestazione tabella colorata
\newcommand{\thd}[1]{\textbf{\color{white}#1}}
\newcommand{\rowh}{\rowcolor{fnDark}}
\newcommand{\rowa}{\rowcolor{fnLight}}
% ---------- Caption ----------------------------------------------------
\usepackage{caption}
\captionsetup{font=small,labelfont={bf,color=fnTeal},labelsep=period}
% ---------- TikZ / pgfplots -------------------------------------------
\usepackage{tikz}
\usetikzlibrary{arrows.meta,positioning,calc,shapes.geometric,shapes.misc,
fit,backgrounds,chains,decorations.pathreplacing,decorations.markings,
matrix,shadows.blur}
\usepackage{pgfplots}
\pgfplotsset{compat=1.17}
\usepackage{tikz-timing}
% stili di blocco riusabili
\tikzset{
fnblock/.style={draw=fnBlue,fill=fnLight,rounded corners=2pt,
minimum height=9mm,minimum width=24mm,align=center,font=\small,
inner sep=4pt,line width=0.7pt},
fnblockT/.style={fnblock,draw=fnTeal,fill=fnLight2},
fnblockD/.style={fnblock,draw=fnDark,fill=fnDark,text=white},
fnblockA/.style={fnblock,draw=fnAmber,fill=white},
fnreg/.style={draw=fnGrey,fill=white,minimum height=8mm,align=center,
font=\footnotesize,inner sep=3pt},
fnstate/.style={draw=fnBlue,fill=fnLight,circle,minimum size=13mm,
align=center,font=\scriptsize,line width=0.7pt},
fnarrow/.style={-{Stealth[length=2.6mm]},line width=0.8pt,draw=fnDark},
fnarrowT/.style={-{Stealth[length=2.6mm]},line width=0.8pt,draw=fnTeal},
fnbus/.style={-{Stealth[length=3mm]},line width=1.6pt,draw=fnBlue},
fnlbl/.style={font=\scriptsize\itshape,fill=white,inner sep=1pt,text=fnGrey}
}
% ---------- varie ------------------------------------------------------
\newcommand{\reg}[1]{\texttt{\textbf{#1}}}
\newcommand{\sig}[1]{\texttt{#1}}
\newcommand{\op}[1]{\texttt{\color{fnBlue}#1}}
\newcommand{\PASS}{\textcolor{fnGreen}{\textbf{PASS}}}
\newcommand{\FAIL}{\textcolor{fnRed}{\textbf{FAIL}}}
\newcommand{\OK}{\textcolor{fnGreen}{\textbf{OK}}}
\newcommand{\code}[1]{\texttt{#1}}
\usepackage{enumitem}
\setlist{noitemsep,topsep=2pt,leftmargin=1.4em}
\usepackage[hidelinks,colorlinks=true,linkcolor=fnBlue,urlcolor=fnTeal,
citecolor=fnBlue]{hyperref}
@@ -1,54 +0,0 @@
\thispagestyle{plain}
\noindent
\begin{tikzpicture}
\node[fill=fnDark,text=white,rounded corners=2pt,inner sep=6pt,
minimum width=\textwidth,anchor=west]
{\large\bfseries Pinout summary --- scope and honesty note};
\end{tikzpicture}
\vspace{6pt}
\noindent
{\footnotesize
V2's top-level module, \code{neural\_multiprocessor.v}, has been
synthesized and placed\&routed \textbf{unconstrained}
(\code{nextpnr-ecp5 --lpf-allow-unconstrained}) throughout this project's
own real-toolchain characterization: every Fmax/resource number in this
datasheet is real and measured, but \textbf{no ball-by-ball pin
assignment (\code{.lpf}) has been generated for V2's top level in this
revision}. Unlike V1's own pinout chapter (which reports a real,
\code{iodb.json}-verified ball map from a constrained place\&route run),
this chapter reports what is \textbf{honestly known} and nothing
invented.
}
\vspace{6pt}
\begin{fnnote}[What is real and reusable]
V2's PSRAM-facing pins (\code{psram\_a}, \code{psram\_dq},
\code{psram\_ce\_n/oe\_n/we\_n/lb\_n/ub\_n/zz\_n}) drive the exact same,
real, unmodified V1 backend chain (\code{memory\_interface.v} $\to$
\code{psram\_controller.v}) as V1's own \code{spi\_neuron\_top}. If V2 is
deployed on the same board, \textbf{V1's own real, verified ball
assignment for these signals (ch.~10 of the V1 datasheet) applies
unchanged} --- the controller was never touched, so its pin requirements
did not change either.
\end{fnnote}
\begin{fnwarn}[What is NOT yet real]
The node-registration bus (\code{reg\_valid}, \code{reg\_node\_id},
\code{reg\_required}, \code{reg\_producer\_ids}, \code{reg\_x\_base},
\code{reg\_w\_base}, \code{reg\_n\_tiles}, \code{reg\_result\_addr},
\code{reg\_ready}) has no assigned physical pins in this revision: every
V2 measurement to date drove this bus directly from a Verilator
testbench or an unconstrained synthesis top-level, never through a real
host-facing SPI (or other) interface with its own placed pinout. Framing
this bus as a real, deployable host interface (analogous to V1's SPI
Mode~0 slave) is explicitly \textbf{future work} --- see
ch.~\ref{ch:roadmap}.
\end{fnwarn}
\vspace{4pt}
\noindent
{\footnotesize\color{fnGrey}
Logical (not physical) port list and field widths: ch.~\ref{ch:regs}
(``Register-level interface''). Real PSRAM signal reuse and board
wiring: ch.~\ref{ch:hw}.\par}
@@ -1,55 +0,0 @@
\chapter{Top-level module}
\label{ch:toplevel}
\section{\texttt{neural\_multiprocessor.v}}
The real, hardware-facing top level: \code{dataflow\_core.v} (Dependency
Manager $+$ Neural Director $+$ \code{N\_SLOTS}$\times$(Memory Manager $+$
Neural Processor) $+$ Activation Cache) with its \code{N\_SLOTS}$+$1
Memory Backend Interface ports funneled through \code{slot\_mem\_arbiter.v}
down to the real, unmodified V1 PSRAM chain
(\code{memory\_interface.v} $\to$ \code{psram\_controller.v}).
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.2cm} C{1.6cm} Y}
\toprule
\rowh \thd{Port} & \thd{Dir} & \thd{Width} & \thd{Function} \\
\midrule
\code{clk}, \code{rst} & IN & 1 & System clock, synchronous reset. \\
\rowa \code{reg\_valid} & IN & 1 & Node registration request (ch.~\ref{ch:host}). \\
\code{reg\_ready} & OUT & 1 & This node id's table slot is \code{EMPTY}. \\
\rowa \code{reg\_node\_id} & IN & $\lceil\log_2\text{N\_NODES}\rceil$ & Node id. \\
\code{reg\_required} & IN & $\lceil\log_2(\text{MAX\_DEPS}{+}1)\rceil$ & Producer count. \\
\rowa \code{reg\_producer\_ids} & IN & \code{MAX\_DEPS}$\times\lceil\log_2\text{N\_NODES}\rceil$ & Packed producer id list. \\
\code{reg\_x\_base}, \code{reg\_w\_base}, \code{reg\_result\_addr} & IN & \code{ADDR\_WIDTH} each & Job descriptor addresses. \\
\rowa \code{reg\_n\_tiles} & IN & 16 & Tile count. \\
\code{psram\_a} & OUT & \code{ADDR\_WIDTH} & PSRAM address bus (real V1 controller, unmodified). \\
\rowa \code{psram\_dq} & INOUT & \code{PSRAM\_DATA\_WIDTH} & PSRAM bidirectional data bus. \\
\code{psram\_ce\_n}, \code{psram\_oe\_n}, \code{psram\_we\_n}, \code{psram\_lb\_n}, \code{psram\_ub\_n}, \code{psram\_zz\_n} & OUT & 1 each & PSRAM control, identical to V1's own real, verified signal set. \\
\bottomrule
\end{tabularx}
\begin{fnnote}[No \texttt{int8\_memory\_access.v} in this datapath]
Earlier milestones instantiated V1's \code{int8\_memory\_access.v}
between the arbiter and \code{memory\_interface.v}. Post word-burst
rewrite (ch.~\ref{ch:mem}), it is no longer instantiated here --- the
file itself is untouched (still frozen V1); V2 simply reuses one layer
lower in the same frozen stack.
\end{fnnote}
\section{Internal hierarchy}
\noindent\code{neural\_multiprocessor.v}
\begin{itemize}[leftmargin=2.4em]
\footnotesize
\item \code{u\_dataflow\_core} : \code{dataflow\_core.v}
\begin{itemize}
\item \code{u\_dep\_mgr} : \code{dependency\_manager.v}
\item \code{u\_director} : \code{neural\_director.v}
\item \code{GEN\_SLOT[0..N\_SLOTS-1]}: \code{memory\_manager.v} $+$ \code{neural\_processor.v}
\begin{itemize}
\item[--] \code{u\_prefetch} : \code{prefetch\_engine.v} (weights only, word-level)
\end{itemize}
\item \code{u\_activation\_cache} : \code{activation\_cache.v}
\end{itemize}
\item \code{u\_arbiter} : \code{slot\_mem\_arbiter.v} (\code{N\_SLOTS}$+$1 ports)
\item \code{u\_memif} : \code{memory\_interface.v} (frozen V1)
\item \code{u\_psram\_ctrl} : \code{psram\_controller.v} (frozen V1)
\end{itemize}
-184
View File
@@ -1,184 +0,0 @@
% ======================================================================
% FPGA-Neural Datasheet -- preamble / stile
% ======================================================================
\usepackage[T1]{fontenc}
\usepackage[utf8]{inputenc}
\usepackage[english]{babel}
\usepackage{helvet}
\renewcommand{\familydefault}{\sfdefault}
\usepackage{courier}
\usepackage{microtype}
\usepackage[a4paper,top=2.4cm,bottom=2.3cm,left=2.2cm,right=2.2cm,headheight=15pt]{geometry}
\usepackage[table]{xcolor}
\usepackage{graphicx}
\usepackage{booktabs}
\usepackage{tabularx}
\usepackage{longtable}
\usepackage{array}
\usepackage{ltablex}
\keepXColumns
\usepackage{multirow}
\usepackage{multicol}
\usepackage{enumitem}
\usepackage{amsmath}
\usepackage{amssymb}
\usepackage{ragged2e}
% ---------- Palette ----------------------------------------------------
\definecolor{fnDark}{HTML}{0B2E4F} % blu profondo (primario)
\definecolor{fnBlue}{HTML}{15629B} % blu medio
\definecolor{fnTeal}{HTML}{0E8F8A} % accento teal
\definecolor{fnAmber}{HTML}{C9761B} % accento ambra
\definecolor{fnRed}{HTML}{B22C34} % fail / warning
\definecolor{fnGreen}{HTML}{2E7D32} % pass / ok
\definecolor{fnGrey}{HTML}{5B6B78}
\definecolor{fnLight}{HTML}{EEF3F7} % sfondo chiaro
\definecolor{fnLight2}{HTML}{E2ECF3}
\definecolor{fnRule}{HTML}{9FB4C4}
\definecolor{codebg}{HTML}{F5F7F9}
\definecolor{codekw}{HTML}{15629B}
\definecolor{codecom}{HTML}{5B6B78}
\definecolor{codestr}{HTML}{0E8F8A}
% ---------- Titoli -----------------------------------------------------
\usepackage{titlesec}
\titleformat{\chapter}[display]
{\normalfont\bfseries\color{fnDark}}
{\filright\Large\color{fnTeal}CHAPTER \thechapter}
{6pt}
{\Huge\filright}
[\vspace{2pt}{\color{fnRule}\titlerule[1.3pt]}]
\titlespacing*{\chapter}{0pt}{6pt}{18pt}
\titleformat{\section}
{\normalfont\large\bfseries\color{fnDark}}{\thesection}{0.6em}{}
\titleformat{\subsection}
{\normalfont\bfseries\color{fnBlue}}{\thesubsection}{0.6em}{}
\titleformat{\subsubsection}
{\normalfont\bfseries\color{fnGrey}}{\thesubsubsection}{0.6em}{}
\titlespacing*{\section}{0pt}{12pt}{4pt}
% ---------- Header / footer -------------------------------------------
\usepackage{fancyhdr}
\pagestyle{fancy}
\fancyhf{}
\renewcommand{\headrulewidth}{0.6pt}
\renewcommand{\footrulewidth}{0.4pt}
\renewcommand{\headrule}{\hbox to\headwidth{\color{fnRule}\leaders\hrule height \headrulewidth\hfill}}
\renewcommand{\footrule}{\hbox to\headwidth{\color{fnRule}\leaders\hrule height \footrulewidth\hfill}}
\renewcommand{\chaptermark}[1]{\markboth{#1}{}}
\fancyhead[L]{\small\color{fnDark}\textbf{FPGA-Neural}}
\fancyhead[R]{\footnotesize\color{fnGrey}\nouppercase{\leftmark}}
\fancyfoot[L]{\small\color{fnGrey}Datasheet~\textbf{Rev.\,\datasheetrev}}
\fancyfoot[C]{\small\color{fnGrey}manvalan/FPGA-Neural}
\fancyfoot[R]{\small\color{fnGrey}\thepage}
\fancypagestyle{plain}{\fancyhf{}%
\fancyfoot[L]{\small\color{fnGrey}Datasheet~\textbf{Rev.\,\datasheetrev}}%
\fancyfoot[C]{\small\color{fnGrey}manvalan/FPGA-Neural}%
\fancyfoot[R]{\small\color{fnGrey}\thepage}%
\renewcommand{\headrulewidth}{0pt}}
% ---------- tcolorbox --------------------------------------------------
\usepackage[most]{tcolorbox}
\tcbuselibrary{skins,breakable}
% Box "nota"
\newtcolorbox{fnnote}[1][Note]{
enhanced, breakable, colback=fnLight, colframe=fnTeal,
boxrule=0.4pt, left=8pt, right=8pt, top=5pt, bottom=5pt,
fonttitle=\bfseries\color{white}, coltitle=white,
attach boxed title to top left={xshift=6pt,yshift=-3pt},
boxed title style={colback=fnTeal,boxrule=0pt,arc=1pt}, title={#1}}
% Box "attenzione"
\newtcolorbox{fnwarn}[1][Warning]{
enhanced, breakable, colback=fnLight, colframe=fnAmber,
boxrule=0.4pt, left=8pt, right=8pt, top=5pt, bottom=5pt,
fonttitle=\bfseries\color{white}, coltitle=white,
attach boxed title to top left={xshift=6pt,yshift=-3pt},
boxed title style={colback=fnAmber,boxrule=0pt,arc=1pt}, title={#1}}
% Box "registro/parametro"
\newtcolorbox{fnspec}[1][Specification]{
enhanced, breakable, colback=white, colframe=fnBlue,
boxrule=0.7pt, left=8pt, right=8pt, top=5pt, bottom=5pt, arc=1.5pt,
fonttitle=\bfseries\color{white}, coltitle=white,
attach boxed title to top left={xshift=6pt,yshift=-3pt},
boxed title style={colback=fnBlue,boxrule=0pt,arc=1pt}, title={#1}}
% ---------- listings (Verilog) ----------------------------------------
\usepackage{listings}
\lstdefinestyle{verilog}{
language=Verilog,
backgroundcolor=\color{codebg},
basicstyle=\ttfamily\scriptsize,
keywordstyle=\color{codekw}\bfseries,
commentstyle=\color{codecom}\itshape,
stringstyle=\color{codestr},
numbers=left, numberstyle=\tiny\color{fnGrey}, numbersep=7pt,
showstringspaces=false, breaklines=true, frame=leftline,
framerule=1.2pt, rulecolor=\color{fnTeal},
xleftmargin=12pt, framexleftmargin=10pt, tabsize=2,
morekeywords={logic,always_ff,always_comb,localparam,signed,genvar,generate,endgenerate}
}
\lstset{style=verilog}
% ---------- Tabelle ----------------------------------------------------
\newcolumntype{L}[1]{>{\raggedright\arraybackslash}p{#1}}
\newcolumntype{C}[1]{>{\centering\arraybackslash}p{#1}}
\newcolumntype{R}[1]{>{\raggedleft\arraybackslash}p{#1}}
\newcolumntype{Y}{>{\raggedright\arraybackslash}X}
\renewcommand{\arraystretch}{1.25}
\arrayrulecolor{fnRule}
% intestazione tabella colorata
\newcommand{\thd}[1]{\textbf{\color{white}#1}}
\newcommand{\rowh}{\rowcolor{fnDark}}
\newcommand{\rowa}{\rowcolor{fnLight}}
% ---------- Caption ----------------------------------------------------
\usepackage{caption}
\captionsetup{font=small,labelfont={bf,color=fnTeal},labelsep=period}
% ---------- TikZ / pgfplots -------------------------------------------
\usepackage{tikz}
\usetikzlibrary{arrows.meta,positioning,calc,shapes.geometric,shapes.misc,
fit,backgrounds,chains,decorations.pathreplacing,decorations.markings,
matrix,shadows.blur}
\usepackage{pgfplots}
\pgfplotsset{compat=1.17}
\usepackage{tikz-timing}
% stili di blocco riusabili
\tikzset{
fnblock/.style={draw=fnBlue,fill=fnLight,rounded corners=2pt,
minimum height=9mm,minimum width=24mm,align=center,font=\small,
inner sep=4pt,line width=0.7pt},
fnblockT/.style={fnblock,draw=fnTeal,fill=fnLight2},
fnblockD/.style={fnblock,draw=fnDark,fill=fnDark,text=white},
fnblockA/.style={fnblock,draw=fnAmber,fill=white},
fnreg/.style={draw=fnGrey,fill=white,minimum height=8mm,align=center,
font=\footnotesize,inner sep=3pt},
fnstate/.style={draw=fnBlue,fill=fnLight,circle,minimum size=13mm,
align=center,font=\scriptsize,line width=0.7pt},
fnarrow/.style={-{Stealth[length=2.6mm]},line width=0.8pt,draw=fnDark},
fnarrowT/.style={-{Stealth[length=2.6mm]},line width=0.8pt,draw=fnTeal},
fnbus/.style={-{Stealth[length=3mm]},line width=1.6pt,draw=fnBlue},
fnlbl/.style={font=\scriptsize\itshape,fill=white,inner sep=1pt,text=fnGrey}
}
% ---------- varie ------------------------------------------------------
\newcommand{\reg}[1]{\texttt{\textbf{#1}}}
\newcommand{\sig}[1]{\texttt{#1}}
\newcommand{\op}[1]{\texttt{\color{fnBlue}#1}}
\newcommand{\PASS}{\textcolor{fnGreen}{\textbf{PASS}}}
\newcommand{\FAIL}{\textcolor{fnRed}{\textbf{FAIL}}}
\newcommand{\OK}{\textcolor{fnGreen}{\textbf{OK}}}
\newcommand{\code}[1]{\texttt{#1}}
\usepackage{enumitem}
\setlist{noitemsep,topsep=2pt,leftmargin=1.4em}
\usepackage[hidelinks,colorlinks=true,linkcolor=fnBlue,urlcolor=fnTeal,
citecolor=fnBlue]{hyperref}