Files
FPGA-Neural/hardware/v1/docs/DatasheetLatex/chapters/05-memoria.tex
T
2026-09-06 15:02:25 +02:00

188 lines
10 KiB
TeX

\chapter{Memory subsystem}
\label{ch:mem}
\section{Memory chain}
The compute engine works with addresses and data at the \emph{byte} level (INT8), while
the PSRAM is a 16-bit word device. Three cascaded modules realize the conversion and
the physical access:
\begin{center}
\begin{tikzpicture}[font=\scriptsize,node distance=8mm]
\node[fnblockD,minimum width=30mm,minimum height=12mm](nm){byte-level master\\{\scriptsize \code{neuron\_memory} / \code{spi\_engine} / \code{layer\_sequencer}}};
\node[fnblockT,right=10mm of nm,minimum width=28mm,minimum height=12mm](ia){\code{int8\_memory\_access}\\{\scriptsize byte $\leftrightarrow$ 16-bit word}};
\node[fnblock,right=10mm of ia,minimum width=26mm,minimum height=12mm](mi){\code{memory\_interface}\\{\scriptsize IDLE/WAIT FSM}};
\node[fnblockA,below=9mm of mi,minimum width=26mm,minimum height=12mm](pc){\code{psram\_controller}\\{\scriptsize async 70\,ns physical bus}};
\node[fnblock,left=10mm of pc,minimum width=26mm,minimum height=12mm](ps){PSRAM\\{\scriptsize 8\,MB 4M$\times$16}};
\draw[fnbus] (nm)--node[fnlbl,above]{req/wr/addr}(ia);
\draw[fnbus] (ia)--node[fnlbl,above]{16-bit}(mi);
\draw[fnbus] (mi)--(pc);
\draw[fnbus] (pc)--node[fnlbl,above]{DQ/A/ctrl}(ps);
\end{tikzpicture}
\end{center}
\section{\texttt{int8\_memory\_access} --- byte/word conversion}
Converts the INT8 interface (byte address) into the word interface. The byte address is
divided by two (\code{addr>>1}) to obtain the word address; the least significant bit
selects the byte:
\begin{itemize}
\item \code{addr[0]=0} $\to$ low byte: \code{lb\_n=0}, \code{ub\_n=1}, data on DQ[7:0];
\item \code{addr[0]=1} $\to$ high byte: \code{lb\_n=1}, \code{ub\_n=0}, data on DQ[15:8].
\end{itemize}
On read it extracts the correct byte from \code{mem\_rdata}. The FSM has two states
(IDLE, WAIT) and returns \code{ready} as a one-cycle pulse.
\section{\texttt{memory\_interface} --- handshake}
Two-state FSM that serializes a single transaction: in IDLE, on the \code{req} request,
it latches \code{wr/addr/wdata/lb\_n/ub\_n} and emits a one-cycle \code{mem\_req} pulse
toward the controller; in WAIT it waits for \code{mem\_ready}, captures \code{rdata} on
read and asserts \code{ready}. It guarantees the ``one transaction at a time'' contract.
\section{\texttt{psram\_controller} --- physical bus}
Asynchronous parallel PSRAM bus controller, with support for the chip's read
\textbf{page mode} (\S~\ref{sec:pagemode}). The main state machine is:
\begin{center}
\begin{tikzpicture}[font=\scriptsize]
\node[fnstate](init) at (0,0){INIT};
\node[fnstate](idle) at (3.2,0){IDLE};
\node[fnstate](read) at (7,2.7){READ};
\node[fnstate](popen) at (11,2.7){PAGE\\OPEN};
\node[fnstate](write) at (7,-2.7){WRITE};
\node[fnstate](ww) at (11,-2.7){WRITE\\WAIT};
\draw[fnarrow] (init)--node[fnlbl,above]{INIT\_CYCLES + CR load}(idle);
\draw[fnarrow] (idle)--node[fnlbl,above,sloped]{req \& !wr}(read);
\draw[fnarrow] (idle)--node[fnlbl,below,sloped]{req \& wr}(write);
\draw[fnarrow] (read)--node[fnlbl,above]{ready}(popen);
\draw[fnarrowT] (popen) to[bend left=25] node[fnlbl,below]{req \& !wr}(read);
\draw[fnarrow] (popen) to[bend right=20] node[fnlbl,above,sloped]{req \& wr}(write);
\draw[fnarrow] (popen) to[out=-100,in=15,looseness=1.15] node[fnlbl,pos=0.55]{tCEM timeout}(idle);
\draw[fnarrow] (write)--node[fnlbl,above]{ACCESS\_CYCLES}(ww);
\draw[fnarrow] (ww) to[out=160,in=-70] node[fnlbl,pos=0.5,left]{ready}(idle);
\end{tikzpicture}
\end{center}
From INIT the controller automatically goes through a configuration-register load
sub-sequence (\code{STATE\_CR\_INIT}, 4 steps) before reaching IDLE for the first
time --- see \S~\ref{sec:pagemode}. The PAGE~OPEN~$\to$~WRITE transition
(bottom-right arrow) internally passes through two transit micro-states,
\code{STATE\_PAGE\_CLOSE} and \code{STATE\_PAGE\_REOPEN} (one cycle each): the
first forces CE\#/OE\# high for at least one cycle before the controller starts
driving the data bus, avoiding contention with the PSRAM's still-active output
($\geq t_{HZ}$); the second restarts the already-latched transaction exactly as
IDLE would. They are not drawn as separate nodes to keep the figure readable.
\subsection{Timing}
\begin{fnspec}[Timing formulas]
$\text{ACCESS\_CYCLES}=\lceil (70\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(random-access latency, $t_{AA}$/$t_{RC}$ = 70~ns)\\[3pt]
$\text{PAGE\_CYCLES}=\lceil (20\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(same-page continuation, $t_{APA}$/$t_{PC}$ = 20~ns)\\[3pt]
$\text{INIT\_CYCLES}=150\times \text{CLK\_FREQ\_MHZ}$ \quad
(power-up initialization, $t_{PU}$ = 150~\textmu s)\\[3pt]
$\text{PAGE\_TIMEOUT\_CYCLES}=\lceil (6000\times \text{CLK\_FREQ\_MHZ})/1000\rceil$ \quad
(automatic page close, safety margin under $t_{CEM}$ = 8~\textmu s)
\end{fnspec}
The data bus is tri-state driven: \code{psram\_dq = dq\_oe ? dq\_out : Z}. On read
\code{dq\_oe=0}; on write \code{dq\_oe=1} during the \code{we\_n} pulse. A WRITE\_WAIT
state keeps \code{ce\_n/lb\_n/ub\_n} active for the final hold before release.
\begin{fnwarn}[This is not QSPI]
This is a classic asynchronous-SRAM interface, \textbf{not} QSPI: most commercial
serial/QSPI ``PSRAM'' parts are not compatible with this controller without a rewrite.
See ch.~\ref{ch:hw} for the recommended part (parallel ISSI).
\end{fnwarn}
\section{Read page mode}
\label{sec:pagemode}
The recommended chip (ch.~\ref{ch:hw}) is ``asynchronous/\textbf{page mode}'': once
an initial random access at $t_{AA}$~=~70~ns has been done, further reads inside the
same 16-word page (address bits above \code{A[3]} unchanged) only cost
$t_{APA}$/$t_{PC}$~=~20~ns, because CE\#/OE\# stay asserted and only the address bus
changes. Page mode is \textbf{disabled by default} at power-up (bit~7 of the
configuration register, CR~=~\texttt{0x0070} by default) and must be explicitly
enabled.
\begin{itemize}
\item \textbf{Enable at boot}: right after INIT, the controller runs the
datasheet's ``software-access sequence'' (2 dummy reads + 2 writes, \texttt{0x0000}
unlock then real CR \texttt{0x00F0} = default with the Page bit set) at the
chip's highest address --- it reuses exactly the same READ/WRITE logic as every
other transaction, so it goes through the same timing checks.
\item \textbf{Page bursts}: after a READ the controller no longer closes CE\#/OE\#
(PAGE~OPEN state). A following read in the same page only waits PAGE\_CYCLES; a
read crossing into a different page still avoids a CE\# toggle but pays a full
ACCESS\_CYCLES for that one word (any change at \code{A[4]} or above requires a
new $t_{AA}$). A counter closes the page before the $t_{CEM}$ limit with a
safety margin.
\item \textbf{Only a WRITE closes the page.} Changes to \code{lb\_n}/\code{ub\_n}
do \emph{not} close it: \code{int8\_memory\_access} alternates these signals on
nearly every access (byte-granular access over the 16-bit bus), so treating them
as a close condition --- the first implementation attempt --- made the real
workload \emph{slower}, not faster (measured: 53.25$\to$61.25 cycles/edge on
\code{graph\_engine}'s gather); removed, corrected to 53.25$\to$37.53
cycles/edge (bandwidth +42\%, \S~\ref{sec:bandwidth}).
\end{itemize}
\begin{fnwarn}[No benefit without a sequential pattern]
Page mode only speeds up accesses that stay in the same page (or nearly) while the
controller is waiting for a new request with the page still open. Isolated,
scattered accesses (a random address every time) still pay a full ACCESS\_CYCLES,
plus a small close/reopen overhead if preceded by a WRITE or a $t_{CEM}$ timeout:
it is not a universal win, it depends on the caller's access pattern.
\end{fnwarn}
Real Fmax (\code{nextpnr-ecp5}, ch.~\ref{ch:impl}) on the integrated
\code{spi\_neuron\_top} system with Type~\#2 enabled: \textbf{75.73~MHz} at
\code{PARALLEL}=2 (was 55.59~MHz before page mode was added) and
\textbf{65.13~MHz} at \code{PARALLEL}=8, both still FAIL against the 80~MHz
target but not regressed. The critical path stays, in both cases, entirely
inside \code{u\_graph\_engine.u\_neuron} (the \code{mac8}/\code{neuron\_parallel}
accumulate chain, ch.~\ref{ch:impl}) --- \code{psram\_controller} never appears
in the critical path despite page mode's resource growth.
\section{Address map and conventions}
The addressing space is \code{ADDR\_WIDTH}=23~bits (\emph{byte} address), for a full
8~MB. The regions do not have hardwired addresses: their bases are registers set by the
host via \op{SET\_BASE} (single-layer path) or read from the descriptor table
(multi-layer path).
\begin{tabularx}{\textwidth}{L{3.2cm} L{3.4cm} Y}
\toprule
\rowh \thd{Region} & \thd{Base} & \thd{Content / convention} \\
\midrule
Input $X$ & \code{x\_base} & Shared input vector, read once per invocation. \\
\rowa Weights $W$ & \code{w\_base} & Neuron-major: weights of neuron $n$ at \code{w\_base + n*N\_INPUTS} bytes. \\
Bias & \code{bias\_addr} & One byte per neuron: bias of neuron $n$ at \code{bias\_addr + n}. \\
\rowa Descriptor table & \code{table\_base} & \code{N\_LAYERS} 11-byte entries (ch.~\ref{ch:seq}). \\
Ping-pong buffers A/B & \code{buf\_a\_base} / \code{buf\_b\_base} & Intermediate outputs between layers. \\
\bottomrule
\end{tabularx}
\subsection{PSRAM physical addressing}
The recommended PSRAM is 4M$\times$16 (8~MB), which requires a 22-bit word address
(A0--A21). \code{int8\_memory\_access} computes \code{addr>>1}, turning the 23-bit byte
address into a 22-bit word address that maps exactly onto A0--A21; bit~22 of
\code{psram\_a} is therefore always 0 and 22 real address lines remain on the PCB.
\section{Bandwidth}
\label{sec:bandwidth}
Measured on \code{graph\_engine}'s edge-list gather (ch.~\ref{ch:grafo}), by
difference between two graph sizes to isolate the per-edge cost from the fixed
per-neuron overhead (\code{sim/graph\_engine\_bandwidth\_tb.v}):
\begin{tabularx}{\textwidth}{L{5.2cm} Y Y Y}
\toprule
\rowh \thd{} & \thd{Before (no page mode)} & \thd{After (page mode)} & \thd{$\Delta$} \\
\midrule
Cycles/edge & 53.25 & 37.53 & $-29.5\%$ \\
\rowa Bandwidth @80\,MHz & 6.01\,MB/s & 8.53\,MB/s & $+41.9\%$ \\
Bandwidth @16\,MHz\textsuperscript{*} & 1.20\,MB/s & 1.71\,MB/s & $+41.9\%$ \\
\bottomrule
\end{tabularx}
\textsuperscript{*}recommended real oscillator (ch.~\ref{ch:hw}).
The model still remains memory-bound by construction: \code{neuron\_memory} reads
$X$ once and re-reads $W$/bias for each neuron (ch.~\ref{ch:seq}), one neuron at a
time; page mode reduces the per-byte cost of a sequential access, it does not
eliminate the access pattern itself.