191 lines
9.0 KiB
TeX
191 lines
9.0 KiB
TeX
\chapter[Graph network (Type \#2)]{Two-level configuration: graph network (Type \#2)}
|
|
\label{ch:grafo}
|
|
|
|
\section{Two network types}
|
|
The engine exposes two \emph{network types} selectable by the host, with the same start
|
|
command dispatching to the correct engine:
|
|
|
|
\begin{itemize}
|
|
\item \textbf{Type \#1 --- classic network (dense).} Layers with neurons per layer, fully
|
|
connected between consecutive layers. It is the \code{layer\_sequencer} path
|
|
(ch.~\ref{ch:seq}), started by \op{RUN\_NETWORK}. Connections are \emph{implicit by
|
|
position}: nothing is enumerated, only the weights are defined, addressed as
|
|
\code{w\_base + k*n\_inputs + j}.
|
|
\item \textbf{Type \#2 --- arbitrary graph (sparse).} Starting from the input neuron ids,
|
|
each neuron's connections up to the output are defined through a per-neuron \emph{sparse
|
|
edge-list}. Connections are \emph{explicit by enumeration}: each connection is an edge
|
|
\code{(src\_id, weight)}; if it is not in the list, it does not exist.
|
|
\end{itemize}
|
|
|
|
\begin{fnnote}[The difference in one line]
|
|
Dense: you define the \emph{weights} by position in a matrix. Graph: you define each
|
|
\emph{connection} as an edge \code{(src\_id, weight)} in a per-neuron list. The two
|
|
descriptor tables share the same 11-byte format but different fields; the
|
|
\code{net\_type} register tells the engine which interpretation to use.
|
|
\end{fnnote}
|
|
|
|
\section{Global activation buffer}
|
|
Type \#2 introduces an \textbf{activation buffer} indexed by \emph{signal id}, one INT8
|
|
byte per id, implemented in \textbf{on-chip \code{DP16KD} block RAM}
|
|
(\code{rtl/act\_buffer.v}). Ids \code{0..N\_in-1} are the inputs; each neuron writes its
|
|
own output into its own id. The source gather reads from here with \emph{single-cycle
|
|
random access}: this is what makes the graph cheap, because it is the access that PSRAM
|
|
(70~ns, sequential) could not accelerate.
|
|
|
|
\begin{fnspec}[V1 sizing]
|
|
\code{N\_TOTAL}=4096 signals, 16-bit id (room to 65\,536 without changing the format).
|
|
Buffer = 4~KB, i.e. 2 \code{DP16KD} blocks out of 108. The real constraint becomes the
|
|
PSRAM edge capacity ($\approx$2\,M edges at 4~B), not block RAM.
|
|
\end{fnspec}
|
|
|
|
\section{Feed-forward DAG and the \texttt{src\_id < out\_id} rule}
|
|
The graph is a feed-forward DAG: every connection points to an \textbf{already-computed}
|
|
id (\code{src\_id < out\_id}). Neurons are processed in ascending id order, so that when a
|
|
neuron is computed all its sources are ready in the buffer. Cycles and recurrence are out
|
|
of scope for V1. The rule is checked at two levels: by the host assembler (compile time)
|
|
and by a runtime guard in \code{graph\_engine} (\code{STATUS.err}), in the same philosophy
|
|
as the elaboration guard on \code{N\_INPUTS \% PARALLEL}.
|
|
|
|
\section{Data formats}
|
|
Both descriptors are 11~bytes/entry, MSB-first, at \code{table\_base}.
|
|
|
|
\subsection{Type \#2 descriptor (graph)}
|
|
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.6cm} Y}
|
|
\toprule
|
|
\rowh \thd{Field} & \thd{Bytes} & \thd{Meaning} \\
|
|
\midrule
|
|
\code{conn\_ptr} & 3 & Byte address in PSRAM of the neuron's edge block. \\
|
|
\rowa \code{n\_conn} & 2 & Real connections (pre-padding). \\
|
|
\code{out\_id} & 2 & Id into which the neuron's output is written. \\
|
|
\rowa \code{activation} & 1 & \code{ACT\_RELU} / \code{ACT\_NONE} (low 2 bits). \\
|
|
\code{bias} & 1 & Neuron bias (INT8). \\
|
|
\rowa \code{reserved} & 2 & 0. \\
|
|
\midrule
|
|
\rowh \thd{Total} & \thd{11} & entries in ascending \code{out\_id} order \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
|
|
\subsection{Graph edge (4~bytes, aligned)}
|
|
\begin{tabularx}{\textwidth}{L{3.4cm} C{1.6cm} Y}
|
|
\toprule
|
|
\rowh \thd{Field} & \thd{Bytes} & \thd{Meaning} \\
|
|
\midrule
|
|
\code{src\_id} & 2 & Source id (uint16 BE). \\
|
|
\rowa \code{weight} & 1 & Weight (INT8). \\
|
|
\code{reserved} & 1 & 0 (4-byte alignment). \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
|
|
\begin{fnnote}[Padding to \texttt{PARALLEL}]
|
|
An arbitrary \code{n\_conn} is not a multiple of \code{PARALLEL}: the neuron's edge-list
|
|
is padded up to the multiple with \textbf{zero-weight} edges (waste
|
|
$\le$\code{PARALLEL}$-1$ per neuron). This keeps the datapath and its guard intact.
|
|
\end{fnnote}
|
|
|
|
\section{\texttt{graph\_engine} --- graph engine}
|
|
\code{rtl/graph\_engine.v} orchestrates Type \#2 \textbf{reusing \code{neuron\_parallel}
|
|
unmodified}, as \code{neuron\_memory} does for the dense case. Key difference: between the
|
|
two modes only the \emph{X addressing} changes. In Type \#1 the input is contiguous
|
|
(\code{x\_base + i}); in Type \#2 it is a gather (\code{act\_buf[src\_id]}). The arithmetic
|
|
core is untouched.
|
|
|
|
\begin{center}
|
|
\begin{tikzpicture}[font=\scriptsize,node distance=4mm,start chain=going below,
|
|
every node/.style={on chain,fnblock,minimum width=52mm}]
|
|
\node[fnblockA]{\code{COPY\_INPUTS}: PSRAM \code{x\_base} $\to$ \code{act\_buf[0..N\_in-1]}};
|
|
\node{\code{READ\_DESC}: descriptor of neuron k};
|
|
\node{\code{READ\_EDGES}: stream edges + gather \code{act\_buf[src\_id]}};
|
|
\node{\code{START\_N} / \code{WAIT\_N}: group of \code{PARALLEL} $\to$ \code{neuron\_parallel}};
|
|
\node[fnblockT]{\code{WRITE\_ACT}: y $\to$ \code{act\_buf[out\_id]}};
|
|
\node{next neuron (id order)};
|
|
\node[fnblockD]{\code{WRITE\_OUTPUTS}: last \code{n\_out} $\to$ PSRAM \code{out\_base}};
|
|
\foreach \i [count=\j from 2] in {1,...,6}
|
|
\draw[fnarrow] (chain-\i) -- (chain-\j);
|
|
\end{tikzpicture}
|
|
\end{center}
|
|
|
|
The outputs are the \textbf{last \code{n\_out}} ids: in a DAG with the
|
|
\code{src\_id < out\_id} ordering the output neurons (sinks, not reused as sources)
|
|
naturally end up with the highest ids. At the end \code{graph\_engine} copies these
|
|
\code{n\_out} bytes into a PSRAM region at \code{out\_base}, which the host reads back with
|
|
\op{READ\_RAM}.
|
|
|
|
\section{Type \#2 opcodes and registers}
|
|
The type is selected with a new opcode; \op{RUN\_NETWORK} dispatches on the
|
|
\code{net\_type} register (details in ch.~\ref{ch:spi}).
|
|
|
|
\begin{tabularx}{\textwidth}{L{2.6cm} L{3.4cm} Y}
|
|
\toprule
|
|
\rowh \thd{Opcode / sel} & \thd{Name} & \thd{Function} \\
|
|
\midrule
|
|
\op{0x11} & SET\_NET\_TYPE & \code{type(1B)}: \code{0x01}=dense (\#1), \code{0x02}=graph (\#2). Default after \op{RESET}=dense. \\
|
|
\rowa \code{SET\_BASE sel 9} & num\_neurons\_graph & Number of graph neurons (uint16). \\
|
|
\code{SET\_BASE sel 10} & n\_out & Number of output ids (uint16). \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
|
|
\begin{fnnote}[Zero regression on Type \#1]
|
|
With \code{net\_type=dense} (the default value after \op{RESET}) the \#1 path is
|
|
bit-identical to before: \op{RUN\_NETWORK} keeps its \code{num\_layers(1B)} payload and the
|
|
framing of the existing opcodes does not change.
|
|
\end{fnnote}
|
|
|
|
\section{Occupancy (Type \#2 enabled)}
|
|
Yosys synthesis of the full \code{spi\_neuron\_top} system with Type \#2 enabled
|
|
(\code{PARALLEL}=2):
|
|
|
|
\begin{tabularx}{\textwidth}{L{4.6cm} Y}
|
|
\toprule
|
|
\rowh \thd{Resource} & \thd{Use} \\
|
|
\midrule
|
|
\code{DP16KD} (block RAM) & 2 (activation buffer) \\
|
|
\rowa \code{MULT18X18D} (DSP) & 4 (2 \code{neuron\_memory} + 2 \code{graph\_engine}) \\
|
|
LUT4 & 2619 \\
|
|
\rowa TRELLIS\_FF & 2467 \\
|
|
\code{\$\_TBUF\_} (PSRAM bus) & 16 \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
The device (108 \code{DP16KD}, 72 DSP, $\approx$44k LUT/FF) stays well below saturation:
|
|
Type \#2 adds a complete mode at a contained resource cost. LUT4/TRELLIS\_FF grew from an
|
|
earlier measurement (2367/2406) because of the PSRAM page mode added to the controller
|
|
(ch.~\ref{ch:mem}, \S~5.5) --- under 6\% utilization, no practical impact.
|
|
|
|
\section{Gather bandwidth (measured)}
|
|
The per-edge gather cost was \textbf{isolated} by building two structurally-identical
|
|
graphs with different edge counts and differencing the cycles: the subtraction cancels the
|
|
fixed per-neuron overhead and leaves the edge cost alone.
|
|
|
|
\begin{fnspec}[Per-edge cost]
|
|
\textbf{37.53 cycles/edge} with PSRAM page mode enabled (ch.~\ref{ch:mem}, \S~5.5) ---
|
|
\textbf{53.25 cycles/edge} without it (pre-page-mode baseline, consistent with theory:
|
|
4~bytes/edge $\times$ $\approx$13 cycles/byte over async PSRAM $\approx$52). At 80~MHz:
|
|
$\approx$2.13\,M edges/s ($\approx$8.5~MB/s, +42\% vs. baseline); at the real 16~MHz
|
|
clock: $\approx$426\,k edges/s ($\approx$1.71~MB/s).
|
|
\end{fnspec}
|
|
|
|
Page-mode read (roadmap G7, ch.~\ref{ch:roadmap}) has been implemented and measured: the
|
|
gather's sequential access benefits directly, cutting the per-edge cost by 29.5\%
|
|
(53.25$\to$37.53 cycles/edge). Each edge still pays \code{int8\_memory\_access}'s
|
|
byte-granular access (4 bytes/edge); page mode reduces the cost of each sequential byte,
|
|
not the number of accesses.
|
|
|
|
\section{\texttt{netasm} host assembler}
|
|
Readable network configuration needs no dedicated FPGA logic: a pseudo-assembly is
|
|
compiled \emph{on the host} (\code{tools/netasm/}) into the exact bytes of the tables and
|
|
edges, then loaded with \op{WRITE\_RAM}. The assembler validates at compile time
|
|
(\code{src\_id < out\_id}, \code{N\_TOTAL} bounds, padding to \code{PARALLEL}),
|
|
complementing the runtime guard.
|
|
|
|
\begin{lstlisting}[language=,caption={Pseudo-assembly example (graph)},basicstyle=\ttfamily\scriptsize]
|
|
NET graph
|
|
INPUTS 4 ; ids 0..3
|
|
NEURON n4 relu bias=2
|
|
CONN 0 w=5
|
|
CONN 1 w=-3
|
|
NEURON n5 none bias=0
|
|
CONN n4 w=2 ; symbolic reference to n4's output
|
|
CONN 2 w=7
|
|
OUTPUT n5
|
|
END
|
|
\end{lstlisting}
|