Was a separate, untracked directory (DataSheet/) outside the repo. Renamed to lowercase and moved in as hardware/v2/docs/datasheet/, with its own .gitignore for LaTeX build byproducts (compiled PDFs stay tracked, .aux/.log/.toc/etc do not). Now versioned and shares this repo's own remote instead of living untracked on disk. Content: IT+EN LaTeX chapter sources, reference manufacturer PDFs, and compiled datasheet PDFs including the 2026-09-07 SDRAM upgrade addendum (AS4C32M16SB-7BIN part/pinout/timing) in the v2-en chapters. Note: hardware/v2/docs/DatasheetLatex/ (and the v1 sibling) is a separate, already-tracked, differently-structured LaTeX document that predates this move -- left untouched, not merged, since its chapter set and content differ and merging was not requested. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013xXuuRUWZScuo1DeYJxs3v
188 lines
9.1 KiB
TeX
188 lines
9.1 KiB
TeX
\chapter[ECP5 implementation \& benchmarks]{ECP5 implementation, real benchmark campaign \& measured results}
|
|
\label{ch:impl2}
|
|
|
|
\section{Flow and verification discipline}
|
|
Every number in this chapter is labelled \textsc{Theoretical},
|
|
\textsc{Simulated}, \textsc{Post-P\&R measured}, or \textsc{Derived}
|
|
(a combination of two real measurements, e.g.\ cycles $\div$ real Fmax).
|
|
No result is invented, approximated to look better, or reported without
|
|
a matching real measurement.
|
|
|
|
\begin{tabularx}{\textwidth}{L{5.0cm} C{3.0cm} Y}
|
|
\toprule
|
|
\rowh \thd{Verification stage} & \thd{Outcome} & \thd{Covers} \\
|
|
\midrule
|
|
RTL simulation (Verilator 5.050) & \PASS & bit-exact correctness vs.\ a software golden model \\
|
|
\rowa ECP5 synthesis (Yosys \code{synth\_ecp5}) & \PASS, 0 problems & synthesizability, resource mapping \\
|
|
Place\&route (real \code{nextpnr-ecp5}) & \PASS at \code{N\_SLOTS}$\le$2 & LUT/FF/DSP, real timing \\
|
|
\rowa Full benchmark campaign & 24/24 bit-exact & 6 workloads $\times$ 4 \code{N\_SLOTS} configurations \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
|
|
\begin{fnnote}[Verilator, not Icarus, for V2]
|
|
Two independent Icarus Verilog v13.0 scheduling defects were found and
|
|
reproduced on minimal repros during V2's own M1 milestone (a
|
|
task/scope-entry desync and a spurious condition evaluation, both
|
|
edge-parity dependent) --- Verilator gives correct results on the same
|
|
repros. V1's own certification (performed separately, with Icarus) was
|
|
unaffected, since its own testbenches already avoided the trigger
|
|
pattern by convention; this is flagged honestly, not glossed over.
|
|
\end{fnnote}
|
|
|
|
\section{V1 vs.\ V2 --- final comparison}
|
|
Both systems full-system (not isolated modules), same
|
|
PARALLEL/P\_IN=8, same real, unmodified V1 PSRAM chain.
|
|
|
|
\begin{tabularx}{\textwidth}{L{3.6cm} C{2.6cm} C{2.6cm} Y}
|
|
\toprule
|
|
\rowh \thd{Metric} & \thd{V1} & \thd{V2 (N\_SLOTS=2)} & \thd{Class} \\
|
|
\midrule
|
|
Fmax & 68.65~MHz (\FAIL) & \textbf{87.72~MHz} (\PASS) & Post-P\&R \\
|
|
\rowa LUT (Total LUT4s) & 8907 & 4359 & Post-P\&R \\
|
|
FF (Total DFFs) & 4900 & 3924 & Post-P\&R \\
|
|
\rowa DSP (MULT18X18D) & 16 & 16 & Post-P\&R \\
|
|
BRAM (DP16KD) & 2 & 0 & Post-P\&R \\
|
|
\rowa cycles/neuron (1 neuron, 8 inputs, real PSRAM) & 209 & 166 & Simulated \\
|
|
Real wall-clock speedup vs.\ V1 & 1.00$\times$ & \textbf{2.6$\times$} & Derived \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
\begin{center}\footnotesize\itshape\color{fnGrey}
|
|
V1's own figures are its already-certified, frozen baseline (not
|
|
re-measured this session); V2's figures are real, current measurements
|
|
including both post-campaign optimizations.\end{center}
|
|
|
|
\begin{fnnote}[Where the win comes from --- and where it does not]
|
|
V2's advantage comes from a faster pipeline and a higher achievable
|
|
clock, \textbf{not} primarily from the multi-processor concurrency the
|
|
architecture was built to add. That concurrency's own real payoff, given
|
|
the single-PSRAM-port memory subsystem, is much smaller than a naive
|
|
\code{N\_SLOTS}$\times$\code{P\_IN} calculation would suggest ---
|
|
\S\ref{sec:scaling}.
|
|
\end{fnnote}
|
|
|
|
\section{Real \texttt{N\_SLOTS} sweep --- Fmax and resources}
|
|
Full system, real place\&route, both memory optimizations active
|
|
(word-burst \S\ref{sec:burstimpl} $+$ activation cache
|
|
\S\ref{sec:cacheimpl}).
|
|
|
|
\begin{tabularx}{\textwidth}{C{1.6cm} C{2.2cm} C{1.6cm} C{1.6cm} C{1.6cm} C{1.6cm} Y}
|
|
\toprule
|
|
\rowh \thd{N\_SLOTS} & \thd{Fmax} & \thd{LUT4} & \thd{FF} & \thd{DSP} & \thd{BRAM} & \thd{80\,MHz} \\
|
|
\midrule
|
|
1 & 131.79~MHz & 2760 & 2405 & 8/72 & 0 & \PASS \\
|
|
\rowa 2 & \textbf{87.72~MHz} & 4359 & 3924 & 16/72 & 0 & \PASS (\textbf{recommended}) \\
|
|
4 & 65.01~MHz & 9158 & 7986 & 32/72 & 0 & \FAIL \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
|
|
\subsection{Fmax versus N\_SLOTS}
|
|
\begin{center}
|
|
\begin{tikzpicture}
|
|
\begin{axis}[
|
|
width=0.68\textwidth,height=6.0cm,
|
|
xlabel={\footnotesize N\_SLOTS}, ylabel={\footnotesize Fmax [MHz]},
|
|
xtick={1,2,4}, ymin=50,ymax=145, ytick={50,65,80,95,110,125,140},
|
|
tick label style={font=\scriptsize}, label style={font=\footnotesize},
|
|
grid=major, grid style={fnRule!40},
|
|
legend style={font=\scriptsize,at={(0.5,-0.28)},anchor=north,legend columns=2}]
|
|
\addplot[fnBlue,mark=square*,thick,mark options={fill=fnBlue}]
|
|
coordinates {(1,152.44)(2,133.58)(4,112.07)};
|
|
\addlegendentry{word-burst only}
|
|
\addplot[fnRed,mark=*,thick,mark options={fill=fnRed}]
|
|
coordinates {(1,131.79)(2,87.72)(4,65.01)};
|
|
\addlegendentry{$+$ activation cache}
|
|
\draw[fnAmber,dashed,thick] (axis cs:1,80)--(axis cs:4,80);
|
|
\node[font=\scriptsize,text=fnAmber] at (axis cs:3.3,74){80 MHz target};
|
|
\end{axis}
|
|
\end{tikzpicture}
|
|
\end{center}
|
|
\begin{center}\footnotesize\itshape\color{fnGrey}
|
|
The activation cache's own real Fmax cost grows much faster with
|
|
N\_SLOTS than the arbiter-widening cost alone --- a single shared
|
|
resource with N\_SLOTS request ports and an unpipelined,
|
|
broadcast-capable hit-check.\end{center}
|
|
|
|
\section{Real parallel scaling}
|
|
\label{sec:scaling}
|
|
Not assumed --- computed from real cycle counts, largest workload
|
|
(256 independent neurons sharing one input vector).
|
|
|
|
\begin{tabularx}{\textwidth}{C{1.6cm} C{2.6cm} C{2.0cm} C{2.6cm} C{2.0cm}}
|
|
\toprule
|
|
\rowh \thd{N\_SLOTS} & \thd{Speedup(N)} & \thd{Efficiency} & \thd{PSRAM utilization} & \thd{Real wall-clock speedup vs.\ N=1} \\
|
|
\midrule
|
|
1 & 1.00$\times$ & 100\% & 55.5--71.8\% & 1.00$\times$ \\
|
|
\rowa 2 & 1.06$\times$ & 53\% & $\approx$90\% & 0.99$\times$ (a wash) \\
|
|
4 & 1.06$\times$ & 27\% & $\approx$90\% & 0.79$\times$ (\emph{slower}) \\
|
|
8 & 1.06$\times$ & 13\% & $\approx$90\% & --- \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
\begin{center}\footnotesize\itshape\color{fnGrey}
|
|
Pre-optimization figures, isolating the real scaling behavior from the
|
|
two memory optimizations' own effect (\S\ref{sec:burstimpl}--\ref{sec:cacheimpl}).\end{center}
|
|
|
|
\begin{fnwarn}[The central, measured finding]
|
|
Real cycle-count speedup from \code{N\_SLOTS}=1 to \code{N\_SLOTS}=8 is
|
|
essentially flat (1.05--1.06$\times$) for sustained, memory-bound
|
|
workloads --- the single shared PSRAM port saturates at
|
|
$\approx$90\% utilization regardless of \code{N\_SLOTS}$\ge$2. Once real
|
|
Fmax degradation is also folded in, \code{N\_SLOTS}=4 measures
|
|
\emph{slower} in real wall-clock time than \code{N\_SLOTS}=1. More
|
|
hardware parallelism made this workload class worse, not better,
|
|
because the bottleneck was never compute.
|
|
\end{fnwarn}
|
|
|
|
\section{Memory optimization \#1 --- word-level burst reads}
|
|
\label{sec:burstimpl}
|
|
See ch.~\ref{ch:mem}, \S\ref{sec:burst}, for the full rationale. Real,
|
|
measured single-job cycle reduction: $-$49\% (1 tile), $-$54\% (3
|
|
tiles), $-$56\% (5 tiles). Combined real wall-clock speedup on the
|
|
256-neuron sustained workload: 2.24--2.37$\times$ across every
|
|
\code{N\_SLOTS} tested, at negligible real Fmax cost.
|
|
|
|
\section{Memory optimization \#2 --- shared activation cache}
|
|
\label{sec:cacheimpl}
|
|
See ch.~\ref{ch:mem}, \S\ref{sec:cache}. A further real 1.66--2.00$\times$
|
|
cycle reduction on top of optimization~\#1, at a real, steep Fmax cost
|
|
that makes \code{N\_SLOTS}=4 fail 80\,MHz outright.
|
|
|
|
\begin{tabularx}{\textwidth}{C{1.6cm} C{2.4cm} C{2.4cm} C{2.4cm}}
|
|
\toprule
|
|
\rowh \thd{N\_SLOTS} & \thd{Wall-clock, baseline} & \thd{Wall-clock, final} & \thd{Total real speedup} \\
|
|
\midrule
|
|
1 & 5118.1~$\mu$s & 1324.9~$\mu$s & \textbf{3.86$\times$} \\
|
|
\rowa 2 & 5169.5~$\mu$s & 2113.9~$\mu$s & \textbf{2.45$\times$} (recommended) \\
|
|
4 & 6498.7~$\mu$s & 2842.6~$\mu$s (Fmax fails) & 2.29$\times$ but a real regression vs.\ \#1 alone \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
|
|
\section{Bottleneck analysis}
|
|
\begin{tabularx}{\textwidth}{L{2.8cm} Y}
|
|
\toprule
|
|
\rowh \thd{Candidate} & \thd{Verdict, with real evidence} \\
|
|
\midrule
|
|
\textbf{Memory (PSRAM port)} & \textbf{The real bottleneck.} $\approx$90\% utilization at N\_SLOTS$\ge$2; real compute-to-memory-wait ratio on the order of 1:170--1:220. \\
|
|
\rowa Compute (Neural Processor) & Not the bottleneck --- the pipeline is idle most of the time waiting for data. \\
|
|
Arbiter overhead & Real but small: a mandatory 1-cycle pending-latch (correctness, not choice) plus modest Fmax cost ($-$6\% at N\_SLOTS=2, word-burst alone). \\
|
|
\rowa Director/dependency logic & Not the bottleneck --- zero lost/duplicated jobs, no queueing backlog observed; a real \emph{fairness} issue exists (\S\ref{ch:sched}) but does not limit throughput. \\
|
|
DSP/LUT/FF availability & Not the bottleneck at N\_SLOTS$\le$4 --- all well under budget; DSP would eventually bind at N\_SLOTS=9 (64/72), never reached in practice since the memory bottleneck dominates first. \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
|
|
\section{Limitations, honestly stated}
|
|
\begin{itemize}
|
|
\item V1's own memory-utilization/stall figures were not re-measured
|
|
this session (V1 is frozen); only its already-certified numbers
|
|
are used for comparison.
|
|
\item No clean per-cycle split between ``processor computing'' and
|
|
``processor waiting for memory'' exists in the current
|
|
instrumentation --- reported figures use tile-delivery-rate
|
|
proxies, not an exact split.
|
|
\item Power/energy: \textbf{NOT MEASURED} --- no ECP5 power estimator
|
|
(\code{ecppower}, \code{icepower}, or equivalent) is available in
|
|
this project's toolchain; no value is invented in its place.
|
|
\item \code{N\_SLOTS}=8 was not re-measured full-system (real PSRAM
|
|
chain) after either memory optimization --- only
|
|
\code{dataflow\_core.v} alone, pre-optimization (92.63~MHz).
|
|
\end{itemize}
|