Warning
Work in progress (WIP): This repository is under active development. Microkernels, benchmark drivers, and analysis tools are actively worked on. Anything might change at anytime without notice.
A high-performance General Matrix Multiply (DGEMM) microbenchmark and exploration suite built on top of MIPPv2 (MyIntrinsics++ v2).
This suite evaluates DGEMM microkernels written in C++ with MIPPv2 across 9 distinct CPU microarchitectures : x86_64 AVX2 & AVX-512, AArch64 NEON, and RISC-V Vector (RVV).
Note
Origins & context: This benchmark was originally written as part of a research internship at LIP6 (Sorbonne Université), focusing on the design, implementation, and benchmarking of RISC-V Vector 1.0 (RVV) support for MIPPv2 on the SpacemiT X100 architecture and AVX2 on the Intel Skylake architecture (because it's the one in my laptop).
The benchmark evaluates single-thread, compute-bound double-precision GEMM throughput against the hardware theoretical peak (
Across all target ISA families (x86_64 AVX2 & AVX-512, ARM NEON, RISC-V Vector), some microkernels achieve >90% of theoretical peak performance (reaching up to >99% on some uarchs):
CRR microkernels use a column-major packed
| SIMD | Microarchitecture | CPU / SoC Model | Freq (GHz) | Peak FLOP/cyc | Comp | Peak GFLOP/s | DGEMM FLOP/cyc | % of Peak |
|---|---|---|---|---|---|---|---|---|
| NEON | Apple M1 Firestorm | M1 Ultra (P-core) | 3.0 | 16.0 | GCC | 48.3 | 15.90 | 99.4 % |
| AVX-512 | AMD Zen 5 Strix Point | Ryzen AI 9 HX 370 | 5.1 | 16.0 | GCC | 81.0 | 15.88 | 99.3 % |
| NEON | Cortex-A76 | Raspberry Pi 5 | 2.4 | 8.0 | Clang | 18.9 | 7.87 | 98.4 % |
| AVX-512 | AMD Zen 4 | Ryzen 9 7900X | 5.4 | 16.0 | GCC | 84.4 | 15.64 | 97.7 % |
| RVV 1.0 | SpacemiT X100 | K3 SoC (VLEN=256) | 2.4 | 8.0 | Clang | 18.5 | 7.70 | 96.2 % |
| AVX2 | Intel Skylake | Core i5-6200U | 2.3 | 16.0 | Clang | 35.0 | 15.21 | 95.0 % |
| AVX2 | Intel Meteor Lake | Redwood Cove P-Core | 5.1** | 16.0 | Clang | 74.4 | 14.59 | 91.2 % |
| RVV 1.0 | SpacemiT A100 | K3 SoC (VLEN=1024) | 2.0 | 8.0 | GCC | 13.2 | 6.62 | 82.7 % |
| RVV 1.0 | SpacemiT X60 | BPI-F3 SoC (VLEN=256) | 1.6 | 8.0 | Clang* | 7.7 | 4.79 | 59.9 % |
Important
Intel Meteor Lake frequency characteristics (burst vs. sustained):
- 5.10 GHz is the peak single-core turbo frequency of the Redwood Cove P-core.
- Under sustained heavy compute loads, thermal and power limits typically throttle the core frequency to ~4.70 GHz.
- For our reported champion microkernel measurements, we verified via hardware frequency monitoring that the core sustained the full 5.10 GHz burst clock during the benchmark runs.
- i.e we made sure that the core is actually running @ 5.1GhZ for the microkernels we expected to perform well. But not for all of them.
Note
Zen 5 Strix Point Datapath Note:
Server/desktop Zen 5 incorporates dual native 512-bit FMA units (
To interpret the reported throughput figures accurately, several methodological and architectural boundaries of the current benchmark suite must be highlighted:
-
Isolated Microkernel Scope (In-Cache / Single-Panel):
- Benchmarks evaluate the innermost
$K$ -loop on pre-packed micropanels$A$ ($M_R \times K$ , column-major) and$B$ ($K \times N_R$ , row-major) accumulating into an in-register tile ($M_R \times N_R$ ) and writing to$C$ . - Overheads of the 5-loop cache-blocking hierarchy (GotoBLAS / BLIS:
$J_C, P_C, I_C, J_R, I_R$ ) and dynamic matrix repacking into memory buffers ($M_C \times K_C$ ,$K_C \times N_C$ ) are not included. End-to-end DGEMM performance within BLIS is part of ongoing integration work.
- Benchmarks evaluate the innermost
-
Canonical GEMM Regime & Throughput Formulation:
- Reported peak throughputs strictly evaluate the compute-bound canonical form
$\alpha = 1.0, \beta = 0.0$ ($C \leftarrow AB$ ), where throughput is computed as$\text{GFLOP/s} = \frac{2 \times M \times N \times K}{\text{Time (seconds)} \times 10^9}$ . Non-canonical operations ($\beta \neq 0.0$ ) introduce load-scale-accumulate overheads on$C$ that are supported functionally but not evaluated for peak throughput.
- Reported peak throughputs strictly evaluate the compute-bound canonical form
-
Tile Geometry & Edge Handling (Tails / Fringes):
- Peak pipeline saturation requires dimensions
$M, N$ to be exact multiples of the register tile$M_R, N_R$ and depth$K$ to align with loop unrolling (typically 4). While all microkernels include functional scalar/vector remainder handling for unit-test validation, fringe handling in production BLAS is typically offloaded to dedicated edge kernels or zero-padded buffers.
- Peak pipeline saturation requires dimensions
-
Single-Thread & Precision Scope:
- All measurements are strictly single-thread IEEE-754 double precision (
double/ FP64). Multi-threaded scaling and memory bus contention are not covered in this phase.
- All measurements are strictly single-thread IEEE-754 double precision (
-
Clock Frequency Baseline & Empirical FPU Ceilings:
- Efficiencies are computed against the architectural theoretical ceiling (
$\text{FLOP/cycle} \times \text{Frequency}$ ) using fixed nominal or verified sustained turbo frequencies. Direct empirical FPU pipeline ceiling measurements (generalized fromx100_uarch_exp/broadcast_bench.cpp) will be integrated directly into the suite to isolate pure FMA pipeline saturation on each machine.
- Efficiencies are computed against the architectural theoretical ceiling (
To ensure empirical reproducibility, benchmarks are compiled with modern toolchains across all evaluated platforms. The table below lists the exact compiler versions and ISA configurations on each test machine:
| Microarchitecture | Platform Tag | SoC / Model | ISA / Vector Unit | g++ Version |
clang++ Version |
|---|---|---|---|---|---|
| SpacemiT X100 | x100 |
SpacemiT K3 | RVV 1.0 (256-bit, VLEN=256) | 15.2.0 (Bianbu 15.2.0-16ubuntu1bb3) | 21.1.8 (Ubuntu 21.1.8-6ubuntu1) |
| SpacemiT A100 | a100 |
SpacemiT K3 | RVV 1.0 (1024-bit via ai) |
15.2.0 (Bianbu 15.2.0-16ubuntu1bb3) | 21.1.8 (Ubuntu 21.1.8-6ubuntu1) |
| SpacemiT X60 | x60 |
Banana Pi BPI-F3 | RVV 1.0 (256-bit, VLEN=256) | 15.2.0 (RISCstar 15.2-r1) | 21.1.8 (cross-built on X100)* |
| Raspberry Pi 5 | rpi5 |
Broadcom BCM2712 (Cortex-A76) | ARMv8-A NEON (128-bit) | 15.2.0 (Ubuntu 15.2.0-16ubuntu1) | 21.1.8 (Ubuntu 21.1.8-6ubuntu1) |
| Apple M1 | m1 |
Apple M1 Ultra (Firestorm) | ARMv8-A NEON (128-bit) | 16.1.1 (Red Hat 16.1.1-1) | 22.1.5 (Fedora 22.1.5-1.fc44) |
| AMD Zen 4 | zen4 |
Ryzen 9 7900X | x86_64 AVX-512 | 13.3.0 (Ubuntu 13.3.0-6ubuntu2) | 18.1.3 (Ubuntu 18.1.3-1ubuntu1) |
| AMD Zen 5 | zen5 |
Ryzen AI 9 HX 370 | x86_64 AVX-512 (256-bit DP) | 13.3.0 (Ubuntu 13.3.0-6ubuntu2) | 18.1.3 (Ubuntu 18.1.3-1ubuntu1) |
| Intel Meteor Lake | meteorlake |
Core Ultra (Redwood Cove P-core) | x86_64 AVX2 / FMA | 13.3.0 (Ubuntu 13.3.0-6ubuntu2) | 18.1.3 (Ubuntu 18.1.3-1ubuntu1) |
| Intel Skylake | skylake |
Core i5-6200U (local laptop) | x86_64 AVX2 / FMA | 14.2.1 (GCC 14.2.1 20250405) | 21.1.8 (Clang 21.1.8) |
Note
* SpacemiT X60 Clang Support: The native distribution on the BPI-F3 board includes Bianbu Clang 18.1.8, which lacks functional RVV 1.0 intrinsic support. Because the X60 and X100 share the exact same vector ISA (-march=rv64gcv_zvl256b -mrvv-vector-bits=zvl) and access the shared NFS filesystem on the Dalek cluster, X60 Clang binaries are compiled on the X100 node with Ubuntu Clang 21.1.8 and then executed natively on the X60 hardware.
The executive comparison below summarizes peak double-precision efficiency achieved across all 9 target CPU microarchitectures under the CRR layout (GotoBLAS / BLIS column-major packed
Detailed individual microarchitecture breakdown plots and RRR comparisons can be generated locally via the plotting suite (see Plotting & Performance Analysis below).
- C++20 Compiler:
g++(>= 12.0) orclang++(>= 15.0). - CMake: version 3.20 or newer.
- Python: 3.9+ with
matplotlib,pandas,numpy. - MIPPv2 Headers: Obtain the headers from the MIPP develop branch or download from MIPP Releases.
cmake -B build -S . \
-DCMAKE_BUILD_TYPE=Release \
-DMIPPV2_INCLUDE_DIR=/path/to/mipp/include
cmake --build build -j$(nproc)Unit tests use Catch2 v3 (automatically fetched if not present on system):
cmake -B build -S . \
-DCMAKE_BUILD_TYPE=Release \
-DMIPPV2_INCLUDE_DIR=/path/to/mipp/include \
-DGEMMBENCH_BUILD_TESTS=ON
cmake --build build -j$(nproc)
ctest --test-dir build --output-on-failureBenchmarks are driven by ./run_benchmarks.sh <platform> [options] [compilers...].
Always specify -o <uarch>_<simd> to ensure non-colliding, standardized result files:
# Benchmark Intel Skylake (AVX2) with GCC and Clang
./run_benchmarks.sh avx2 -o skylake_avx2 gcc clang
# Benchmark Intel Meteor Lake (AVX2) on specific matrix sizes
./run_benchmarks.sh avx2 -o meteorlake_avx2 -s 32 -s 64 -s 96 -s 128 gcc clang
# Benchmark AMD Zen 4 (AVX-512) pinned to core 0
./run_benchmarks.sh avx512 -o zen4_avx512 -c 0 gcc clang
# Benchmark SpacemiT X100 with clean rebuild
./run_benchmarks.sh rvv_x100 -o x100_rvv --rebuild clang
# Benchmark CRR kernels only (BLIS-like K sweep on native MR×NR tile)
./run_benchmarks.sh avx512 --crr -o zen4_avx512 gcc clangThe analytical suite plot_results.py reads the hierarchical hardware specification in uarch_config.json, scans the benchmark datasets, identifies the champion kernel for each microarchitecture, and outputs comparison plots:
# Generate all RRR plots from reference datasets
python3 plot_results.py --input-dir gemm-bench-results --output plots
# Filter specific architectures or SIMD extensions
python3 plot_results.py --input-dir gemm-bench-results --uarch zen4 zen5 m1
python3 plot_results.py --input-dir gemm-bench-results --simd avx512 neon
# Override clock frequencies (in GHz) for dynamic boost clock analysis
python3 plot_results.py --input-dir gemm-bench-results --freq zen4=5.4,m1=3.0plot_results_crr.py analyzes CRR micro-kernel benchmark results with automatic champion selection and FLOP/cycle efficiency. It can optionally generate side-by-side comparisons against RRR baselines:
# Generate all CRR plots (auto-detects gemm-bench-results/crr/)
python3 plot_results_crr.py --output plots_crr --png
# With side-by-side CRR vs RRR comparison
python3 plot_results_crr.py --output plots_crr --compare-rrr gemm-bench-results --png
# Filter specific architectures
python3 plot_results_crr.py --uarch zen4 m1 x100 --output plots_crrplots/(RRR):cross_uarch_efficiency_comparison.svg: Multi-architecture comparison in % of theoretical peak.cross_uarch_flop_cycle_comparison.svg: Multi-architecture comparison in DP FLOP/cycle.<simd>_<uarch>_overview.svg: Per-architecture detailed breakdown.
plots_crr/(CRR):cross_uarch_crr_efficiency_comparison.svg: CRR efficiency across architectures.cross_uarch_crr_flop_cycle_comparison.svg: CRR FLOP/cycle comparison.cross_uarch_crr_vs_rrr_comparison.svg: CRR vs RRR speedup (with--compare-rrr).<simd>_<uarch>_crr_overview.svg: Per-architecture CRR breakdown.
This end-to-end tutorial demonstrates how to implement, register, test, benchmark, and analyze a new GEMM microkernel across the entire pipeline.
Microkernels are implemented in include/microkernels/GemmUKernel_opti_rrr.hpp (RRR layout) or include/microkernels/GemmUKernel_opti_crr.hpp (CRR layout). Exploratory kernels go in include/microkernels/GemmUKernel_explo.hpp.
-
Choose Register Tile Geometry (
$M_R \times N_R$ ):- In elements:
$M_R$ rows$\times$ $N_R$ columns, where$N_R = N_V \times VL$ ($N_V$ vectors wide,$VL = \text{vlen}$ ). - Accumulator vector registers required =
$M_R \times N_V = M_R \times \lceil N_R / VL \rceil$ . - For example, a
$2 \times 2VL$ tile ($M_R=2, N_V=2$ ): 4 accumulator registers. - Total register budget
$\approx (M_R \times N_V) + N_V + \min(M_R, 2)$ operands. - Ensure the total budget fits within the architectural register file (16 registers on AVX2, 32 on AVX-512 / NEON / RVV) with zero spilling. On RVV, account for register grouping: physical registers =
$\text{LMUL} \times \text{logical registers}$ .
- In elements:
Tip
-
Register Pressure vs. Loop Unrolling: When unrolling loops (e.g. along
$K$ ), declaring more intermediate variables than available architectural registers is often viable: modern compilers perform live-range analysis and can interleave register reuse without spilling to the stack. Maybe check the assembly though. -
MIPPv2 LMUL Abstraction for Power-of-Two Tiles: For configurations with power-of-two vector widths (
$N_V \in {2, 4, 8}$ , e.g.$2 \times 2VL$ ,$4 \times 2VL$ ,$8 \times 4VL$ ), you can leverage MIPPv2'slmultemplate parameter (set0<T, lmul>(),load<T, lmul>(),fmadd()). This produces concise, portable code across architectures: on RVV it maps directly to hardware vector register groups, while on fixed-width ISAs (AVX2, AVX-512, NEON) MIPPv2 automatically emulates LMUL by unrolling native SIMD registers under a unified API.
- Inner Loop Pattern:
// In include/microkernels/GemmUKernel_opti_rrr.hpp (inside template <typename T> class GemmUKernel): template <int lmul = 1, bool FastPath = false> static inline void gemm_mippv2_myukernel_core(const PackedRowMajor<T> &__restrict A, const PackedRowMajor<T> &__restrict B, PackedRowMajor<T> &__restrict C, T alpha = T{1}, T beta = T{0}) { using namespace mipp; constexpr size_t VL = N<T, lmul>(); constexpr size_t MR = 2; constexpr size_t NR = 2 * VL; for (size_t i = 0; i < A.rows; i += MR) { for (size_t j = 0; j < B.cols; j += NR) { /* * 2x2VL register tile * * c00 c01 * c10 c11 */ auto c00 = set0<T, lmul>(); auto c01 = set0<T, lmul>(); auto c10 = set0<T, lmul>(); auto c11 = set0<T, lmul>(); const auto a0_ptr = A[i + 0]; const auto a1_ptr = A[i + 1]; for (size_t k = 0; k < A.cols; ++k) { const auto b0 = load<T, lmul>(B[k] + j); const auto b1 = load<T, lmul>(B[k] + j + VL); const auto a0 = set1<T, lmul>(a0_ptr[k]); const auto a1 = set1<T, lmul>(a1_ptr[k]); c00 = fmadd(a0, b0, c00); c01 = fmadd(a0, b1, c01); c10 = fmadd(a1, b0, c10); c11 = fmadd(a1, b1, c11); } const auto c0_ptr = C[i + 0] + j; const auto c1_ptr = C[i + 1] + j; Epilogue::store<T, lmul, FastPath>(c0_ptr, c00, alpha, beta); Epilogue::store<T, lmul, FastPath>(c0_ptr + VL, c01, alpha, beta); Epilogue::store<T, lmul, FastPath>(c1_ptr, c10, alpha, beta); Epilogue::store<T, lmul, FastPath>(c1_ptr + VL, c11, alpha, beta); } } } // Dispatch wrappers (inside GemmUKernel<T>): template <int lmul = 1> static inline void gemm_mippv2_myukernel(const PackedRowMajor<T> &__restrict A, const PackedRowMajor<T> &__restrict B, PackedRowMajor<T> &__restrict C) { gemm_mippv2_myukernel_core<lmul, true>(A, B, C); } template <int lmul = 1> static inline void gemm_mippv2_myukernel(const PackedRowMajor<T> &__restrict A, const PackedRowMajor<T> &__restrict B, PackedRowMajor<T> &__restrict C, T alpha, T beta) { if (__builtin_expect(alpha == T{1} && beta == T{0}, 1)) { gemm_mippv2_myukernel_core<lmul, true>(A, B, C); } else { gemm_mippv2_myukernel_core<lmul, false>(A, B, C, alpha, beta); } }
-
Add the enum identifier in
UKernelType:enum class UKernelType { // ... mippv2_myukernel, };
-
Add a
KernelDescriptortokernelTable:{UKernelType::mippv2_myukernel, "mippv2_myukernel", RRR, 64, 4, 3},- Specify layout:
RRR(Row-major$A, B, C$ ) orCRR(Col-major$A$ , Row-major$B, C$ ). -
MRandNR_vecdefine the register tile geometry for auto-sizing. - Specify minimal tile size: typically
64.
- Specify layout:
Add the kernel dispatch branch in the switch (cfg.kernel) statement in src/main.cpp:
case UKernelType::mippv2_myukernel:
BENCH_KERNEL_FASTPATH(
gemm.gemm_mippv2_myukernel(A, B, C),
gemm.gemm_mippv2_myukernel(A, B, C, cfg.alpha, cfg.beta));Add validation test cases for testGemmSuite() and for testGemmSuiteAlphaBeta():
// In testGemmSuite():
TEST_KERNEL("MIPPv2 MyUKernel",
gemm.gemm_mippv2_myukernel(A, B, C_test));
// In testGemmSuiteAlphaBeta():
SECTION("mippv2_myukernel") {
resetMatrix(C_test, C_init);
gemm.gemm_mippv2_myukernel(A, B, C_test, p.alpha, p.beta);
checkMatrixEqual(C_ref, C_test);
}Run ctest --test-dir build --output-on-failure to verify correctness and numerical precision.
In run_benchmarks.sh, add the kernel name to UNIVERSAL_KERNELS:
mippv2_myukernelExecute the benchmark specifying the output prefix and kernel:
./run_benchmarks.sh avx2 -o myarch_avx2 -k mippv2_myukernel gcc clangIn uarch_config.json, add or update the target entry under its SIMD family (e.g., "avx2"):
"avx2": {
"myarch": {
"name": "My Architecture",
"cpu_model": "Custom Core",
"frequency_ghz": 3.5,
"peak_flop_per_cycle": 16.0,
"expected_best_kernel": "mippv2_myukernel",
"csv_pattern": "*gemm_results_*myarch*_{compiler}.csv"
}
}Run python3 plot_results.py --input-dir gemm-bench-results/rrr to generate updated cross-architecture graphs.
Current Focus & Roadmap (Prioritized TODO):
- Comparison against production BLAS (BLIS & OpenBLAS)*: Benchmark MIPPv2 microkernels against real BLAS libraries across all platforms. To do so the repo will need to be extended to a proper GEMM benchmark suite. And not just implement microkernels.
- Single-Precision Support (SGEMM / FP32): Extend the benchmark suite to evaluate FP32 as well.
-
Comparison with Hand-Written Native Intrinsics: Implement intrinsics baselines (native RVV
riscv_vector.h, AVX-512immintrin.h, NEONarm_neon.h) to measure the exact abstraction overhead introduced by MIPPv2. -
Cross-SIMD Abstraction Comparisons: Benchmark against alternative SIMD abstraction frameworks, like Google Highway, EVE, and
std::simd. -
Non-Canonical GEMM Evaluation: Benchmark throughput and epilogue overhead for non-canonical forms (
$\beta \neq 0.0$ ,$\alpha \neq 1.0$ ) to quantify the cost of$C$ tile reloads and scaling.
Note
BLIS/OpenBLAS for RVV: Ongoing vendor and community work on RVV-optimized BLAS kernels reportedly outperforms current upstream public releases of BLIS and OpenBLAS, but these patches are not yet fully merged/streamlined in mainline distributions. Benchmarking methodology should evaluate both upstream releases and patched branches.
All of this is a lot of work. And I'm theoretically on vacations. I'm pretty sure I will do it eventually. But idk when.
- Adrien Cassagne (@kouchy) for access to the single-board computers (SBCs) and nodes on the Dalek Cluster ("Dalek: An Unconventional and Energy-Aware Heterogeneous Cluster", arXiv:2508.10481).
- SpacemiT, for the SpacemiT K3 Technical Whitepaper.
- camel-cdr, for the RVV Benchmark Results project.
- The authors of "Great Expectations: Benchmarking the Real-World Performance of RVV 1.0 in HPC" (arXiv:2608.28097).
Generative AI assistants (Google Gemini and OpenAI ChatGPT) were utilized during the development of this repository to accelerate boilerplate tasks and infrastructure setup, specifically:
- Mechanically generating C++ microkernel implementations across varying register tile geometries (
$M_R \times N_R$ ) and accumulator layouts. - Automation scripts (shell runners and Python data parsing / plotting pipelines).
- Benchmark driver CLI scaffolding, argument handling, and boilerplate wiring in C++.
- Build system configuration and test suite plumbing.
The foundational DGEMM microkernel designs tuned for Intel Skylake (AVX2 + FMA) and the SpacemiT K3 X100 core are the result of multiple weeks of work and reflection on the underlying microarchitectures and platforms. And the best X100 microkernels are, I think, some of the best results of my internship.
For subsequent target microarchitectures, the methodology was empirical, the exploration workflow relied on pre-established invariants:
-
Layout Exploration: Both the
$RRR$ format (Row-major$A, B, C$ ) and the$CRR$ format (Column-major$A$ , Row-major$B, C$ ) are implemented. CRR uses the GotoBLAS / BLIS packing convention where$A$ is stored contiguously along$M_R$ with leading dimension$= M_R$ , matching the inner kernel's access pattern. CRR microkernels achieve comparable or higher peak efficiency than RRR on most microarchitectures (up to 99.3% on M1). -
Inner-Loop Access Patterns: Preserving either the scalar broadcast pattern (
mipp::set1+mipp::fmadd) or the indexed FMA pattern (mipp::fmaddi) on matrix$A$ .
Under these, the exploration focused on sweeping register block dimensions (
- How many accumulators can I have at most?
- How many accumulators do I want?
- How many accumulators do I need at least?
In practice, it often boiled down to telling Gemini to implement a bunch of tile sizes I thought might yield good results. Then I ran the benchmarks on the target hardware and iterated over the ones that performed well "Gemini please make gemm_mippv2_firestorm_mr8_nr2_fmaddi using gemm_mippv2_firestorm_mr6_nr4_fmaddi as reference" and so on. I feel like the most inefficient part of the automation was sometimes the agent sitting between the keyboard and the chair. Oh well.