Skip to content

Repository files navigation

GEMM Benchmark suite across modern microarchitectures using MIPPv2

Warning

Work in progress (WIP): This repository is under active development. Microkernels, benchmark drivers, and analysis tools are actively worked on. Anything might change at anytime without notice.

A high-performance General Matrix Multiply (DGEMM) microbenchmark and exploration suite built on top of MIPPv2 (MyIntrinsics++ v2).

This suite evaluates DGEMM microkernels written in C++ with MIPPv2 across 9 distinct CPU microarchitectures : x86_64 AVX2 & AVX-512, AArch64 NEON, and RISC-V Vector (RVV).

Note

Origins & context: This benchmark was originally written as part of a research internship at LIP6 (Sorbonne Université), focusing on the design, implementation, and benchmarking of RISC-V Vector 1.0 (RVV) support for MIPPv2 on the SpacemiT X100 architecture and AVX2 on the Intel Skylake architecture (because it's the one in my laptop).


Microarchitectural Performance & Peak Efficiency

The benchmark evaluates single-thread, compute-bound double-precision GEMM throughput against the hardware theoretical peak ($\text{FLOP/cycle} \times \text{Frequency}$).

Across all target ISA families (x86_64 AVX2 & AVX-512, ARM NEON, RISC-V Vector), some microkernels achieve >90% of theoretical peak performance (reaching up to >99% on some uarchs):

CRR Layout (Col-Major A, Row-Major B, C) — GotoBLAS / BLIS packing

CRR microkernels use a column-major packed $A$ panel (leading dimension $= M_R$), matching the access pattern of the GotoBLAS / BLIS macro-kernel inner loop. The benchmark evaluates a single register tile ($M_R \times N_R$) across panel depth $K \in {32, 64, 128, 256, 512, 1024}$.

SIMD Microarchitecture CPU / SoC Model Freq (GHz) Peak FLOP/cyc Comp Peak GFLOP/s DGEMM FLOP/cyc % of Peak
NEON Apple M1 Firestorm M1 Ultra (P-core) 3.0 16.0 GCC 48.3 15.90 99.4 %
AVX-512 AMD Zen 5 Strix Point Ryzen AI 9 HX 370 5.1 16.0 GCC 81.0 15.88 99.3 %
NEON Cortex-A76 Raspberry Pi 5 2.4 8.0 Clang 18.9 7.87 98.4 %
AVX-512 AMD Zen 4 Ryzen 9 7900X 5.4 16.0 GCC 84.4 15.64 97.7 %
RVV 1.0 SpacemiT X100 K3 SoC (VLEN=256) 2.4 8.0 Clang 18.5 7.70 96.2 %
AVX2 Intel Skylake Core i5-6200U 2.3 16.0 Clang 35.0 15.21 95.0 %
AVX2 Intel Meteor Lake Redwood Cove P-Core 5.1** 16.0 Clang 74.4 14.59 91.2 %
RVV 1.0 SpacemiT A100 K3 SoC (VLEN=1024) 2.0 8.0 GCC 13.2 6.62 82.7 %
RVV 1.0 SpacemiT X60 BPI-F3 SoC (VLEN=256) 1.6 8.0 Clang* 7.7 4.79 59.9 %

Important

Intel Meteor Lake frequency characteristics (burst vs. sustained):

  • 5.10 GHz is the peak single-core turbo frequency of the Redwood Cove P-core.
  • Under sustained heavy compute loads, thermal and power limits typically throttle the core frequency to ~4.70 GHz.
  • For our reported champion microkernel measurements, we verified via hardware frequency monitoring that the core sustained the full 5.10 GHz burst clock during the benchmark runs.
  • i.e we made sure that the core is actually running @ 5.1GhZ for the microkernels we expected to perform well. But not for all of them.

Note

Zen 5 Strix Point Datapath Note: Server/desktop Zen 5 incorporates dual native 512-bit FMA units ($32\text{ DP FLOP/cycle}$), AMD's mobile Strix Point SoC implements dual 256-bit physical datapaths (double-pumped 512-bit FMA, identical to Zen 4), yielding a physical ceiling of $16\text{ DP FLOP/cycle}$. If you own a Zen5 with 512-bit native FMA units and don't know what to do with it, I'd be very happy if you provided benchmarks results on the existing microkernels :o .

Methodological Scope & Current Limitations

To interpret the reported throughput figures accurately, several methodological and architectural boundaries of the current benchmark suite must be highlighted:

  1. Isolated Microkernel Scope (In-Cache / Single-Panel):
    • Benchmarks evaluate the innermost $K$-loop on pre-packed micropanels $A$ ($M_R \times K$, column-major) and $B$ ($K \times N_R$, row-major) accumulating into an in-register tile ($M_R \times N_R$) and writing to $C$.
    • Overheads of the 5-loop cache-blocking hierarchy (GotoBLAS / BLIS: $J_C, P_C, I_C, J_R, I_R$) and dynamic matrix repacking into memory buffers ($M_C \times K_C$, $K_C \times N_C$) are not included. End-to-end DGEMM performance within BLIS is part of ongoing integration work.
  2. Canonical GEMM Regime & Throughput Formulation:
    • Reported peak throughputs strictly evaluate the compute-bound canonical form $\alpha = 1.0, \beta = 0.0$ ($C \leftarrow AB$), where throughput is computed as $\text{GFLOP/s} = \frac{2 \times M \times N \times K}{\text{Time (seconds)} \times 10^9}$. Non-canonical operations ($\beta \neq 0.0$) introduce load-scale-accumulate overheads on $C$ that are supported functionally but not evaluated for peak throughput.
  3. Tile Geometry & Edge Handling (Tails / Fringes):
    • Peak pipeline saturation requires dimensions $M, N$ to be exact multiples of the register tile $M_R, N_R$ and depth $K$ to align with loop unrolling (typically 4). While all microkernels include functional scalar/vector remainder handling for unit-test validation, fringe handling in production BLAS is typically offloaded to dedicated edge kernels or zero-padded buffers.
  4. Single-Thread & Precision Scope:
    • All measurements are strictly single-thread IEEE-754 double precision (double / FP64). Multi-threaded scaling and memory bus contention are not covered in this phase.
  5. Clock Frequency Baseline & Empirical FPU Ceilings:
    • Efficiencies are computed against the architectural theoretical ceiling ($\text{FLOP/cycle} \times \text{Frequency}$) using fixed nominal or verified sustained turbo frequencies. Direct empirical FPU pipeline ceiling measurements (generalized from x100_uarch_exp/broadcast_bench.cpp) will be integrated directly into the suite to isolate pure FMA pipeline saturation on each machine.

Hardware Platforms & Compiler Toolchains

To ensure empirical reproducibility, benchmarks are compiled with modern toolchains across all evaluated platforms. The table below lists the exact compiler versions and ISA configurations on each test machine:

Microarchitecture Platform Tag SoC / Model ISA / Vector Unit g++ Version clang++ Version
SpacemiT X100 x100 SpacemiT K3 RVV 1.0 (256-bit, VLEN=256) 15.2.0 (Bianbu 15.2.0-16ubuntu1bb3) 21.1.8 (Ubuntu 21.1.8-6ubuntu1)
SpacemiT A100 a100 SpacemiT K3 RVV 1.0 (1024-bit via ai) 15.2.0 (Bianbu 15.2.0-16ubuntu1bb3) 21.1.8 (Ubuntu 21.1.8-6ubuntu1)
SpacemiT X60 x60 Banana Pi BPI-F3 RVV 1.0 (256-bit, VLEN=256) 15.2.0 (RISCstar 15.2-r1) 21.1.8 (cross-built on X100)*
Raspberry Pi 5 rpi5 Broadcom BCM2712 (Cortex-A76) ARMv8-A NEON (128-bit) 15.2.0 (Ubuntu 15.2.0-16ubuntu1) 21.1.8 (Ubuntu 21.1.8-6ubuntu1)
Apple M1 m1 Apple M1 Ultra (Firestorm) ARMv8-A NEON (128-bit) 16.1.1 (Red Hat 16.1.1-1) 22.1.5 (Fedora 22.1.5-1.fc44)
AMD Zen 4 zen4 Ryzen 9 7900X x86_64 AVX-512 13.3.0 (Ubuntu 13.3.0-6ubuntu2) 18.1.3 (Ubuntu 18.1.3-1ubuntu1)
AMD Zen 5 zen5 Ryzen AI 9 HX 370 x86_64 AVX-512 (256-bit DP) 13.3.0 (Ubuntu 13.3.0-6ubuntu2) 18.1.3 (Ubuntu 18.1.3-1ubuntu1)
Intel Meteor Lake meteorlake Core Ultra (Redwood Cove P-core) x86_64 AVX2 / FMA 13.3.0 (Ubuntu 13.3.0-6ubuntu2) 18.1.3 (Ubuntu 18.1.3-1ubuntu1)
Intel Skylake skylake Core i5-6200U (local laptop) x86_64 AVX2 / FMA 14.2.1 (GCC 14.2.1 20250405) 21.1.8 (Clang 21.1.8)

Note

* SpacemiT X60 Clang Support: The native distribution on the BPI-F3 board includes Bianbu Clang 18.1.8, which lacks functional RVV 1.0 intrinsic support. Because the X60 and X100 share the exact same vector ISA (-march=rv64gcv_zvl256b -mrvv-vector-bits=zvl) and access the shared NFS filesystem on the Dalek cluster, X60 Clang binaries are compiled on the X100 node with Ubuntu Clang 21.1.8 and then executed natively on the X60 hardware.


Performance Visualizations

The executive comparison below summarizes peak double-precision efficiency achieved across all 9 target CPU microarchitectures under the CRR layout (GotoBLAS / BLIS column-major packed $A$ panel), grouped by SIMD family:

GEMMBench Cross-Microarchitecture CRR Efficiency Comparison

Detailed individual microarchitecture breakdown plots and RRR comparisons can be generated locally via the plotting suite (see Plotting & Performance Analysis below).


Prerequisites & Building

1. Requirements

  • C++20 Compiler: g++ (>= 12.0) or clang++ (>= 15.0).
  • CMake: version 3.20 or newer.
  • Python: 3.9+ with matplotlib, pandas, numpy.
  • MIPPv2 Headers: Obtain the headers from the MIPP develop branch or download from MIPP Releases.

2. Standalone Build

cmake -B build -S . \
  -DCMAKE_BUILD_TYPE=Release \
  -DMIPPV2_INCLUDE_DIR=/path/to/mipp/include

cmake --build build -j$(nproc)

3. Build and Run Unit Tests

Unit tests use Catch2 v3 (automatically fetched if not present on system):

cmake -B build -S . \
  -DCMAKE_BUILD_TYPE=Release \
  -DMIPPV2_INCLUDE_DIR=/path/to/mipp/include \
  -DGEMMBENCH_BUILD_TESTS=ON

cmake --build build -j$(nproc)
ctest --test-dir build --output-on-failure

Running Benchmarks

Benchmarks are driven by ./run_benchmarks.sh <platform> [options] [compilers...].

Always specify -o <uarch>_<simd> to ensure non-colliding, standardized result files:

# Benchmark Intel Skylake (AVX2) with GCC and Clang
./run_benchmarks.sh avx2 -o skylake_avx2 gcc clang

# Benchmark Intel Meteor Lake (AVX2) on specific matrix sizes
./run_benchmarks.sh avx2 -o meteorlake_avx2 -s 32 -s 64 -s 96 -s 128 gcc clang

# Benchmark AMD Zen 4 (AVX-512) pinned to core 0
./run_benchmarks.sh avx512 -o zen4_avx512 -c 0 gcc clang

# Benchmark SpacemiT X100 with clean rebuild
./run_benchmarks.sh rvv_x100 -o x100_rvv --rebuild clang

# Benchmark CRR kernels only (BLIS-like K sweep on native MR×NR tile)
./run_benchmarks.sh avx512 --crr -o zen4_avx512 gcc clang

Plotting & Performance Analysis

RRR Plots

The analytical suite plot_results.py reads the hierarchical hardware specification in uarch_config.json, scans the benchmark datasets, identifies the champion kernel for each microarchitecture, and outputs comparison plots:

# Generate all RRR plots from reference datasets
python3 plot_results.py --input-dir gemm-bench-results --output plots

# Filter specific architectures or SIMD extensions
python3 plot_results.py --input-dir gemm-bench-results --uarch zen4 zen5 m1
python3 plot_results.py --input-dir gemm-bench-results --simd avx512 neon

# Override clock frequencies (in GHz) for dynamic boost clock analysis
python3 plot_results.py --input-dir gemm-bench-results --freq zen4=5.4,m1=3.0

CRR Plots

plot_results_crr.py analyzes CRR micro-kernel benchmark results with automatic champion selection and FLOP/cycle efficiency. It can optionally generate side-by-side comparisons against RRR baselines:

# Generate all CRR plots (auto-detects gemm-bench-results/crr/)
python3 plot_results_crr.py --output plots_crr --png

# With side-by-side CRR vs RRR comparison
python3 plot_results_crr.py --output plots_crr --compare-rrr gemm-bench-results --png

# Filter specific architectures
python3 plot_results_crr.py --uarch zen4 m1 x100 --output plots_crr

Output Figures

  • plots/ (RRR):
    • cross_uarch_efficiency_comparison.svg: Multi-architecture comparison in % of theoretical peak.
    • cross_uarch_flop_cycle_comparison.svg: Multi-architecture comparison in DP FLOP/cycle.
    • <simd>_<uarch>_overview.svg: Per-architecture detailed breakdown.
  • plots_crr/ (CRR):
    • cross_uarch_crr_efficiency_comparison.svg: CRR efficiency across architectures.
    • cross_uarch_crr_flop_cycle_comparison.svg: CRR FLOP/cycle comparison.
    • cross_uarch_crr_vs_rrr_comparison.svg: CRR vs RRR speedup (with --compare-rrr).
    • <simd>_<uarch>_crr_overview.svg: Per-architecture CRR breakdown.

Developer Tutorial: Adding a New Microkernel

This end-to-end tutorial demonstrates how to implement, register, test, benchmark, and analyze a new GEMM microkernel across the entire pipeline.

Step 1: Implement the Kernel in C++

Microkernels are implemented in include/microkernels/GemmUKernel_opti_rrr.hpp (RRR layout) or include/microkernels/GemmUKernel_opti_crr.hpp (CRR layout). Exploratory kernels go in include/microkernels/GemmUKernel_explo.hpp.

  1. Choose Register Tile Geometry ($M_R \times N_R$):
    • In elements: $M_R$ rows $\times$ $N_R$ columns, where $N_R = N_V \times VL$ ($N_V$ vectors wide, $VL = \text{vlen}$).
    • Accumulator vector registers required = $M_R \times N_V = M_R \times \lceil N_R / VL \rceil$.
    • For example, a $2 \times 2VL$ tile ($M_R=2, N_V=2$): 4 accumulator registers.
    • Total register budget $\approx (M_R \times N_V) + N_V + \min(M_R, 2)$ operands.
    • Ensure the total budget fits within the architectural register file (16 registers on AVX2, 32 on AVX-512 / NEON / RVV) with zero spilling. On RVV, account for register grouping: physical registers = $\text{LMUL} \times \text{logical registers}$.

Tip

  • Register Pressure vs. Loop Unrolling: When unrolling loops (e.g. along $K$), declaring more intermediate variables than available architectural registers is often viable: modern compilers perform live-range analysis and can interleave register reuse without spilling to the stack. Maybe check the assembly though.
  • MIPPv2 LMUL Abstraction for Power-of-Two Tiles: For configurations with power-of-two vector widths ($N_V \in {2, 4, 8}$, e.g. $2 \times 2VL$, $4 \times 2VL$, $8 \times 4VL$), you can leverage MIPPv2's lmul template parameter (set0<T, lmul>(), load<T, lmul>(), fmadd()). This produces concise, portable code across architectures: on RVV it maps directly to hardware vector register groups, while on fixed-width ISAs (AVX2, AVX-512, NEON) MIPPv2 automatically emulates LMUL by unrolling native SIMD registers under a unified API.
  1. Inner Loop Pattern:
    // In include/microkernels/GemmUKernel_opti_rrr.hpp (inside template <typename T> class GemmUKernel):
    template <int lmul = 1, bool FastPath = false>
    static inline void
    gemm_mippv2_myukernel_core(const PackedRowMajor<T> &__restrict A,
                               const PackedRowMajor<T> &__restrict B,
                               PackedRowMajor<T> &__restrict C,
                               T alpha = T{1}, T beta = T{0}) {
      using namespace mipp;
    
      constexpr size_t VL = N<T, lmul>();
    
      constexpr size_t MR = 2;
      constexpr size_t NR = 2 * VL;
    
      for (size_t i = 0; i < A.rows; i += MR) {
        for (size_t j = 0; j < B.cols; j += NR) {
          /*
           * 2x2VL register tile
           *
           * c00 c01
           * c10 c11
           */
    
          auto c00 = set0<T, lmul>();
          auto c01 = set0<T, lmul>();
    
          auto c10 = set0<T, lmul>();
          auto c11 = set0<T, lmul>();
    
          const auto a0_ptr = A[i + 0];
          const auto a1_ptr = A[i + 1];
    
          for (size_t k = 0; k < A.cols; ++k) {
            const auto b0 = load<T, lmul>(B[k] + j);
            const auto b1 = load<T, lmul>(B[k] + j + VL);
    
            const auto a0 = set1<T, lmul>(a0_ptr[k]);
            const auto a1 = set1<T, lmul>(a1_ptr[k]);
    
            c00 = fmadd(a0, b0, c00);
            c01 = fmadd(a0, b1, c01);
    
            c10 = fmadd(a1, b0, c10);
            c11 = fmadd(a1, b1, c11);
          }
    
          const auto c0_ptr = C[i + 0] + j;
          const auto c1_ptr = C[i + 1] + j;
          Epilogue::store<T, lmul, FastPath>(c0_ptr, c00, alpha, beta);
          Epilogue::store<T, lmul, FastPath>(c0_ptr + VL, c01, alpha, beta);
    
          Epilogue::store<T, lmul, FastPath>(c1_ptr, c10, alpha, beta);
          Epilogue::store<T, lmul, FastPath>(c1_ptr + VL, c11, alpha, beta);
        }
      }
    }
    
    // Dispatch wrappers (inside GemmUKernel<T>):
    template <int lmul = 1>
    static inline void
    gemm_mippv2_myukernel(const PackedRowMajor<T> &__restrict A,
                          const PackedRowMajor<T> &__restrict B,
                          PackedRowMajor<T> &__restrict C) {
      gemm_mippv2_myukernel_core<lmul, true>(A, B, C);
    }
    
    template <int lmul = 1>
    static inline void
    gemm_mippv2_myukernel(const PackedRowMajor<T> &__restrict A,
                          const PackedRowMajor<T> &__restrict B,
                          PackedRowMajor<T> &__restrict C,
                          T alpha, T beta) {
      if (__builtin_expect(alpha == T{1} && beta == T{0}, 1)) {
        gemm_mippv2_myukernel_core<lmul, true>(A, B, C);
      } else {
        gemm_mippv2_myukernel_core<lmul, false>(A, B, C, alpha, beta);
      }
    }

Step 2: Register in include/microkernels/GemmUKernel.h

  1. Add the enum identifier in UKernelType:

    enum class UKernelType {
        // ...
        mippv2_myukernel,
    };
  2. Add a KernelDescriptor to kernelTable:

    {UKernelType::mippv2_myukernel, "mippv2_myukernel", RRR, 64, 4, 3},
    • Specify layout: RRR (Row-major $A, B, C$) or CRR (Col-major $A$, Row-major $B, C$).
    • MR and NR_vec define the register tile geometry for auto-sizing.
    • Specify minimal tile size: typically 64.

Step 3: Wire in src/main.cpp

Add the kernel dispatch branch in the switch (cfg.kernel) statement in src/main.cpp:

case UKernelType::mippv2_myukernel:
  BENCH_KERNEL_FASTPATH(
      gemm.gemm_mippv2_myukernel(A, B, C),
      gemm.gemm_mippv2_myukernel(A, B, C, cfg.alpha, cfg.beta));

Step 4: Add Unit Tests in tests/GemmUKernelTest.cpp

Add validation test cases for $C = A \times B$ in testGemmSuite() and for $C = \alpha (A \times B) + \beta C$ in testGemmSuiteAlphaBeta():

// In testGemmSuite():
TEST_KERNEL("MIPPv2 MyUKernel",
            gemm.gemm_mippv2_myukernel(A, B, C_test));

// In testGemmSuiteAlphaBeta():
SECTION("mippv2_myukernel") {
  resetMatrix(C_test, C_init);
  gemm.gemm_mippv2_myukernel(A, B, C_test, p.alpha, p.beta);
  checkMatrixEqual(C_ref, C_test);
}

Run ctest --test-dir build --output-on-failure to verify correctness and numerical precision.

Step 5: Shell Driver Integration (run_benchmarks.sh)

In run_benchmarks.sh, add the kernel name to UNIVERSAL_KERNELS:

    mippv2_myukernel

Step 6: Benchmark Execution

Execute the benchmark specifying the output prefix and kernel:

./run_benchmarks.sh avx2 -o myarch_avx2 -k mippv2_myukernel gcc clang

Step 7: Configure in uarch_config.json & Plot

In uarch_config.json, add or update the target entry under its SIMD family (e.g., "avx2"):

"avx2": {
  "myarch": {
    "name": "My Architecture",
    "cpu_model": "Custom Core",
    "frequency_ghz": 3.5,
    "peak_flop_per_cycle": 16.0,
    "expected_best_kernel": "mippv2_myukernel",
    "csv_pattern": "*gemm_results_*myarch*_{compiler}.csv"
  }
}

Run python3 plot_results.py --input-dir gemm-bench-results/rrr to generate updated cross-architecture graphs.


Current Focus & Roadmap (Prioritized TODO):

  1. Comparison against production BLAS (BLIS & OpenBLAS)*: Benchmark MIPPv2 microkernels against real BLAS libraries across all platforms. To do so the repo will need to be extended to a proper GEMM benchmark suite. And not just implement microkernels.
  2. Single-Precision Support (SGEMM / FP32): Extend the benchmark suite to evaluate FP32 as well.
  3. Comparison with Hand-Written Native Intrinsics: Implement intrinsics baselines (native RVV riscv_vector.h, AVX-512 immintrin.h, NEON arm_neon.h) to measure the exact abstraction overhead introduced by MIPPv2.
  4. Cross-SIMD Abstraction Comparisons: Benchmark against alternative SIMD abstraction frameworks, like Google Highway, EVE, and std::simd.
  5. Non-Canonical GEMM Evaluation: Benchmark throughput and epilogue overhead for non-canonical forms ($\beta \neq 0.0$, $\alpha \neq 1.0$) to quantify the cost of $C$ tile reloads and scaling.

Note

BLIS/OpenBLAS for RVV: Ongoing vendor and community work on RVV-optimized BLAS kernels reportedly outperforms current upstream public releases of BLIS and OpenBLAS, but these patches are not yet fully merged/streamlined in mainline distributions. Benchmarking methodology should evaluate both upstream releases and patched branches.

All of this is a lot of work. And I'm theoretically on vacations. I'm pretty sure I will do it eventually. But idk when.


Acknowledgments


Generative AI Disclosure

Generative AI assistants (Google Gemini and OpenAI ChatGPT) were utilized during the development of this repository to accelerate boilerplate tasks and infrastructure setup, specifically:

  • Mechanically generating C++ microkernel implementations across varying register tile geometries ($M_R \times N_R$) and accumulator layouts.
  • Automation scripts (shell runners and Python data parsing / plotting pipelines).
  • Benchmark driver CLI scaffolding, argument handling, and boilerplate wiring in C++.
  • Build system configuration and test suite plumbing.

Engineering Methodology & Kernel Design

The foundational DGEMM microkernel designs tuned for Intel Skylake (AVX2 + FMA) and the SpacemiT K3 X100 core are the result of multiple weeks of work and reflection on the underlying microarchitectures and platforms. And the best X100 microkernels are, I think, some of the best results of my internship.

For subsequent target microarchitectures, the methodology was empirical, the exploration workflow relied on pre-established invariants:

  • Layout Exploration: Both the $RRR$ format (Row-major $A, B, C$) and the $CRR$ format (Column-major $A$, Row-major $B, C$) are implemented. CRR uses the GotoBLAS / BLIS packing convention where $A$ is stored contiguously along $M_R$ with leading dimension $= M_R$, matching the inner kernel's access pattern. CRR microkernels achieve comparable or higher peak efficiency than RRR on most microarchitectures (up to 99.3% on M1).
  • Inner-Loop Access Patterns: Preserving either the scalar broadcast pattern (mipp::set1 + mipp::fmadd) or the indexed FMA pattern (mipp::fmaddi) on matrix $A$.

Under these, the exploration focused on sweeping register block dimensions ($M_R \times N_R$) that made sense given the uarch characteristics. Register file size, number of FMA pipelines/execution units and FMA latency being the three things that answer the questions :

  • How many accumulators can I have at most?
  • How many accumulators do I want?
  • How many accumulators do I need at least?

In practice, it often boiled down to telling Gemini to implement a bunch of tile sizes I thought might yield good results. Then I ran the benchmarks on the target hardware and iterated over the ones that performed well "Gemini please make gemm_mippv2_firestorm_mr8_nr2_fmaddi using gemm_mippv2_firestorm_mr6_nr4_fmaddi as reference" and so on. I feel like the most inefficient part of the automation was sometimes the agent sitting between the keyboard and the chair. Oh well.

Releases

Packages

Contributors

Languages