Skip to content

arm64: SME2 GEMM and level-3 kernels for Apple M4 (VORTEXM4) - #6074

Open
tesch1 wants to merge 1 commit into
OpenMathLib:developfrom
tesch1:sme2-kernels
Open

tesch1 wants to merge 1 commit into
OpenMathLib:developfrom
tesch1:sme2-kernels

Conversation

@tesch1

@tesch1 tesch1 commented Sep 30, 2026 •

Copy link
Copy Markdown

Supersedes #6072. That PR ran SYMM, SYRK, SYR2K, TRMM and TRSM on the SME GEMM kernel through hooks in interface/{symm,syrk,syr2k,trsm}.c that included a header from kernel/arm64/, and it added an M4-tuned size threshold to interface/gemm.c. This PR does the same work the usual OpenBLAS way: SME2 kernels for the level-3 driver, set in KERNEL.VORTEXM4 and param.h, with no changes to the interface files. The thread cap for this target is in num_cpu_avail() (common_thread.h).

Changes (one commit)

  1. SME2 GEMM kernel
    sme_sgemm_kernel / sme_dgemm_kernel run a blocked SME2 GEMM (after Deng et al., arXiv:2512.21473) on SME2 cores with a 512-bit streaming vector length, checked at run time. The existing SME kernels stay as the fallback. Problems below 3000 multiply-adds run in plain loops inside the kernel, so interface/gemm.c is unchanged. cpuid detects the M4 Pro, and later Apple cores with FEAT_SME, as VORTEXM4.
  2. Shared SME2 code: sme2_gemm_tile.h (ZA tile access and C block transfers) and sme2_gemm_detect.h (compiler and run-time checks) serve both the direct GEMM and the packed kernels.
  3. SME2 level-3 driver kernels
    • param.h: GEMM_UNROLL_M 64 (one row of ZA tiles) and GEMM_UNROLL_N of one streaming vector (16 fp32, 8 fp64).
    • gemm_kernel_sme2.c: the GEMM kernel, and with TRMMKERNEL the TRMM kernel. C accumulates in ZA with one tile row for each column of C.
    • trsm_kernel_sme2.c: LN, LT, RN and RT, with the panel walk of the SVE kernels. Each C block takes the update, then the triangular solve in ZA.
    • M-side copies (gemm, symm, trmm, trsm), with a dense tail panel as in the SVE kernels. The N side uses the generic copies, which the driver requires on that side.
    • kernel/generic/laswp_ncopy_16.c and neg_tcopy_64.c: the make build needs these widths and cannot override them.
    • Without SME2 at run time, or in the compiler, the kernels run plain loops on the same layout. GCC builds keep the NEON kernels.
  4. Blocking (measured): fp32 P = Q = 1024, fp64 P = Q = 512.
  5. CMake fix for builds without single precision
    With BUILD_SINGLE=OFF, kernel/CMakeLists.txt took the A-side TRSM copies from generic/*_${SGEMM_UNROLL_M}.c even where the KERNEL file sets TRSMCOPY*_M. The main CMake loop and the make build already honor these overrides.
  6. One level-3 thread per SME unit
    M4 has one SME unit per performance cluster, and the cores of the cluster share it. With one thread per core (12 on M4 Pro, E-cores included), the threads contend for the units: SYMM ran slower with 12 threads than with 1, and fp64 ran slower than the NEON kernels of develop. num_cpu_avail(level) now caps levels 3 (level-3 BLAS) and 4 (LAPACK) at the number of units (hw.perflevel0), for real precision on VORTEXM4 only. The level argument was not used before. The interface files are unchanged. Complex, bfloat16 and half precision, levels 1 and 2, and all other targets are unchanged.

Compared with #6072

What is different:

Where #6072 was faster (M4 Pro, GFLOPS, n = 256 / 1024 / 4096 unless noted):

threads develop #6072 this PR
SGEMM, n = 8 / 16 / 24 1 2 / 13 / 34 20 / 68 / 89 9 / 37 / 56
SSYRK 12 170 / 436 / 669 895 / 2213 / 2488 735 / 1797 / 2164
SSYRK 1 82 / 106 / 111 833 / 1453 / 1404 618 / 1363 / 1289
STRMM 1 84 / 106 / 112 523 / 1205 / 1266 295 / 938 / 1290
DSYMM 12 146 / 239 / 387 408 / 833 / 843 431 / 718 / 709
DSYMM 1 52 / 56 / 56 337 / 414 / 421 294 / 363 / 372
DSYRK 12 97 / 184 / 328 331 / 794 / 799 242 / 548 / 639
DSYRK 1 45 / 53 / 56 289 / 413 / 409 227 / 363 / 348
DTRMM 12 150 / 205 / 397 241 / 667 / 778 250 / 593 / 697
DTRMM 1 45 / 55 / 56 230 / 376 / 402 136 / 304 / 357

Where this PR is faster:

threads develop #6072 this PR
SSYMM 12 235 / 460 / 754 942 / 2544 / 2812 1197 / 2794 / 2952
STRSM 12 191 / 348 / 678 231 / 800 / 1606 412 / 1501 / 2382
STRSM 1 62 / 95 / 108 229 / 653 / 1017 229 / 701 / 1159
STRMM 12 235 / 451 / 785 507 / 1830 / 2414 498 / 1796 / 2557
DTRSM 12 134 / 198 / 402 114 / 384 / 639 211 / 545 / 684
SPOTRF, SGETRF, DPOTRF, DGETRF, n = 4096 12 346, 364, 170, 207 NEON (not measured) 1554, 859, 558, 467

SGEMM and DGEMM at n >= 32 use the same SME2 kernel in both PRs, so their speed is the same.

Files outside kernel/arm64

  • param.h and cmake/prebuild.cmake: the VORTEXM4 entries only.
  • kernel/generic/: two new files.
  • kernel/CMakeLists.txt: the BUILD_SINGLE=OFF fix (item 5).
  • common_thread.h: the thread cap (item 6). It is a no-op except for real precision on VORTEXM4 with clang on macOS.
  • cpuid_arm64.c, README, CONTRIBUTORS.

Testing (Apple M4 Pro, Apple clang 21, gfortran 16)

  • Fortran BLAS tests, CBLAS tests and utest: all pass. This covers the static make build, DYNAMIC_ARCH=1 (the runtime selects vortexm4), CMake (ctest 14/14, and 5/5 with BUILD_SINGLE=OFF), and GCC (NEON kernels).
  • LAPACK testing: 5,438,659 tests, 0 numerical errors. xeigtstc < cec.in and xeigtstz < zec.in stop with a gfortran input error ("Bad repeat count" at zchkee.F:1288) while they read their input, before any BLAS call.
  • A randomized check of GEMM, SYMM, SYRK, SYR2K, TRMM and TRSM against naive references: all transpose, side and uplo cases, alpha/beta including beta = 0 with NaN in C, sizes 0 to 1500, 1, 2 and 12 threads. No failures.
  • The new copy routines were compared element by element with the generic ones: at width 2 against the generic *_2.c files, and at width 64 against simple reference versions with random offsets.

Performance (M4 Pro, GFLOPS, n = 4096, column major)

Default thread count (12):

develop this PR
SGEMM 1241 3275
SSYMM 754 2952
SSYRK 669 2164
STRMM 785 2557
STRSM 678 2382
SPOTRF 346 1554
SGETRF 364 859
DGEMM 338 888
DSYMM 387 709
DSYRK 328 639
DTRMM 397 697
DTRSM 402 684
DPOTRF 170 558
DGETRF 207 467

One thread: SSYMM 1455, SSYRK 1289, STRMM 1290, STRSM 1159, DSYMM 372, DSYRK 348, DTRMM 357, DTRSM 347. develop is at 105-112 (fp32) and 52-56 (fp64) for these routines.

Notes

  • Measured on M4 Pro only. M5 and later have a different SME and need their own measurements; see the comment on sme_level3_threads() in common_thread.h.
  • Small problems that gained from more than one thread per SME unit lose some with the cap: SSYRK at n = 256 952 -> 735, SPOTRF at n = 1024 688 -> 580.

AI assistance

AI coding tools (Claude Code) helped to prepare this PR. The changes were reviewed, tested and benchmarked as described above.
No decompilation or reverse engineering was used. The only sources were the paper by Deng et al. (arXiv:2512.21473), MTGEMM-A (https://github.com/tesch1/mtgemm-a), the existing OpenBLAS kernels and drivers (the level-3 kernels follow the generic and SVE kernels), and Arm's public documentation of the SME instructions and intrinsics.

SME2 kernels for real single and double precision on cores with SME2 and a 512-bit streaming vector length
(checked at run time), built with clang for VORTEXM4. GCC builds keep the NEON kernels.

- sme_sgemm_kernel / sme_dgemm_kernel (the direct SME path of interface/gemm.c) run a blocked SME2 GEMM
  after Deng et al., arXiv:2512.21473: analytical blocking, ZA transposition of A, online packing of B,
  multi-vector micro-kernels, one thread per SME unit. Problems below 3000 multiply-adds run in plain loops
  in the kernel. The existing SME kernels stay as the fallback.
- Level-3 driver kernels (gemm_kernel_sme2.c with the TRMM kernel, trsm_kernel_sme2.c) with GEMM_UNROLL_M 64
  and GEMM_UNROLL_N one streaming vector, and A-panel copies for GEMM, SYMM, TRMM and TRSM, so that SYMM,
  SYRK, SYR2K, TRMM, TRSM, GEMMT and the LAPACK routines on the driver run on SME2. TRSM updates each C block
  and solves it in ZA. Without SME2 at run time the kernels run plain loops on the same layout.
- param.h: VORTEXM4 unroll and blocking (fp32 P = Q = 1024, fp64 P = Q = 512); cmake/prebuild.cmake the same
  unroll for cross builds; kernel/generic laswp_ncopy_16.c and neg_tcopy_64.c for the new unroll values.
- kernel/CMakeLists.txt: builds without single precision use the TRSMCOPY*_M overrides of the KERNEL file.
- num_cpu_avail() caps levels 3 and 4 at one thread per SME unit (per performance cluster, hw.perflevel0)
  for real precision on VORTEXM4; with a thread per core the kernels contend for the shared units.
- cpuid: M4 Pro and later Apple cores with FEAT_SME detect as VORTEXM4.

M4 Pro, GFLOPS at n = 4096, default threads, develop -> this change:
  s: gemm 1241 -> 3275, symm 754 -> 2952, syrk 669 -> 2164, trmm 785 -> 2557, trsm 678 -> 2382,
     potrf 346 -> 1554, getrf 364 -> 859
  d: gemm 338 -> 888, symm 387 -> 709, syrk 328 -> 639, trmm 397 -> 697, trsm 402 -> 684,
     potrf 170 -> 558, getrf 207 -> 467

Measured on M4 Pro only; M5 and later have a different SME and need their own measurements.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant