Conversation
SME2 kernels for real single and double precision on cores with SME2 and a 512-bit streaming vector length
(checked at run time), built with clang for VORTEXM4. GCC builds keep the NEON kernels.
- sme_sgemm_kernel / sme_dgemm_kernel (the direct SME path of interface/gemm.c) run a blocked SME2 GEMM
after Deng et al., arXiv:2512.21473: analytical blocking, ZA transposition of A, online packing of B,
multi-vector micro-kernels, one thread per SME unit. Problems below 3000 multiply-adds run in plain loops
in the kernel. The existing SME kernels stay as the fallback.
- Level-3 driver kernels (gemm_kernel_sme2.c with the TRMM kernel, trsm_kernel_sme2.c) with GEMM_UNROLL_M 64
and GEMM_UNROLL_N one streaming vector, and A-panel copies for GEMM, SYMM, TRMM and TRSM, so that SYMM,
SYRK, SYR2K, TRMM, TRSM, GEMMT and the LAPACK routines on the driver run on SME2. TRSM updates each C block
and solves it in ZA. Without SME2 at run time the kernels run plain loops on the same layout.
- param.h: VORTEXM4 unroll and blocking (fp32 P = Q = 1024, fp64 P = Q = 512); cmake/prebuild.cmake the same
unroll for cross builds; kernel/generic laswp_ncopy_16.c and neg_tcopy_64.c for the new unroll values.
- kernel/CMakeLists.txt: builds without single precision use the TRSMCOPY*_M overrides of the KERNEL file.
- num_cpu_avail() caps levels 3 and 4 at one thread per SME unit (per performance cluster, hw.perflevel0)
for real precision on VORTEXM4; with a thread per core the kernels contend for the shared units.
- cpuid: M4 Pro and later Apple cores with FEAT_SME detect as VORTEXM4.
M4 Pro, GFLOPS at n = 4096, default threads, develop -> this change:
s: gemm 1241 -> 3275, symm 754 -> 2952, syrk 669 -> 2164, trmm 785 -> 2557, trsm 678 -> 2382,
potrf 346 -> 1554, getrf 364 -> 859
d: gemm 338 -> 888, symm 387 -> 709, syrk 328 -> 639, trmm 397 -> 697, trsm 402 -> 684,
potrf 170 -> 558, getrf 207 -> 467
Measured on M4 Pro only; M5 and later have a different SME and need their own measurements.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Supersedes #6072. That PR ran SYMM, SYRK, SYR2K, TRMM and TRSM on the SME GEMM kernel through hooks in
interface/{symm,syrk,syr2k,trsm}.cthat included a header fromkernel/arm64/, and it added an M4-tuned size threshold tointerface/gemm.c. This PR does the same work the usual OpenBLAS way: SME2 kernels for the level-3 driver, set inKERNEL.VORTEXM4andparam.h, with no changes to the interface files. The thread cap for this target is innum_cpu_avail()(common_thread.h).Changes (one commit)
sme_sgemm_kernel/sme_dgemm_kernelrun a blocked SME2 GEMM (after Deng et al., arXiv:2512.21473) on SME2 cores with a 512-bit streaming vector length, checked at run time. The existing SME kernels stay as the fallback. Problems below 3000 multiply-adds run in plain loops inside the kernel, sointerface/gemm.cis unchanged. cpuid detects the M4 Pro, and later Apple cores with FEAT_SME, as VORTEXM4.sme2_gemm_tile.h(ZA tile access and C block transfers) andsme2_gemm_detect.h(compiler and run-time checks) serve both the direct GEMM and the packed kernels.param.h:GEMM_UNROLL_M64 (one row of ZA tiles) andGEMM_UNROLL_Nof one streaming vector (16 fp32, 8 fp64).gemm_kernel_sme2.c: the GEMM kernel, and withTRMMKERNELthe TRMM kernel. C accumulates in ZA with one tile row for each column of C.trsm_kernel_sme2.c: LN, LT, RN and RT, with the panel walk of the SVE kernels. Each C block takes the update, then the triangular solve in ZA.kernel/generic/laswp_ncopy_16.candneg_tcopy_64.c: the make build needs these widths and cannot override them.With
BUILD_SINGLE=OFF,kernel/CMakeLists.txttook the A-side TRSM copies fromgeneric/*_${SGEMM_UNROLL_M}.ceven where the KERNEL file setsTRSMCOPY*_M. The main CMake loop and the make build already honor these overrides.M4 has one SME unit per performance cluster, and the cores of the cluster share it. With one thread per core (12 on M4 Pro, E-cores included), the threads contend for the units: SYMM ran slower with 12 threads than with 1, and fp64 ran slower than the NEON kernels of develop.
num_cpu_avail(level)now caps levels 3 (level-3 BLAS) and 4 (LAPACK) at the number of units (hw.perflevel0), for real precision on VORTEXM4 only. Thelevelargument was not used before. The interface files are unchanged. Complex, bfloat16 and half precision, levels 1 and 2, and all other targets are unchanged.Compared with #6072
What is different:
interface/{symm,syrk,syr2k,trsm}.c, which includedkernel/arm64/sme_level3.h. That header split each routine recursively into calls to the direct SME GEMM kernel. This PR changes no interface file. The routines go through the normal level-3 driver with SME2 kernels, so LAPACK routines on the driver (GETRF, POTRF, TRTRI, ...) run on SME2 too. In arm64 SME2: faster SGEMM/DGEMM, and SYMM/SYRK/SYR2K/TRMM/TRSM on the SME GEMM kernel #6072 they still used the NEON kernels, because its hooks were in the BLAS interface only.interface/gemm.c. It sent GEMMs below about 20^3 multiply-adds to the NEON driver, and it also applied to complex precision and to other SME cores. This PR puts a limit of 3000 multiply-adds (about 14^3) inside the real SME2 kernel instead, with plain loops below it.num_cpu_avail()instead.Where #6072 was faster (M4 Pro, GFLOPS, n = 256 / 1024 / 4096 unless noted):
Where this PR is faster:
SGEMM and DGEMM at n >= 32 use the same SME2 kernel in both PRs, so their speed is the same.
Files outside kernel/arm64
param.handcmake/prebuild.cmake: the VORTEXM4 entries only.kernel/generic/: two new files.kernel/CMakeLists.txt: theBUILD_SINGLE=OFFfix (item 5).common_thread.h: the thread cap (item 6). It is a no-op except for real precision on VORTEXM4 with clang on macOS.cpuid_arm64.c, README, CONTRIBUTORS.Testing (Apple M4 Pro, Apple clang 21, gfortran 16)
DYNAMIC_ARCH=1(the runtime selects vortexm4), CMake (ctest 14/14, and 5/5 withBUILD_SINGLE=OFF), and GCC (NEON kernels).xeigtstc < cec.inandxeigtstz < zec.instop with a gfortran input error ("Bad repeat count" at zchkee.F:1288) while they read their input, before any BLAS call.*_2.cfiles, and at width 64 against simple reference versions with random offsets.Performance (M4 Pro, GFLOPS, n = 4096, column major)
Default thread count (12):
One thread: SSYMM 1455, SSYRK 1289, STRMM 1290, STRSM 1159, DSYMM 372, DSYRK 348, DTRMM 357, DTRSM 347. develop is at 105-112 (fp32) and 52-56 (fp64) for these routines.
Notes
sme_level3_threads()incommon_thread.h.AI assistance
AI coding tools (Claude Code) helped to prepare this PR. The changes were reviewed, tested and benchmarked as described above.
No decompilation or reverse engineering was used. The only sources were the paper by Deng et al. (arXiv:2512.21473), MTGEMM-A (https://github.com/tesch1/mtgemm-a), the existing OpenBLAS kernels and drivers (the level-3 kernels follow the generic and SVE kernels), and Arm's public documentation of the SME instructions and intrinsics.