You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
On Apple M4 (SME2, 512-bit streaming vector length), develop (63d7f22) leaves most of the SME unit unused:
SGEMM/DGEMM:sme_[sd]gemm_kernel ([WIP] Add ARMv9.2 SME GEMM kernels ported from vlovero's project #5971) runs at about 0.6x of Accelerate on one thread, and on one
thread whatever the thread count, while the M4 Pro has two SME units (one per performance cluster).
fp32 geometric means over the 24 workloads of Deng et al. (arXiv:2512.21473), row-major / column-major, GFLOPS:
Accelerate
develop
SGEMM, 1 thread
1127 / 1185
707 / 766
SGEMM, default threads
2364 / 2445
702 / 766
DGEMM, 1 thread
321 / 344
275 / 283
SYMM, SYRK, SYR2K, TRMM, TRSM run through the level-3 driver on NEON kernels: about 100-115 GFLOPS in fp32
and 53-58 in fp64 at n = 1024 on one thread, against 791-1758 and 265-433 for Accelerate. The SME kernel cannot
simply be plugged into the driver, because SYMM/TRMM share the GEMM unroll sizes (where [WIP] Arm®v9-A architecture SME2 SGEMM kernels #5011 stopped).
Detection: the M4 Pro (hw.cpufamily 0x17d5b93a) is not recognised by getarch, so a build without TARGET falls back to ARMV8 without any SME code.
On Apple M4 (SME2, 512-bit streaming vector length),
develop(63d7f22) leaves most of the SME unit unused:SGEMM/DGEMM:
sme_[sd]gemm_kernel([WIP] Add ARMv9.2 SME GEMM kernels ported from vlovero's project #5971) runs at about 0.6x of Accelerate on one thread, and on onethread whatever the thread count, while the M4 Pro has two SME units (one per performance cluster).
fp32 geometric means over the 24 workloads of Deng et al. (arXiv:2512.21473), row-major / column-major, GFLOPS:
SYMM, SYRK, SYR2K, TRMM, TRSM run through the level-3 driver on NEON kernels: about 100-115 GFLOPS in fp32
and 53-58 in fp64 at n = 1024 on one thread, against 791-1758 and 265-433 for Accelerate. The SME kernel cannot
simply be plugged into the driver, because SYMM/TRMM share the GEMM unroll sizes (where [WIP] Arm®v9-A architecture SME2 SGEMM kernels #5011 stopped).
Detection: the M4 Pro (
hw.cpufamily0x17d5b93a) is not recognised bygetarch, so a build withoutTARGETfalls back to ARMV8 without any SME code.Measured on an M4 Pro, macOS 27, Apple clang 21; harness and raw data in https://github.com/tesch1/mtgemm-a.
Prepared with AI coding assistance.