Skip to content

et-backend: F32 vecdot GEMV + matrix-engine GEMM - #29

Draft
RehanQasim-dev wants to merge 5 commits into
aifoundry-org:etfrom
RehanQasim-dev:upstream-f32-gemv-gemm
Draft

et-backend: F32 vecdot GEMV + matrix-engine GEMM#29
RehanQasim-dev wants to merge 5 commits into
aifoundry-org:etfrom
RehanQasim-dev:upstream-f32-gemv-gemm

Conversation

@RehanQasim-dev

@RehanQasim-dev RehanQasim-dev commented Jul 23, 2026

Copy link
Copy Markdown

Overview

Improves the ET backend's F32 MUL_MAT path for both decode (GEMV) and prefill (GEMM):

GEMV (decode, N <= 2): Stripes output elements across every hart of all 32 shires instead of blocking work into 16-element chunks, which only filled 8 shires for a typical decode GEMV. Adds a register-resident f32 row-dot helper with L2 prefetch for the streamed weight row, and stages the reused B activation vector into per-shire L2 SCP. Also fixes the matrix-engine GEMV dispatch check, which compared src1->ne[0] instead of src1->ne[1] and so never actually caught the decode case, sending it to the matrix engine kernel where it stalled on a single/near-single-column matmul.

GEMM (prefill, N > 2): Adds a double-buffered producer/consumer F32 matrix-engine kernel — hart 1 transposes weights into double-buffered L2 SCP while hart 0 runs tensor-engine compute — with a weight-reuse path and software prefetch for both weights and activations.

Dispatch now routes N <= 2 to the vecdot GEMV kernel and N > 2 to this matrix-engine GEMM kernel.

Additional information

Performance (Llama-3.2-1B-Instruct F32, ET-SoC-1):

Prefill t/s

N et optimized speedup
100 210.11 185.72 0.88x
220 299.64 344.39 1.15x
512 345.19 526.19 1.52x
700 320.69 429.39 1.34x
900 327.94 480.64 1.47x

Verified with llama-bench on ET-SoC-1 hardware (Llama-3.2-1B-Instruct F32), comparing this branch ("optimized") against unmodified et ("et") at the same prompt sizes used in the Q4_0/Q8_0 matrix-engine PRs.

Note: N=100 is a regression (0.88x), reproduced twice (4 repetitions each). The new weight-reuse matrix-engine kernel appears to have per-call setup overhead that doesn't amortize until N is large enough — gains are consistent and substantial from N=220 upward, but the smallest prefill size is currently worse than baseline et. Flagging for review before merge; may need a higher matrix-engine dispatch threshold for F32 than the current N > 2, or further tuning of the small-N case in the kernel itself.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES - the optimization strategies and design decisions are my own. AI assisted with understanding the hardware reference manual, some pieces of code implementation and guided debugging. I have thoroughly reviewed the code.

RehanQasim-dev and others added 5 commits July 30, 2026 09:22
Adds the L2-SCP hart-to-hart counter primitives (already used
locally by the stock mul_mat_f16_matrix_engine.c and
mul_mat_Q4_0_matrix_engine.c kernels) to the shared platform.h so
new kernels can use them too, and switches both stock kernels to
the shared copy instead of their own private ones to avoid a
duplicate-definition conflict.
Stripe output elements across every hart of all 32 shires instead of
blocking work into 16-element chunks, which only filled 8 shires for a
typical decode GEMV. Adds a register-resident f32 row-dot helper with
L2 prefetch for the streamed weight row, and stages the reused B
activation vector into per-shire L2 SCP.

Also fixes the matrix-engine GEMV dispatch check, which compared
src1->ne[0] instead of src1->ne[1] and so never actually caught the
n=1 decode case, sending it to the matrix engine kernel where it
stalls on a single-column matmul.

Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
(cherry picked from commit b62eede)
Adds a double-buffered producer/consumer F32 matrix-engine kernel
(hart 1 transposes weights into double-buffered L2 SCP while hart 0
runs tensor-engine compute) with a weight-reuse path and software
prefetch for both weights and activations, and wires MUL_MAT dispatch
so N <= 2 uses the vecdot GEMV kernel and N > 2 uses this
matrix-engine GEMM kernel.

Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
(cherry picked from commit 637d9a9)
Kernel source (mul_mat_f32.c, mul_mat_f32_matrix_engine.c) confirmed
byte-identical to the et-perf-debug branch, so no kernel porting needed here —
only the dispatch threshold was stale. Re-measured independently on this branch
(no graph-profiler/perf-counter instrumentation here) to confirm the crossover.
Llama-3.2-1B-Instruct-f32.gguf, ET_DEVICES=0, llama-bench -n 0:

N | vec_dot | matrix-engine (forced) | final (threshold=7):
 2 |  9.46 |  5.25 |  9.46
 3 | 11.61 |  7.66 | 11.61
 4 | 13.11 | 10.01 | 13.11
 5 | 14.39 | 12.66 | 14.39
 6 | 15.11-15.15 (r=5) | 14.98-15.10 (r=5) | 15.11
 7 | 15.71-15.75 (r=5) | 17.64-17.73 (r=5) | 17.74
 8 | 16.18 | 20.79 | 20.39
 9 | 16.53 | 23.02 | 22.90
10 | 16.90 | 25.26 | 25.65
20 |   n/a |   n/a | 42.93
30 |   n/a |   n/a | 63.02
50 |   n/a |   n/a | 100.36
70 |   n/a |   n/a | 134.90
100 |   n/a |   n/a | 182.84
120 |   n/a |   n/a | 215.30

Result: N=7, identical to the value measured on et-perf-debug. Final column
tracks the better of the two forced curves. Correctness verified via
test-backend-ops (202/202) and llama-cli coherence.

(cherry picked from commit c1b6ac3e9209960483f3702700cbc54a0229691f)
The comment said N <= 2 but the dispatch threshold above it was
tuned to N >= 7 in the previous commit, so this fallback now also
covers N=3..6.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant