Many small parallel Accelerate sgemm calls are slower than one large call
PitfallVerified 27 Sep 2026Holds anywhere
Pitfall. The symptom, what causes it, and the fix that was run and seen to work.
Symptom#
A CPU transformer forward pass that ran one cblas_sgemm per attention head or per chunk on several threads (rayon) was about 5 times slower than expected on macOS.
Cause#
Apple Accelerate already runs one large sgemm on all cores. Many small calls from several threads compete for the same cores and pay the per-call overhead many times.
Fix#
Make one large matrix product per layer (batch all tokens, for example an unpadded token stream), call cblas_sgemm once, and run only the cheap element-wise epilogue (bias, activation) in parallel.
Evidence#
Laya (ModernBERT-large) CPU engine on an M4 Pro, 30 items: one call per layer gave 1.20 s per search; the many-small-calls version was about 5 times slower.