Ray's Knowledge Base

candle compiles Metal kernels on the first GPU run of each binary

PitfallVerified 27 Sep 2026Holds anywhere
Pitfall. The symptom, what causes it, and the fix that was run and seen to work.

Symptom#

The first GPU (Metal) run of a candle model in a new binary takes about 10 s; later runs of the same binary take about 0.3 s. When the first call runs under a short timeout (2 s), it is killed every time and the GPU never gets fast.

Cause#

candle builds its Metal kernels from source at runtime, and macOS caches the compiled result per binary. A copy of the same binary at another path was slow again (10.00 s), while the original binary's third run took 0.27 s. A process killed before the compile finishes leaves no cache.

Fix#

Warm the GPU once with no time limit (for example in an install or fetch step: load the model and score one item), record that this binary is warm, and use the CPU engine until then. Rebuilding or moving the binary needs a new warm-up.

Evidence#

Laya reranker on an M4 Pro: first GPU run 9.61 s, second 0.27 s; a copied binary 10.00 s; after a warm-up step, rkb search used the GPU in 0.26 s.