candle compiles Metal kernels on the first GPU run of each binary
Symptom#
The first GPU (Metal) run of a candle model in a new binary takes about 10 s; later runs of the same binary take about 0.3 s. When the first call runs under a short timeout (2 s), it is killed every time and the GPU never gets fast.
Cause#
candle builds its Metal kernels from source at runtime, and macOS caches the compiled result per binary. A copy of the same binary at another path was slow again (10.00 s), while the original binary's third run took 0.27 s. A process killed before the compile finishes leaves no cache.
Fix#
Warm the GPU once with no time limit (for example in an install or fetch step: load the model and score one item), record that this binary is warm, and use the CPU engine until then. Rebuilding or moving the binary needs a new warm-up.
Evidence#
Laya reranker on an M4 Pro: first GPU run 9.61 s, second 0.27 s; a copied binary 10.00 s; after a warm-up step, rkb search used the GPU in 0.26 s.