granite-docling-258M runs at about 1 token per second through transformers on Apple Silicon
Symptom#
ibm-granite/granite-docling-258M loaded with AutoModelForImageTextToText (transformers 5.17, torch 2.14) and generate() takes 467 s for one PDF page on MPS (641 tokens, about 1 token/s). The same run on the CPU had not finished after 10 minutes. A 258M-parameter model should be far faster.
Cause#
Not found. The transformers path for this Idefics3-based model is very slow on this machine on both MPS (bfloat16) and CPU (float32).
Fix#
Use the MLX build with mlx-vlm:
, =
=
= .
The page image was rendered with pdftoppm -r 144 -png. The Image processor also needs torchvision installed for the transformers path, or it fails with Missing optional dependencies: torchvision.
Evidence#
rait, 2026-09-26, Apple M4 Pro, mlx-vlm 0.7.3: the page that took 467 s on MPS took 7.8 s with MLX, and its DocTags held the same table cells and caption. 245 pages averaged 7.4 s per page.