Ray's Knowledge Base

granite-docling-258M runs at about 1 token per second through transformers on Apple Silicon

PitfallVerified 28 Sep 2026Holds os: macos
Pitfall. The symptom, what causes it, and the fix that was run and seen to work.

Symptom#

ibm-granite/granite-docling-258M loaded with AutoModelForImageTextToText (transformers 5.17, torch 2.14) and generate() takes 467 s for one PDF page on MPS (641 tokens, about 1 token/s). The same run on the CPU had not finished after 10 minutes. A 258M-parameter model should be far faster.

Cause#

Not found. The transformers path for this Idefics3-based model is very slow on this machine on both MPS (bfloat16) and CPU (float32).

Fix#

Use the MLX build with mlx-vlm:

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
model, proc = load("ibm-granite/granite-docling-258M-mlx")
prompt = apply_chat_template(proc, model.config, "Convert this page to docling.", num_images=1)
text = generate(model, proc, prompt, ["page.png"], max_tokens=8192, temperature=0.0, verbose=False).text

The page image was rendered with pdftoppm -r 144 -png. The Image processor also needs torchvision installed for the transformers path, or it fails with Missing optional dependencies: torchvision.

Evidence#

rait, 2026-09-26, Apple M4 Pro, mlx-vlm 0.7.3: the page that took 467 s on MPS took 7.8 s with MLX, and its DocTags held the same table cells and caption. 245 pages averaged 7.4 s per page.