Model
https://huggingface.co/ibm-granite/granite-docling-258M
Architecture
Reported architecture is Idefics3ForConditionalGeneration (model_type: idefics3), composed of:
- Vision encoder:
idefics3_vision — a SigLIP2-style ViT (512px images, 16px patches, 768 hidden size, 12 layers/heads).
- Connector: a pixel-shuffle downsampler (
scale_factor: 4) followed by a linear projection into the text hidden size.
- Text decoder: a small
llama-type decoder (576 hidden size, 30 layers).
- Checkpoint dtype is
bfloat16.
This is a multimodal (image+text) model — the vision tower and connector only run when image inputs are present; text-only prompts skip them.
Relation to existing Spyre support
- The vision encoder's patch embedding is
Conv2dLayer, the same vLLM layer class already given an OOT Spyre implementation for Pixtral (spyre_inference/custom_ops/conv.py).
- The vision encoder's self-attention is bidirectional with no KV cache, the same shape of attention already supported for pooling/embedding encoder models (BERT, RoBERTa, BGE, granite-embedding) and for Gemma-4's vision tower.
- The pixel-shuffle connector's reshape/permute sequence is architecturally analogous to Pixtral's patch-merger, which needed a Spyre-specific fix for its own reshape/permute geometry.
- The text decoder is a standard dense Llama-family architecture, the same family already exercised on Spyre via existing Llama/Granite text checkpoints.
No implementation approach is prescribed here — filing this to track enabling the model.
Model
https://huggingface.co/ibm-granite/granite-docling-258M
Architecture
Reported architecture is
Idefics3ForConditionalGeneration(model_type: idefics3), composed of:idefics3_vision— a SigLIP2-style ViT (512px images, 16px patches, 768 hidden size, 12 layers/heads).scale_factor: 4) followed by a linear projection into the text hidden size.llama-type decoder (576 hidden size, 30 layers).bfloat16.This is a multimodal (image+text) model — the vision tower and connector only run when image inputs are present; text-only prompts skip them.
Relation to existing Spyre support
Conv2dLayer, the same vLLM layer class already given an OOT Spyre implementation for Pixtral (spyre_inference/custom_ops/conv.py).No implementation approach is prescribed here — filing this to track enabling the model.