Skip to content

Enable ibm-granite/granite-docling-258M #1037

Description

@gkumbhat

Model

https://huggingface.co/ibm-granite/granite-docling-258M

Architecture

Reported architecture is Idefics3ForConditionalGeneration (model_type: idefics3), composed of:

  • Vision encoder: idefics3_vision — a SigLIP2-style ViT (512px images, 16px patches, 768 hidden size, 12 layers/heads).
  • Connector: a pixel-shuffle downsampler (scale_factor: 4) followed by a linear projection into the text hidden size.
  • Text decoder: a small llama-type decoder (576 hidden size, 30 layers).
  • Checkpoint dtype is bfloat16.

This is a multimodal (image+text) model — the vision tower and connector only run when image inputs are present; text-only prompts skip them.

Relation to existing Spyre support

  • The vision encoder's patch embedding is Conv2dLayer, the same vLLM layer class already given an OOT Spyre implementation for Pixtral (spyre_inference/custom_ops/conv.py).
  • The vision encoder's self-attention is bidirectional with no KV cache, the same shape of attention already supported for pooling/embedding encoder models (BERT, RoBERTa, BGE, granite-embedding) and for Gemma-4's vision tower.
  • The pixel-shuffle connector's reshape/permute sequence is architecturally analogous to Pixtral's patch-merger, which needed a Spyre-specific fix for its own reshape/permute geometry.
  • The text decoder is a standard dense Llama-family architecture, the same family already exercised on Spyre via existing Llama/Granite text checkpoints.

No implementation approach is prescribed here — filing this to track enabling the model.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions