andreagemelli's picture
Fix generation from the published artifact; rewrite model card
2c4010c verified
|
Raw
History Blame Contribute Delete
5.31 kB
metadata
language:
  - en
  - fr
  - de
  - es
  - it
  - pl
license: apache-2.0
tags:
  - vision-language-model
  - multimodal
  - visual-question-answering
  - image-captioning
  - vlm
base_model:
  - OpenGVLab/InternViT-300M-448px-V2_5
  - PleIAs/Baguettotron

baguettotron-internvit-alignment

Baguettotron-VLM is an open, fully-reproducible, multilingual Vision-Language Model in the sub-1B parameter class. It extends PleIAs/Baguettotron β€” a 321M text-only reasoning SLM β€” with visual capabilities via the InternViT-300M-448px-V2.5 vision encoder and a lightweight MLP projector, for a total of ~628M parameters. It inherits six European languages (EN, FR, DE, ES, IT, PL) from the Baguettotron backbone.

Two checkpoints are published:

  • baguettotron-internvit-alignment β€” the projector-only warmup. It describes images, and nothing more.
  • baguettotron-vision-vqa β€” instruction-tuned on top of it. It goes past plain description and follows visual instructions, so prefer it for a richer chat experience.

Apache 2.0. Both live in the Baguettotron-VLM collection. Source: github.com/andreagemelli/baguettotron-vlm Β· Write-up: andreagemelli.me/posts/baguettotron-vlm

Alignment checkpoint. Only the MLP projector was trained (on LLaVA-CC3M-Pretrain-595K) β€” the ViT and the Baguettotron LLM are unmodified base weights. It writes short captions and nothing more. Published so the alignment stage can be reproduced; for actual use take baguettotron-vision-vqa.

Architecture

Image (448Γ—448)
  β†’ InternViT-300M-448px-V2.5 (304M, frozen) β†’ 1024 tokens Γ— 1024d
  β†’ Pixel unshuffle (factor=2)                β†’ 256 tokens Γ— 4096d
  β†’ MLP projector (2-layer, ~2.7M)            β†’ 256 tokens Γ— 576d
  β†’ Interleave with text tokens
  β†’ Baguettotron (321M, Llama arch, 80L, h=576)

Total: ~628M parameters

Usage

pip install "transformers>=4.56,<5" torch pillow timm einops accelerate
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
from PIL import Image

model = AutoModelForImageTextToText.from_pretrained(
    "andreagemelli/baguettotron-internvit-alignment",
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="auto",  # also tested on Apple Silicon (mps) and CPU
)
processor = AutoProcessor.from_pretrained(
    "andreagemelli/baguettotron-internvit-alignment",
    trust_remote_code=True,
)

image = Image.open("photo.jpg").convert("RGB")
inputs = processor(
    messages=[{"role": "user", "content": "<image>\nDescribe the image concisely."}],
    image=image,
)
inputs = {k: v.to(model.device) for k, v in inputs.items() if v is not None}

print(model.chat(**inputs))

Example output

Greedy, prompt Describe the image concisely. β€” verbatim output. Images from COCO val2017.

output
a cat is sleeping on the couch
the bear is a good friend.
a sign for a stop
the bus is a red double - decoration

Chat template

Trained on short image captions with no <think> traces. The processor emits a bare assistant prefix (<|im_start|>assistant\n) and the model completes the caption directly. Keep prompts simple ("Describe the image").

Limitations

  • Resolution ceiling. One 448Γ—448 crop β†’ 256 visual tokens puts document text at roughly 2–4 px/char. OCR, charts and documents are out of reach by architecture, not by budget. Neither published checkpoint reads text in an image.
  • Hallucinations, especially on fine-grained or text-heavy questions.
  • Multilingual capability is inherited, not verified. The backbone covers six languages; the VLM was never evaluated on non-English benchmarks.

Tested against transformers 4.57. Newer major versions may need adjustments.

Contributions and suggestions are very welcome β€” issues, PRs, and ideas for better data mixes, training recipes, or evaluation setups are all appreciated. Open an issue or PR on the GitHub repo.

Citation

If you use or extend Baguettotron-VLM in your research, please cite it:

@misc{gemelli2026baguettotronvlm,
  title        = {Baguettotron-VLM: An Open, Reproducible, Multilingual Sub-1B Vision-Language Model},
  author       = {Gemelli, Andrea},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/andreagemelli/baguettotron-vision-vqa}},
  note         = {Code: \url{https://github.com/andreagemelli/baguettotron-vlm}}
}

License

Apache 2.0 β€” see the GitHub repo.