robotflowlabs 's Collections

ANIMA VLM

INT4 vision-language models for robotic scene understanding. Qwen2.5-VL for visual QA and grounding.