How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("image-text-to-text", model="NIyueeE/Qwen3.5-0.8B-cocreator")
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    },
]
pipe(text=messages)
# Load model directly
from transformers import AutoProcessor, AutoModelForMultimodalLM

processor = AutoProcessor.from_pretrained("NIyueeE/Qwen3.5-0.8B-cocreator")
model = AutoModelForMultimodalLM.from_pretrained("NIyueeE/Qwen3.5-0.8B-cocreator", device_map="auto")
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    },
]
inputs = processor.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=40)
print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:]))
Quick Links

Qwen3.5-0.8B-cocreator

Qwen3.5-0.8B fine-tuned on the CoCreator Driving Scene dataset for driving scene causal understanding.

Model Details

  • Base model: Qwen/Qwen3.5-0.8B
  • Dataset: NIyueeE/cocreator-driving-scene - 1,227 driving scene samples, each with multi-frame video and causal text descriptions
  • Fine-tuning method: QLoRA (4-bit) via Unsloth
  • Vision: Native multimodal (image+text)

Training

Platform

Google Colab (colab.research.google.com) with NVIDIA A100-SXM4-40GB.

Training Log

Unsloth 2026.5.5: Fast Qwen3_5 patching. Transformers: 5.5.0.
NVIDIA A100-SXM4-40GB. Num GPUs = 1. Max memory: 39.494 GB.
Torch: 2.10.0+cu128. CUDA: 8.0. CUDA Toolkit: 12.8. Triton: 3.6.0
Bfloat16 = TRUE. FA [Xformers = 0.0.35. FA2 = False]

Num examples = 1,227 | Num Epochs = 7 | Total steps = 50
Batch size per device = 128 | Gradient accumulation steps = 1
Total batch size (128 x 1 x 1) = 128
Trainable parameters = 13,181,952 of 866,167,872 (1.52% trained)

Loss Curve

Training loss

Training Script

See finetune_cocreator_coclab.ipynb for the complete fine-tuning notebook.

Hyperparameters

Parameter Value
LoRA r 16
LoRA alpha 16
LoRA dropout 0
Target modules all-linear
Fine-tuned layers vision + language + attention + MLP
Optimizer adamw_8bit
Learning rate 5e-5 (cosine schedule)
Max steps 50
Epochs 7
Gradient checkpointing unsloth
Resolution 800×450 (resized)

Usage

from transformers import AutoModel, AutoTokenizer
import torch

model = AutoModel.from_pretrained(
    "NIyueeE/Qwen3.5-0.8B-cocreator",
    torch_dtype=torch.bfloat16,
    trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained(
    "NIyueeE/Qwen3.5-0.8B-cocreator",
    trust_remote_code=True,
)

Intended Use

This model is fine-tuned for driving scene causal understanding. It takes multi-frame driving images as input and generates causal relationship text descriptions. The primary use case is as a feature extractor in the ReCogDrive autonomous driving VLA training pipeline.

License

Apache 2.0

Downloads last month
6
Safetensors
Model size
0.9B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train NIyueeE/Qwen3.5-0.8B-cocreator