Molmo2-4B Pairwise Judge (Direct, merged)

Standalone BF16 Molmo2-4B for pairwise judging of AI-generated videos. Given a user's preference questionnaire, a generation prompt, and two candidate videos, it emits exactly one token: 1, 2, or Tie.

Merged from LoRA vlm_preference_judge/outputs/molmo2_pairwise_judge/checkpoint-3450 (rank 16, alpha 32, dropout 0.05; targets att_proj, attn_out, ff_proj, ff_out) into the local Molmo2-4B base. Direct-answer format (no rating CoT).

Loading

import torch
from transformers import AutoModelForImageTextToText, AutoProcessor

model_id = "theblackcat102/Molmo2-4B-Pairwise-Judge-Direct"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="auto",
)

Use the Molmo2 chat-template with type="video" content. Constrain or post-process the answer to 1, 2, or Tie.

Prompt format

user: judge instructions + questionnaire + Generation task: t2v/i2v. Prompt used to generate both videos: "..." + Video 1: + video + Video 2: + video + Which video would this person prefer: 1, 2, or Tie? Answer with exactly one token. assistant: 1 | 2 | Tie

Downloads last month
30
Safetensors
Model size
5B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for theblackcat102/Molmo2-4B-Pairwise-Judge-Direct

Finetuned
(14)
this model