You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

MER2026 Track 3 (MER-Prefer) โ€” LoRA adapters

LoRA adapters for our 5th-place system on MER2026 Track 3, plus the full single-model zoo behind the paper's negative results.

Every number in the paper is verifiable without these weights โ€” the per-sample scores are in the GitHub repo and verify_paper.py recomputes 34 checks on CPU. The adapters are here for anyone who wants to regenerate those scores from video, or build on the models.

Adapters

subfolder checkpoint role Holdout-99 Test
v3 350 base recipe, no subtitles 76.75 89.72
bt 240 Bradleyโ€“Terry pairwise objective 82.81 90.25
MT590 590 individual-emotion multi-task, ฯ=0.5 89.90 89.46
r70 620 individual-emotion multi-task, ฯ=0.7 85.83 89.19
all3 1170 three-task mixture 88.90 90.51
s220 220 seed variant, subtitle recipe 78.77 88.66
s480 480 +474 majority-voted pairs 86.86 89.46
vlora 270 +vision-tower LoRA 82.81 89.46
cons 230 swap-consistency regularizer 80.79 89.18
synth 920 +synthetic preference pairs 82.81 86.82
vl 390 cross-backbone probe (Qwen2.5-VL base) 81.74 86.82

vl adapts Qwen2.5-VL, not Qwen2.5-Omni. Every other adapter is over Qwen/Qwen2.5-Omni-7B (Thinker), LoRA r=16 ฮฑ=32 on the language layers only โ€” 40.4 M trainable parameters, 0.45 % of the backbone. The vision and audio towers are frozen, which the paper argues is why input-side diversification fails to decorrelate members.

Checkpoint indices are only meaningful at the training length in the paper (mixed dataset 13,573 examples, 3 epochs โ‰ˆ 2,547 optimizer steps at accumulation 16). all3 in particular must be checkpoint 1170: that is the checkpoint that produced the Stage-1 scores the slot-averaging result rests on. A different checkpoint is a different model and the gains do not transfer.

The submitted system

Both official submissions use the same recipe โ€” four slots at equal weight, with two models averaged into one vote in the third slot:

SLOTS = [['v3'], ['bt'], ['MT590', 'all3'], ['r70']]

all3 as a separate fifth vote contributes exactly +0.0000, because it near-duplicates MT590 (soft-score correlation 0.969). Folded into MT590's vote it lowers that vote's variance without adding a decision direction, and gains on both boards: 91.5612 โ†’ 91.8240 (Stage 1), 69.6208 โ†’ 69.6388 (Stage 2). This is score-space averaging, not weight-space model souping.

Usage

from transformers import Qwen2_5OmniThinkerForConditionalGeneration
from peft import PeftModel

base = Qwen2_5OmniThinkerForConditionalGeneration.from_pretrained(
    "Qwen/Qwen2.5-Omni-7B", torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(base, "<user>/<repo>", subfolder="MT590")

Scoring is a single forward pass, not generation: append the shared "a" prefix so the next position is the decision token, take a binary softmax over the 1 / 2 token ids, then cancel position bias with a swapped second pass:

p_a1 = 0.5 * p_fwd + 0.5 * (1 - p_bwd)

Decision rule as submitted:

p_a1 > 0.5  -> a1
p_a1 < 0.5  -> a2
p_a1 = 0.5  -> a1 iff p_fwd >= 0.5

The tie branch is structural: when the forward and swapped passes agree exactly, p_a1 is exactly 0.5 by construction. See code/01_inference_engine.py in the GitHub repo for the exact implementation, including the MT590 exception (its submission left ties at a2).

Training data

Official MER2026 data only: EmoPrefer-Data (574 pairs), EmoPrefer-Data-V2 (2,096), the official subtitle file, and the Track-1 individual-emotion set (9,395 clips). Not redistributed here โ€” obtain it from the MER2026 organizers.

Limitations

Single task, single backbone family, one competition. The decorrelation principle the paper argues for is an empirical regularity at this task and scale, with a proposed mechanism (frozen perception towers), not a theorem. See the Limitations section of the paper.

Citation

@inproceedings{zhou2026emopref,
  title     = {Parameter-Efficient Post-Training for Multimodal Emotional
               Preference Prediction: A Top-5 System, the Limits of Improving
               It, and a Holdout That Fails to Predict the Test},
  author    = {Zhou, Bojian},
  booktitle = {Proceedings of the Workshop on Multimodal, Robust, and Affective
               Computing (MRAC) at ACM Multimedia},
  year      = {2026}
}

Adapters inherit the licence of their base model (Qwen2.5-Omni / Qwen2.5-VL).

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for bojianzhou/mer2026-track3-lora

Adapter
(59)
this model