Instructions to use bojianzhou/mer2026-track3-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use bojianzhou/mer2026-track3-lora with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
MER2026 Track 3 (MER-Prefer) โ LoRA adapters
LoRA adapters for our 5th-place system on MER2026 Track 3, plus the full single-model zoo behind the paper's negative results.
- ๐ Paper: Parameter-Efficient Post-Training for Multimodal Emotional Preference Prediction. MRAC 2026 @ ACM Multimedia.
- ๐ป Code & per-sample scores: (https://github.com/zhoubojian-stevenchow/mer2026-track3-reproduction)
- ๐ Official result: 91.8240 (Stage 1) / 69.6388 (Stage 2), average 80.7314, 5th of 21
Every number in the paper is verifiable without these weights โ the
per-sample scores are in the GitHub repo and verify_paper.py recomputes 34
checks on CPU. The adapters are here for anyone who wants to regenerate those
scores from video, or build on the models.
Adapters
| subfolder | checkpoint | role | Holdout-99 | Test |
|---|---|---|---|---|
v3 |
350 | base recipe, no subtitles | 76.75 | 89.72 |
bt |
240 | BradleyโTerry pairwise objective | 82.81 | 90.25 |
MT590 |
590 | individual-emotion multi-task, ฯ=0.5 | 89.90 | 89.46 |
r70 |
620 | individual-emotion multi-task, ฯ=0.7 | 85.83 | 89.19 |
all3 |
1170 | three-task mixture | 88.90 | 90.51 |
s220 |
220 | seed variant, subtitle recipe | 78.77 | 88.66 |
s480 |
480 | +474 majority-voted pairs | 86.86 | 89.46 |
vlora |
270 | +vision-tower LoRA | 82.81 | 89.46 |
cons |
230 | swap-consistency regularizer | 80.79 | 89.18 |
synth |
920 | +synthetic preference pairs | 82.81 | 86.82 |
vl |
390 | cross-backbone probe (Qwen2.5-VL base) | 81.74 | 86.82 |
vl adapts Qwen2.5-VL, not Qwen2.5-Omni. Every other adapter is over
Qwen/Qwen2.5-Omni-7B (Thinker), LoRA r=16 ฮฑ=32 on the language layers only
โ 40.4 M trainable parameters, 0.45 % of the backbone. The vision and audio
towers are frozen, which the paper argues is why input-side diversification
fails to decorrelate members.
Checkpoint indices are only meaningful at the training length in the paper
(mixed dataset 13,573 examples, 3 epochs โ 2,547 optimizer steps at accumulation
16). all3 in particular must be checkpoint 1170: that is the checkpoint
that produced the Stage-1 scores the slot-averaging result rests on. A different
checkpoint is a different model and the gains do not transfer.
The submitted system
Both official submissions use the same recipe โ four slots at equal weight, with two models averaged into one vote in the third slot:
SLOTS = [['v3'], ['bt'], ['MT590', 'all3'], ['r70']]
all3 as a separate fifth vote contributes exactly +0.0000, because it
near-duplicates MT590 (soft-score correlation 0.969). Folded into MT590's
vote it lowers that vote's variance without adding a decision direction, and
gains on both boards: 91.5612 โ 91.8240 (Stage 1), 69.6208 โ 69.6388
(Stage 2). This is score-space averaging, not weight-space model souping.
Usage
from transformers import Qwen2_5OmniThinkerForConditionalGeneration
from peft import PeftModel
base = Qwen2_5OmniThinkerForConditionalGeneration.from_pretrained(
"Qwen/Qwen2.5-Omni-7B", torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(base, "<user>/<repo>", subfolder="MT590")
Scoring is a single forward pass, not generation: append the shared "a" prefix
so the next position is the decision token, take a binary softmax over the 1 /
2 token ids, then cancel position bias with a swapped second pass:
p_a1 = 0.5 * p_fwd + 0.5 * (1 - p_bwd)
Decision rule as submitted:
p_a1 > 0.5 -> a1
p_a1 < 0.5 -> a2
p_a1 = 0.5 -> a1 iff p_fwd >= 0.5
The tie branch is structural: when the forward and swapped passes agree exactly,
p_a1 is exactly 0.5 by construction. See code/01_inference_engine.py in the
GitHub repo for the exact implementation, including the MT590 exception (its
submission left ties at a2).
Training data
Official MER2026 data only: EmoPrefer-Data (574 pairs), EmoPrefer-Data-V2 (2,096), the official subtitle file, and the Track-1 individual-emotion set (9,395 clips). Not redistributed here โ obtain it from the MER2026 organizers.
Limitations
Single task, single backbone family, one competition. The decorrelation principle the paper argues for is an empirical regularity at this task and scale, with a proposed mechanism (frozen perception towers), not a theorem. See the Limitations section of the paper.
Citation
@inproceedings{zhou2026emopref,
title = {Parameter-Efficient Post-Training for Multimodal Emotional
Preference Prediction: A Top-5 System, the Limits of Improving
It, and a Holdout That Fails to Predict the Test},
author = {Zhou, Bojian},
booktitle = {Proceedings of the Workshop on Multimodal, Robust, and Affective
Computing (MRAC) at ACM Multimedia},
year = {2026}
}
Adapters inherit the licence of their base model (Qwen2.5-Omni / Qwen2.5-VL).
- Downloads last month
- -
Model tree for bojianzhou/mer2026-track3-lora
Base model
Qwen/Qwen2.5-Omni-7B