Mizar-3B: Audio Understanding with RT-OPD

Paper · Mizar family · GitHub code · Mizar-159M · Mizar-3B

Mizar-3B is the Ke-based 3B member of the Mizar audio-language model family. This repository contains its LoRA adapter from the main experiment in Reward-Tilted On-Policy Distillation for Acoustic Grounding in Audio-Language Models. It needs the pinned KE-Team/Ke-Omni-R-3B base. It is not a standalone base model. The model accepts audio and a multiple-choice question and generates text.

Mizar model family

Model Model weights Code Foundation / training
Mizar-159M KaiyangLi/Mizar-159M Mizar_159M CED-Small + SmolLM2-135M; three-stage audio-language training
Mizar-3B KaiyangLi/Mizar-3B RT-OPD Ke-Omni-R-3B + RT-OPD; released as a LoRA adapter

These models share the Mizar family name and audio-understanding focus. They use different foundations and training recipes. Mizar-3B is the name of the released Ke-based model; RT-OPD is its distillation method. The Qwen profile in this repository remains a separate RT-OPD experiment.

Which checkpoint is this?

Seed 85, fixed final step 626, selected post hoc by the highest Macro-3 among the five main-experiment seeds (82–86). This release choice does not replace the paper's five-seed mean and is not an independent test-set estimate.

Result MMAU full (9,000) MMAR (1,000) ADQA-cl (1,577) Macro-3
This seed 85 adapter 72.7778 61.1000 57.0704 63.6494
Paper: five-seed mean 72.7222 60.0800 56.4490 63.0837 ± 0.3592

All numbers are percentages; ± is sample standard deviation of per-seed Macro-3. Seed 86 has the highest MMAU alone (73.1222%), but seed 85 has the highest Macro-3. MMAU uses the official hidden-label scoring service; mini is a separate split. main_results.json records all ten Ke/Qwen runs and evidence hashes.

Evaluation inputs and scoring code

Loading and running

Use the matching code and pinned Python 3.10 environment from the RT-OPD release.

git clone https://github.com/KaiyangLi1992/RT-OPD.git
cd RT-OPD
bash ke/environment/setup.sh
ke/.venv/bin/hf auth login
ke/.venv/bin/python tools/infer.py \
  --adapter KaiyangLi/Mizar-3B \
  --audio /absolute/path/example.wav \
  --question "Which sound is audible?" \
  --choices "A dog barking" "A piano playing"

The helper downloads the pinned base, loads its Thinker component with Transformers 4.52.4, attaches this adapter with PEFT 0.19.1, and runs greedy text generation. The teacher is unnecessary at inference. On GPUs without native BF16 support, --dtype float16 is a convenience option; this is not the paper's frozen vLLM benchmark pipeline. Follow docs/EVALUATION.md for paper evaluation. For immutable runs, pass the HF commit ID to --revision.

Training recipe

Frozen Ke-Omni-R 7B teacher; fresh Ke-Omni-R-3B student; 10,000 fixed training rows; two epochs / 626 steps; global batch 32 (4 GPUs × microbatch 4 × accumulation 2); learning rate 7.5e-5; LoRA rank 64, alpha 128, dropout 0.05; CE + 0.25 × gated reverse KL; target alpha 1.0; student rollouts at temperature 1.0, top-p 0.95, top-k 64, maximum 96 new tokens. The teacher scores the same student prefix with and without audio; there is no separate reference model or donor audio in this main method.

Files and provenance

  • adapter_model.safetensors: original, unmodified trained adapter bytes.
  • adapter_config.json: original PEFT settings with the base path replaced by its public model ID and pinned revision for portability.
  • base_model.json: exact base revision and per-file SHA-256 values.
  • CHECKPOINT_PROVENANCE.json: original and published identities, selection policy, and original configuration hash. Optimizer/RNG state is not needed for inference and is not part of this inference release.
  • SHA256SUMS.json: checksums of published payload files.

The model is for research on audio understanding. Dataset and base-model terms remain applicable; see the source repositories linked in the code's data guide. This package does not include hidden MMAU gold labels.

Frozen data manifests

reproducibility/frozen_data.tar.gz contains the exact 10,000-row training manifest, frozen teacher gate, valid vocabulary, benchmark manifests and audio hashes shared by Ke and Qwen. The code automatically downloads this archive at a pinned revision before preparing models/audio. The archive contains no audio or model weights and no hidden MMAU test labels. Upstream dataset terms apply.

License

The adapter weights and the files authored for this release are licensed under the BSD 3-Clause Clear License. The adapter requires its base model, KE-Team/Ke-Omni-R-3B, which is fine-tuned from Qwen/Qwen2.5-Omni-3B; using the adapter with that base remains subject to the base models' terms, including the Qwen Research License. Dataset manifests are not relicensed and keep their upstream terms. See NOTICE.

Downloads last month
31
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KaiyangLi/Mizar-3B

Adapter
(1)
this model

Collection including KaiyangLi/Mizar-3B

Paper for KaiyangLi/Mizar-3B