--- license: other license_name: qwen-research license_link: LICENSE base_model: Qwen/Qwen2.5-Omni-3B base_model_relation: finetune library_name: transformers pipeline_tag: audio-text-to-text tags: - audio - audio-question-answering - audio-classification - candidate-scoring - qwen2_5_omni - safetensors --- # AudioJev **Direct audio decisions with order-calibrated candidate probabilities. Built with Qwen.** AudioJev takes an audio waveform, a natural-language question, and a list of candidate descriptions. It returns a probability distribution over those candidates using the next-token logits of their position labels. This release contains the **random-derangement SKL (RD-SKL) model with lambda = 0.5 and training seed = 20261001**, fine-tuned from [Qwen2.5-Omni-3B](https://huggingface.co/Qwen/Qwen2.5-Omni-3B). ## Checkpoint | Setting | Value | |---|---| | Model name | AudioJev | | Base model | Qwen/Qwen2.5-Omni-3B | | Training seed | 20261001 | | SKL weight | 0.5 | | Training schedule | Semantic: 1,024 updates; joint: 1,024; general: 2,048 | | Total optimizer updates | 4,096 | | Released checkpoint | Final general-stage step 2,048 | | Checkpoint selection | Fixed final step | | Training objective | Mean of two candidate cross-entropies + 0.5 × aligned mean symmetric KL | | Weight format | Original FP32 Safetensors shards, approximately 18.8 GB | | Recommended inference dtype | BF16 | | Candidate labels | `0`–`9`, then `A`–`Z` (2–36 candidates) | The audio encoder and language-model decision path were trained without a LoRA adapter. The visual branch was excluded from training; audio generation is disabled. This repository contains the complete saved Thinker checkpoint and its processor/tokenizer files. The weight shards are byte-identical to the evaluated checkpoint. For each training example, RD-SKL supervises the original candidate order and a random derangement. The second distribution is aligned by candidate identity before applying symmetric KL. Inference uses one forward pass in the supplied candidate order. ## Inference Python inference, the HTTP service, installation dependencies, and runnable examples are maintained in the standalone **AudioJev-Inference** project. This model repository distributes the model weights, tokenizer/processor assets, model card, and license files. After installing AudioJev-Inference, start the service with: ```bash audiojev-serve --model shlv/AudioJev --device cuda:0 ``` The Python API uses `AudioJev("shlv/AudioJev")`. Follow the AudioJev-Inference README for authentication, local/offline model loading, API requests, and deployment. The service downloads the model from this repository; it does not require a training-project checkout. AudioJev decisions use candidate-label logits at the final input position, normalized over the supplied candidates. The service documents its candidate-order convention separately from the original-order evaluation below. ## Evaluation The following scores are for **this single checkpoint in the original candidate order**, not the mean over three training seeds. | Benchmark | Questions | Accuracy | |---|---:|---:| | MMAU full (1,000 public + 9,000 officially scored hidden questions) | 10,000 | 69.41% | | MMAR | 1,000 | 56.30% | | MMSU evaluated short-audio subset | 3,931 | 62.45% | The experiment's source audit found 119 hidden MMAU questions with audio-source overlap against general-stage fitting/development data; the official full score retains all questions. That score should be interpreted with this overlap in mind. ## Scope AudioJev is a closed-set decision model: probabilities are normalized over the candidates supplied by the caller. Provide a suitable fallback explicitly if the application requires one. Order consistency is the training objective; it does not establish that a reported probability equals an empirical correctness rate, and changing candidate order can still change a prediction. The released model is intended for audio-conditioned classification and question answering. Open-ended generation, vision performance, speech generation, and live turn-taking deployment have not been established by this release. Training audio and benchmark answer data are not included. ## License and attribution This model is derived from Qwen2.5-Omni-3B and carries the **Qwen Research License Agreement**, supplied in [LICENSE](LICENSE). The upstream agreement grants use for research or evaluation and requires a separate license from Alibaba Cloud for commercial use. See [Notice](Notice) for attribution and a description of modified files. **Built with Qwen.** The original pretrained weights were modified by AudioJev fine-tuning with RD-SKL, lambda 0.5, seed 20261001. This release does not grant rights beyond the applicable upstream terms.