Audio-Text-to-Text
Transformers
Safetensors
qwen2_5_omni
text-to-audio
audio
audio-question-answering
audio-classification
candidate-scoring
Instructions to use shlv/AudioJev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use shlv/AudioJev with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("shlv/AudioJev") model = AutoModelForMultimodalLM.from_pretrained("shlv/AudioJev", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 4,822 Bytes
25bc587 d506a7b 25bc587 4e0323d 25bc587 d506a7b 4e0323d 25bc587 4e0323d 25bc587 4e0323d 25bc587 4e0323d 25bc587 d506a7b 25bc587 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 | ---
license: other
license_name: qwen-research
license_link: LICENSE
base_model: Qwen/Qwen2.5-Omni-3B
base_model_relation: finetune
library_name: transformers
pipeline_tag: audio-text-to-text
tags:
- audio
- audio-question-answering
- audio-classification
- candidate-scoring
- qwen2_5_omni
- safetensors
---
# AudioJev
**Direct audio decisions with order-calibrated candidate probabilities. Built with Qwen.**
AudioJev takes an audio waveform, a natural-language question, and a list of candidate descriptions. It returns a probability distribution over those candidates using the next-token logits of their position labels. This release contains the **random-derangement SKL (RD-SKL) model with lambda = 0.5 and training seed = 20261001**, fine-tuned from [Qwen2.5-Omni-3B](https://huggingface.co/Qwen/Qwen2.5-Omni-3B).
## Checkpoint
| Setting | Value |
|---|---|
| Model name | AudioJev |
| Base model | Qwen/Qwen2.5-Omni-3B |
| Training seed | 20261001 |
| SKL weight | 0.5 |
| Training schedule | Semantic: 1,024 updates; joint: 1,024; general: 2,048 |
| Total optimizer updates | 4,096 |
| Released checkpoint | Final general-stage step 2,048 |
| Checkpoint selection | Fixed final step |
| Training objective | Mean of two candidate cross-entropies + 0.5 × aligned mean symmetric KL |
| Weight format | Original FP32 Safetensors shards, approximately 18.8 GB |
| Recommended inference dtype | BF16 |
| Candidate labels | `0`–`9`, then `A`–`Z` (2–36 candidates) |
The audio encoder and language-model decision path were trained without a LoRA adapter. The visual branch was excluded from training; audio generation is disabled. This repository contains the complete saved Thinker checkpoint and its processor/tokenizer files. The weight shards are byte-identical to the evaluated checkpoint.
For each training example, RD-SKL supervises the original candidate order and a random derangement. The second distribution is aligned by candidate identity before applying symmetric KL. Inference uses one forward pass in the supplied candidate order.
## Inference
Python inference, the HTTP service, installation dependencies, and runnable examples are maintained in the standalone **AudioJev-Inference** project. This model repository distributes the model weights, tokenizer/processor assets, model card, and license files.
After installing AudioJev-Inference, start the service with:
```bash
audiojev-serve --model shlv/AudioJev --device cuda:0
```
The Python API uses `AudioJev("shlv/AudioJev")`. Follow the AudioJev-Inference README for authentication, local/offline model loading, API requests, and deployment. The service downloads the model from this repository; it does not require a training-project checkout.
AudioJev decisions use candidate-label logits at the final input position, normalized over the supplied candidates. The service documents its candidate-order convention separately from the original-order evaluation below.
## Evaluation
The following scores are for **this single checkpoint in the original candidate order**, not the mean over three training seeds.
| Benchmark | Questions | Accuracy |
|---|---:|---:|
| MMAU full (1,000 public + 9,000 officially scored hidden questions) | 10,000 | 69.41% |
| MMAR | 1,000 | 56.30% |
| MMSU evaluated short-audio subset | 3,931 | 62.45% |
The experiment's source audit found 119 hidden MMAU questions with audio-source overlap against general-stage fitting/development data; the official full score retains all questions. That score should be interpreted with this overlap in mind.
## Scope
AudioJev is a closed-set decision model: probabilities are normalized over the candidates supplied by the caller. Provide a suitable fallback explicitly if the application requires one. Order consistency is the training objective; it does not establish that a reported probability equals an empirical correctness rate, and changing candidate order can still change a prediction.
The released model is intended for audio-conditioned classification and question answering. Open-ended generation, vision performance, speech generation, and live turn-taking deployment have not been established by this release. Training audio and benchmark answer data are not included.
## License and attribution
This model is derived from Qwen2.5-Omni-3B and carries the **Qwen Research License Agreement**, supplied in [LICENSE](LICENSE). The upstream agreement grants use for research or evaluation and requires a separate license from Alibaba Cloud for commercial use. See [Notice](Notice) for attribution and a description of modified files.
**Built with Qwen.** The original pretrained weights were modified by AudioJev fine-tuning with RD-SKL, lambda 0.5, seed 20261001. This release does not grant rights beyond the applicable upstream terms.
|