File size: 4,822 Bytes
25bc587
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d506a7b
25bc587
 
 
4e0323d
25bc587
d506a7b
4e0323d
 
25bc587
 
4e0323d
25bc587
 
4e0323d
25bc587
4e0323d
25bc587
 
 
 
 
 
 
 
 
 
 
d506a7b
25bc587
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
---
license: other
license_name: qwen-research
license_link: LICENSE
base_model: Qwen/Qwen2.5-Omni-3B
base_model_relation: finetune
library_name: transformers
pipeline_tag: audio-text-to-text
tags:
- audio
- audio-question-answering
- audio-classification
- candidate-scoring
- qwen2_5_omni
- safetensors
---

# AudioJev

**Direct audio decisions with order-calibrated candidate probabilities. Built with Qwen.**

AudioJev takes an audio waveform, a natural-language question, and a list of candidate descriptions. It returns a probability distribution over those candidates using the next-token logits of their position labels. This release contains the **random-derangement SKL (RD-SKL) model with lambda = 0.5 and training seed = 20261001**, fine-tuned from [Qwen2.5-Omni-3B](https://huggingface.co/Qwen/Qwen2.5-Omni-3B).

## Checkpoint

| Setting | Value |
|---|---|
| Model name | AudioJev |
| Base model | Qwen/Qwen2.5-Omni-3B |
| Training seed | 20261001 |
| SKL weight | 0.5 |
| Training schedule | Semantic: 1,024 updates; joint: 1,024; general: 2,048 |
| Total optimizer updates | 4,096 |
| Released checkpoint | Final general-stage step 2,048 |
| Checkpoint selection | Fixed final step |
| Training objective | Mean of two candidate cross-entropies + 0.5 × aligned mean symmetric KL |
| Weight format | Original FP32 Safetensors shards, approximately 18.8 GB |
| Recommended inference dtype | BF16 |
| Candidate labels | `0`–`9`, then `A`–`Z` (2–36 candidates) |

The audio encoder and language-model decision path were trained without a LoRA adapter. The visual branch was excluded from training; audio generation is disabled. This repository contains the complete saved Thinker checkpoint and its processor/tokenizer files. The weight shards are byte-identical to the evaluated checkpoint.

For each training example, RD-SKL supervises the original candidate order and a random derangement. The second distribution is aligned by candidate identity before applying symmetric KL. Inference uses one forward pass in the supplied candidate order.

## Inference

Python inference, the HTTP service, installation dependencies, and runnable examples are maintained in the standalone **AudioJev-Inference** project. This model repository distributes the model weights, tokenizer/processor assets, model card, and license files.

After installing AudioJev-Inference, start the service with:

```bash
audiojev-serve --model shlv/AudioJev --device cuda:0
```

The Python API uses `AudioJev("shlv/AudioJev")`. Follow the AudioJev-Inference README for authentication, local/offline model loading, API requests, and deployment. The service downloads the model from this repository; it does not require a training-project checkout.

AudioJev decisions use candidate-label logits at the final input position, normalized over the supplied candidates. The service documents its candidate-order convention separately from the original-order evaluation below.

## Evaluation

The following scores are for **this single checkpoint in the original candidate order**, not the mean over three training seeds.

| Benchmark | Questions | Accuracy |
|---|---:|---:|
| MMAU full (1,000 public + 9,000 officially scored hidden questions) | 10,000 | 69.41% |
| MMAR | 1,000 | 56.30% |
| MMSU evaluated short-audio subset | 3,931 | 62.45% |

The experiment's source audit found 119 hidden MMAU questions with audio-source overlap against general-stage fitting/development data; the official full score retains all questions. That score should be interpreted with this overlap in mind.

## Scope

AudioJev is a closed-set decision model: probabilities are normalized over the candidates supplied by the caller. Provide a suitable fallback explicitly if the application requires one. Order consistency is the training objective; it does not establish that a reported probability equals an empirical correctness rate, and changing candidate order can still change a prediction.

The released model is intended for audio-conditioned classification and question answering. Open-ended generation, vision performance, speech generation, and live turn-taking deployment have not been established by this release. Training audio and benchmark answer data are not included.

## License and attribution

This model is derived from Qwen2.5-Omni-3B and carries the **Qwen Research License Agreement**, supplied in [LICENSE](LICENSE). The upstream agreement grants use for research or evaluation and requires a separate license from Alibaba Cloud for commercial use. See [Notice](Notice) for attribution and a description of modified files.

**Built with Qwen.** The original pretrained weights were modified by AudioJev fine-tuning with RD-SKL, lambda 0.5, seed 20261001. This release does not grant rights beyond the applicable upstream terms.